Petroleum drilling big data processing method and system

Through wavelet transformation and machine learning methods, the noise interference of sensor signals in oil drilling is identified and eliminated, and data anomalies caused by extreme downhole environments are solved, and the reliability and decision-making efficiency of drilling data are improved.

CN120561689AInactive Publication Date: 2025-08-29HEBEI PETROLEUM VOCATIONAL & TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510690771.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-08-29
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

During oil drilling, the extreme underground environment causes sensor package to rupture, circuit aging or signal drift, resulting in data loss, jumping or noise interference, affecting drilling safety and efficiency.

Method used

The high-frequency signal and low-frequency signal are separated by wavelet transformation, the timing/frequency domain correlation between sensor signals and environmental parameters are analyzed, the abnormal type is judged using the random forest/LSTM model, and calibration instructions or shutdown checks are triggered based on the confidence and hazard level to optimize the data correction strategy.

Benefits of technology

Effectively identify and eliminate noise interference, improve data accuracy, avoid safety accidents, and improve drilling decision-making accuracy and system adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561689A_ABST
    Figure CN120561689A_ABST
Patent Text Reader

Abstract

The invention discloses a petroleum drilling big data processing method and system, and belongs to the field of data mining, and the method comprises the following steps: carrying out the wavelet transformation of a petroleum drilling original signal, separating a high-frequency signal from a low-frequency signal, extracting a characteristic frequency band energy ratio, and judging whether noise interference detection exists or not; synchronously analyzing the time sequence / frequency domain correlation between the sensor signal and the environmental parameter, and judging an abnormal source based on a preset threshold value; training a random forest / LSTM model based on historical data, inputting signal statistics and environmental parameters, and outputting an anomaly type; compared with the prior art, the method has the beneficial effects that whether data abnormity is caused by environmental parameters or not is judged by analyzing the time sequence / frequency domain correlation between sensor signals and the environmental parameters, a calibration instruction or shutdown check is triggered when the data abnormity occurs, and a subsequent data correction strategy is synchronously optimized; and safety accidents caused by abnormal data are avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data mining, and in particular relates to a method and system for processing oil drilling big data. Background Art

[0002] Oil drilling data acquisition relies on real-time monitoring of key parameters such as downhole temperature, pressure, drilling speed, and torque using various sensors, providing a basis for decision-making regarding drilling safety and efficiency. However, extreme downhole environments (such as temperatures exceeding 200°C, pressures exceeding 100 MPa, H₂S corrosion, and mechanical vibration) can easily lead to sensor package cracking, circuit aging, or signal drift, resulting in data loss, data jumps, or noise interference. For example, piezoelectric sensors can generate harmonic distortion in vibrating environments, while thermocouples are prone to zero drift in high-temperature gradients.

[0003] Abnormal data in oil drilling may lead to misjudgment of drilling parameters, causing safety accidents such as blowouts and formation leakage, and at the same time interfere with formation pressure assessment and drilling fluid density design. In intelligent drilling systems, abnormal data will also affect the accuracy of automated decision-making and increase the risk of unplanned drilling stoppages, which requires improvement. Summary of the Invention

[0004] Based on this, it is necessary to provide a method and system for processing oil drilling big data to address the above problems.

[0005] The embodiment of the present invention is implemented as follows: a method for processing oil drilling big data, the method comprising the following steps:

[0006] Perform wavelet transform on the original oil drilling signal to separate high-frequency signals (20-50g corresponds to the frequency band: 50Hz-200Hz, and the noise signal also falls in this frequency band) and low-frequency signals (pressure / temperature band: 0-10Hz). Then, the energy proportion of the characteristic frequency band is extracted to determine whether there is noise interference detection.

[0007] Synchronously analyze the time series / frequency domain correlation between sensor signals (high-frequency vibration noise signals and low-frequency effective signals) and environmental parameters (mechanical stress, H2S concentration, EMI spectrum), and determine the root cause of the anomaly (deformation / corrosion / electromagnetic interference) based on preset thresholds (such as stress-drift slope, corrosion-signal distortion, and EMI-noise energy);

[0008] Based on historical data, a random forest / LSTM model is trained, which inputs signal statistics (mean, variance, kurtosis) and environmental parameters (H2S concentration, vibration RMS), and outputs anomaly types (corrosion / vibration / EMI).

[0009] Based on the anomaly confidence level (e.g., probability >90%) and hazard level (e.g., pressure drift rate), a calibration instruction or shutdown inspection is triggered, and subsequent data correction strategies are simultaneously optimized.

[0010] In one embodiment, the present invention provides a method for processing oil drilling big data, wherein the step of performing wavelet transform on the original oil drilling signal, separating the high-frequency signal from the low-frequency signal, extracting the energy proportion of the characteristic frequency band, and determining whether there is noise interference detection specifically includes:

[0011] According to the non-stationary characteristics of the drilling signal, select the original Daubechies wavelet signal with excellent time-frequency localization characteristics (such as db6), determine the number of decomposition layers (such as 5 layers), and decompose the original signal step by step into high-frequency (detail coefficients) and low-frequency (approximate coefficients) according to the frequency band;

[0012] Reconstruct the high-frequency band (50-200Hz, corresponding to detail coefficients D1-D3) and the low-frequency band (0-10Hz, corresponding to approximation coefficient A5), and calculate the signal energy of each frequency band (energy = sum of squares of signal amplitudes);

[0013] The low-frequency energy ratio under historical normal operating conditions (e.g., >70%) is counted, and the high-frequency and low-frequency energy ratio is calculated in real time. If the high-frequency energy ratio suddenly increases (e.g., >60%) and continuously exceeds the threshold, it is determined that vibration noise has flooded the valid signal area (e.g., time period t1-t2), interfering with detection.

[0014] In one embodiment, the present invention provides a method for processing oil drilling big data, wherein the step of synchronously analyzing the time series / frequency domain correlation between sensor signals and environmental parameters and determining the root cause of the abnormality based on a preset threshold specifically includes:

[0015] Calculate the mean / variance using a sliding window to establish a normal fluctuation range for the pressure / temperature signal (e.g., ±5%). If the signal continuously deviates from the baseline and exceeds the threshold (e.g., >10%), it is determined to be a drift anomaly caused by package deformation.

[0016] Compare the correlation between pH value, H2S concentration and electrode response current. If the current drop or nonlinear fluctuation is positively correlated with the concentration of the corrosive product, it is determined that the data anomaly is caused by electrode corrosion.

[0017] Collect electromagnetic radiation intensity (e.g., 30MHz-1GHz) and sensor signal spectrum. If the noise peak coincides with the electromagnetic interference frequency band (e.g., harmonics from a variable-frequency drilling rig), it is determined that the data anomaly is caused by electromagnetic interference (EMI). Adaptive notch filtering can be used to eliminate interference.

[0018] In one embodiment, the present invention provides a method for processing oil drilling big data, wherein the steps of training a random forest / LSTM model based on historical data, inputting signal statistics and environmental parameters, and outputting anomaly types specifically include:

[0019] Filter valid samples from historical data, remove periods of communication interruption or sensor failure, align signal statistics (mean, variance, kurtosis) with environmental parameters (H2S concentration, vibration RMS), and construct a dataset labeled with anomaly types (corrosion / vibration / EMI);

[0020] For time series data (such as LSTM) or static features (such as random forest), the dataset is divided into training, validation, and test sets, and parameters are optimized (number of LSTM hidden layer nodes, random forest tree depth), using K-fold cross-validation to prevent overfitting;

[0021] Based on the random forest / LSTM model, the system outputs multiple classification probabilities (such as corrosion 80%, EMI 15%, and vibration 5%) and sets a confidence threshold (such as the highest category probability > 70%). When the confidence threshold is reached, the system outputs the anomaly type. If the probability distribution is dispersed (such as corrosion 45%, EMI 40%) and the confidence threshold is not reached, the system outputs a multi-factor coupling anomaly.

[0022] In one embodiment, the present invention provides a method for processing oil drilling big data, wherein the step of synchronously analyzing the time series / frequency domain correlation between sensor signals and environmental parameters and determining the root cause of the abnormality based on a preset threshold further includes:

[0023] Calculate the residual (e.g., difference) between multiple sensor data with the same parameters. If the residual exceeds a threshold (e.g., ±5% of the range), mark it as an abnormal node.

[0024] For missing / abnormal parameters (such as H2S concentration), Kalman filtering or LSTM time series prediction is used to generate alternative values ​​and annotate confidence levels. The corrosion rate model and stress-strain equation are combined to constrain the rationality of the infilled values ​​(for example, H2S concentration cannot be negative).

[0025] Periodically send self-test instructions (such as electrode impedance test and pressure zero point calibration) to diagnose the sensor hardware status; record the sensor's cumulative working time and the number of extreme working conditions, predict the sensor's remaining life based on the Weibull distribution model, and remind you to replace it in advance.

[0026] In one embodiment, the present invention provides a petroleum drilling big data processing system, comprising:

[0027] The interference elimination module is used to perform wavelet transform on the original oil drilling signal, separate the high-frequency signal (20-50g corresponds to the frequency band: 50Hz-200Hz, and the noise signal also falls in this frequency band) and the low-frequency signal (pressure / temperature band: 0-10Hz), extract the energy proportion of the characteristic frequency band, and determine whether there is noise interference detection;

[0028] The abnormality root cause judgment module is used to synchronously analyze the time series / frequency domain correlation between sensor signals (high-frequency vibration noise signals and low-frequency effective signals) and environmental parameters (mechanical stress, H2S concentration, EMI spectrum), and determine the abnormality root cause (deformation / corrosion / electromagnetic interference) based on preset thresholds (such as stress-drift slope, corrosion-signal distortion, and EMI-noise energy);

[0029] Anomaly type output module, used to train random forest / LSTM models based on historical data, inputs signal statistics (mean, variance, kurtosis) and environmental parameters (H2S concentration, vibration RMS), and outputs anomaly types (corrosion / vibration / EMI);

[0030] The exception handling module is used to trigger calibration instructions or shutdown inspections based on the exception confidence level (such as probability >90%) and hazard level (such as pressure drift rate), and simultaneously optimize subsequent data correction strategies.

[0031] In one embodiment, the present invention provides a petroleum drilling big data processing system, wherein the interference elimination module includes:

[0032] The frequency band decomposition unit is used to select the original Daubechies wavelet signal with excellent time-frequency localization characteristics (such as db6) according to the non-stationary characteristics of the drilling signal, determine the number of decomposition layers (such as 5 layers), and decompose the original signal into high-frequency (detail coefficients) and low-frequency (approximate coefficients) step by step according to the frequency band;

[0033] The high-frequency and low-frequency reconstruction unit is used to reconstruct the high-frequency band (50-200Hz, corresponding to detail coefficients D1-D3) and the low-frequency band (0-10Hz, corresponding to approximation coefficient A5), and calculate the signal energy of each frequency band (energy = sum of squares of signal amplitudes);

[0034] The interference judgment unit is used to count the low-frequency energy proportion under historical normal operating conditions (e.g., >70%) and calculate the high-frequency and low-frequency energy ratio in real time. If the high-frequency energy proportion suddenly increases (e.g., >60%) and continuously exceeds the threshold, it is determined that vibration noise has submerged the valid signal area (e.g., time period t1-t2), and interference detection is initiated.

[0035] In one embodiment, the present invention provides a petroleum drilling big data processing system, wherein the abnormality root cause determination module includes:

[0036] The package deformation determination unit is used to calculate the mean / variance using a sliding window to establish a normal fluctuation range for the pressure / temperature signal (e.g., ±5%). If the signal continuously deviates from the baseline and exceeds a threshold (e.g., >10%), it is determined to be a drift anomaly caused by package deformation.

[0037] The corrosion determination unit is used to compare the correlation between pH value, H2S concentration and electrode response current. If the current sudden drop or nonlinear fluctuation is positively correlated with the concentration of corrosive substances, it is determined that the data anomaly is caused by electrode corrosion.

[0038] The electromagnetic interference determination unit is used to collect electromagnetic radiation intensity (such as 30MHz-1GHz) and sensor signal spectrum. If the noise peak coincides with the electromagnetic interference frequency band (such as the harmonics of a variable-frequency drilling rig), it is determined to be a data anomaly caused by electromagnetic interference (EMI). Adaptive notch filtering can be used to eliminate interference.

[0039] In one embodiment, the present invention provides a petroleum drilling big data processing system, wherein the abnormality type output module includes:

[0040] The training set construction unit is used to filter valid samples from historical data, eliminate periods of communication interruption or sensor failure, align signal statistics (mean, variance, kurtosis) with environmental parameters (H2S concentration, vibration RMS), and construct a data set labeled with anomaly types (corrosion / vibration / EMI);

[0041] The training set partitioning unit is used to divide the dataset into training, validation, and test sets for time series data (such as LSTM) or static features (such as random forest), optimize parameters (such as the number of hidden layer nodes in LSTM and the depth of random forest trees), and use K-fold cross-validation to prevent overfitting.

[0042] The judgment output unit is used to output multi-classification probabilities (such as corrosion 80%, EMI 15%, and vibration 5%) based on the random forest / LSTM model. It sets a confidence threshold (such as the highest category probability > 70%). When the confidence threshold is reached, the abnormality type is output. If the probability distribution is dispersed (such as corrosion 45%, EMI 40%), and the confidence threshold is not reached, a multi-factor coupling abnormality is output.

[0043] In one embodiment, the present invention provides a petroleum drilling big data processing system, wherein the abnormality root cause determination module further includes:

[0044] The abnormal data detection unit is used to calculate the residual (such as the difference) between the data of multiple sensors with the same parameters. If the residual exceeds the threshold (such as ±5% of the range), it is marked as an abnormal node;

[0045] Abnormal data replaces units. For missing / abnormal parameters (such as H2S concentration), Kalman filtering or LSTM time series prediction is used to generate replacement values ​​and annotate confidence levels. Combined with the corrosion rate model and stress-strain equation, the rationality of the filled-in values ​​is constrained (for example, H2S concentration cannot be negative).

[0046] The abnormal advance reminder unit is used to periodically send self-test instructions (such as electrode impedance test and pressure zero point calibration) to diagnose the sensor hardware status; record the sensor's cumulative working time and the number of extreme working conditions, predict the remaining life of the sensor based on the Weibull distribution model, and remind you to replace it in advance.

[0047] Compared with the existing technology, the beneficial effect of the present invention is: the present invention determines whether the data anomaly is caused by environmental parameters by analyzing the time / frequency domain correlation between sensor signals and environmental parameters, and triggers calibration instructions or shutdown inspections when the data is abnormal, and simultaneously optimizes subsequent data correction strategies to avoid safety accidents caused by abnormal data. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 A flow chart of a method for processing oil drilling big data provided by an embodiment of the present invention.

[0049] Figure 2 A schematic diagram of a process for noise interference detection and judgment provided by an embodiment of the present invention.

[0050] Figure 3 A schematic diagram of a process for determining the root cause of data anomalies provided by an embodiment of the present invention.

[0051] Figure 4 A schematic diagram of the process of outputting exception types provided by an embodiment of the present invention.

[0052] Figure 5 A schematic diagram of the process of replacing abnormal parameters provided by an embodiment of the present invention.

[0053] Figure 6 A schematic diagram of an oil drilling big data processing system provided by an embodiment of the present invention.

[0054] Figure 7 A schematic diagram of an interference elimination module provided in an embodiment of the present invention.

[0055] Figure 8 This is a schematic diagram of the first part of the abnormality root cause determination module provided in an embodiment of the present invention.

[0056] Figure 9 A schematic diagram of an exception type output module provided in an embodiment of the present invention.

[0057] Figure 10 This is a schematic diagram of the second part of the abnormality root cause determination module provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0058] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0059] In one embodiment, Figure 1 As shown, a method for processing oil drilling big data includes the following steps:

[0060] Step S1: Perform wavelet transform on the original oil drilling signal to separate the high-frequency signal (20-50g, where g is vibration acceleration, corresponding to the frequency band: 50Hz-200Hz, and noise signals also fall within this frequency band) from the low-frequency signal (pressure / temperature band: 0-10Hz). The energy proportion of the characteristic frequency band is extracted to determine whether there is noise interference detection.

[0061] Step S2: Synchronously analyze the time series / frequency domain correlation between sensor signals (high-frequency vibration noise signals and low-frequency effective signals) and environmental parameters (mechanical stress, H2S concentration, EMI spectrum), and determine the root cause of the anomaly (deformation / corrosion / electromagnetic interference) based on preset thresholds (such as stress-drift slope, corrosion-signal distortion, and EMI-noise energy).

[0062] Step S3: Train a random forest / LSTM model based on historical data, input signal statistics (mean, variance, kurtosis), environmental parameters (H2S concentration, vibration RMS), and output anomaly type (corrosion / vibration / EMI).

[0063] In step S4, based on the abnormality confidence (e.g., probability > 90%) and hazard level (e.g., pressure drift rate), a calibration instruction or shutdown inspection is triggered, and subsequent data correction strategies are optimized simultaneously.

[0064] Step S1 leverages the frequency-domain decomposition properties of the wavelet transform to prioritize the separation of high-frequency vibration noise from low-frequency valid signals, addressing the insensitivity of traditional time-domain analysis to non-stationary signals and establishing a clean data foundation for subsequent detection. Step S2 introduces cross-domain correlation analysis between environmental parameters and sensor signals, overcoming the limitations of single-signal threshold detection. An anomaly root cause inference network is constructed through multi-dimensional physical correlations such as stress-drift and corrosion-distortion. Step S3 integrates historical experience with real-time features based on machine learning. Leveraging the interpretability of random forests and the time series modeling capabilities of LSTMs, probabilistic identification of anomaly types under complex operating conditions is achieved. Step S4, using a dual-criteria decision-making mechanism based on confidence and criticality, converts diagnostic results into graded response actions. This not only avoids unplanned downtime caused by misjudgments but also enhances system adaptability through continuous optimization of data correction strategies. These four steps, combined, significantly improve the reliability and decision-making efficiency of condition monitoring in the oil drilling industry.

[0065] In one embodiment, Figure 2 As shown, a method for processing oil drilling big data, said step S1, performing wavelet transform on the original oil drilling signal, separating the high-frequency signal and the low-frequency signal, extracting the energy proportion of the characteristic frequency band, and judging whether there is noise interference detection step, specifically includes:

[0066] Step S11: Based on the non-stationary characteristics of the drilling signal, a Daubechies wavelet original signal with excellent time-frequency localization characteristics (such as db6) is selected, the number of decomposition layers (such as 5 layers) is determined, and the original signal is decomposed step by step into high frequencies (detail coefficients) and low frequencies (approximation coefficients) according to the frequency band.

[0067] Step S12: reconstruct the high frequency band (50-200 Hz, corresponding to detail coefficients D1-D3) and the low frequency band (0-10 Hz, corresponding to approximation coefficient A5), and calculate the signal energy of each frequency band (energy = sum of squares of signal amplitudes);

[0068] Step S13: Count the historical low-frequency energy proportion under normal operating conditions (e.g., >70%), and calculate the high-frequency and low-frequency energy ratio in real time. If the high-frequency energy proportion suddenly increases (e.g., >60%) and continuously exceeds the threshold, it is determined that vibration noise has submerged the valid signal area (e.g., time period t1-t2), and interference detection is performed.

[0069] In step S11, the Daubechies wavelet (such as db6) is selected and the number of decomposition layers (5 layers) is determined. Due to its tight support and vanishing moment characteristics, it can retain the signal mutation characteristics (such as vibration impact) while separating the frequency band boundaries of noise and valid signals through multi-scale decomposition (high frequency D1-D3 corresponds to 50-200Hz, and low frequency A5 corresponds to 0-10Hz). In step S12, the complex time domain signal is converted into quantifiable frequency domain features by reconstructing the high-frequency and low-frequency components and calculating the energy ratio (such as low-frequency energy >70% is a normal baseline), solving the detection problem of effective signal submersion caused by noise energy penetration. In step S13, a dynamic energy threshold judgment criterion is established based on historical statistics (such as high-frequency sudden increase >60%). The suddenness of vibration noise (such as time period t1-t2) is captured through real-time energy ratio monitoring, avoiding the missed detection of transient interference caused by traditional fixed thresholds.

[0070] In one embodiment, Figure 3 As shown, a method for processing oil drilling big data, said step S2, synchronously analyzing the time series / frequency domain correlation between sensor signals and environmental parameters, and determining the root cause of the abnormality based on a preset threshold, specifically includes:

[0071] Step S21: Calculate the mean / variance using a sliding window to establish a normal fluctuation range for the pressure / temperature signal (e.g., ±5%). If the signal continuously deviates from the baseline and exceeds a threshold (e.g., >10%), it is determined to be a drift anomaly caused by package deformation.

[0072] Step S22: comparing the correlation between pH value, H2S concentration and electrode response current. If the current sudden drop or nonlinear fluctuation is positively correlated with the concentration of the corrosive product, it is determined that the data anomaly is caused by electrode corrosion.

[0073] In step S23, electromagnetic radiation intensity (e.g., 30 MHz-1 GHz) and sensor signal spectrum are collected. If the noise peak coincides with the electromagnetic interference frequency band (e.g., harmonics from a variable-frequency drilling rig), it is determined that the data anomaly is caused by electromagnetic interference (EMI). Adaptive notch filtering can be used to eliminate interference.

[0074] Step S21 uses sliding window statistics (mean / variance) to dynamically define the normal fluctuation baseline (e.g., ±5%) of the pressure / temperature signal. By monitoring the signal's continuous deviation from the threshold (e.g., >10%) and combining it with mechanical stress parameters, it locks in the drift anomaly caused by package deformation, solving the problem of insufficient sensitivity of traditional threshold methods to gradual failures (e.g., metal fatigue deformation). Step S22 focuses on the chemical corrosion mechanism. By establishing a nonlinear correlation between H2S concentration and electrode response current (e.g., a sudden current drop and the accumulation of corrosive products change synchronously), it uses the differences in electrochemical characteristics to identify data distortion caused by electrode corrosion, thus avoiding the misjudgment of latent corrosion by single signal analysis. Step S23 targets the frequency domain coupling characteristics of electromagnetic interference. By comparing the electromagnetic radiation spectrum (30MHz-1GHz) with the sensor noise peak distribution, it locates the interference source in a specific frequency band (e.g., variable frequency drilling rig harmonics), combines it with adaptive notch filtering to achieve interference suppression, and breaks through the bottleneck of time domain filtering in suppressing non-stationary EMI.

[0075] In one embodiment, Figure 4 As shown, a method for processing oil drilling big data, said step S3, training a random forest / LSTM model based on historical data, inputting signal statistics and environmental parameters, and outputting anomaly types, specifically includes:

[0076] Step S31: Filter valid samples from historical data, remove periods of communication interruption or sensor failure, align signal statistics (mean, variance, kurtosis) with environmental parameters (H2S concentration, vibration RMS), and construct a dataset labeled with anomaly types (corrosion / vibration / EMI);

[0077] Step S32: For time series data (such as LSTM) or static features (such as random forest), the data set is divided into training set, validation set, and test set, and the parameters are optimized (number of LSTM hidden layer nodes, random forest tree depth), and K-fold cross-validation is used to prevent overfitting;

[0078] In step S33, the random forest / LSTM model outputs multi-classification probabilities (e.g., corrosion 80%, EMI 15%, vibration 5%), sets a confidence threshold (e.g., highest category probability > 70%), and outputs the abnormality type when the confidence threshold is reached. If the probability distribution is dispersed (e.g., corrosion 45%, EMI 40%) and the confidence threshold is not reached, a multi-factor coupling abnormality is output.

[0079] Step S31 constructs a labeled dataset with strong physical correlation by screening valid samples and aligning multi-source data (signal statistics, environmental parameters), solving the problem of model training bias caused by data fragmentation and noisy labels in industrial scenarios; step S32 adopts cross-validation and parameter optimization strategies (such as adjusting the number of LSTM hidden layer nodes) to balance the model's temporal feature capture capability and generalization performance, avoiding overfitting caused by dynamic changes in drilling conditions; step S33 quantifies the uncertainty of the diagnostic results through multi-classification probability output and confidence threshold mechanism (such as the highest category probability >70%), which can not only make quick decisions in the case of a single dominant anomaly (such as corrosion 80%), but also identify multi-factor coupling scenarios (such as corrosion 45%, EMI 40%) to prevent the risk of misjudgment.

[0080] In one embodiment, Figure 5 As shown, a method for processing oil drilling big data, said step S2, synchronously analyzing the time series / frequency domain correlation between sensor signals and environmental parameters, and determining the root cause of the abnormality based on a preset threshold, further includes:

[0081] Step S24, calculating the residual (such as the difference) between the multi-sensor data of the same parameter, and marking it as an abnormal node if the residual exceeds a threshold (such as ±5% of the range);

[0082] Step S25: For missing / abnormal parameters (such as H2S concentration), use Kalman filtering or LSTM time series prediction to generate replacement values ​​and annotate the confidence level. Combined with the corrosion rate model and stress-strain equation, constrain the rationality of the filled-in value (for example, H2S concentration cannot be negative).

[0083] Step S26: Periodically send self-test instructions (such as electrode impedance test and pressure zero point calibration) to diagnose the sensor hardware status; record the sensor's cumulative working time and the number of extreme working conditions, predict the sensor's remaining life based on the Weibull distribution model, and remind you to replace it in advance.

[0084] If the environmental parameter sensor fails or communication interference causes data distortion, the correlation analysis of steps S21, S22, and S23 will draw incorrect conclusions based on incorrect input. This situation needs to be avoided, so steps S24 and S25 are designed; at the same time, step S26 is involved to remind you to replace the sensor in advance to avoid affecting the acquisition of oil drilling data.

[0085] In one embodiment, Figure 6 As shown, a petroleum drilling big data processing system includes:

[0086] Interference elimination module 1 is used to perform wavelet transform on the original oil drilling signal, separate the high-frequency signal (20-50g corresponds to the frequency band: 50Hz-200Hz, and the noise signal also falls in this frequency band) and the low-frequency signal (pressure / temperature band: 0-10Hz), extract the energy proportion of the characteristic frequency band, and determine whether there is noise interference detection;

[0087] Abnormal root cause determination module 2 is used to synchronously analyze the time series / frequency domain correlation between sensor signals (high-frequency vibration noise signals and low-frequency effective signals) and environmental parameters (mechanical stress, H2S concentration, EMI spectrum), and determine the abnormal root cause (deformation / corrosion / electromagnetic interference) based on preset thresholds (such as stress-drift slope, corrosion-signal distortion, and EMI-noise energy);

[0088] Anomaly type output module 3 is used to train a random forest / LSTM model based on historical data. It inputs signal statistics (mean, variance, kurtosis) and environmental parameters (H2S concentration, vibration RMS) and outputs the anomaly type (corrosion / vibration / EMI).

[0089] The exception handling module 4 is used to trigger a calibration instruction or shutdown inspection based on the exception confidence (such as probability > 90%) and hazard level (such as pressure drift rate), and simultaneously optimize the subsequent data correction strategy.

[0090] Abnormal handling module 4 implements abnormal response by building a dual-factor "confidence-criticality" decision engine and a dynamic optimization closed loop. First, the confidence level of the diagnostic result is determined based on the multi-classification probability output by abnormality type output module 3 (e.g., corrosion probability 80%), combined with a preset confidence threshold (e.g., >90%). This is then mapped to a hazard level matrix (e.g., low / medium / high risk) based on real-time parameters (e.g., pressure drift rate, H2S concentration growth rate). When both high confidence (probability >90%) and high risk (e.g., pressure drift rate >5% / min) are met, a shutdown command is automatically triggered and fault location information is delivered. If confidence meets the requirements but the criticality is moderate (e.g., vibration RMS value exceeds the limit but there is no sudden pressure change), online calibration (e.g., zero drift compensation) is initiated. For suspected abnormalities with insufficient confidence (e.g., corrosion probability 45% + EMI 40%), multi-sensor cross-validation and LSTM time series prediction are activated to dynamically correct the judgment results. The feedback mechanism is used to simultaneously record misjudgment cases and correction effects, and reinforcement learning is used to optimize threshold settings (such as adaptive adjustment of hazard level critical values) and data filling strategies (such as weighted fusion Kalman filtering and physical model output).

[0091] In one embodiment, Figure 7 As shown, an oil drilling big data processing system, the interference elimination module 1 includes:

[0092] The frequency band decomposition unit 11 is used to select a Daubechies wavelet original signal (such as db6) with excellent time-frequency localization characteristics based on the non-stationary characteristics of the drilling signal, determine the number of decomposition layers (such as 5 layers), and decompose the original signal into high frequencies (detail coefficients) and low frequencies (approximate coefficients) step by step according to the frequency band.

[0093] The high-frequency and low-frequency reconstruction unit 12 is used to reconstruct the high-frequency band (50-200 Hz, corresponding to detail coefficients D1-D3) and the low-frequency band (0-10 Hz, corresponding to approximation coefficient A5), and calculate the signal energy of each frequency band (energy = sum of squares of signal amplitudes);

[0094] The interference judgment unit 13 is used to count the low-frequency energy proportion under historical normal operating conditions (e.g., >70%) and calculate the high-frequency and low-frequency energy ratio in real time. If the high-frequency energy proportion suddenly increases (e.g., >60%) and continuously exceeds the threshold, it is determined that the vibration noise has submerged the effective signal area (e.g., time period t1-t2), and interference detection is performed.

[0095] The number of decomposition layers of the frequency band decomposition unit 11 is not fixed. For example, according to the difference in high-frequency noise intensity, the number of decomposition layers can be adjusted to 4-6 layers for dynamic switching to avoid frequency band aliasing or information loss caused by the preset number of layers, thereby improving the frequency band separation accuracy and noise suppression effect.

[0096] Interference judgment unit 13 relies solely on a single threshold (e.g., >60%) for high-frequency energy percentage to determine noise interference, making it susceptible to interference from transient pulses or non-vibration noise (e.g., electromagnetic transients). Electromagnetic interference is an interference item in abnormality root cause judgment module 2, leading to misjudgments. Therefore, interference judgment unit 13 can further integrate multi-dimensional feature fusion criteria: ① Introducing time-domain statistics (e.g., high-frequency kurtosis to detect impulsive noise); ② Combining energy gradients of adjacent frequency bands (e.g., the D3 / D4 energy ratio to distinguish between continuous vibration and transient interference); and ③ Designing a threshold for the energy change rate within a sliding window (e.g., an increase >30% within 10 seconds). By combining time-frequency domain features with dynamic trend analysis, the misjudgment rate of a single indicator can be reduced, enhancing robustness in complex noise scenarios.

[0097] In one embodiment, Figure 8 As shown, in an oil drilling big data processing system, the abnormality root cause judgment module 2 includes:

[0098] The package deformation determination unit 21 is used to calculate the mean / variance through a sliding window to establish a normal fluctuation range of the pressure / temperature signal (e.g., ±5%). If the signal continuously deviates from the baseline and exceeds a threshold (e.g., >10%), it is determined to be a drift anomaly caused by package deformation.

[0099] The corrosion determination unit 22 is used to compare the correlation between pH value, H2S concentration and electrode response current. If the current sudden drop or nonlinear fluctuation is positively correlated with the concentration of corrosive substances, it is determined that the data anomaly is caused by electrode corrosion;

[0100] The electromagnetic interference determination unit 23 is used to collect electromagnetic radiation intensity (e.g., 30 MHz-1 GHz) and sensor signal spectrum. If the noise peak coincides with the electromagnetic interference frequency band (e.g., harmonics of a variable-frequency drilling rig), it is determined to be a data anomaly caused by electromagnetic interference (EMI). Adaptive notch filtering can be used to eliminate interference.

[0101] The package deformation determination unit 21 uses a fixed threshold (e.g., ±5%) to determine signal drift, without considering the dynamic coupling effects of parameters such as mechanical stress and temperature. A multi-parameter joint probability model (e.g., a copula function) can be further constructed to dynamically adjust the threshold based on the coordinated changes in stress, temperature, and drift rate. For example, metal expansion under high-temperature conditions can cause baseline offset, necessitating the introduction of a temperature compensation factor (e.g., threshold = 5% × (1 + 0.02ΔT)). Simultaneously, stress sensor data should be combined to verify that the drift is consistent with the physical laws of deformation. This can prevent misjudgment of complex failures by a single signal threshold and improve the adaptability of deformation detection to different operating conditions.

[0102] In one embodiment, Figure 9 As shown, in a petroleum drilling big data processing system, the abnormality type output module 3 includes:

[0103] The training set construction unit 31 is used to filter valid samples from historical data, eliminate periods of communication interruption or sensor failure, align signal statistics (mean, variance, kurtosis) with environmental parameters (H2S concentration, vibration RMS), and construct a data set labeled with anomaly types (corrosion / vibration / EMI);

[0104] The training set partitioning unit 32 is used to divide the data set into training set, validation set, and test set for time series data (such as LSTM) or static features (such as random forest), optimize parameters (such as the number of hidden layer nodes of LSTM and the depth of random forest trees), and use K-fold cross-validation to prevent overfitting;

[0105] The judgment output unit 33 is used to output multi-classification probabilities (such as corrosion 80%, EMI 15%, and vibration 5%) based on the random forest / LSTM model, set a confidence threshold (such as the highest category probability > 70%), and output the abnormality type when the confidence threshold is reached; if the probability distribution is dispersed (such as corrosion 45%, EMI 40%), and the confidence threshold is not reached, a multi-factor coupling abnormality is output.

[0106] Judgment output unit 33 only outputs the probability of anomaly type and lacks a physical explanation of the diagnostic results. An interpretable AI module could be added: ① Using SHAP values ​​(Shapley Additive Explanations) to quantify the contribution of signal statistics (such as kurtosis) and environmental parameters (H2S concentration) to the classification results; ② Integrating domain knowledge bases (such as corrosion rate equations and EMI propagation models) to generate a multi-level root cause chain (e.g., H2S concentration > 50 ppm → electrode corrosion probability increases → current nonlinearity increases). Outputting visual reports (such as decision path diagrams and key parameter time series comparisons) can help users quickly locate the source of the fault and verify the logical rationality, thereby enhancing the credibility of human-machine collaborative decision-making and the operational guidance value.

[0107] In one embodiment, Figure 10 As shown, in a petroleum drilling big data processing system, the abnormality root cause judgment module 2 also includes:

[0108] An abnormal data detection unit 24 is used to calculate the residual (such as the difference) between the data of multiple sensors with the same parameters. If the residual exceeds a threshold (such as ±5% of the range), it is marked as an abnormal node;

[0109] Abnormal data replacement unit 25: For missing / abnormal parameters (such as H2S concentration), use Kalman filtering or LSTM time series prediction to generate replacement values ​​and annotate confidence levels. Combined with the corrosion rate model and stress-strain equation, the rationality of the filled-in values ​​is constrained (for example, H2S concentration cannot be negative).

[0110] The abnormal advance warning unit 26 is used to periodically send self-test instructions (such as electrode impedance test, pressure zero point calibration) to diagnose the sensor hardware status; record the sensor's cumulative working time and the number of extreme working conditions, predict the remaining life of the sensor based on the Weibull distribution model, and remind you to replace it in advance.

[0111] The abnormal early warning unit 26 integrates the sensor's cumulative operating time, the frequency of extreme operating conditions (such as overtemperature and overvibration), and real-time health indicators (such as electrode impedance growth rate and zero-point drift) to construct a Weibull distribution parameter estimation model. Using maximum likelihood or Bayesian inference, the shape parameter β (reflecting the failure mode; for example, β < 1 indicates early failure) and the scale parameter η (characteristic life) are calibrated based on historical failure data. Next, the real-time operating load spectrum (such as rain flow counts of the vibration root mean square value) is introduced to calculate the cumulative damage degree. The η value is dynamically corrected using the Miner linear damage rule. For example, for every 10% increase in high-frequency vibration cumulative damage, η is reduced to 90% of its original value. Furthermore, cross-platform degradation data from similar sensors is integrated (transfer learning) to optimize parameter estimation bias under small sample sizes. During prediction, the remaining reliability R(t) is calculated using the Weibull distribution reliability function. When R(t) is lower than a set threshold (e.g., 5%) or the predicted remaining life, such as the mean time to failure (MTTF) = ηΓ (1+1 / β), is less than the maintenance cycle, Γ is a gamma function, triggering a graded warning (e.g., a yellow warning with 30 days remaining and a red warning with 7 days remaining), the spare parts inventory system is linked to automatically generate a replacement work order, achieving a seamless closed loop of life prediction and maintenance decision-making.

[0112] It should be understood that, although the various steps in the flow chart of each embodiment of the present invention are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence according to the order indicated by the arrows. Unless otherwise specified herein, the execution of these steps is not strictly limited in order, and these steps can be performed in other orders. Moreover, at least a portion of the steps in each embodiment may include a plurality of sub-steps or a plurality of stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these sub-steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0113] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0114] The above-described embodiments merely illustrate several implementations of the present invention, and while their descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make numerous variations and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.

[0115] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

[0116] In addition, it should be understood that although this specification is described in terms of implementation methods, not every implementation method contains only one independent technical solution. This narrative method of the specification is only for the sake of clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other implementation methods that can be understood by those skilled in the art.

Claims

1. A method for processing oil drilling big data, characterized in that: The oil drilling big data processing method includes the following steps: Perform wavelet transform on the original oil drilling signal to separate high-frequency signals from low-frequency signals, extract the energy proportion of characteristic frequency bands, and determine whether there is noise interference detection; Synchronously analyze the time series / frequency domain correlation between sensor signals and environmental parameters, and determine the root cause of abnormalities based on preset thresholds; Train a random forest / LSTM model based on historical data, input signal statistics and environmental parameters, and output anomaly types; Based on the anomaly confidence and hazard level, calibration instructions or shutdown inspections are triggered, and subsequent data correction strategies are optimized simultaneously.

2. The oil drilling big data processing method according to claim 1, characterized in that: The step of performing wavelet transform on the original oil drilling signal, separating the high-frequency signal from the low-frequency signal, extracting the energy proportion of the characteristic frequency band, and determining whether there is noise interference detection specifically includes: According to the non-stationary characteristics of drilling signals, the original Daubechies wavelet signal with excellent time-frequency localization characteristics is selected, the number of decomposition layers is determined, and the original signal is decomposed into high frequency and low frequency levels according to the frequency band. Reconstruct the high frequency band and the low frequency band, and calculate the signal energy of each frequency band respectively; Statistics are collected on the historical low-frequency energy ratio under normal operating conditions, and the high-frequency and low-frequency energy ratios are calculated in real time. If the high-frequency energy ratio suddenly increases and continuously exceeds the threshold, it is determined that vibration noise has submerged the effective signal area, interfering with detection.

3. The oil drilling big data processing method according to claim 1, characterized in that: The step of synchronously analyzing the time series / frequency domain correlation between the sensor signal and the environmental parameter and determining the root cause of the abnormality based on a preset threshold specifically includes: The mean / variance is calculated using a sliding window to establish the normal fluctuation range of the pressure / temperature signal. If the signal continuously deviates from the baseline and exceeds the threshold, it is determined to be a drift anomaly caused by package deformation. Compare the correlation between pH value, H2S concentration and electrode response current. If the current drop or nonlinear fluctuation is positively correlated with the concentration of the corrosive product, it is determined that the data anomaly is caused by electrode corrosion. The electromagnetic radiation intensity and sensor signal spectrum are collected. If the noise peak coincides with the electromagnetic interference frequency band, it is determined that the data is abnormal due to electromagnetic interference.

4. The oil drilling big data processing method according to claim 1, characterized in that: The steps of training the random forest / LSTM model based on historical data, inputting signal statistics and environmental parameters, and outputting anomaly types specifically include: Filter valid samples from historical data, eliminate periods of communication interruption or sensor failure, align signal statistics with environmental parameters, and construct a dataset with annotated anomaly types. For time series data or static features, the dataset is divided into training set, validation set, and test set, and parameters are optimized. K-fold cross-validation is used to prevent overfitting. Based on the random forest / LSTM model, the system outputs multi-classification probabilities and sets a confidence threshold. When the confidence threshold is reached, the system outputs the abnormality type. If the probability distribution is dispersed and the confidence threshold is not reached, the system outputs a multi-factor coupling abnormality.

5. The oil drilling big data processing method according to claim 1, characterized in that: The step of synchronously analyzing the time series / frequency domain correlation between the sensor signal and the environmental parameter and determining the root cause of the abnormality based on a preset threshold value further includes: Calculate the residual between multiple sensor data with the same parameters. If the residual exceeds the threshold, mark it as an abnormal node; For missing / abnormal parameters, use Kalman filtering or LSTM time series prediction to generate alternative values ​​and annotate the confidence level; combine the corrosion rate model and stress-strain equation to constrain the rationality of the filled-in values; Periodically send self-test instructions to diagnose the sensor hardware status; record the sensor's cumulative working time and the number of extreme working conditions, predict the sensor's remaining life based on the Weibull distribution model, and remind you to replace it in advance.

6. A big data processing system for oil drilling, characterized in that: The oil drilling big data processing system includes: Interference elimination module, used to perform wavelet transform on the original oil drilling signal, separate high-frequency signals from low-frequency signals, extract the energy proportion of characteristic frequency bands, and determine whether there is noise interference detection; The abnormality root cause judgment module is used to synchronously analyze the time series / frequency domain correlation between sensor signals and environmental parameters, and determine the abnormality root cause based on the preset threshold; The anomaly type output module is used to train the random forest / LSTM model based on historical data, input signal statistics and environmental parameters, and output the anomaly type; The exception handling module is used to trigger calibration instructions or shutdown inspections based on the exception confidence and hazard level, and simultaneously optimize subsequent data correction strategies.

7. The oil drilling big data processing system according to claim 6, characterized in that: The interference elimination module includes: The frequency band decomposition unit is used to select the original Daubechies wavelet signal with excellent time-frequency localization characteristics according to the non-stationary characteristics of the drilling signal, determine the number of decomposition layers, and decompose the original signal into high frequency and low frequency levels according to the frequency band; The high-frequency and low-frequency reconstruction unit is used to reconstruct the high-frequency band and the low-frequency band and calculate the signal energy of each frequency band respectively; The interference judgment unit is used to count the proportion of low-frequency energy under historical normal operating conditions and calculate the high-frequency and low-frequency energy ratio in real time. If the proportion of high-frequency energy suddenly increases and continuously exceeds the threshold, it is determined that vibration noise has submerged the effective signal area and interference detection is initiated.

8. The oil drilling big data processing system according to claim 6, characterized in that: The abnormality root cause judgment module includes: The package deformation determination unit is used to calculate the mean / variance through a sliding window to establish the normal fluctuation range of the pressure / temperature signal. If the signal continuously deviates from the baseline and exceeds the threshold, it is determined to be a drift anomaly caused by package deformation. The corrosion determination unit is used to compare the correlation between pH value, H2S concentration and electrode response current. If the current sudden drop or nonlinear fluctuation is positively correlated with the concentration of corrosive substances, it is determined that the data anomaly is caused by electrode corrosion. The electromagnetic interference determination unit is used to collect the electromagnetic radiation intensity and sensor signal spectrum. If the noise peak coincides with the electromagnetic interference frequency band, it is determined that the data is abnormal due to electromagnetic interference.

9. The oil drilling big data processing system according to claim 6, characterized in that: The exception type output modules include: The training set construction unit is used to filter valid samples from historical data, eliminate periods of communication interruption or sensor failure, align signal statistics with environmental parameters, and construct a data set with annotated anomaly types; The training set partitioning unit is used to divide the data set into training set, validation set, and test set for time series data or static features, optimize parameters, and use K-fold cross-validation to prevent overfitting; The judgment output unit is used to output multi-classification probabilities based on the random forest / LSTM model, set the confidence threshold, and output the abnormality type when the confidence threshold is reached; if the probability distribution is dispersed and the confidence threshold is not reached, a multi-factor coupling abnormality is output.

10. The oil drilling big data processing system according to claim 6, characterized in that: The abnormality root cause judgment module also includes: Anomaly data detection unit, used to calculate the residual between multi-sensor data with the same parameters. If the residual exceeds the threshold, it is marked as an abnormal node; Abnormal data replaces units. For missing / abnormal parameters, Kalman filtering or LSTM time series prediction is used to generate replacement values ​​and annotate confidence levels. The corrosion rate model and stress-strain equation are combined to constrain the rationality of the filled-in values. The abnormal advance reminder unit is used to periodically send self-test instructions to diagnose the sensor hardware status; record the sensor's cumulative working time and the number of extreme working conditions, predict the remaining life of the sensor based on the Weibull distribution model, and remind you to replace it in advance.