Underground geophysical prospecting sampling data quality monitoring method and system
By combining real-time data acquisition with deep learning, the problem of low efficiency in controlling the quality of downhole geophysical sampling data was solved, enabling efficient and accurate anomaly identification and early warning, thus ensuring the continuity and safety of downhole geophysical operations.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-17
- Publication Date
- 2026-04-10
AI Technical Summary
The current quality control of downhole geophysical sampling data mainly relies on manual sampling, which leads to low efficiency and easy omission of abnormal data, making it difficult to achieve efficient and accurate quality monitoring in the case of massive amounts of data.
It employs real-time data acquisition and preprocessing, extracts time-series features, combines traditional quality evaluation rules with a deep learning time-series anomaly detection model, identifies suspected abnormal data through a dual verification mechanism, distinguishes between equipment failures and environmental interference anomalies, triggers a three-level early warning mechanism, and supports model iterative optimization and threshold updates.
It achieves full coverage and real-time monitoring of downhole geophysical sampling data, improves monitoring efficiency and accuracy, ensures no abnormal data is missed, provides accurate data support, provides a reliable basis for geological interpretation and resource assessment, and reduces the safety risks of downhole operations.
Smart Images

Figure CN121834589A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of downhole geophysical prospecting, and in particular to a downhole geophysical prospecting sampling data quality monitoring method and system. BACKGROUND
[0002] Downhole geophysical prospecting generally refers to a method of placing geophysical exploration instruments wholly or partially in a borehole or tunnel to excite and observe geophysical fields. During downhole geophysical prospecting, data sampling operations are usually performed. Downhole geophysical prospecting sampling data is the core basis for geological structure interpretation and resource reserve evaluation, and the data quality directly determines the reliability of geophysical prospecting results. Therefore, quality monitoring operations are performed on the sampling data.
[0003] The existing downhole geophysical prospecting sampling data quality control mainly relies on a manual sampling mode. Technical personnel judge whether the data is abnormal by observing waveforms and calculating amplitude fluctuations based on traditional geophysical prospecting data quality evaluation rules.
[0004] However, this method usually increases the labor intensity of workers when facing a large amount of sampling data, resulting in low sampling efficiency and easy omission of abnormal data. Therefore, it is necessary to use an intelligent downhole geophysical prospecting sampling data quality monitoring method with high efficiency and accuracy. SUMMARY
[0005] The present application aims to provide a downhole geophysical prospecting sampling data quality monitoring method and system to solve the problems in the background art.
[0006] To achieve the above-mentioned purpose, the present application provides the following technical solution: a downhole geophysical prospecting sampling data quality monitoring method, comprising the following specific steps: Step 1: Real-time acquisition of downhole geophysical prospecting sampling data and pre-processing of the sampling data; Step 2: Extraction of time sequence features of the sampling data in Step 1, wherein the time sequence features at least include waveform features and amplitude features; Step 3: Preliminary screening and identification of the time sequence features extracted in Step 2 based on traditional geophysical prospecting data quality evaluation rules, and determination of whether it is suspected abnormal data in combination with a pre-set static threshold; Step 4: If it is suspected abnormal data, input the suspected abnormal data into a deep learning time sequence anomaly detection model, and determine the abnormal data by comparing the output result of the deep learning time sequence anomaly detection model with a dynamic threshold; Step 5: Type identification of the determined abnormal data to distinguish between equipment failure type abnormality and environmental interference type abnormality, and triggering of a corresponding level of early warning; Step Six: Store the monitoring data and identification results throughout the entire process, and periodically initiate model iteration optimization and threshold updates to provide data quality inspection and testing support for downhole geophysical exploration operations.
[0007] Preferably, the sampling data in step one specifically includes waveform data, amplitude data, seismic wave data, electromagnetic induction data, and time-series correlation information. The sampling data acquisition frequency in step one is consistent with the sampling frequency of the geophysical equipment. During the acquisition process, the integrity of data transmission is verified in real time. When the data transmission delay exceeds 500ms or the data packet loss rate exceeds 3%, a data transmission anomaly prompt is triggered, and a backup transmission link is activated. The sampling data preprocessing in step one includes denoising and data completion. Specifically, a wavelet threshold denoising algorithm is used to remove Gaussian white noise and impulse noise. Missing data is completed using linear interpolation. The completion ratio of a single batch of data is ≤10%. When the missing ratio exceeds 10%, it is directly marked as data integrity anomaly and included in the suspected abnormal data.
[0008] Preferably, the time-series features in step two also include frequency features, phase features, and data continuity features. After the time-series features are extracted in step two, each feature is standardized to remove abnormal feature points that exceed the effective value threshold. The effective value threshold is the mean of each feature in historical normal sampling data ± 3 times the standard deviation.
[0009] Preferably, the traditional geophysical data quality evaluation rules in step three include signal-to-noise ratio (SNR) evaluation rules, amplitude stability rules, and data integrity rules. The preset static thresholds in step three are SNR ≥ 20dB, amplitude fluctuation ≤ 15%, and data missing rate ≤ 5%. If any threshold corresponding to any rule is not met, it is judged as suspected abnormal data. The SNR evaluation rule is obtained by calculating the ratio of the peak signal value to the peak noise value. The amplitude stability rule is obtained by calculating the ratio of the difference between the maximum and minimum amplitude values of a single batch of data to the average amplitude value.
[0010] Preferably, the deep learning temporal anomaly detection model in step four is a fusion model of bidirectional LSTM and attention mechanism. When training the deep learning temporal anomaly detection model, a training set is constructed using historical normal sampling data and labeled abnormal data, and the model parameters are optimized through five-fold cross-validation. The dynamic threshold in step four is dynamically adjusted based on the loss function value during the model training process, and the initial value is 1.2 times the mean of the loss function.
[0011] Preferably, the real-time adjustment formula for the dynamic threshold in step four is as follows:
[0012] in, This is the current dynamic threshold. is an initial dynamic threshold value, is an adjustment coefficient, and the value range is 0.05-0.1, is a model loss value of the current batch data, is an average loss value of 100 batches of data.
[0013] Preferably, the abnormal type identification in step five is identified by constructing a feature matching library, which contains sensor failure, transmission link failure and other device failure feature templates, and temperature mutation, electromagnetic interference and other environmental interference feature templates. If the similarity between the abnormal data and a certain type of feature template is ≥85%, it is determined as the corresponding type of abnormality.
[0014] Preferably, the corresponding level of early warning in step five includes first-level warning, second-level warning and third-level warning. The first-level warning corresponds to device failure class serious abnormality, the abnormality degree of the device failure class serious abnormality is ≥90%, the second-level warning corresponds to device failure class general abnormality or environmental interference class serious abnormality, wherein the abnormality degree of the device failure class general abnormality ranges from ≥60% to <90%, and the abnormality degree of the environmental interference class serious abnormality is ≥85%, and the third-level warning corresponds to environmental interference class general abnormality, the abnormality degree of the environmental interference class general abnormality is <85%. The abnormality degree is quantitatively calculated by the feature difference between abnormal data and normal data. The early warning information includes abnormal occurrence time, channel number, abnormal type, abnormality degree and disposal suggestion. The actions after the first-level warning, second-level warning and third-level warning are triggered respectively as follows: If the first-level warning is triggered, real-time sound and light warning is performed, and sampling is suspended, and the response time is ≤1 minute; If the second-level warning is triggered, a pop-up window warning is performed, and the response time is ≤5 minutes; If the third-level warning is triggered, an abnormal log is recorded, and continuous monitoring is performed, and the response time is ≤15 minutes.
[0015] Preferably, the storage of monitoring whole process data in step six specifically includes original sampling data, preprocessed data, extracted time sequence features, suspected abnormal data, abnormal data, model output results and warning records, and a distributed database is used for storage, supporting multi-dimensional retrieval according to time, channel and abnormal type. The periodic starting of model iteration optimization and threshold value updating in step six specifically includes starting the iteration training of the deep learning time sequence anomaly detection model when the amount of new data reaches 20% of the historical training data amount, and synchronously updating the model parameters, dynamic threshold value and preset static threshold value in the traditional evaluation rule.
[0016] A downhole geophysical prospecting sampling data quality monitoring system, characterized in that it comprises: A data acquisition and processing module, the data acquisition module is used for real-time acquisition of downhole geophysical prospecting sampling data, and pre-processing of the sampling data; A data feature extraction module, the data feature extraction module is used for extracting the time sequence features of the sampling data; A traditional rule preliminary screening identification module, the traditional rule preliminary screening identification module is used for preliminary screening identification of the extracted time sequence features based on a traditional geophysical prospecting data quality evaluation rule, and combining a preset static threshold value, to determine whether it is suspected abnormal data; A deep learning anomaly detection module, the deep learning anomaly detection module is used for inputting the suspected abnormal data into a deep learning time sequence anomaly detection model, and comparing the output result of the deep learning time sequence anomaly detection model with a dynamic threshold value to determine abnormal data; An anomaly identification module, the anomaly identification module is used for type identification of the determined abnormal data, to distinguish between equipment fault type anomalies and environmental interference type anomalies, and trigger a corresponding level of early warning; A data storage module, the data storage module is used for storing monitoring full-process data and identification results, and periodically starting model iteration optimization and threshold value updating.
[0017] The technical effects and advantages of the present application are as follows: (1) The present application realizes full-coverage and real-time monitoring of downhole geophysical prospecting sampling data by using multi-channel real-time acquisition and full-quantity intelligent monitoring process, which is conducive to completely solving the problems of incomplete coverage and low efficiency of traditional manual sampling inspection, reducing the labor intensity of workers, improving the monitoring efficiency, ensuring that there is no abnormal omission of massive sampling data, providing support for the continuity of geophysical prospecting operation, and having high efficiency and accuracy; (2) The present application is set in a way that the dual verification mechanism of traditional geophysical prospecting quality rule preliminary screening and deep learning time sequence anomaly detection model is matched, which is conducive to complementary advantages of field professional knowledge and intelligent algorithm, replaces the traditional fixed threshold value with poor adaptability and single deep learning, improves the accuracy and reliability of data identification, and provides accurate data for geological interpretation and resource evaluation; (3) The present application is set in a way that the feature matching library and the abnormal degree quantification are matched, which is conducive to realizing accurate differentiation of abnormal types and quantitative evaluation of abnormal degrees, solving the problem of deviation in disposal direction caused by traditional manual sampling inspection which can only judge abnormalities and cannot locate the cause, improving the accuracy of abnormal type identification, and facilitating targeted equipment maintenance or environmental intervention by workers, which is conducive to preventing operation stagnation and secondary quality problems caused by blind investigation; (4) The application is advantageous to realize the graded response and rapid access of abnormal risks by using the setting mode of the three-level early warning mechanism and the response time threshold value, solve the traditional monitoring and early warning ambiguity and response lag, and is advantageous to make the early warning information pass through the local sound and light, remote platform and short message push at the same time, ensure the timely disposal of serious abnormalities, and significantly reduce the safety risk and data loss of downhole geophysical exploration operation. BRIEF DESCRIPTION OF DRAWINGS
[0018] Fig. 1 It is a downhole geophysical exploration sampling data quality monitoring method flowchart of the application; Fig. 2 It is a downhole geophysical exploration sampling data quality monitoring system block diagram of the application; Fig. 3 It is a suspected abnormal data judgment logic diagram of the application. DETAILED DESCRIPTION
[0019] The technical solutions in the embodiments of the application will be clearly and completely described below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the protection scope of the application.
[0020] The application provides a downhole geophysical exploration sampling data quality monitoring method as shown in Figs. 1-3 The method comprises the following specific steps: Step one: real-time acquisition of downhole geophysical exploration sampling data, and pre-processing of the sampling data; Step two: extraction of the time sequence characteristics of the sampling data in step one, wherein the time sequence characteristics at least include waveform characteristics and amplitude characteristics; Step three: preliminary screening and identification of the time sequence characteristics extracted in step two based on the conventional geophysical data quality evaluation rules, and combination with a pre-set static threshold value to determine whether it is suspected abnormal data; Step four: if it is suspected abnormal data, the suspected abnormal data is input into a deep learning time series anomaly detection model, and the result output by the deep learning time series anomaly detection model is compared with a dynamic threshold to determine abnormal data. The construction process of the deep learning time series anomaly detection model includes constructing a network structure that fuses bidirectional LSTM (Bi-LSTM) and multi-head attention mechanism. The Bi-LSTM layer contains 2-4 hidden layers, each hidden layer has 128-512 neurons, the dropout rate is set to 0.2-0.3, the number of heads of the multi-head attention mechanism is set to 4-8 heads, the output dimension is consistent with the output dimension of the Bi-LSTM layer, the forget gate, input gate and output gate weights of the Bi-LSTM layer are initialized as Xavier normal distribution, the bias term is initialized as 0, and the multi-head attention score is the fusion feature output by the fully connected layer after splicing the single-head attention scores; a convolution feature extraction layer is added at the input end of the Bi-LSTM layer, 3 1D convolution kernels (kernel sizes are 3, 5 and 7 respectively) are used to extract local time series features, the convolution step is 1, the padding method is same, and the activation function is ReLU; a fully connected layer is added between the multi-head attention mechanism layer and the output layer, the number of neurons of the fully connected layer is 64-256, BatchNorm batch normalization processing is used, the output layer uses a Sigmoid activation function, and the output range is [0, 1] abnormal probability value. The training process of the deep learning time series anomaly detection model includes collecting nearly 3 years of historical geophysical sampling data, screening out 90% of normal data and 10% of labeled abnormal data (including equipment failure and environmental interference labeling), and dividing them into training set, validation set and test set according to the ratio of 7:2:1; the model is trained using the Adam optimizer, the initial learning rate is 0.001, the cosine annealing learning rate decay strategy is used, the minimum learning rate is 0.0001, the training rounds are 200-500 rounds, and the batch size is set to 32-128; the loss function uses the weighted sum of cross-entropy loss function and mean square error loss function (the weight proportions are 0.7 and 0.3 respectively), the model parameters are optimized through 5-fold cross-validation, when the validation set loss value does not decrease for 30 consecutive rounds and is higher than 1.1 times of the current optimal loss value, the training is stopped and the optimal model parameters are saved, and the initial value T_0 of the dynamic threshold is determined based on the optimal loss value of the validation set during the model training process. Step five: type identification is performed on the determined abnormal data to distinguish between equipment failure type abnormality and environmental interference type abnormality, and trigger the corresponding level of early warning; Step six: store the monitoring whole process data and identification results, and periodically start model iteration optimization and threshold update to provide data quality inspection and detection support for downhole geophysical exploration.
[0021] Further, the sampling data in step one specifically includes waveform data, amplitude data, seismic wave data, electromagnetic induction data, and time sequence correlation information. The sampling data acquisition frequency in step one is consistent with the sampling frequency of the geophysical prospecting equipment. The data transmission integrity is verified in real time during the acquisition process. When the data transmission delay exceeds 500 ms or the data packet loss rate exceeds 3%, an abnormal data transmission prompt is triggered, and a backup transmission link is started. The pre-processing of the sampling data in step one includes denoising and data completion. Specifically, the wavelet threshold denoising algorithm is used to remove Gaussian white noise and impulse noise. The wavelet threshold denoising algorithm decomposes the noisy data into wavelet coefficients of different scales through wavelet transform, reconstructs the wavelet coefficients corresponding to the noise after threshold processing, and realizes noise removal. The linear interpolation method is used to complete the missing data. The linear interpolation method completes the missing values through the linear relationship of the effective data before and after the missing point. It is suitable for continuous / discrete missing situations in geophysical prospecting sampling data. The single-batch data completion ratio is ≤10%. When the missing ratio exceeds 10%, the data is directly marked as data integrity anomaly and included in the suspected abnormal data, which is beneficial to provide preliminary screening basis for data quality inspection and detection of downhole geophysical prospecting operation.
[0022] Further, the time sequence features in step two also include frequency features, phase features, and data continuity features. After the time sequence features are extracted in step two, each feature is standardized to remove abnormal feature points that exceed the effective value threshold of the feature. The effective value threshold of the feature is the mean ± 3 times the standard deviation of each feature in the historical normal sampling data.
[0023] Specifically, the traditional geophysical data quality evaluation rules in step three include signal-to-noise ratio evaluation rules, amplitude stability rules, and data integrity rules. The signal-to-noise ratio evaluation rule refers to the ratio of effective signal to noise in geophysical sampling data, reflecting the signal purity of the data, and is the core rule for judging whether the data is disturbed by noise. The amplitude stability rule refers to the fluctuation degree of the amplitude of geophysical sampling data in the time dimension, reflecting whether the amplitude abnormally fluctuates due to equipment failure (such as sensor drift) or environmental interference (such as vibration). The data integrity rule refers to the missing degree of geophysical sampling data, reflecting whether there are interruptions and packet loss problems in the data acquisition / transmission process. The pre-set static threshold in step three is signal-to-noise ratio ≥ 20 dB, amplitude fluctuation amplitude ≤ 15%, and data missing rate ≤ 5%. If any of the threshold values corresponding to the rules does not meet the standard, it is determined as suspected abnormal data. The signal-to-noise ratio evaluation rule is obtained by calculating the ratio of the signal peak value to the noise peak value. The amplitude stability rule is obtained by calculating the difference between the maximum and minimum values of the amplitude of single-batch data as a proportion of the amplitude mean value.
[0024] Specifically, the deep learning time series anomaly detection model in step four is a fusion model of bidirectional LSTM and attention mechanism. The deep learning time series anomaly detection model is trained using historical normal sampling data and labeled abnormal data to construct a training set, and the model parameters are optimized through five-fold cross-validation. The dynamic threshold in step four is dynamically adjusted based on the loss function value during model training. The initial value is 1.2 times the average value of the loss function. Five-fold cross-validation is performed by dividing the data set into five mutually exclusive subsets. Four subsets are used for training and one subset is used for validation. The final integrated result selects the optimal model parameters. The specific operation steps are as follows: data set preprocessing collect labeled historical geophysical data (90% normal data + 10% abnormal data), divide the total training set (70%) and independent test set (10%) according to the ratio of 7:2:1, and randomly and uniformly divide the total training set into five equal subsets; model parameter candidate set determination, for the deep learning model (Bi-LSTM + multi-head attention), determine the core parameters to be optimized and the candidate range: Bi-LSTM hidden layer neuron number: [128, 256, 512]; number of multi-head attention heads: [4, 6, 8]; initial learning rate of Adam optimizer: [0.0005, 0.001, 0.002]; Dropout rate: [0.2, 0.25, 0.3]; cycle training and validation, for each candidate parameter, perform 5 rounds of training-validation; parameter performance evaluation, calculate the average validation set loss value and average abnormal recognition accuracy of each candidate parameter in 5 rounds of validation, and select the highest parameter combination as the optimal parameter; if there are multiple parameter groups with similar indicators, prefer the parameter with lower model complexity (such as smaller neuron number); model final verification, use the independent test set to verify the performance of the optimal parameters, if the test set accuracy is greater than or equal to the preset threshold, the final model parameters are determined; otherwise, return to step two, adjust the parameter candidate set and re-verify, the constraint condition is that during the training process, if the validation set loss value does not decrease for 30 consecutive rounds and is higher than 1.1 times the current optimal loss value, terminate the training of this group of parameters and switch to the next candidate parameter.
[0025] Further, the real-time adjustment formula of the dynamic threshold in step four is as follows:
[0026] wherein, is the current dynamic threshold, is the initial dynamic threshold, is the adjustment coefficient, and the value range is 0.05-0.1, is the model loss value of the current batch of data, is the average loss value of 100 batches of data.
[0027] Further, the abnormal type identification in step five is identified by constructing a feature matching library. The feature matching library includes sensor failure, transmission link failure and other device failure feature templates, as well as temperature mutation, electromagnetic interference and other environmental interference feature templates. If the similarity between the abnormal data and a certain type of feature template is greater than or equal to 85%, it is determined that the corresponding type of abnormality occurs. The specific steps of constructing the feature matching library are as follows: build an abnormal type classification framework, first determine the classification system of the core abnormality in the downhole geophysical prospecting scene, ensure that the main abnormality is covered, and the first-level classification is divided into two categories: device failure class abnormality and environmental interference class abnormality. The second-level classification under the device failure class abnormality includes sensor failure (such as sensitivity drift and failure), transmission link failure (such as packet loss and interruption), and power supply module failure (such as unstable voltage). The second-level classification under the environmental interference class abnormality includes temperature mutation (such as a change in downhole temperature of more than 5°C per minute), electromagnetic interference (such as interference generated by the start and stop of downhole electrical equipment), and vibration interference (such as disturbance caused by drilling operations). Collect and pre-process abnormal feature template data, collect historical abnormal data from downhole geophysical prospecting operations in the past 3-5 years, and require at least 1000 valid samples for each type of second-level abnormality. Each sample is a time series containing 1024-4096 data points, and all data must be clearly labeled with the abnormal type, occurrence time and working condition background. At the same time, collect the working condition data when the abnormality occurs, including sensor calibration records, downhole temperature, electromagnetic intensity and equipment operating status, etc., to exclude irrelevant interference factors. Then, perform the same preprocessing operation as the real-time monitoring process on the collected abnormal data to ensure that the feature dimensions of the template features and the real-time monitoring data remain consistent. Then, remove invalid samples, such as samples with multiple abnormal superpositions, labeling errors and data missing rates exceeding 10%. Finally, at least 800 valid samples of each type of second-level abnormality are retained. Then, extract the core features and construct standardized templates. For each type of valid sample of the second-level abnormality, extract the common basic features and unique differentiated features, and then integrate them to form standardized templates. Extract the basic features, and all types of abnormalities need to extract the following six common features: amplitude mean, amplitude fluctuation amplitude, signal-to-noise ratio, main frequency value, phase shift and data missing rate. Calculate the mean and standard deviation of each feature for all valid samples of each type of abnormality to reflect the overall performance and fluctuation range of the basic features of this type of abnormality. Extract differentiated features, supplement specialized features according to the unique performance of different types of abnormalities and quantify them. For example, sensor failure needs to extract amplitude drift rate and waveform distortion, electromagnetic interference needs to extract high-frequency burr density and signal-to-noise ratio decay rate, and temperature mutation needs to extract main frequency shift rate. These features can accurately distinguish different types of abnormalities. Then, integrate the features to form standardized templates, integrate the basic features (including mean and standard deviation range) and differentiated features (including mean and standard deviation range) of each type of second-level abnormality together, and arrange them in a unified format to clearly present the feature rules of this type of abnormality.The template validity is verified again, similarity verification is performed first, 20% of the valid samples of each type of anomaly are selected as a verification set, the similarity between each verification set sample and the corresponding anomaly template is calculated, and the similarity of the present solution is required to reach 85% or more for a successful match. The verification pass rate of each type of anomaly template is calculated, that is, the proportion of the number of samples that match successfully to the total number of samples in the verification set, which needs to reach 90% or more to be qualified. If it does not meet the standard, go back to the previous step to supplement samples or extract features again, and then cross-verify, calculate the similarity between a certain type of anomaly sample and other types of anomaly templates, and require that the cross-similarity be no more than 30% to avoid confusion between different anomaly templates and ensure the accuracy of classification and identification. Finally, store the feature matching library and dynamically update it. The feature templates are stored in a distributed database, which supports fast retrieval by primary classification and secondary classification, and associates corresponding abnormal handling suggestions in the template, such as the handling suggestion for sensor failure, which is to immediately stop the machine and calibrate the sensor, to facilitate synchronous pushing during subsequent early warning. A dynamic updating mechanism is established to ensure that the template adapts to changes in working conditions. In terms of incremental updating, real-time monitoring of each newly added 100 sets of labeled abnormal data extracts the features of these data, and updates the mean and standard deviation of the corresponding template using a weighted average method, with historical data weight accounting for 80% and new data weight accounting for 20%. In terms of iterative updating, when the number of newly added abnormal samples reaches 20% of the original sample size in the library, the full process of feature extraction, template quantization, and effectiveness verification is re-executed to optimize the performance of the template. In terms of new types, if an unknown anomaly is monitored that has a similarity of less than 85% with all existing templates, collect 500 groups of samples of this type of anomaly and supplement them to the abnormal classification system to build a new feature template, and realize the construction of the feature matching library.
[0028] Further, the corresponding level of early warning in step five includes first-level early warning, second-level early warning, and third-level early warning. The first-level early warning corresponds to a device fault class serious anomaly, the anomaly degree of the fault class serious anomaly is ≥ 90%, the second-level early warning corresponds to a device fault class general anomaly or an environmental interference class serious anomaly, wherein the anomaly degree of the fault class general anomaly ranges from ≥ 60% to < 90%, and the anomaly degree of the environmental interference class serious anomaly is ≥ 85%, and the third-level early warning corresponds to an environmental interference class general anomaly, the anomaly degree of the environmental interference class general anomaly is < 85%. The anomaly degree is calculated by quantifying the feature difference between the abnormal data and the normal data. The early warning information includes the time of anomaly occurrence, the channel number, the anomaly type, the anomaly degree, and the handling suggestion. The actions after the first-level early warning, the second-level early warning, and the third-level early warning are as follows: If the first-level early warning is triggered, real-time sound and light early warning is performed, and sampling is suspended, and the response time is ≤ 1 minute; If the second-level early warning is triggered, a pop-up window early warning is performed, and the response time is ≤ 5 minutes; If the third-level early warning is triggered, an abnormal log is recorded, and continuous monitoring is performed, and the response time is ≤15 minutes.
[0029] Specifically, the storage of the monitoring full-process data in step six specifically includes original sampling data, pre-processed data, extracted time sequence features, suspected abnormal data, abnormal data, model output results, and early warning records, and a distributed database is used for storage, supporting multi-dimensional retrieval according to time, channel, and abnormal type, and the periodic starting of model iteration optimization and threshold updating in step six specifically includes starting the iteration training of the deep learning time sequence anomaly detection model when the amount of new data reaches 20% of the historical training data, synchronously updating the model parameters, the dynamic threshold, and the preset static threshold in the traditional evaluation rule, after the iteration training, using A / B testing to verify the performance, when the abnormal identification accuracy of the new model is improved by ≥5% compared with the original model, the misjudgment rate is reduced by ≥3%, and the response delay is reduced by ≥10%, the original model is replaced and put into real-time monitoring, and at the same time, the historical model version and threshold parameters are saved for tracing and comparison.
[0030] A downhole geophysical sampling data quality monitoring system, comprising: A data acquisition and processing module, the data acquisition module is used for real-time acquisition of downhole geophysical sampling data, and pre-processing of the sampling data; A data feature extraction module, the data feature extraction module is used for extracting time sequence features of the sampling data; A traditional rule preliminary screening identification module, the traditional rule preliminary screening identification module is used for preliminary screening identification of the extracted time sequence features based on a traditional geophysical data quality evaluation rule, and in combination with a preset static threshold, judging whether it is suspected abnormal data; A deep learning anomaly detection module, the deep learning anomaly detection module is used for inputting the suspected abnormal data into a deep learning time sequence anomaly detection model, and determining abnormal data through comparison of the deep learning time sequence anomaly detection model output result and the dynamic threshold; An abnormal identification module, the abnormal identification module is used for type identification of the determined abnormal data, distinguishing between device fault type abnormality and environmental interference type abnormality, and triggering early warning of corresponding levels; A data storage module, the data storage module is used for storing monitoring full-process data and identification results, and periodically starting model iteration optimization and threshold updating.
[0031] As an implementation, the traditional rule preliminary screening identification module relies on preset static thresholds (such as signal-to-noise ratio ≥ 20 dB, amplitude fluctuation amplitude ≤ 15%, and data missing rate ≤ 5%) to preliminarily judge the suspected abnormal data. However, the downhole geophysical environment is dynamic and changeable, for example, temperature fluctuation, electromagnetic interference or equipment aging may cause the distribution of data characteristics to deviate, and the static threshold is difficult to adapt to these changes in real time, which may cause misjudgment or missed judgment. Specifically, the static threshold is set based on historical normal data, but in the long-term operation, environmental factors or equipment state evolution may make the threshold invalid, reducing the accuracy of preliminary screening. This defect not only affects the input quality of the subsequent deep learning model, but also may increase the risk of false alarm, weakening the reliability of the system.
[0032] To overcome this defect, the embodiment introduces a dynamic adaptive threshold adjustment mechanism on the basis of the traditional rule preliminary screening identification module and the data storage module. The mechanism automatically updates the static threshold by analyzing the historical data stream in real time, ensuring that it is synchronized with environmental changes. Specifically, the data storage module regularly archives historical normal sampling data and abnormal records, and the traditional rule preliminary screening identification module integrates a lightweight machine learning algorithm (such as sliding window statistics or online clustering) to calculate the dynamic distribution of data characteristics. For example, every certain time interval (such as 24 hours), the system uses recent data (such as the last 1000 batches) in the data storage module to recalculate the mean and standard deviation of each feature (such as signal-to-noise ratio and amplitude fluctuation), and adjusts the static threshold to the range of mean ± 2 times standard deviation to cover normal fluctuations. At the same time, through the feedback of the abnormal identification module, misjudgment cases after threshold adjustment are identified to further optimize the adjustment frequency and amplitude.
[0033] As an implementation, the deep learning anomaly detection module adopts a fusion model of bidirectional LSTM and attention mechanism, relying on historical training data for anomaly detection. However, the model has fixed parameters after training, and when new types of anomalies (such as unforeseen equipment failures or environmental disturbances) occur downhole, the model may not be able to effectively identify them, leading to the risk of missed detection. The document indicates that the model is optimized through five-fold cross-validation, but lacks an online learning mechanism, which cannot absorb new data in real time, which may limit its adaptability in long-term deployment. Especially in the complex downhole geophysical environment, the type of anomaly may evolve with the upgrading of technology or changes in working conditions, and the static model is difficult to cover all scenarios.
[0034] To address this shortcoming, the present embodiment introduces an incremental learning framework based on existing deep learning anomaly detection modules, anomaly identification modules, and data storage modules. This framework allows the model to continuously learn new anomaly patterns during operation without having to retrain the entire model. In terms of specific implementation, the data storage module not only stores historical data, but also records newly detected anomaly cases and their annotations (type identification results through the anomaly identification module) in real time. The deep learning anomaly detection module integrates online learning algorithms, such as using elastic weight consolidation or gradient descent optimization, to fine-tune model parameters using new data periodically (e.g., every week). At the same time, the anomaly identification module is responsible for verifying the authenticity of new anomalies to ensure the quality of incremental learning. For example, when the system detects an unknown anomaly with a similarity to existing feature templates below 85%, it will be temporarily marked as "to-be-confirmed anomaly" and added to the training set after being confirmed through manual review or multi-sensor fusion.
[0035] Through the incremental learning mechanism, the model can dynamically adapt to unknown anomalies, improving the generalization ability of the system. Specifically, the data storage module provides the data foundation, the anomaly identification module provides the annotation feedback, and the deep learning anomaly detection module performs model updates. This not only reduces the manual intervention in model maintenance, but also ensures the accuracy of the system in long-term operation. In addition, incremental learning can be combined with dynamic threshold adjustment to form a closed-loop optimization, avoiding model obsolescence.
[0036] Finally, it should be noted that the above-described only for the preferred embodiments of the present application, and not for the purpose of limiting the present application, although the foregoing detailed description of the present application, for those skilled in the art, it still can be modified, or part of the technical features of the equivalent replacement, within the spirit and principles of the present invention, any modification, equivalent replacement, improvement, etc., should be included within the scope of the present application.
Claims
1. A method for monitoring the quality of downhole geophysical sampling data, characterized in that: include: Step 1: Collect downhole geophysical sampling data in real time and preprocess the sampling data; Step 2: Extract the temporal features of the sampled data from Step 1. The temporal features include at least waveform features and amplitude features. Step 3: Based on traditional geophysical data quality evaluation rules, the time-series features extracted in Step 2 are initially screened and identified, and combined with preset static thresholds, it is determined whether they are suspected abnormal data. Step 4: If the data is suspected to be abnormal, input the suspected abnormal data into the deep learning temporal anomaly detection model, and determine the abnormal data by comparing the output of the deep learning temporal anomaly detection model with the dynamic threshold. Step 5: Identify the types of the identified abnormal data, distinguish between equipment failure anomalies and environmental interference anomalies, and trigger the corresponding level of warning; Step Six: Store the monitoring data and identification results throughout the entire process, and periodically initiate model iteration optimization and threshold updates to provide data quality inspection and testing support for downhole geophysical exploration operations.
2. The method for monitoring the quality of downhole geophysical sampling data according to claim 1, characterized in that: The sampling data in step one specifically includes waveform data, amplitude data, seismic wave data, electromagnetic induction data, and time-series correlation information. The sampling data acquisition frequency in step one is consistent with the sampling frequency of the geophysical equipment. During the acquisition process, the integrity of data transmission is verified in real time. When the data transmission delay exceeds 500ms or the data packet loss rate exceeds 3%, a data transmission anomaly prompt is triggered, and a backup transmission link is activated. The sampling data preprocessing in step one includes denoising and data completion. Specifically, a wavelet threshold denoising algorithm is used to remove Gaussian white noise and impulse noise. Missing data is completed using linear interpolation. The completion ratio of a single batch of data is ≤10%. When the missing ratio exceeds 10%, it is directly marked as data integrity anomaly and included in the suspected anomaly data.
3. The method for monitoring the quality of downhole geophysical sampling data according to claim 1, characterized in that: The time-series features in step two also include frequency features, phase features, and data continuity features. After the time-series features are extracted in step two, each feature is standardized to remove abnormal feature points that exceed the effective value threshold. The effective value threshold is the mean of each feature in historical normal sampling data ± 3 times the standard deviation.
4. The method for monitoring the quality of downhole geophysical sampling data according to claim 1, characterized in that: The traditional geophysical data quality evaluation rules in step three include signal-to-noise ratio (SNR) evaluation rules, amplitude stability rules, and data integrity rules. The preset static thresholds in step three are SNR ≥ 20dB, amplitude fluctuation ≤ 15%, and data missing rate ≤ 5%. If any threshold corresponding to any rule is not met, the data is judged as suspected abnormal data. The SNR evaluation rule is obtained by calculating the ratio of the peak signal value to the peak noise value. The amplitude stability rule is obtained by calculating the ratio of the difference between the maximum and minimum amplitude values of a single batch of data to the average amplitude value.
5. The method for monitoring the quality of downhole geophysical sampling data according to claim 1, characterized in that: The deep learning temporal anomaly detection model in step four is a fusion model of bidirectional LSTM and attention mechanism. When training the deep learning temporal anomaly detection model, a training set is constructed using historical normal sampling data and labeled abnormal data, and the model parameters are optimized through five-fold cross-validation. The dynamic threshold in step four is dynamically adjusted based on the loss function value during the model training process, and the initial value is 1.2 times the mean of the loss function.
6. The method for monitoring the quality of downhole geophysical sampling data according to claim 1, characterized in that: The real-time adjustment formula for the dynamic threshold in step four is as follows: in, This is the current dynamic threshold. As the initial dynamic threshold, This is an adjustment factor, and its value ranges from 0.05 to 0.
1. This represents the model loss value for the current batch of data. This represents the average loss value for 100 batches of data.
7. The method for monitoring the quality of downhole geophysical sampling data according to claim 1, characterized in that: The anomaly type identification in step five is performed by constructing a feature matching library. The feature matching library includes feature templates for sensor failures, transmission link failures, and other equipment failures, as well as feature templates for temperature mutations, electromagnetic interference, and other environmental interferences. If the similarity between abnormal data and a certain type of feature template is ≥85%, it is determined to be an anomaly of the corresponding type.
8. The method for monitoring the quality of downhole geophysical sampling data according to claim 1, characterized in that: The corresponding warning levels in step five include Level 1, Level 2, and Level 3 warnings. Level 1 warnings correspond to severe equipment malfunctions with an anomaly severity of ≥90%. Level 2 warnings correspond to general equipment malfunctions or severe environmental interference, where the anomaly severity ranges from ≥60% to <90%, and the severe environmental interference anomaly severity is ≥85%. Level 3 warnings correspond to general environmental interference anomalies with an anomaly severity of <85%. The anomaly severity is calculated through a metric of the difference between abnormal and normal data. The warning information includes the anomaly occurrence time, channel number, anomaly type, anomaly severity, and handling recommendations. The actions triggered by Level 1, Level 2, and Level 3 warnings are as follows: If a Level 1 warning is triggered, a real-time audible and visual warning will be issued, and sampling will be suspended, with a response time of ≤1 minute. If a Level 2 alert is triggered, a pop-up alert will be displayed, and the response time will be ≤5 minutes. If a Level 3 alert is triggered, an anomaly log will be recorded and continuous monitoring will be conducted, with a response time of ≤15 minutes.
9. The method for monitoring the quality of downhole geophysical sampling data according to claim 8, characterized in that: The data storage monitoring process in step six specifically includes raw sampled data, preprocessed data, extracted time-series features, suspected abnormal data, abnormal data, model output results, and early warning records. It is stored in a distributed database and supports multi-dimensional retrieval by time, channel, and anomaly type. The periodic initiation of model iteration optimization and threshold update in step six specifically includes starting the iterative training of the deep learning time-series anomaly detection model when the amount of new data reaches 20% of the historical training data, and synchronously updating the model parameters, dynamic thresholds, and preset static thresholds in traditional evaluation rules.
10. A downhole geophysical sampling data quality monitoring system, characterized in that, include: The data acquisition and processing module is used to acquire downhole geophysical sampling data in real time and preprocess the sampling data. The data feature extraction module is used to extract the temporal features of the sampled data; The traditional rule-based preliminary screening and identification module is used to perform preliminary screening and identification on the extracted time-series features based on traditional geophysical data quality evaluation rules, and to determine whether the data is suspected to be abnormal based on a preset static threshold. A deep learning anomaly detection module is used to input suspected abnormal data into a deep learning temporal anomaly detection model, and to determine abnormal data by comparing the output of the deep learning temporal anomaly detection model with a dynamic threshold. An anomaly identification module is used to identify the type of determined abnormal data, distinguish between equipment failure anomalies and environmental interference anomalies, and trigger corresponding level of early warning. The data storage module is used to store the monitoring process data and identification results, and to periodically initiate model iteration optimization and threshold updates.
Citation Information
Cited By
A power distribution terminal security protection method and system based on data encryption
CN122293325A