Intelligent separation method for high-dimensional multi-component characteristics of signal-noise superimposed wave field in well seismic

By improving the inherent timescale decomposition method and tree-structured ensemble model, the problem of signal-noise field separation in well seismic exploration is solved, achieving signal-noise separation in high-dimensional feature space, improving the signal-to-noise ratio and resolution of well DAS data, and supporting subsequent inversion and interpretation.

CN117310804BActive Publication Date: 2026-07-24JILIN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JILIN UNIVERSITY
Filing Date
2023-10-09
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In well seismic exploration, the data obtained by fiber optic distributed acoustic sensor technology is mixed with high-energy and diverse types of noise, which seriously interferes with the complex uplink and downlink wave field, affecting the accuracy of inversion imaging and stratigraphic interpretation. Existing methods are unable to effectively separate the signal-noise wave field.

Method used

An improved intrinsic timescale decomposition method is adopted, combined with multi-component, multi-dimensional time-frequency feature attribute space and tree ensemble model, to construct a high-integration classification framework. The improved intrinsic timescale decomposition method is used to obtain distortion-free and lightly aliased multi-component decomposition results, and the signal-noise feature points are discriminated and classified in high-dimensional space to achieve the separation of signal and noise.

Benefits of technology

It effectively improves the signal-to-noise ratio, fidelity, and resolution of DAS data in wells, ensuring high accuracy of seismic records and supporting subsequent inversion and interpretation work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117310804B_ABST
    Figure CN117310804B_ABST
Patent Text Reader

Abstract

The present application relates to a kind of well seismic signal-noise superposition wave field high-dimensional multi-component feature intelligent separation method, belong to geophysical oil and gas resource exploration field.Establish time-domain multi-component decomposition result serves different dimension differentiation characteristic property expression;Six-dimensional characteristic attribute space based on multi-component signal-noise component is established, and the mathematical characteristics of complex signal wave and multiple types of noise in DAS record, such as direct wave and reflected wave, are revealed;Dual-stage highly integrated framework with tree model as base learner is set up to complete the feature point category determination separation task in high-dimensional attribute space, so as to accurately and efficiently realize low-damage preservation of signal wave field and maximum suppression of noise wave field.The method is flexible, accurately preserves the amplitude characteristics and frequency energy components of DAS two-dimensional exploration record signal field in well, the mathematical basis is reliable and reliable, and the feature is strongly interpretable, which provides an important basis for thin and fine layer detection, oil and gas reservoir exploitation and other practical applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of geophysical oil and gas resource exploration, and particularly to the processing of seismic exploration data. It proposes an intelligent separation method for high-dimensional multi-component features of signal-noise superimposed wavefields in well seismic data ("-" represents "AND"), which is applicable to the separation of complex signal-noise superimposed wavefields in high-precision well seismic exploration records in the Tahe Oilfield. Background Technology

[0002] Fiber optic distributed acoustic sensor (DAS) technology is a novel data acquisition technology that has been widely used in well seismic exploration and oil and gas monitoring in recent years, offering significant advantages such as wide detection range, high accuracy, and low cost. In practical applications of well seismic exploration in the Tarim Basin, this technology, leveraging the high-temperature and high-pressure resistant properties of optical fiber, has replaced conventional electronic geophones, enabling high-precision exploration in deep and even ultra-deep wells. However, due to the harsh monitoring environment in deep wells and interference and diffraction issues occurring in the optical fiber, the currently acquired well DAS data is contaminated with high-energy noise of various types, severely interfering with the complex uplink and downlink wave fields and directly affecting the accuracy of subsequent operations such as inversion imaging and formation interpretation. Therefore, signal-to-noise separation to improve the resolution and signal-to-noise ratio of well DAS data is imperative. Currently, commonly used methods for seismic random noise removal include several categories: time-frequency domain filtering algorithms based on median filtering and bandpass filtering; threshold denoising algorithms based on sparse domain transforms such as wavelets and shear waves; decomposition denoising algorithms based on empirical mode decomposition; and deep learning network framework methods that have emerged in recent years. These methods all have limitations in practical applications: time-frequency filtering methods based on filters are simple to implement but prone to leakage and attenuation of effective signal components; sparse domain transforms are more suitable for suppressing high-frequency, Gaussian noise; decomposition algorithms suffer from the common problem of mode aliasing; and the interpretability of black-box models in deep learning still needs further research.

[0003] Signal decomposition is a method of decomposing a signal into a series of wavelets according to a certain strategy. After decomposition, the wavelets usually have different characteristics in the time and frequency domains, which is a prerequisite for subsequent data processing and analysis. Considering the severe signal-noise field aliasing in well seismic records obtained by fiber optic distributed acoustic sensing (DAS) technology, as well as the diverse types and uneven distribution of noise, the high energy performance of which can submerge weak reflected signals and interfere with the characterization of strong reflected wave noise, the high difficulty in noise removal, and the easy attenuation of signal amplitude-frequency characteristics during noise suppression, this paper presents some challenges.

[0004] Studies have shown that seismic signals exhibit non-stationary characteristics such as time-varying spectrum and space-varying wavenumber, while inherent time-scale decomposition algorithms are a class of adaptive time-frequency analysis methods suitable for non-stationary signals. This method involves setting a baseline extraction operator. and rotation extraction operator To extract the non-stationary signal X respectively t Single-channel linear trend L t and the forward component H t This decomposes the non-stationary signal into a series of monotonic positive rotation components (PRCs) between extreme points. The above process mainly consists of the following four steps:

[0005] 1) Calculate the original non-stationary signal X t extreme points

[0006]

[0007] Where τ k X represents the extreme point. k / L k Representing X t / L t The corresponding extreme values ​​in the signal, where α represents the gain control parameter (0 < α < 1).

[0008] 2) Use a linear shrinkage transformation to fit the extreme points obtained in step 1), and the curves between the extreme points should show a monotonic trend.

[0009] The baseline signal L can be obtained by connecting the curves between all extreme points. t .

[0010]

[0011] 3) According to L t The forward component H can be obtained. t .

[0012]

[0013] 4) Place L t Using the original signal again, repeat the above steps until the iteration stopping condition is met. After n decompositions, the decomposed wavelet components PRC1-PRC are obtained. n .

[0014] The distribution of time-domain extrema points of a series of wavelet components obtained after the seismic signal is decomposed by its inherent time scale is from dense to sparse, and the corresponding frequency distribution in the frequency domain is from high to low. Different decomposed components correspond to the characterization information of different types of signals or noise. While preserving the morphological characteristics and inherent instantaneous amplitude of time-domain signals and noise, the frequency and relative phase information are obtained. This is the basis for obtaining the time-frequency multidimensional characteristics of various traveling waves and noise fields in the DAS record in the well, and it is also a prerequisite for the separation of various superimposed wave fields. Summary of the Invention

[0015] The purpose of this invention is to provide an intelligent separation method for high-dimensional multi-component features of signal-noise superimposed wavefields in well seismic data. This method solves the problems of low recording quality and difficulty in separating signal-noise multi-wavefield superposition in well seismic data obtained using distributed fiber optic sensor technology, which are detrimental to subsequent inversion imaging and structural interpretation operations. Guided by multi-component, multi-dimensional time-frequency feature attribute space, this invention constructs a highly integrated tree-like model classification framework to separate complex signal-noise wavefield superposition features, effectively improving the signal-to-noise ratio (SNR) of well DAS data and truly meeting the "three high requirements" (high SNR, high fidelity, and high resolution) of well DAS data, thus better serving subsequent work stages such as inversion and interpretation. This algorithm first obtains the distortion-free, lightly aliased multi-component decomposition results of DAS records in wells based on an improved intrinsic time-scale decomposition method. Then, it spans a six-dimensional feature attribute space using multiple time- and frequency-domain feature factors as coordinate vectors, mapping low-dimensional challenges to a higher-dimensional space for resolution. Finally, it constructs a comprehensive strong model based on a decision tree structure, integrating the initial and secondary stages, for the discrimination and classification of feature points of different signal-to-noise types in the higher-dimensional space, thereby achieving the separation and independence of the corresponding time-domain signal and complex noise components. This invention effectively demonstrates the differences in the characteristic representation of different signal waves and multiple types of superimposed noise in well DAS data using a multi-component high-dimensional feature attribute space. It provides a sufficient and reliable statistical and mathematical basis for signal-to-noise field analysis and modeling from a theoretical perspective. It constructs a complex tree-like decision integration model in a high-dimensional space for multi-component, multi-dimensional feature separation tasks, solving the difficult problem of low-dimensional signal-to-noise field superposition based on dimensionality-upgrading mapping, and effectively improving the signal-to-noise ratio, fidelity, and resolution of well DAS data.

[0016] The above-mentioned objective of this invention is achieved through the following technical solution:

[0017] A smart method for separating high-dimensional multi-component features of signal-noise superimposed wavefields in well seismic boreholes includes the following steps:

[0018] Step (1) Acquisition of 2D seismic exploration records in wells using DAS: After transmitting lasers into the optical fibers deployed in the well using distributed optical fiber sensing technology, stress sensing and optical signal transmission during the propagation of seismic waves are integrated using optical fibers; the optical signal is received at the ground receiving end and demodulated and converted into an electrical signal to obtain the 2D seismic exploration records in the well.

[0019] Step (2) Establishment of multi-component high-dimensional time-frequency feature attribute space:

[0020] 2.1 Establishing multi-component decomposition components using an improved intrinsic timescale decomposition method

[0021] The linear shrinkage transformation method in the inherent time-scale decomposition is replaced with a cubic spline interpolation algorithm, which uses a smooth curve to fit all extreme points in series, thereby effectively improving the problems of wavelet spikes, distortion, and signal discontinuity after decomposition. On the other hand, a sieving process is added in each decomposition process, and small iteration loops are nested in large iterations to refine the coarse components obtained from the large iterations, further reducing the overlap range of information at different time scales and reducing the degree of frequency component aliasing. At the same time, based on the characteristics of uneven noise distribution intensity and large energy differences in DAS records in wells, a flexible iteration termination condition based on standard deviation and similarity measurement is proposed. That is, the Manhattan distance with clear physical meaning is used to measure the difference between the baseline signal and the original PRC components before and after each sieving iteration of the seismic record data, and the standard deviation (SD) of Formula 1 is used as the standard for measuring the smoothness of the distance curve to determine the number of iteration terminations, avoiding the use of a uniform standard as an inappropriate or unsuitable measure under different data performance conditions.

[0022]

[0023] Where x represents signal data, The mean of the data is represented by N, and the signal length is N.

[0024] The mathematical process of the improved intrinsic time-scale decomposition algorithm is described as follows:

[0025]

[0026] Where PRC(t) is the decomposition component obtained in each iteration, h is the number of components obtained by decomposition, l represents the number of sieving iterations, and r(t) and r*(t) represent the residual components remaining after decomposition.

[0027] 2.2 Multi-dimensional time and frequency domain feature factors span a six-dimensional feature attribute space

[0028] Six-dimensional time-frequency independent domain features are extracted from each component using time-frequency domain features, and a six-dimensional feature attribute space is constructed based on these features. In this six-dimensional feature attribute space, the decomposition results of the DAS record in the well are characterized and described by six-dimensional feature factors, further amplifying the differences in the performance of signal and noise components that are not easily observed in low-dimensional space, thus providing a multi-dimensional spatial data basis for the final separation of different categories of components. The feature factors used include four time-domain features: Kurtosis, Peak factor, Impulse factor, and Clearance factor (Equations 3-6), and two frequency-domain features: root mean square frequency (RMSF) and root variation frequency (RVF) (Equations 7 and 8). These features respectively extract time-domain information including the amplitude fluctuation and peak distribution characteristics of each wavelet component of the seismic data, as well as frequency-domain information including the main frequency band range and frequency energy fluctuation.

[0029]

[0030]

[0031]

[0032]

[0033]

[0034]

[0035] Among them, f k For a signal to contain frequency components, S k Let K be the signal power spectral energy, and K be the total number of frequency components. max(·) is the operation to find the maximum value of the signal.

[0036] Step (3) Output of probability prediction results of tree ensemble model: Formulate a two-stage ensemble large-batch decision tree model to build a comprehensive strong learning model with high precision and high accuracy in multi-dimensional space to complete the classification and discrimination task of signal-noise multi-component feature points. The model construction process is mainly divided into the following two stages:

[0037] Initial ensemble stage: The learner structure based on the large batch decision tree model is trained and generated differently using feedback training and self-sampling training methods respectively. Then, according to the classification or regression task, the appropriate voting method or averaging method is selected to integrate them into multiple strong integrated models with different emphases. In this process, the instability of using only a single learner in the classification process is effectively avoided. At the same time, the integration approach that takes into account both accuracy and diversity indicators greatly improves the final prediction accuracy of the single learner, thereby constructing a tree-like decision model with high classification accuracy.

[0038] Secondary integration stage: Considering the randomness of noise interference and the differences between shot gather records in actual exploration, to improve the generalization performance of the model as much as possible, a secondary integration is carried out on the basis of the primary integration according to the meta-learning strategy, so as to output a highly integrated and highly applicable classification learning model with decision tree as the basic structure. It not only has the ability to accurately distinguish fuzzy multi-component feature point data of DAS records in wells, but also can fully cope with various types of interference in actual DAS data.

[0039] Step (4) Separation of signal and noise multi-overlapping wave fields: After establishing multi-component and multi-dimensional feature attribute spaces in sequence, the constructed tree-like integrated model is used to classify a large number of feature sample points mapped from the DAS data in the well to the high-dimensional space; the feature categories are divided into signal and noise categories, and the output of the integrated model is determined as the most likely category probability prediction for a certain feature point: the prediction with a probability greater than 50% is the signal category, and the prediction with a probability less than 50% is the noise category; the feature corresponding components that are judged to be the signal category are retained and added together, while the feature corresponding components that are judged to be the noise category are discarded; through the above operations, the signal and noise wave fields in the DAS record in the well can be separated.

[0040] The beneficial effects of this invention are as follows: Based on the inherent time-scale decomposition, this invention proposes an improved algorithm with flexible iteration conditions, no distortion, and light aliasing, providing a fast and efficient means to achieve the initial independence of the signal-noise wave field in well DAS data; furthermore, based on multi-component decomposition, a feature factor statistical extraction method is used to map the two-dimensional spatiotemporal domain seismic data record to a high-dimensional time-frequency domain feature attribute space. While providing a reliable theoretical characteristic difference characterization for signal waves such as direct waves, strong and weak reflection waves, and various types of noise such as background random noise, fading noise, and speckled optical noise, this invention also provides a strong mathematical foundation for the further separation of signal-noise category feature points under the guidance of the high-dimensional feature space. Considering the difficulty in linear discrimination of high-dimensional features, this invention performs two high-level integrations based on the decision tree structure. While ensuring the high precision and accuracy of the integrated model, it is as compatible as possible with the diversity of the base model. This ensures that the six-dimensional feature points of the complex superimposed noise field with non-uniformity and different energy distributions in the DAS record in the well have a strong ability to distinguish them in actual engineering. It can efficiently complete the signal-noise field separation task, ensure the high requirements of seismic records, and lay a solid foundation for subsequent processing and interpretation. Attached Figure Description

[0041] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate the invention and are used to explain it, but do not constitute an undue limitation of the invention.

[0042] Figure 1 This image shows a comparison of the data before and after using cubic spline interpolation to replace the linear contraction transform in the inherent timescale decomposition. Part a) shows the 55Hz seismic single-wavelet structure simulated using the Ricker wavelet (top) and its noisy data after adding Gaussian white noise interference (bottom). Part b) shows the results of fitting the seismic wave data (circled by rectangles in part a) using both linear contraction transform and cubic spline interpolation. Part c) compares the PRC components decomposed using the two fitting methods. Arrows indicate signal distortion caused by noise interference under linear contraction transform.

[0043] Figure 2 This section compares the sieving process before and after incorporating a new iteration stopping criterion based on Manhattan distance and standard deviation. The arrows indicate the waveform comparison of the decomposed components before and after the sieving iteration, while the dashed line represents the baseline signal obtained after the sieving iteration process compared to the initial PRC. n The Manhattan distance curve trend between components is shown, with solid dots indicating the standard deviation of the current Manhattan curve fluctuation and rectangles indicating the signal-to-noise aliasing.

[0044] Figure 3The images show the time-domain waveform decomposition results and frequency-domain energy spectral density analysis curves of simulated seismic data with added actual fading noise after decomposition using the inherent time-scale decomposition and its improved algorithm. Part a) includes the time-domain waveforms (left) and energy spectral density curves (right) of the simulated clean seismic signal, actual fading noise, and synthesized noisy data. Circles indicate noise jumps. Part b) shows the multi-component time-domain waveforms (left) and energy spectral density (right) obtained from the inherent time-scale decomposition. Arrows indicate missing signal valleys, and circles indicate signal morphological distortion. Part c) shows the multi-component time-domain waveforms (left) and energy spectral density (right) obtained from the improved algorithm decomposition. Arrows indicate the decomposed high-frequency noise components, and rectangles mark the observable signal morphology. The blue vertical lines in all energy spectral density curves indicate the location of the mean frequency.

[0045] Figure 4 Figure 1 shows the histograms of the six-dimensional eigenvalue distributions of the intrinsic time-scale decomposition and its improved method. Figure a) shows the histograms of the time-domain (left) and frequency-domain (right) eigenvalue distributions of the simulated clean seismic signal, actual fading noise, and noisy data, respectively. Figure b) shows the histograms of the time-domain (left) and frequency-domain (right) eigenvalue distributions of the components obtained by the intrinsic time-scale decomposition. Figure c) shows the histograms of the time-domain (left) and frequency-domain (right) eigenvalue distributions of the components obtained by the improved decomposition algorithm. In the figures, the time-domain features are 1: kurtosis, 2: peak factor, 3: impulse factor, and 4: margin factor; the frequency-domain features are 5: root mean square frequency and 6: frequency standard deviation. Rectangles indicate easily confused categories with similar time-domain and frequency-domain features; ellipses indicate frequency-domain eigenvalue distributions that contradict the time-domain feature categories; and light-colored labels indicate components where arbitrary signal wavelet morphology can be clearly observed.

[0046] Figure 5 This is a schematic diagram of the highly integrated tree model structure with primary and secondary stages used in this invention. DT is the abbreviation for decision tree.

[0047] Figure 6 This is a record of actual DAS exploration data in a well. Rectangles A and B delineate the recording areas for different processing requirements. Arrow 1 marks the direct wave, arrow 2 marks the reflected wave, arrow 3 marks the weak converted wave, double arrows mark fading noise, and ellipses mark optical spot noise.

[0048] Figure 7This section presents the results of separating the signal and noise wavefields from actual well DAS exploration data. Part a) shows the signal wavefield recovery results using five methods: CITD-TS model, CITD-fixed component preservation, ITD-TS model, wavelet thresholding, and RPCA. Part b) shows the noise wavefield separation results using the corresponding algorithms. ITD and CITD are abbreviations for Inherent Time Scale Decomposition and its improved algorithms, respectively; TS is an abbreviation for Tree Integrated Model; and RPCA is an abbreviation for Principal Component Analysis. In part a), single arrows indicate the comparison of signal wave amplitude preservation, and white arrows indicate the unremoved overlapping portion of fading noise in the signal wavefield. In part b), single arrows indicate signal residue in the noise field, ellipses indicate optical spot-type noise in the noise field, and double arrows indicate fading-type noise in the noise field. Detailed Implementation

[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0050] See Figures 1 to 7As shown, this invention discloses an intelligent separation method for complex wavefields with signal-noise superposition in well seismic exploration data under multi-component high-dimensional feature attribute space. To address the severe interference of complex superimposed noise fields on the wavefields of various types of seismic signals obtained by distributed optical fiber acoustic sensor (DAS) technology in wells, this invention first designs an improved method suitable for processing DAS data in wells, based on the distortion and aliasing drawbacks of inherent time-scale decomposition algorithms. Furthermore, considering the uneven distribution and varying intensity of the noise field, a flexible screening iteration stopping condition based on the stationarity of the distance metric curve is proposed, establishing time-domain multi-component decomposition results to serve the expression of differentiated characteristic properties in different dimensions. Secondly, a six-dimensional feature attribute space based on multi-component signal-noise components is established using time and frequency domain feature factors as extraction methods. This maps the difficult separation problem of low-dimensional signal-noise wavefields to a high-dimensional signal-noise feature differentiation attribute space for resolution, revealing the mathematical characteristics of complex signal waves such as direct waves and reflected waves, as well as various types of noise in DAS records, providing sufficient theoretical and statistical basis for wavefield analysis and separation. Finally, a two-stage highly integrated framework with a tree-like model as the base learner is established to complete the feature point category determination and separation task in the high-dimensional attribute space, thereby accurately and efficiently achieving low-damage preservation of the signal wavefield and maximum suppression of the noise wavefield. The method proposed in this invention flexibly and accurately preserves the amplitude characteristics and frequency energy components of the DAS two-dimensional exploration record signal field in the well. It has a solid and reliable mathematical foundation and strong interpretability of features, providing important basis for practical applications such as thin and fine strata detection and oil and gas reservoir development.

[0051] The present invention presents an intelligent separation method for high-dimensional multi-component features of the signal-noise superimposed wavefield in well seismic imaging. This method employs a suitable signal decomposition method to provide positive feedback to address practical application challenges, establishing preliminary independence between the signal and noise fields. This provides reliable prior knowledge on the time or frequency scales for subsequent processing and analysis. Based on an improved intrinsic time-scale decomposition algorithm, the invention constructs a high-dimensional attribute space for time and frequency domain feature factors. A tree-structured integrated model, guided by this high-dimensional feature attribute space, establishes probability thresholds for the classification results of seismic signal fluctuations and various types of random noise. This fully demonstrates the differences in characteristic performance between direct waves, primary reflection waves, and multiple reflection waves detected by DAS technology and background noise, speckled noise, and fading noise in terms of time-domain morphology and frequency-domain energy distribution. This provides reliable mathematical and statistical basis for signal-noise field analysis and highly interpretable theoretical characteristics for signal-noise field separation, serving as a favorable foundation for subsequent inversion imaging analysis and stratigraphic interpretation. The specific steps are as follows:

[0052] Step (1): Acquisition of DAS 2D seismic exploration records in the well:

[0053] In vertical seismic profiling applications, distributed optical fiber sensing (DAS) technology integrates stress sensing and optical signal transmission during seismic wave propagation by transmitting laser light into optical fibers deployed in the well. The optical signal is received at the ground receiver and, after demodulation and conversion into an electrical signal, a two-dimensional seismic record from the well is obtained. Due to the inherent characteristics of optical fibers, well records obtained using DAS technology have a higher sampling frequency (2500Hz), better spatiotemporal resolution, a wider detection range, and smaller inter-track spacing (1m) compared to those obtained using conventional electronic geophones. This makes them more advantageous for discovering fine-grained formations, thin-layer structures, and small oil and gas reservoirs. The dominant frequency range of seismic signals obtained using DAS technology is between 40Hz and 150Hz.

[0054] However, due to the effects of high temperature and high pressure environments downhole on optical fibers, as well as interference during electronic instrument detection, various types of complex noise are mixed into the downhole DAS exploration data of the Tahe Oilfield, such as ringing noise, optical spot noise, fading noise, and background noise. These noises are strong and diverse in nature, severely interfering with the continuity of direct and reflected waves and potentially completely drowning out weak signals, greatly reducing the resolution of seismic exploration records and having a very adverse impact on subsequent stratigraphic inversion. Therefore, how to separate the noise wavefield without changing the characteristics of the up and down wavefields is an important research topic.

[0055] Step (2) Establishment of multi-component high-dimensional time-frequency feature attribute space:

[0056] Step 2.1: Establish multi-component decomposition components using an improved intrinsic timescale decomposition method.

[0057] In the actual decomposition of seismic records, while inherent timescale decomposition has advantages such as high speed, ease of implementation, and the inclusion of accurate time information in the decomposed wavelets, it also has a series of inapplicability issues due to the limitations of the algorithm itself. For example, the decomposed wavelet spikes and severe signal distortion caused by linear contraction transformation under noise interference, and the signal-noise aliasing problem between wavelets caused by incomplete decomposition, etc.

[0058] To address the above problems, this invention replaces the linear shrinkage transformation method in the inherent time-scale decomposition with a cubic spline interpolation algorithm, and uses a smooth curve to cascade and fit all extreme points, thereby effectively improving the problems of wavelet spikes, distortion, and signal discontinuity after decomposition. On the other hand, a sieving process is added to each decomposition process, with small iteration loops nested within large iterations to refine the coarse components obtained from the large iterations, further reducing the overlap range of information at different time scales and alleviating the degree of frequency component aliasing. At the same time, based on the characteristics of uneven noise distribution intensity and large energy differences in DAS records in wells, a flexible iteration termination condition based on standard deviation and similarity measurement is proposed. That is, the Manhattan distance (i.e., street distance) with clear physical meaning is used to measure the difference between the baseline signal and the original PRC components before and after each sieving iteration of the seismic record data, and the standard deviation (Equation 1) is used as the standard for measuring the smoothness of the distance curve to determine the number of iteration terminations, so as to avoid using a uniform standard as an inappropriate or unsuitable measure under different data performance conditions.

[0059]

[0060] Where x is a signal Let N be the signal mean and N be the signal length.

[0061] The mathematical process of the improved intrinsic timescale decomposition algorithm can be described as follows:

[0062]

[0063] Where h is the number of components obtained from the decomposition, l represents the number of sieving iterations, and r(t) and r*(t) represent the residual components remaining after decomposition.

[0064] Compared to inherent time-scale decomposition, the introduction of cubic spline interpolation makes the seismic signal morphology in the wavelet components obtained from DAS decomposition in wellbore more complete and distortion-free, and the extracted signal component characteristics are closer to the actual seismic wavelet characteristics. Furthermore, because component aliasing is suppressed and components at different time scales are separated, the signal and noise superimposed wavefield components are more independent, which is beneficial for the final signal-noise wavefield separation operation. The re-established iteration stopping criterion during the screening process is proposed to address the uneven mixing of multiple noise energy in DAS data, effectively solving the drawback of the inherent time-scale decomposition algorithm's lack of flexibility in practical engineering applications and obtaining better decomposition results.

[0065] Step 2.2: Multi-dimensional time and frequency domain feature factors span a six-dimensional feature attribute space.

[0066] After the seismic data is decomposed into multiple time-domain components using an improved inherent time-scale algorithm, this invention further utilizes time-frequency domain features to extract six-dimensional time- and frequency-independent domain features from each component, and uses this as a basis to span a six-dimensional feature attribute space. In this high-dimensional feature space, the decomposed multiple-component results of the DAS record in the well are characterized and described by six-dimensional feature factors, further amplifying the differences in the performance of signal and noise components that are not easily observed in the low-dimensional space, thus providing a multi-dimensional spatial data foundation for the final separation of different categories of components. The feature factors used include four time-domain features (Equations 6-9) such as kurtosis, peak factor, impulse factor, and margin factor, and two frequency-domain features (Equations 10 and 11) such as root mean square frequency and frequency standard deviation. These features respectively extract time-domain information including the amplitude fluctuation and peak distribution characteristics of each wavelet component of the seismic data, as well as frequency-domain information such as the main frequency band range and frequency energy fluctuation.

[0067]

[0068]

[0069]

[0070]

[0071]

[0072] Among them, f k For a signal to contain frequency components, S k Let K be the signal power spectral energy, and K be the total number of frequency components.

[0073] The above-mentioned features represent the multifaceted and multidimensional relationships between seismic signals and various types of noise, even between different types of seismic waves and noise fields. They provide theoretical data support for differentiated signal-to-noise multi-component analysis, possessing at least one dimension of distinguishable characteristics applicable to the separation of different classes. This allows challenging problems that are difficult to solve in low dimensions to be addressed in a differentiated feature attribute space. Research shows that all six types of features are applicable to the time-frequency information expression of borehole seismic exploration data.

[0074] Step (3) Output of probability prediction results from the tree-based ensemble model:

[0075] Although the differences between feature points corresponding to seismic signals and different types of noise tend to be significant in the constructed high-dimensional feature attribute space, it is difficult to establish simple and usable linearly separable criteria or partitionable hyperplanes for feature classification and subsequent feature separation operations in high-dimensional space. To address the challenge of linearly inseparable multi-component feature points in a six-dimensional feature attribute space, this invention further proposes a two-stage integrated large-batch decision tree model implementation scheme to construct a comprehensive strong learning model that achieves high precision and accuracy in classifying and discriminating multi-component feature points in multi-dimensional space. This successfully solves the application problem of linearly inseparable feature points in high-dimensional space and avoids overfitting and instability defects that easily occur in single-model classification by using an integrated approach. The above model construction process mainly consists of the following two stages:

[0076] i. Initial Ensemble Stage: The learner structure based on the large-batch decision tree model is differentiated through two training methods: a highly dependent feedback training approach and a highly diverse self-sampling training approach. Depending on the classification or regression task, appropriate voting or averaging methods are selected to ensemble the learners into multiple strongly integrated models with different emphases. This process effectively avoids the instability inherent in using only a single learner during classification, while significantly improving the final prediction accuracy of individual learners by simultaneously considering accuracy and diversity metrics, thus constructing a tree-structured decision model with high classification accuracy.

[0077] ii. Secondary integration stage: Considering the randomness of noise interference and the differences between shot gather records in actual exploration, and to improve the generalization performance of the model as much as possible, a secondary integration is further implemented on the basis of the primary integration according to the meta-learning strategy. The output is a highly integrated and highly applicable classification learning model with a decision tree as the basic structure. It not only has the ability to accurately distinguish fuzzy multi-component feature point data of DAS records in wells, but also can fully cope with various types of interference in actual DAS data.

[0078] Step (4) Separation of signal and noise multi-overlapping wave fields

[0079] After establishing a multi-component, multi-dimensional feature attribute space, a pre-constructed tree-structured integrated model is used to classify a large number of feature sample points mapped from the well DAS data to the high-dimensional space. The feature categories are tentatively set as signal and noise, and the output of the integrated model is determined to be a probability prediction of the category bias for a given feature point: a prediction with a probability greater than 50% is for the signal category, and a prediction with a probability less than 50% is for the noise category. The feature components classified as signal are retained and added together, while those classified as noise are discarded. Through these operations, the signal and noise fields in the well DAS record can be separated. The retained signal components, after addition and combination, should contain high-amplitude, high-energy direct wave components, primary reflection waves with slightly reduced amplitude and dominant frequency compared to the direct wave, low-amplitude, low-dominant-frequency, weak-energy multiple reflection waves, and transmitted and converted waves appearing in the shallow and mid-deep layers of the record. The components to be removed should include at least the severe signal interference from speckled optical noise, long-period irregular fading noise, etc. After wave field separation, the clear and interference-free spatiotemporal morphology of the signal can be directly used for the study of different types of wave field characteristics and modeling, while the noise field can be traced back to its causes through property analysis, so as to avoid the causes and reduce interference during the instrument detection and acquisition stage.

[0080] Example:

[0081] First, let's take simulated well DAS seismic single-channel data mixed with background random noise and fading noise as an example ( Figure 3 This demonstrates how to separate multiple types of signal wavefields from dual-noise fields using the aforementioned technical process. A single channel of data consists of 1200 sampling points at a sampling frequency of 2500Hz and a sampling interval of 0.4ms. Due to the high-precision detection characteristics of DAS technology, well-drilled DAS records should contain more clearly defined and complex uplink and downlink wavefields with higher dominant frequencies compared to surface seismic exploration. Therefore, the single channel record should include... Figure 3 Part a) shows the three types of signal wave patterns simulated by the Rick wavelet: direct wave signal, first reflection wave signal, and multiple reflection wave signal. Based on the propagation law of seismic waves, the amplitude and dominant frequency of these three types of signal waves are assumed to decrease sequentially: the simulated amplitude of the direct wave is 3 and the dominant frequency is 75Hz; the simulated amplitude of the first reflection wave is 2.4 and the dominant frequency is 65Hz; the simulated amplitude of multiple reflection wave 1 is 1 and the dominant frequency is 50Hz; and the simulated amplitude of multiple reflection wave 2 is 0.5 and the dominant frequency is 45Hz.

[0082] Add to the above simulated earthquake clean signal Figure 3The fading noise actually collected is shown in section a). In practical applications, background noise is uniformly distributed at the bottom of the exploration record. Therefore, the added fading noise inevitably contains background noise, forming a superimposed dual noise field. From a time-domain perspective, this dual noise field contains multiple frequency components, including a uniformly distributed high-frequency spike, a large-wave low-frequency pattern, and abrupt peaks as circled in the ellipse. After obtaining simulated noisy data using an additive model as a reference, it can be observed that the noise field significantly interferes with the signal wave field with differences in amplitude and dominant frequency distribution. Figure 3 As shown in the energy spectral density curves, the dominant frequency of the noisy data after actual fading noise interference shifts significantly to lower frequencies compared to the dominant frequency of the pure signal, and the shape of the energy spectral density curve is more constrained by the noise energy spectral density curve. Therefore, accurately and with low damage, separating the signal-noise wave field while preserving the inherent characteristics of various wave fields is a key challenge in separation research.

[0083] 1. Establishment of signal-noise multi-component components based on the improved inherent time-scale decomposition algorithm

[0084] To address the severe signal distortion and component aliasing issues inherent in time-scale decomposition algorithms, this invention improves upon them by employing cubic spline interpolation and establishing a sieving process with flexible iteration stopping conditions. Figure 1 , 2 Each with such Figure 1 Using the simulated Ricker wavelet noisy data with added zero-mean Gaussian white noise across the entire frequency band, as shown in section a), the applicability and feasibility of the above two improvements are illustrated. Replacing the linear shrinkage transformation in the extreme point fitting process with cubic spline interpolation can effectively avoid issues such as... Figure 1 In part b), the arrow indicates the signal unevenness and discontinuity caused by noise interference, thus preventing the decomposed seismic signal from exhibiting complete morphology and no distortion. Figure 1 Part c) further helps to improve the fidelity of the signal wave field after recovery. Figure 2 The process of screening iteration under Manhattan distance iteration conditions, measured by standard deviation, is demonstrated. The iteration stops by measuring the steady trend of the Manhattan distance curve between the baseline signal and the original decomposed components before and after screening iteration using standard deviation. This effectively improves the component aliasing phenomenon at different time scales and flexibly addresses the uneven and highly variable noise distribution in actual DAS records.

[0085] After establishing flexible iteration stopping conditions, the inherent timescale decomposition algorithm is improved. Figure 1 It exhibits excellent frequency division capability in the decomposition of noisy analog single-channel data in part a). For example... Figure 3As shown in section c), this method decomposes noisy data into nine positive components (PRC1-PRC9) with extreme point distribution ranging from dense to sparse and frequency components ranging from high to low. Among them, PRC1 and PRC2 are mainly characterized by high-frequency glitches, with average frequencies above 400Hz; in PRC3-PRC6, the simulated Ricker wavelet wave pattern can be clearly observed, and the strong signal wave (direct wave and first reflection wave) and the weak wave (multiple reflection waves 1 and 2) are not completely in the same component, proving that the method has a detailed frequency division capability with little overlap, and its energy spectral density curve is closer to the trend of the simulated pure signal; PRC7-PRC9 are mainly characterized by large fluctuations and low frequency, with average frequencies all below 50Hz. Based on the multi-component decomposition results of the three different time-frequency manifestations above, the components of the signal and noise fields can be initially and independently classified. For example, the noise field should include high-frequency components represented by PRC1-PRC2 and low-frequency components represented by PRC7-PRC9, while the signal field mainly includes frequency components PRC3-PRC6 with varying intensities and observable wavelet morphology.

[0086] Figure 3 Part b) also presents the unimproved inherent timescale decomposition results. In contrast, the unimproved algorithm's time-domain decomposition results exhibit missing wavelet valleys in the seismic signal, as indicated by the arrows, and weak wavelet waveform distortion enclosed in elliptical frames, significantly reducing the fidelity of the seismic signal after component decomposition. From the energy spectral density distribution curve, the unimproved algorithm lacks strong frequency division capability, such as... Figure 3 In part c), the high-frequency noise components above 400Hz shown in PRC1 and PRC2 were not successfully separated, causing them to continue to mix in the signal wave field and form high-frequency interference. Therefore, the improved intrinsic time-scale algorithm can further reduce aliasing at different frequency scales, obtain more accurate time information with higher resolution, lay a preliminary multi-component foundation for the high-dimensional feature space spanning in the next stage, and is more suitable for the separation of signal and noise fields in well DAS records with complex and diverse frequency components.

[0087] By retaining and adding the components from the decomposition results of the two methods in which the signal wavelet shape is clearly observable, the separated signal wavefield can be obtained. This wavefield is compared with the simulated clean signal, and the separated wavefield results are described by combining the objective time-domain indicators of signal-to-noise ratio (SNR) and root mean square error (RMSE) with the objective frequency-domain indicator of mean frequency, as shown in Table 1 below:

[0088] Table 1. Comparison of data retention results in single wells processed by the inherent time-scale decomposition method before and after improvement.

[0089] Improved inherent timescale decomposition 9.3217 0.1306 75.9578 Inherent timescale decomposition algorithm 8.9786 0.1359 97.2515

[0090] The mean frequency of the pure seismic signal is 72.7339 Hz; the mean frequency of the noisy data is reduced to 48.4472 Hz due to fading noise interference, with a signal-to-noise ratio (SNR) of -1.4420 dB and a mean square error (MSE) of 0.4511. The comparison results further confirm the conclusions above. The improved algorithm demonstrates a more significant quality improvement for noisy signals, exhibiting a higher SNR (9.3217 dB) and a smaller MSE (0.1306). Furthermore, the mean frequency recovered by the improved algorithm is closer to the true frequency of the simulated pure seismic signal.

[0091] 2. Extraction of six-dimensional feature points from high-dimensional attribute space in time and frequency domains using multiple components

[0092] Figure 4 The distribution of multidimensional eigenvalues ​​extracted from the multi-components obtained by the improved intrinsic time-scale decomposition using six time-frequency domain feature factors, including kurtosis, margin, and frequency standard deviation, is presented in histogram format. Figure 4 In terms of the characteristics of simulated clean data, actual fading noise, and aliased noisy data in part a), the clean seismic signal has a significantly higher distribution of time-domain eigenvalues ​​and lower frequency-domain eigenvalue disturbances, while the six-dimensional characteristics of fading noise and noisy data are closer. This indicates that actual fading noise has a strong interference effect on the property characterization of seismic signals, which greatly increases the difficulty of restoring the complete energy field of low-loss or lossless signals.

[0093] In the feature representation process, the margin factor (feature 4) exhibits more drastic numerical fluctuations before and after noise addition compared to the other three types of feature factors. Therefore, it demonstrates superior stability compared to the other three types of features and has better noise sensitivity. The diverse selection of feature factors with different strengths can comprehensively and from multiple perspectives uncover the differences in the time-frequency domain characteristics of the signal and noise fields, thus providing a sufficient data foundation for the separation of complex superimposed multi-wave fields and making the partitioning of different types of feature point values ​​in high-dimensional space highly feasible.

[0094] Based on the preliminary classification results of the signal and noise wave field components (the signal component is light gray), the time and frequency domain eigenvalue distribution characteristics of each component are analyzed: the time domain, frequency domain, or both domain characteristics of the PRC3-PRC5 components obtained by the improved inherent time-scale decomposition algorithm all exhibit obvious signal characteristic tendencies, making them easy to classify as signals in the feature space. This proves that the improved algorithm preserves the seismic signal morphology in the wavelet relatively completely and is identifiable, while retaining less noise interference. Although the seismic wavelet morphology can be observed in the decomposition results of PRC6, its eigenvalue distribution is closer to that of PRC7 and PRC8, and the criteria for classifying it are not as clear as those for PRC3-PRC5. This necessitates the use of feature classification to determine these unclear components, directly determining the accuracy of signal wavefield recovery and the degree of noise wavefield residue, thus becoming a key challenge in the feature classification stage. The unimproved algorithm also suffers from the problem of fuzzy component characteristics. Figure 4 (Part b)). However, apart from this, the components PRC1 and PRC2, which exhibit obvious high-amplitude signal characteristics in the time-domain decomposition results, did not show the clear time and frequency domain feature values ​​they should have presented—their time-domain feature value distribution is closer to the distribution of pure signal values, but their frequency-domain feature value distribution is closer to the distribution of fading noise feature values. The reasons for this contradiction are twofold: first, the failure to decompose high-frequency noise components during the decomposition process in the unimproved algorithm leads to frequency domain feature interference; second, it is also highly correlated with signal distortion caused by linear contraction transform. Both of these causes are effectively improved in the improved decomposition algorithm through the sieving process and cubic spline interpolation, thus providing a clearer and easier-to-distinguish time and frequency feature data foundation.

[0095] 3. Tree-based ensemble model for feature classification and discrimination in high-dimensional space to achieve signal-noise field separation.

[0096] by Figure 4 Based on the results of the multi-component feature difference analysis, in utilizing Figure 5 When the dual-stage highly integrated tree model with the structure shown is used to determine the category of feature points in the high-dimensional feature attribute space, the model should, on the basis of ensuring the correct classification of feature points with obvious signal and noise characteristics, further improve the classification accuracy of fuzzy feature points, so as to minimize the loss of signal energy and the incomplete suppression of noise during the separation of signal and noise wave fields.

[0097] For example Figure 6 The actual well DAS seismic exploration record shown in the figure further confirms the applicability and reliability of the method of the present invention for separating complex signal and noise fields. Figure 6The record, acquired from oil wells in the Tarim Basin of China, consists of 1500 seismic traces acquired at a frequency of 2500 Hz with a sampling duration of 919.6 ms. The uplink and downlink wave fields in this record are complex, with dense primary reflection wave fields severely interfered with by optical noise, significantly reducing signal continuity. Multiple types of noise fields are unevenly distributed, with strong energy noise in some blocks completely obscuring the temporal representation of the signal wave field within those areas. Therefore, suppressing noise fields while simultaneously recovering the signal wave field with high precision and low damage has become a key objective of this current mission.

[0098] Taking regions A and B, where the distribution of signal and noise wavefields differs significantly, as examples, a specific analysis is conducted: Region A mainly includes a direct wavefield, a dense primary reflection wavefield formed by thin geological structures, and a complex noise wavefield formed by fading noise, strong optical spots, and background random noise; Region B mainly consists of a signal wavefield formed by weak primary reflections and weak converted waves, as well as a noise wavefield composed of fading noise and background random noise. Based on the recording characteristics of these two regions, the signal and noise wavefields differ significantly in both their two-dimensional morphological characteristics in the spatiotemporal domain and their energy distribution in the frequency domain. This necessitates that the decomposition algorithm in the technical solution should be able to achieve preliminary independence of the multiple components containing the signal and noise components in the context of non-uniform wavefield distributions between different regions and seismic traces, and that the tree-structured ensemble model should possess a highly efficient ability to distinguish and differentiate high-dimensional spatial feature points.

[0099] Referring to the above technical process Figure 6 The actual well DAS exploration data was processed and... Figure 7 The image shows the separated uplink and downlink wave fields. Figure 7 Part a) and complex noise combined wave field ( Figure 7 (Part b)). For ease of representation, the above technical process method is abbreviated as CITD-TS. Simultaneously, four methods are set up for comparison: the method of retaining fixed components PRC2-PRC5 based on decomposition (CITD), the multi-component multidimensional feature model discrimination method based on the unimproved algorithm (ITD-TS), the wavelet threshold filtering method, and the principal component analysis (RPCA) method. Comparing the separation results of CITD-TS and CITD methods, it can be seen that both methods can effectively separate the signal and noise wave fields. In the signal wave field, the direct wave and reflected wave in region A and the transmitted wave in region B are clearly visible, and the noise wave field also contains the complete forms of fading noise and optical spot noise. However, compared to the former, the CITD method, which only retains fixed components, leads to… Figure 7 The direct wave energy at the point indicated by the arrow in section a) is severely attenuated, and... Figure 7In part b), a significant residual primary reflection wave field is observed, indicating that the signal wave field in this part is still superimposed with the noise wave field, and the separated signal wave field exhibits considerable loss of frequency components and attenuation of amplitude characteristics. In contrast, a feature classification model with high accuracy and high generalization is more sufficient and necessary for preserving real-time signal energy when separating the signal and noise wave fields of complex DAS records in wells. The ITD-TS algorithm and wavelet threshold filtering method have limited ability to separate the signal and noise wave fields, and the superposition between the two is still severe. The RPCA algorithm only successfully separated the patchy optical noise field as shown by the ellipse. For the dense reflection wave signal field and noise wave field in the 100-400ms range on the left side of region A, the separation is incomplete and uneven, and the energy continuity of the signal wave field is still severely disturbed. Based on the above results, it can be concluded that the CITD-TS method is not only correct and highly feasible, but also has superior ability to separate multi-superimposed signal and noise wave fields, and can be better applied to the signal and noise separation of DAS records in wells with different processing requirements and different characteristics.

[0100] In summary, the tree-structured ensemble model for separating signal and noise fields in a multi-component high-dimensional feature attribute space is effective. Starting from the time-frequency domain characteristics of DAS records in wells, this method considers the diverse types and varying spatiotemporal characteristics of signal wavefields, as well as the uneven distribution and strong energy interference of noise fields. It specifically improves upon the shortcomings of traditional time-scale decomposition algorithms, such as incomplete decomposition leading to signal distortion, obtaining time-domain multi-components with less frequency component aliasing and clearer time-scale information. Based on this data, a high-dimensional feature attribute space is constructed. A tree-structured ensemble classification model is designed to meet the specific needs of multi-dimensional feature point category representation, effectively completing the task of multi-dimensional feature point category determination. This allows for accurate and effective separation of signal and noise components even in the case of multiple overlapping wavefields. This invention provides sufficient and reliable theoretical support for the modeling, analysis, and separation of signal and noise multi-wavefields in well-drilled DAS data, and also lays a corresponding data foundation for subsequent inversion and interpretation work.

[0101] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the invention by those skilled in the art. Any modifications, equivalent substitutions, or improvements made to the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for intelligent separation of high-dimensional multi-component features of signal-noise superimposed wavefields in well seismic boreholes, characterized in that: Includes the following steps: Step (1) Obtaining DAS 2D seismic exploration records in the well; Step (2) Establishment of multi-component high-dimensional time-frequency feature attribute space: Step (2.1) establishes the multi-component decomposition components using an improved intrinsic timescale decomposition method: The linear shrinkage transformation method in the inherent time scale decomposition is replaced by the cubic spline difference algorithm, and all extreme points are fitted in series in the form of smooth curves, thereby effectively improving the problems of wavelet spikes, distortion and discontinuity of signal morphology after decomposition; on the other hand, a sieving process is added in each decomposition process, and small iteration loops are nested in the large iteration to refine the coarse components obtained by the large iteration, further reducing the overlap range of information at different time scales and reducing the degree of frequency component aliasing. At the same time, based on the characteristics of uneven noise distribution intensity and large energy difference in the DAS record in the well, a flexible iteration termination condition based on standard deviation and similarity measurement is proposed. That is, the Manhattan distance with clear physical meaning is used to measure the difference between the baseline signal and the original decomposition component before and after each sieving iteration of the seismic record data, and the standard deviation of formula (1) is used as the standard for measuring the smoothness of the distance curve to determine the number of iteration terminations, so as to avoid using a uniform standard as an inappropriate and unsuitable measurement under different data performance conditions. (1); in, Represents signal data, The mean of the data is represented by N, where N is the signal length. The mathematical process of the improved intrinsic time-scale decomposition algorithm is described as follows: (2); Where PRC(t) is the decomposition component obtained in each iteration, h is the number of components obtained by decomposition, l represents the number of sieving iterations, and r(t) and r*(t) represent the residual components remaining after decomposition. Step (2.2) Multi-dimensional time and frequency domain feature factors constitute a six-dimensional feature attribute space: Using time-frequency domain features, six-dimensional time-frequency independent domain features are extracted for each component, and a six-dimensional feature attribute space is constructed based on this. In this six-dimensional feature attribute space, the decomposition of multi-component results of DAS records in wells is characterized and described by six-dimensional feature factors, further amplifying the differences in the performance of signal and noise components that are not easily observed in low-dimensional space, thereby providing a multi-dimensional spatial data basis for the final separation of different categories of components. The feature factors used include four time-domain features: kurtosis, peak factor, impulse factor, and margin factor, Equation (3)-Equation (6), and two frequency-domain features: root mean square frequency and frequency standard deviation, Equation (7) and Equation (8). These features respectively extract time-domain information including the amplitude fluctuation and peak distribution characteristics of each wavelet component of the seismic data, as well as frequency-domain information including the main frequency band range and frequency energy fluctuation. (3); (4); (5); (6); (7); (8); Among them, f k For a signal to contain frequency components, S k Let K be the signal power spectral energy, and K be the total number of frequency components. max(·) is the operation to find the maximum value of the signal; Step (3) Output the probability prediction results of the tree ensemble model; Step (4) Separation of signal and noise multi-overlapping wave fields.

2. The intelligent separation method for high-dimensional multi-component features of the signal-noise superimposed wavefield in well seismic boreholes according to claim 1, characterized in that: The acquisition of the well DAS two-dimensional seismic exploration record in step (1) is as follows: After the distributed optical fiber sensing technology emits a laser into the optical fiber deployed in the well, it uses the optical fiber to realize the integration of stress sensing and optical signal transmission during the propagation of seismic waves; the optical signal is received at the ground receiving end and demodulated and converted into an electrical signal to obtain the well exploration two-dimensional seismic record.

3. The intelligent separation method for high-dimensional multi-component features of the signal-noise superimposed wavefield in well seismic boreholes according to claim 1, characterized in that: The output of the probability prediction result of the tree-like ensemble model in step (3) is: to formulate a two-stage ensemble large-batch decision tree model to construct a comprehensive strong learning model that completes the signal-noise multi-component feature point classification and discrimination task with high precision and high accuracy in multi-dimensional space. The model construction process is mainly divided into the following two stages: Initial ensemble stage: The learner structure based on the large batch decision tree model is trained and generated differently using feedback training and self-sampling training methods respectively. Then, according to the classification or regression task, the appropriate voting method or averaging method is selected to integrate them into multiple strong integrated models with different emphases. In this process, the instability of using only a single learner in the classification process is effectively avoided. At the same time, the integration approach that takes into account both accuracy and diversity indicators greatly improves the final prediction accuracy of the single learner, thereby constructing a tree-like decision model with high classification accuracy. Secondary integration stage: Considering the randomness of noise interference and the differences between shot gather records in actual exploration, to improve the generalization performance of the model as much as possible, a secondary integration is carried out on the basis of the primary integration according to the meta-learning strategy, so as to output a highly integrated and highly applicable classification learning model with decision tree as the basic structure. It not only has the ability to accurately distinguish fuzzy multi-component feature point data of DAS records in wells, but also can fully cope with various types of interference in actual DAS data.

4. The intelligent separation method for high-dimensional multi-component features of the signal-noise superimposed wavefield in well seismic boreholes according to claim 1, characterized in that: The signal-noise multi-overlapping wave field separation in step (4) is as follows: After establishing a multi-component, multi-dimensional feature attribute space in sequence, the constructed tree-like integrated model is used to classify a large number of feature sample points mapped from the DAS data in the well to the high-dimensional space; the feature categories are divided into signal and noise categories, and the output of the integrated model is determined as the most likely category probability prediction for a certain feature point: the prediction with a probability greater than 50% is the signal category, and the prediction with a probability less than 50% is the noise category; the feature corresponding components that are judged to be the signal category are retained and added together, while the feature corresponding components that are judged to be the noise category are discarded; through the above operations, the signal and noise wave fields in the DAS record in the well can be separated.