Sea sand automatic identification method based on multi-source geophysical data fusion
By constructing a dual-channel deep learning model that integrates drilling, single-channel seismic, and shallow seismic profile data, efficient and accurate automatic identification of marine sand was achieved. This solves the problems of low efficiency and low accuracy in marine sand identification in existing technologies, and improves the accuracy and efficiency of marine sand resource identification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- INST OF GEOMECHANICS
- Filing Date
- 2026-01-26
- Publication Date
- 2026-05-12
AI Technical Summary
In existing technologies, sea sand identification methods rely on single seismic data, which is susceptible to noise interference. The identification accuracy of thin sand layers is insufficient, shallow stratigraphic profile data has not been effectively fused, and the sparsity of drilling data limits the effectiveness of model training. As a result, sea sand identification is inefficient and inaccurate, making it difficult to meet the needs of precise exploration in large-scale sea areas.
By acquiring drilling data, single-channel seismic data, and shallow seismic profile data, a dual-channel deep learning model is constructed to perform data fusion and intelligent analysis. The model is trained using lithological labels calibrated in the fused data volume and drilling data to achieve efficient, accurate, and automatic identification of marine sand.
It significantly improves the accuracy and efficiency of marine sand resource identification, solves the problems of inconsistent spatial benchmarks and large scale differences in multi-source data, enhances the ability to identify complex geological features, and significantly improves the mudstone exclusion effect.
Smart Images

Figure CN122017970A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of marine geological exploration, and in particular to an automatic identification method for sea sand based on the fusion of multi-source geophysical data. Background Technology
[0002] Marine sand, as an important marine mineral resource, is widely used in construction, road building, and other fields. Its accurate identification and distribution survey are crucial for the rational development and utilization of this resource. Traditional marine sand identification methods mainly rely on manual analysis of seismic and drilling data, which suffers from low efficiency, strong subjectivity, and significant reliance on experience for accuracy. While existing technologies have attempted to use single seismic data combined with machine learning for lithological identification, single data types have limitations: single-channel seismic data is susceptible to noise interference, resulting in insufficient accuracy in identifying thin sand layers; shallow seismic profile data is mostly used for qualitative analysis and has not been effectively integrated with other data; furthermore, there is a lack of quantitative utilization of the "penetrating mud but not sand" physical characteristic of shallow seismic profiles, leading to incomplete mudstone removal and a high misjudgment rate in marine sand identification. The sparsity of drilling data also limits the effectiveness of model training, making it difficult to meet the needs of accurate marine sand exploration over large areas. Therefore, how to fully integrate the complementarity of multi-source geophysical data, utilize deep learning technology to achieve efficient, accurate, and automatic identification of marine sand, and enhance the mudstone removal effect by leveraging the penetration differences of shallow stratigraphic profiles has become a pressing technical problem to be solved in the field of marine geological exploration. Summary of the Invention
[0003] Therefore, it is necessary to provide an automatic identification method for marine sand based on the fusion of multi-source geophysical data to solve at least one of the above-mentioned technical problems.
[0004] To achieve the above objectives, an automatic identification method for sea sand based on multi-source geophysical data fusion includes the following steps:
[0005] Step S1: Acquire drilling data, single-channel seismic data, and shallow seismic profile data of the target area, and preprocess them to generate a fused data volume;
[0006] Step S2: Based on the fused data volume, construct a dual-channel deep learning model;
[0007] Step S3: Train a dual-channel deep learning model using the lithological labels identified in the fused data volume and drilling data;
[0008] Step S4: Use the trained dual-channel deep learning model to identify the target area, verify the identification results, and generate the sea sand distribution results.
[0009] The beneficial effects of this invention are as follows: Through deep fusion and intelligent analysis of multi-source geophysical data, the accuracy and efficiency of marine sand resource identification are significantly improved. Firstly, by establishing a spatial registration and scale normalization mechanism for drilling data, single-channel seismic data, and shallow seismic profile data, a fused data volume with accurate spatial mapping relationships is constructed. This solves the technical problems of inconsistent spatial benchmarks and large scale differences in multi-source data in traditional methods, providing a reliable data foundation for subsequent analysis. Secondly, an innovative dual-channel deep learning model architecture is designed, optimizing the extraction of seismic time-series features and shallow seismic profile image features respectively. Through a feature fusion layer, multi-dimensional information is effectively integrated, overcoming the limited feature representation capabilities of a single data source and significantly enhancing the model's ability to identify complex geological features. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating the steps of an automatic identification method for sea sand based on the fusion of multi-source geophysical data.
[0011] Figure 2 for Figure 1 A detailed flowchart illustrating the implementation steps of step S2.
[0012] Figure 3 Flowchart for automatic sea sand identification;
[0013] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0014] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0015] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0016] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0017] To achieve the above objectives, please refer to Figures 1 to 3 An automatic identification method for sea sand based on multi-source geophysical data fusion includes the following steps:
[0018] Step S1: Acquire drilling data, single-channel seismic data, and shallow seismic profile data of the target area, and preprocess them to generate a fused data volume;
[0019] Step S2: Based on the fused data volume, construct a dual-channel deep learning model;
[0020] Step S3: Train a dual-channel deep learning model using the lithological labels identified in the fused data volume and drilling data;
[0021] Step S4: Use the trained dual-channel deep learning model to identify the target area, verify the identification results, and generate the sea sand distribution results.
[0022] All specific values involved in this embodiment are exemplary parameters used to clearly illustrate the technical operation process and are not the only limitation of the present invention.
[0023] In one embodiment, the initial step is to acquire the original geophysical and geological data of the target work area. Drilling data is derived from completion reports and digital logging curves of all completed engineering geological boreholes and offshore oil wells within the target area. Specifically, this includes the geographic coordinates, depth, sonic transit time curves, density curves, gamma curves, and lithological columnar sections determined by geological logging for each well point. Lithological labels are represented by numerical codes, where 1 represents sand layers and 0 represents non-sand layers such as mudstone. Single-channel seismic data is acquired from the Boomer or Sparker source and single-channel hydrophone receiving system deployed on the survey vessel. The data format is SEG-Y standard seismic gathers, and each survey line includes geographic coordinates, two-way travel time, and reflected wave amplitude sequences. Shallow seismic profile data is acquired from the parametric array shallow profile system aboard the same survey vessel. The data format is JSF or XTF profile image files, and each survey line includes geographic coordinates, depth, and reflection intensity information. Preprocessing begins with data organization, extracting the top and bottom depths of all labeled sand layers from drilling data, calculating their thickness, and simultaneously recording the corresponding acoustic velocity and resistivity values at each depth point, forming a "Drilling-Sand Layer-Property" index table. For single-channel seismic data preprocessing, professional seismic interpretation software (such as KingdomSuite) is used to perform bandpass filtering on the raw SEG-Y data, filtering out high-frequency noise and low-frequency surge interference, retaining the effective frequency band from 20Hz to 1000Hz. Subsequently, automatic gain control is applied to each seismic channel to balance the reflection amplitudes of shallow, intermediate, and deep layers. For shallow stratigraphic profile data preprocessing, a shallow profile processing system is used to perform TVG time-varying gain compensation and water multiple wave suppression on the original profile, generating a profile image with relatively balanced reflection energy. The core fusion processing uses drilling data as the absolute spatial reference. First, spatial registration is performed: the navigation coordinates from the single-track seismic survey line header file, the navigation coordinates of the shallow seismic profile survey line, and the wellhead coordinates are imported into ArcGIS software and unified into the WGS-84UTM projection coordinate system. Through spatial interpolation and survey line resampling, it is ensured that each depth sampling point has unique and consistent geographic coordinates in three-dimensional space, with a registration accuracy requirement of less than 1 meter in planar position error. Next, scale normalization is performed: based on the sonic velocity in the "Drilling-Sand Layer-Physical Property" index table, the Dix formula is used to generate a time-depth converted velocity field for the neighborhood of each well. This velocity field is used to batch convert the single-track seismic data from the time domain to the depth domain, with the converted depth sampling interval set to 0.1 meters. At the same time, the shallow seismic profile data is also resampled at a depth interval of 0.1 meters. Finally, for each spatially matched point, its data volume contains three strictly aligned data layers: the depth domain amplitude sequence from the single-track seismic data, the depth domain reflection intensity sequence from the shallow seismic profile, and the lithological labels and physical property parameters from the well. This data body is defined as the merged data body.
[0024] The data structure and features are based on the fused data volume generated in step S1. The model adopts a modular architecture, with a computational graph at its core consisting of two independent feature extraction paths and a fusion decision head. The first channel, the seismic time series feature extraction channel, is specifically designed for processing single-channel seismic amplitude sequences in the fused data volume. The input layer of this channel is defined as a one-dimensional vector of length L, where L corresponds to the number of depth sampling points within an analysis window in the fused data volume. For example, L=256 corresponds to a depth range of 25.6 meters. The main structure of this channel consists of four sequentially connected one-dimensional convolutional modules, each containing a one-dimensional convolutional layer, a batch normalization layer, and a ReLU activation function. The first convolutional layer uses 64 filters with a width of 5 to capture wavelet waveform features; subsequent convolutional layers increase the number of filters and decrease the width to extract higher-level abstract reflection patterns. The second channel, the shallow seismic profile feature extraction channel, is specifically designed for processing shallow seismic profile image blocks in the fused data volume. The input layer of this channel is defined as a two-dimensional matrix with a receiving size of H×W, where H represents the number of pixels in the depth direction (aligned with L), and W represents a finite number of channels or feature dimensions in the horizontal direction. The main structure of this channel employs a miniaturized two-dimensional convolutional neural network containing two convolutional pooling combinations. In each combination, a 3x3 convolutional layer is used to extract local reflection structure features, followed by a 2x2 max-pooling layer to reduce spatial dimensionality and increase feature invariance. After extracting high-level features, the output feature maps of the two channels are flattened into one-dimensional feature vectors. These two feature vectors are then fed into the third part of the model—the feature fusion and decision module. This module first concatenates the two vectors along the feature dimension to form a longer joint feature vector. This joint vector is then processed through three fully connected layers, each followed by batch normalization and the ReLU activation function. The output neurons of the last fully connected layer have two numbers, corresponding to the sand and non-sand classes respectively, and output a normalized probability distribution through the Softmax function. The entire model structure is defined through the code implementation of the deep learning framework PyTorch, and all its parameters (weights and biases) are in a pre-training state during initialization.
[0025] The entire fused dataset is divided into training and validation sets. A stratified random sampling strategy based on drilling points is employed to ensure that the ratio of sandstone to mudstone samples in the training and validation sets is consistent with the overall distribution, typically 8:2. Before training, the input data is standardized by subtracting the mean from the single-channel seismic amplitude value and the shallow seismic profile reflection intensity value, and then dividing by its standard deviation. Training is conducted on a server equipped with an NVIDIA RTX 3090 graphics processor using the PyTorch framework. The optimizer is Adam, with an initial learning rate of 0.001. The cross-entropy loss function is used, with the model's predicted class probability distribution and the actual drilling lithology labels as inputs. Training is performed in batches with a batch size of 32. In each training batch, the program randomly selects 32 sample pairs from the training set. Each sample pair contains a single-channel seismic sequence, a corresponding shallow seismic profile image block, and a lithology label at a center point depth. The model forward propagates to obtain 32 prediction results, and the average loss value is calculated. The gradient of the loss function with respect to each model parameter is calculated using the backpropagation algorithm. The optimizer updates the values of all model parameters based on the gradient direction and learning rate, aiming to minimize the loss function value. During training, the penetration constraint of the shallow seismic profile is added as a hard rule. Specifically, after each model makes a prediction on a batch of data, before calculating the final loss, a post-processing logic is added: for each sample point predicted as sand, the signal attenuation of its corresponding original shallow seismic profile below that depth is read from the fused data volume. A penetration depth threshold H is defined, for example, 2.5 meters. If the effective reflection signal of the shallow profile below this point is continuous, clear, and traceable to a depth exceeding H, then the point is determined not to meet the typical characteristics of strong reflection and shallow penetration of sand layers. Regardless of the original predicted probability of the model, the predicted label of this sample point is forcibly corrected to non-sand, and the corrected label is used in the loss calculation. The threshold H is determined based on the maximum thickness of sand layers in known wells in the work area and the typical penetration depth of the shallow profile system in mudstone. The entire training iteration is carried out. After each round, the model performance is evaluated on the validation set. When the validation set loss no longer decreases for 5 consecutive rounds, the early stopping mechanism is triggered, and training terminates the dual-channel deep learning model that can be used for prediction.
[0026] Prepare the target area's identification data. This data needs to undergo the same preprocessing steps as in step S1 to generate a standard-format "target area fusion data volume," but its drilling data is only used for subsequent validation and does not include lithology labels. Divide the target area into a grid (e.g., 10m × 10m), extract the vertical data of each grid point, forming samples with a structure completely consistent with that used during training. Then, load the optimal model parameter file saved in step S3 and set the model to inference mode. Batch input all sample point data from the target area into the model, and perform forward propagation, outputting a probability value for each point belonging to a sand layer. This value ranges from 0 to 1, forming a preliminary marine sand probability volume. Apply a threshold (e.g., 0.5) to the probability volume for binarization to obtain "preliminary marine sand identification data," which is a binary matrix in three-dimensional space, where 1 indicates a prediction of sand and 0 indicates a prediction of non-sand. Validation is divided into quantitative and qualitative parts. Quantitative validation: Within the target area, select reserved wells that did not participate in model training and validation as validation wells. In the preliminary marine sand identification data, all predicted points traversed by the trajectory of each verification well are extracted and compared with the actual lithology columnar section of the drilling data at each depth. Four core indicators are calculated: prediction error of the top boundary depth of the sand layer, prediction error of the bottom boundary depth of the sand layer, accuracy of sand layer identification, and mudstone exclusion rate. The average depth prediction error is no greater than 5% of the actual thickness, and the mudstone exclusion rate is no less than 90%. Qualitative verification: Several representative complete single-channel seismic survey lines and shallow seismic profile lines are selected in the target area, and a comprehensive comparative profile map is generated using professional plotting software (such as MATLAB or Python's Matplotlib library). Under the same depth-horizontal coordinate system, three profiles are drawn side by side from top to bottom: the first line shows the single-channel seismic amplitude color profile, the second line shows the shallow seismic profile grayscale or pseudo-color profile, and the third line shows the area identified as a sand layer in the "preliminary marine sand identification data" in a transparent color block overlay manner. The comparative map was independently interpreted by at least two experienced marine geologists, focusing on checking the continuity of the identified sand bodies in the lateral direction, whether their boundaries coincided with the termination points of seismic reflections or abrupt changes in shallow profile features, and whether there were any isolated anomalies that clearly violated geological laws. Combining quantitative verification indicators with human qualitative interpretation, the preliminary marine sand probability volume underwent local correction and smoothing. For example, sections with depth errors exceeding limits were manually adjusted, and small-scale anomalies deemed unreasonable by the engineers were removed. The corrected 3D identification data volume was then spatially rasterized and mapped. The output "Marine Sand Distribution Results" included at least three maps: a contour map of the burial depth of the top interface of the marine sand layer; a contour map of the thickness of the marine sand layer; and a confidence zoning map of the marine sand identification results. This map divided the work area into high-confidence, medium-confidence, and low-confidence zones based on the predicted probability values and human verification results.
[0027] It is necessary to supplement the data of five wells in a certain sea area, and extract parameters such as the top and bottom burial depth (range 5-30m), thickness (1-5m), density (2.0-2.2g / cm³), and wave velocity (1500-1800m / s) of the sand layer in each well; collect data of 20 single-channel seismic survey lines in the sea area (sampling rate 250Hz, time length 2s); collect data of 20 shallow stratigraphic profiles in the corresponding area (grayscale image format, resolution 512×1024). (2) Data feature analysis: Through well calibration, it is clear that the sand layer in this area exhibits medium to strong amplitude (amplitude value ≥0.6) and continuous reflection axis (continuity ≥80%) in single-channel seismic surveys, while the mudstone exhibits weak reflection (amplitude value <0.3) and chaotic reflection; in the shallow stratigraphic profiles, the sand body has strong reflection energy, penetration depth ≤15m, and missing lower layers in the corresponding area, while the mudstone has a penetration depth ≥30m, weak reflection, and continuous layers in the corresponding area. (3) Spatial registration: Based on the GPS coordinates of the survey line and the well location coordinates, ArcGIS software is used to project the single-channel seismic and shallow seismic profile data to the WGS84 coordinate system. Spatial matching between the well and the survey line is achieved through coordinate correction, and the error is controlled within 2m.
[0028] Please refer to [link / reference needed] for further information. Figure 3 The three graphic squares on the left of the image represent the three types of raw input data: drilling data, single-channel seismic data, and shallow seismic profile data. The model employs two parallel feature extraction channels: a seismic time-series feature channel, specifically processing single-channel seismic time-series data from the fused data volume, and an image profile feature channel, specifically processing shallow seismic profile image data from the fused data volume. The features extracted from both channels are ultimately integrated into a single fusion network layer to form a complete dual-channel deep learning model.
[0029] Preferably, step S1 includes the following steps:
[0030] Step S11: Acquire drilling data, single-channel seismic data, and shallow seismic profile data of the target area, and perform preprocessing;
[0031] Step S12: Using drilling data as a benchmark, analyze single-channel seismic data and shallow seismic profile data to establish the correspondence between lithology and geophysical response;
[0032] Step S13: Based on the correspondence, project the single-channel seismic data and shallow seismic profile data to the same geographic coordinate system to generate spatial registration data;
[0033] Step S14: Convert the single-channel seismic data in the spatial registration data from the time domain to the depth domain to generate the fused data volume.
[0034] In one embodiment, completion reports of all completed geological boreholes within the target sea area are collected, including wellhead latitude and longitude coordinates, digital logging curves for the entire well section, and lithology columnar sections determined by core analysis. Simultaneously, raw seismic data files and shallow profile data files acquired during marine geophysical survey cruises within the same sea area are obtained. The seismic data is a SEG-Y format gather received using a Boomer source and a single-channel hydrophone, recording shot coordinates, two-way travel time, and reflection amplitude. The shallow profile data is an XTF format profile file acquired using a parametric array source, recording survey line coordinates, depth, and reflection intensity. First, these three types of raw data undergo format standardization and noise reduction. Environmental correction is applied to the drilling logging curves, and a 20-800Hz bandpass filter is used on the seismic data to suppress swell and high-frequency noise. TVG gain compensation is used on the shallow profile data to suppress water column multiples.
[0035] Using the well lithology columnar section as the absolute standard, the locations corresponding to the top and bottom depths of the sand layers exposed by the well are determined along the spatially closest single-channel seismic trace and shallow seismic profile. On the single-channel seismic trace, the amplitude, frequency, phase, and waveform envelope characteristic parameters corresponding to the top and bottom reflection interfaces of the sand layers are extracted. On the shallow seismic profile image, the intensity, continuity, and attenuation of the reflection signal from the underlying layers corresponding to the sand layers are extracted. The various physical properties of the sand layers are correlated with the extracted geophysical characteristic parameters and entered into a database to form a "Sand Layer-Seismic Response-Shallow Profile Response" lookup table. This process is repeated for the mudstone layers to establish a comparative feature database.
[0036] The database established in step S12 is invoked to retrieve the original navigation coordinates of all recorded seismic traces and shallow seismic profiles. Using a coordinate transformation library, all coordinates are uniformly converted to Gauss-Kruger projection coordinates in the CGCS2000 coordinate system. Based on the transformed coordinates, spatial interpolation resampling is performed on the single-trace seismic line points and shallow seismic profile points to generate a data point sequence with perfectly matched planar positions and consistent sampling intervals. Each point simultaneously has both a seismic trace index and a shallow seismic profile index, thereby generating a precisely matched spatial registration data index file.
[0037] Extract the corrected drilling sonic transit time curve from step S11 and calculate the layer velocity. Using this velocity, apply the ray tracing time-depth conversion algorithm to each single-channel seismic record in the spatial registration data index file generated in step S13 to convert the seismic reflection time series into a depth series, with the conversion depth interval set to 0.1 meters. Simultaneously, resample the spatially registered shallow seismic profile depth series at 0.1-meter intervals. Merge the converted depth-domain single-channel seismic amplitude series and the resampled shallow seismic profile reflection intensity series according to the spatial registration data index file, generating a data unit for each spatial coordinate point containing a seismic amplitude vector, a shallow seismic profile intensity vector, and a corresponding depth label. The collection of all data units constitutes the fused data volume used for model input.
[0038] It is necessary to supplement this by using drilling data as the calibration benchmark to extract lithological parameters such as the top and bottom burial depth, thickness, density, and wave velocity of sand layers, and establish a "lithology-logging response" correspondence; analyzing the differences in seismic response between sand layers (medium-strong amplitude, continuous reflection axis) and mudstone (weak reflection, disordered) in single-channel seismic data to clarify characteristic parameters such as amplitude, frequency, and phase; utilizing the penetration difference of shallow stratigraphic profiles ("penetrating mud but not sand") to identify the characteristic differences between sand bodies (strong reflection, abrupt changes in penetration depth, missing lower layers) and mudstone (deep penetration, weak reflection, continuous layers). Random noise is added to the seismic data, amplitude scaling, and time-shift transformation are applied to expand the training sample; interpolation is performed on the drilling calibration data to compensate for the lack of sparse well locations.
[0039] Preferably, step S12 includes the following steps:
[0040] Step S121: Based on drilling data, extract spatial location information of sand and mudstone layers from single-channel seismic data and shallow seismic profile data;
[0041] Step S122: Based on spatial location information, statistically analyze the reflection characteristic parameters of single-channel seismic data and the penetration characteristic parameters of shallow seismic profile data to establish the correspondence between lithology and geophysical response.
[0042] In one embodiment, the depths of the top and bottom of sand layers recorded in the drilling lithology column are used as a benchmark to locate corresponding layers in spatially registered single-channel seismic data and shallow seismic profile data. Specifically, the nearest neighbor interpolation method is used to determine the nearest seismic line point based on the spatial distance between the drilling coordinates and the seismic and shallow seismic lines. On the single-channel seismic trace, the position of the seismic reflection phase axis corresponding to the top and bottom interfaces of the sand layers is accurately calibrated according to the time-depth conversion relationship, and its time or depth coordinates are recorded. On the shallow seismic profile image, the position of the reflection phase axis corresponding to the sand layers is calibrated according to the depth correspondence relationship, and its depth coordinates are recorded. The same operation is performed on the mudstone layers, thereby extracting a set of precise spatial location information for the sand layers and mudstone in the two types of geophysical data.
[0043] Based on the extracted spatial location information, characteristic parameters of sandstone and mudstone layers were statistically analyzed in single-channel seismic data and shallow seismic profile data. For single-channel seismic data, within the calibrated sandstone and mudstone sections, reflection characteristic parameters such as dominant frequency, amplitude intensity, instantaneous phase, and waveform similarity of reflected waves were calculated. For shallow seismic profile data, within the calibrated sections, penetration characteristic parameters such as the continuity of the reflection phase axis, reflection intensity, and attenuation gradient of the reflected signal below the section were calculated. The statistically obtained characteristic parameters of sandstone and mudstone layers were compared and analyzed to clarify their differences, and a database of correspondences between seismic reflection characteristics and shallow seismic penetration characteristics was established, indexed by lithology. This database clearly defines the geophysical response models corresponding to different lithologies.
[0044] Preferably, step S13 includes the following steps:
[0045] Step S131: Obtain georeferenced information for single-channel seismic data and shallow seismic profile data in the corresponding relationship;
[0046] Step S132: Based on geographic reference information, perform coordinate transformation on single-channel seismic data and shallow seismic profile data respectively, and generate spatial registration data with a unified geographic coordinate benchmark.
[0047] In one embodiment, the original geographic reference information of the recorded single-track seismic lines and shallow seismic profiles is read from the lithology and geophysical response correspondence database established in step S12. This information contains the original coordinates of each shot point or each recording point, typically in WGS-84 geodetic latitude and longitude format, and includes the coordinate reference system code used.
[0048] Based on the georeferenced information obtained in step S131, the coordinate transformation function in the Geographic Information System (GIS) library is used to perform batch coordinate transformations on the original single-track seismic survey point set and the shallow seismic profile survey point set. The coordinates of all points are uniformly converted from the original WGS-84 geodetic coordinate system to Gauss-Kruger 6-degree zone projection coordinates under the CGCS2000 national geodetic coordinate system. After the transformation, in the unified Cartesian coordinate system, the survey points of both types of data are spatially sorted and resampled according to the new coordinates, ensuring that data points from both the single-track seismic and shallow seismic profiles correspond to each other at the same planar location, thereby generating a spatially registered data index table with strict spatial location matching relationships.
[0049] It should be added that spatial registration: based on the coordinates of the survey line and the well location, single-channel seismic and shallow seismic profile data are projected onto the same geographic coordinate system to achieve precise spatial matching between the well and the seismic survey line, with the error controlled within the meter level.
[0050] Preferably, step S14 includes the following steps:
[0051] Step S141: Obtain drilling logging speed information corresponding to the spatial registration data;
[0052] Step S142: Based on the drilling logging velocity information, perform time-depth conversion processing on the single-channel seismic time data in the spatial registration data to generate depth domain seismic data;
[0053] Step S143: Merge the depth domain seismic data with the shallow seismic profile data in the spatial registration data to generate a fused data volume.
[0054] In one embodiment, sonic logging curve data corresponding to the drilling points listed in the spatial registration data index file generated in step S13 are extracted from the logging data table in the project database. This data contains information on sonic propagation time (time difference) that varies with depth, which is used to establish the conversion relationship between time and depth.
[0055] A time-depth conversion algorithm based on drilling acoustic velocity is adopted. The specific process is as follows: using the acoustic transit time data obtained in step S141, the layer velocity model corresponding to the well is calculated through integration; then, this velocity model is applied to the single-track seismic time series of the corresponding area in the spatial registration data; using ray tracing or direct integration, the two-way travel time of the seismic records is converted into depth values, generating a depth-domain seismic amplitude sequence corresponding to the actual subsurface depth. The relative amplitude relationship remains unchanged during the conversion process.
[0056] The depth-domain seismic data output in step S142 is strictly aligned and merged with the shallow seismic profile reflection intensity data already in the depth domain in the spatial registration data, according to their shared spatial coordinates (planar position and depth). The data cell corresponding to each coordinate point contains the depth-domain seismic amplitude value, the shallow seismic profile reflection intensity value, and its spatial label, ultimately generating a structured three-dimensional fused data volume, which serves as the input for subsequent model training.
[0057] Preferably, step S2 includes the following steps:
[0058] Step S21: Based on the characteristics of single-channel seismic data in the fused data volume, construct a network channel for seismic time series characteristics;
[0059] Step S22: Based on the features of shallow seismic profile data in the fused data volume, construct a network channel for image profile features;
[0060] Step S23: Construct a fusion network layer based on the network channels of image profile features and earthquake time series features to form a complete dual-channel deep learning model.
[0061] In one embodiment, a network channel for seismic time series features is constructed based on the characteristics of single-channel seismic data in the fused data volume. The input layer of this channel is designed to receive a one-dimensional vector of length 256, corresponding to the seismic amplitude sampling sequence within a depth range of 25.6 meters in the fused data volume. The main structure of the channel consists of four one-dimensional convolutional modules connected in series. Each module contains a one-dimensional convolutional layer, a batch normalization layer, and a ReLU activation function. The first convolutional layer uses 64 filters with a width of 5, and the number of filters in subsequent layers increases to 128, while the width decreases to 3, to extract multi-scale time series features from local waveforms to macroscopic sequences.
[0062] Based on the features of shallow seismic profile data in the fused data volume, a network channel for image profile features is constructed. The input layer of this channel is defined as receiving a 256×8 two-dimensional matrix, where 256 represents the depth dimension and 8 represents the feature dimension. The channel employs a two-dimensional convolutional neural network architecture, containing two convolutional-pooling units. Each unit consists of a 3×3 convolutional layer and a 2×2 max-pooling layer. The first unit uses 32 filters, and the second unit uses 64 filters. Convolutional operations extract spatial structural features from the profile, and pooling operations enhance feature invariance.
[0063] The 256-dimensional feature vector output from the seismic time-series feature channel is concatenated with the 512-dimensional feature vector output from the image profile feature channel to form a 768-dimensional joint feature vector. A fusion network layer, consisting of three fully connected layers with 384, 192, and 2 neurons respectively, is then applied to this joint feature vector. The last fully connected layer uses the Softmax activation function to output the classification probabilities of sandy and non-sandy layers. This structure achieves deep fusion of seismic time-series features and image profile features, forming an end-to-end dual-channel deep learning model.
[0064] It should be added that the features extracted from the two channels are fused by concatenation and fully connected layers to generate a multi-dimensional feature matrix containing "lithology-seismicity-penetration".
[0065] Preferably, step S21 includes the following steps:
[0066] Step S211: Perform format checks and data integrity verification on the single-channel seismic time series data in the fused data volume;
[0067] Step S212: Construct a one-dimensional convolutional network structure based on the validated single-channel earthquake time series data to form a network channel for earthquake time series features.
[0068] In one embodiment, the format and data integrity of the single-channel seismic time series data stored in the fused data volume are checked. The checks include verifying whether the sampling rate, number of sampling points, and data format in the data file header conform to the SEG-Y standard; verifying the consistency of data length for each seismic channel and checking for data truncation or anomalous padding values; detecting whether the amplitude values are within the normalized range of [-1, 1], and identifying and marking outlier data points. Simultaneously, the completeness of the correspondence between data indexes and spatial coordinates is checked to ensure that each seismic channel has corresponding geographic coordinate information.
[0069] Based on validated single-channel earthquake time-series data, a one-dimensional convolutional network structure was constructed. The network input layer was designed to receive a one-dimensional time-series data of 1024 sampling points, corresponding to a 2-second earthquake record. The main body of the network contains three one-dimensional convolutional modules, each consisting of a one-dimensional convolutional layer, a batch normalization layer, and a ReLU activation function connected sequentially. The first layer uses 64 filters with a width of 21 and a stride of 2; the second layer uses 128 filters with a width of 11; and the third layer uses 256 filters with a width of 5. Each convolutional layer is followed by a max-pooling layer with a width of 2, and finally, a global average pooling layer outputs a 256-dimensional feature vector, forming the network channel for earthquake time-series features.
[0070] It should be added that the first channel (single-channel seismic feature extraction) adopts a 1D-CNN architecture. The input layer is the amplitude, frequency, and phase temporal features of a single-channel seismic wave (dimension 1×300). It passes through two convolutional layers (convolutional kernel size 3, number 32 and 64) and a pooling layer (pooling kernel size 2) to extract a 64-dimensional sand layer reflection feature vector. An attention mechanism (SE module) is added to strengthen the reflection axis weights corresponding to the drilling calibration sand layers.
[0071] Preferably, step S22 includes the following steps:
[0072] Step S221: Standardize the shallow seismic profile image data in the fused data volume to obtain standardized shallow seismic profile image data;
[0073] Step S222: Based on the format characteristics of standardized shallow seismic profile image data, construct a two-dimensional convolutional network channel containing convolutional and pooling layers.
[0074] In one embodiment, the shallow seismic profile image data stored in the fused data volume is standardized. First, the pixel value matrix of the original image data is read, and its numerical distribution range is detected. A maximum-minimum normalization method is used to linearly transform the pixel values to the [0,1] interval. The specific calculation formula is: (current pixel value - minimum value) / (maximum value - minimum value). Simultaneously, the images are resized to a fixed size of 256×256 pixels, and a bilinear interpolation algorithm is used to maintain the image aspect ratio. The processed data is saved as standardized shallow seismic profile image data.
[0075] Based on the format characteristics of standardized shallow seismic profile image data, a two-dimensional convolutional network channel is constructed. The input layer is designed to receive a single-channel image of 256×256×1. The network architecture contains four convolutional modules, each consisting of a two-dimensional convolutional layer, a batch normalization layer, a ReLU activation function, and a 2×2 max pooling layer. The first layer uses 32 3×3 convolutional kernels with a stride of 1; subsequent layers use 64, 128, and 256 convolutional kernels respectively. Finally, a global average pooling layer outputs a 512-dimensional feature vector, forming the network channel of the image profile features.
[0076] It should be added that the second channel (shallow seismic profile feature extraction): adopts a 2D-CNN architecture, with the input layer being a grayscale image of the shallow seismic profile (dimension 512×1024×1), which passes through two convolutional layers (3×3 kernel size, number 32 and 64) and a pooling layer (2×2 pooling kernel size) to extract a 128-dimensional sand body penetration feature vector.
[0077] Preferably, step S3 includes the following steps:
[0078] Step S31: Use the lithological labels calibrated in the fused data volume and drilling data as input samples to iteratively train the dual-channel deep learning model;
[0079] Step S32: During model training, incorporate the penetration depth threshold constraint from the shallow seismic profile data to correct the model's prediction results.
[0080] In one embodiment, single-channel seismic time series and shallow seismic profile image data contained in the fused data volume are used as input features, and combined with lithological labels (1 for sandy layers and 0 for non-sandy layers) from the drilling data to form training sample pairs. Batch gradient descent is employed, with a batch size of 32, a learning rate of 0.001, and a cross-entropy loss function. During training, a batch of samples is randomly selected from the training set in each iteration, input into the dual-channel deep learning model for forward propagation calculation, and the loss value is calculated after obtaining the predicted probability distribution. The network weight parameters are then updated using the backpropagation algorithm to complete one parameter optimization. The entire training process iterates for 100 rounds, and the model performance is evaluated on the validation set after each round.
[0081] During model training, a penetration depth threshold constraint is applied to the prediction results of each batch. Specifically, when the model predicts a location to be a sand layer, the attenuation characteristics of the reflection signal below the predicted depth are checked in the corresponding shallow seismic profile data. A penetration depth threshold of 2.5 meters is set. If a continuous and clear reflection signal still exists within 2.5 meters below the location, the prediction result is determined to be inconsistent with the physical characteristics of sand layers ("strong reflection, shallow penetration"), and the predicted label of the sample is forcibly corrected to non-sand layer. The corrected label is then used in loss calculation and gradient backpropagation. This constraint is embedded as a hard rule in the training process, ensuring that the model learning process conforms to prior geophysical knowledge.
[0082] Preferably, step S4 includes the following steps:
[0083] Step S41: Input the processed multi-source data of the target area into the trained dual-channel deep learning model to obtain preliminary sea sand identification data;
[0084] Step S42: Compare and analyze the preliminary marine sand identification data with the verification drilling data to generate marine sand distribution results.
[0085] In one embodiment, preprocessed multi-source data of the target area is input into a trained dual-channel deep learning model for batch inference. Specifically, the single-channel seismic time-series data and shallow seismic profile image data of the target area are first standardized according to the same format as the training data to generate standardized input samples. Then, the optimal model parameter file saved during training is loaded, and the model is set to inference mode. The target area data is divided into several batches and sequentially input into the model for forward propagation calculation. The model outputs the probability value of each spatial location belonging to a sand layer, with the probability value ranging from 0 to 1. A threshold of 0.5 is applied to the probability values of all locations for binarization classification, generating preliminary sea sand identification data, where 1 indicates a predicted sand layer and 0 indicates a predicted non-sand layer.
[0086] The preliminary marine sand identification data obtained in step S41 is spatially matched and quantitatively compared with the reserved verification drilling data within the target area. The specific implementation process is as follows: First, based on the geographical coordinates of the verification wells, all predicted points traversed by the drilling trajectory are extracted from the preliminary marine sand identification data. Then, the prediction results are compared with the actual lithology columnar section of the drilling data at each depth, and quantitative indicators such as the prediction error of the top and bottom depths of the sand layers, the accuracy of sand layer identification, and the mudstone exclusion rate are calculated. Based on the comparison results, the identification data is corrected and optimized, ultimately generating a marine sand distribution map containing the distribution range, sand layer thickness, and identification confidence level.
[0087] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.
Claims
1. A method for automatic identification of sea sand based on multi-source geophysical data fusion, characterized in that, Includes the following steps: Step S1: Acquire drilling data, single-channel seismic data, and shallow seismic profile data of the target area, and preprocess them to generate a fused data volume; Step S2: Based on the fused data volume, construct a dual-channel deep learning model; Step S3: Train a dual-channel deep learning model using the lithological labels identified in the fused data volume and drilling data; Step S4: Use the trained dual-channel deep learning model to identify the target area, verify the identification results, and generate the sea sand distribution results.
2. The automatic identification method for sea sand based on multi-source geophysical data fusion according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Acquire drilling data, single-channel seismic data, and shallow seismic profile data of the target area, and perform preprocessing; Step S12: Using drilling data as a benchmark, analyze single-channel seismic data and shallow seismic profile data to establish the correspondence between lithology and geophysical response; Step S13: Based on the correspondence, project the single-channel seismic data and shallow seismic profile data to the same geographic coordinate system to generate spatial registration data; Step S14: Convert the single-channel seismic data in the spatial registration data from the time domain to the depth domain to generate the fused data volume.
3. The automatic identification method for sea sand based on multi-source geophysical data fusion according to claim 2, characterized in that, Step S12 includes the following steps: Step S121: Based on drilling data, extract spatial location information of sand and mudstone layers from single-channel seismic data and shallow seismic profile data; Step S122: Based on spatial location information, statistically analyze the reflection characteristic parameters of single-channel seismic data and the penetration characteristic parameters of shallow seismic profile data to establish the correspondence between lithology and geophysical response.
4. The automatic identification method for sea sand based on multi-source geophysical data fusion according to claim 2, characterized in that, Step S13 includes the following steps: Step S131: Obtain georeferenced information for single-channel seismic data and shallow seismic profile data in the corresponding relationship; Step S132: Based on geographic reference information, perform coordinate transformation on single-channel seismic data and shallow seismic profile data respectively, and generate spatial registration data by unifying the geographic coordinate benchmark.
5. The automatic identification method for sea sand based on multi-source geophysical data fusion according to claim 2, characterized in that, Step S14 includes the following steps: Step S141: Obtain drilling logging speed information corresponding to the spatial registration data; Step S142: Based on the drilling logging velocity information, perform time-depth conversion processing on the single-channel seismic time data in the spatial registration data to generate depth domain seismic data; Step S143: Merge the depth domain seismic data with the shallow seismic profile data in the spatial registration data to generate a fused data volume.
6. The automatic identification method for sea sand based on multi-source geophysical data fusion according to claim 1, characterized in that, Step S2 includes the following steps: Step S21: Based on the characteristics of single-channel seismic data in the fused data volume, construct a network channel for seismic time series characteristics; Step S22: Based on the features of shallow seismic profile data in the fused data volume, construct a network channel for image profile features; Step S23: Construct a fusion network layer based on the network channels of image profile features and earthquake time series features to form a complete dual-channel deep learning model.
7. The automatic identification method for sea sand based on multi-source geophysical data fusion according to claim 6, characterized in that, Step S21 includes the following steps: Step S211: Perform format checks and data integrity verification on the single-channel seismic time series data in the fused data volume; Step S212: Construct a one-dimensional convolutional network structure based on the validated single-channel earthquake time series data to form a network channel for earthquake time series features.
8. The automatic identification method for sea sand based on multi-source geophysical data fusion according to claim 6, characterized in that, Step S22 includes the following steps: Step S221: Standardize the shallow seismic profile image data in the fused data volume to obtain standardized shallow seismic profile image data; Step S222: Based on the format characteristics of standardized shallow seismic profile image data, construct a two-dimensional convolutional network channel containing convolutional and pooling layers.
9. The automatic identification method for sea sand based on multi-source geophysical data fusion according to claim 1, characterized in that, Step S3 includes the following steps: Step S31: Use the lithological labels calibrated in the fused data volume and drilling data as input samples to iteratively train the dual-channel deep learning model; Step S32: During model training, incorporate the penetration depth threshold constraint from the shallow seismic profile data to correct the model's prediction results.
10. The automatic identification method for sea sand based on multi-source geophysical data fusion according to claim 1, characterized in that, Step S4 includes the following steps: Step S41: Input the processed multi-source data of the target area into the trained dual-channel deep learning model to obtain preliminary sea sand identification data; Step S42: Compare and analyze the preliminary marine sand identification data with the verification drilling data to generate marine sand distribution results.