Fault recognition model training method and system and storage medium

Through orthogonal labeling strategy and multi-model collaborative training method, the problems of low labeling efficiency and insufficient accuracy in fault identification are solved, efficient and accurate fault identification is achieved, the labeling cost is reduced and the adaptability of the model in complex geological structures is improved.

CN120669308APending Publication Date: 2025-09-19INST OF ADVANCED TECH UNIV OF SCI & TECH OF CHINA +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510689894.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies have limitations in fault identification due to data quality and subjective bias. Traditional labeling strategies are inefficient and difficult to achieve high-precision fault segmentation in complex geological structures. Synthetic seismic data differs significantly from real data, resulting in insufficient model generalization capabilities.

Method used

An orthogonal labeling strategy is adopted in combination with the mutual guidance training method of the main model and the auxiliary model. By generating auxiliary pseudo-labels and main pseudo-labels, multi-dimensional information is used for collaborative training to reduce the labeling cost and improve the fault recognition accuracy and generalization ability.

Benefits of technology

With less labeled data, more continuous and accurate fault segmentation is achieved, which reduces the cost of manual labeling and shows stronger generalization ability in real-world data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120669308A_ABST
    Figure CN120669308A_ABST
Patent Text Reader

Abstract

The invention discloses a fault recognition model training method and system and a storage medium, and the method comprises the steps: obtaining a seismic data set which comprises a three-dimensional seismic data body and a full-resolution fault mark; inputting the three-dimensional seismic data volume into the main model and inputting slice data of the three-dimensional seismic data volume into the auxiliary model to respectively obtain a main model predicted value and an auxiliary model predicted value; combining the main model predicted value with an orthogonal labeling label to generate an auxiliary pseudo label, and performing first post-processing operation on the main model predicted value to generate an auxiliary mask; combining the auxiliary model prediction value with an orthogonal labeling label to generate a main pseudo label, and performing a second post-processing operation on the auxiliary model prediction value to generate a main mask; and iteratively and circularly carrying out mutual guiding training of the auxiliary model based on auxiliary pseudo label and auxiliary mask guiding training and the main model based on main pseudo label and main mask guiding training on the seismic data set until the obtained trained main model is used for executing a fault identification task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of geophysics and artificial intelligence technology, and in particular to a fault identification model training method, system and storage medium. Background Art

[0002] Oil and natural gas are essential energy and economic pillars of modern society. The exploration and production of these resources require a comprehensive assessment, including resource evaluation, reservoir characterization, and well location optimization. These processes rely on seismic data obtained through reflection measurements to indirectly depict complex subsurface geological structures. Among these geological structures, faults are particularly important. Faults are typically formed under tectonic stress, causing rock formations to fracture and shift, creating discontinuities in the Earth's crust. Faults play a dual role in oil and gas migration—acting as both pathways that facilitate fluid flow and barriers that hinder it. Therefore, faults directly affect reservoir connectivity, entrapment mechanisms, and ultimately the success or failure of exploration and production strategies. Given this importance, accurately identifying faults from seismic data has become a critical step in subsurface geological interpretation.

[0003] To address the limitations of artificial fault interpretation caused by data quality and subjective bias, researchers have proposed various automatic fault detection methods, including phase unwrapping, Hough transform, ant colony tracking, and clustering. However, the accuracy and precision of these traditional fault interpretation algorithms in complex geological structures are still limited. The introduction of deep learning methods has improved the accuracy of fault segmentation and has shown better performance than traditional fault interpretation algorithms. However, they usually require training with voxel-level classification labels in a fully supervised learning environment. However, obtaining fine-grained voxel-level annotations of real seismic data for earthquake fault interpretation is extremely expensive and prone to subjective bias. Therefore, most models are trained on synthetic seismic datasets. However, due to the significant structural differences between synthetic seismic data and real seismic data, synthetic seismic data cannot fully reflect the complexity of real geological structures. Models trained only on synthetic seismic data are often difficult to generalize to complex real-world scenarios.

[0004] Therefore, training effective fault interpretation models with limited annotations in a semi-supervised or weakly supervised environment has become a key research direction. Semi-supervised and weakly supervised learning methods use a small amount of labeled data and a large amount of unlabeled data to train earthquake fault interpretation models, which reduces the dependence on voxel-level labels. This makes it possible to train fault identification models using sparsely labeled real earthquake datasets. However, these semi-supervised and weakly supervised learning methods usually rely on real earthquake datasets labeled using traditional annotation strategies, such as Figure 1(a), (b), and (c) show the subvoxel annotation, slice-by-slice annotation, and sparse point annotation, respectively. The gray cube represents the 3D seismic volume, and the red area represents the annotation area. Figure 1 The subvoxel annotations shown in (a) in Figure 1 lead to significant label redundancy and often overfitting of the labeled regions due to the high similarity between adjacent slices. Figure 1 The slice-by-slice annotation along a single direction shown in (b) improves the segmentation effect on this axis, but ignores the three-dimensional continuity and consistency of the fault structure; Figure 1 The sparse point annotation shown in (c) in the figure may help the model focus on global features, but often ignores the fine fault details in a single slice. Therefore, these traditional annotation strategies are inefficient and have a lot of annotation redundancy. It is difficult to achieve high segmentation performance without a large number of labeled samples.

[0005] And as Figure 1 The new orthogonal labeling strategy shown in (d) only selects a portion of the slices in two orthogonal directions, which is different from Figure 1 Compared with (a) in , it avoids redundant annotations on adjacent slices; Figure 1 Compared with (b), slices in different orthogonal directions provide complementary information and enhance the identification of fault continuity; Figure 1Compared to (c) in [1], the orthogonal strategy provides complete two-dimensional fault slice labels, improving local accuracy. Therefore, exploring fault recognition training based on orthogonal labeling has great potential in reducing labeling workload while improving model performance for complex real-world geological structures. For example, in the related art, the paper "Intelligent Fault Recognition Based on Orthogonal Labeling and Weakly Supervised Learning, Proceedings of the Second Annual Conference on Petroleum Geophysical Exploration of China (Volume 2), Zhang Zheng et al." proposed a fault recognition framework based on orthogonal labeling, seismic image registration, and weakly supervised learning. This approach combines orthogonal labeling and registration, first orthogonally labeling a portion of the data volume, then generating pseudo-labels through registration techniques. Finally, weakly supervised training is performed using both pseudo-labeled and unlabeled data. By combining orthogonal labeling, seismic image registration, and weakly supervised learning, this approach reduces reliance on large amounts of labeled data and improves the network's fault recognition performance for real seismic data. However, this approach focuses on using seismic image registration techniques to generate pseudo-labels to enhance the network's ability to recognize complex faults. However, the seismic image registration process requires pre-training the network, and the pre-training process cannot completely break away from voxel-level labels. The patent application document with publication number CN119251613A proposes a semi-supervised 3D fault recognition method based on sparse annotation. This scheme uses a small amount of two-dimensional slice annotation of some field data, crops and enhances the sparsely labeled data, generates pseudo labels through a teacher-student model structure, and combines unlabeled data for semi-supervised learning to improve the model's adaptability to real data. However, it does not use an orthogonal annotation method and is not an orthogonal annotation strategy. Summary of the Invention

[0006] The technical problem to be solved by the present invention is how to achieve more continuous and accurate fault segmentation while using less annotated data.

[0007] The present invention solves the above technical problems through the following technical means:

[0008] In a first aspect, a fault identification model training method is proposed, the method comprising:

[0009] Acquire a seismic dataset, which includes a 3D seismic data volume and full-resolution fault annotations;

[0010] Inputting the 3D seismic data volume into the main model and inputting the slice data of the 3D seismic data volume into the auxiliary model, respectively obtaining the main model prediction value and the auxiliary model prediction value;

[0011] Combining the primary model predictions with the orthogonal annotation labels to generate auxiliary pseudo-labels, and performing a first post-processing operation on the primary model predictions to generate auxiliary masks;

[0012] Combining the auxiliary model predictions with the orthogonal annotation labels to generate primary pseudo labels, and performing a second post-processing operation on the auxiliary model predictions to generate a primary mask;

[0013] An auxiliary model guided by auxiliary pseudo labels and auxiliary masks and a main model guided by main pseudo labels and main masks are trained iteratively on the seismic dataset until a trained main model is obtained for performing the fault identification task.

[0014] Furthermore, the three-dimensional seismic data volume x i and the full-resolution tomographic annotation y i They are all defined in three orthogonal dimensions, namely time V, survey line I and contact survey line ∈;

[0015] The slice data is the two-dimensional slice data obtained by slicing the three-dimensional seismic data along the direction of the survey line. Or two-dimensional slice data cut in the direction of the contact line

[0016] Furthermore, the main model is a three-dimensional network model based on the V-Net architecture.

[0017] Furthermore, the auxiliary model includes a first sub-model and a second sub-model, and both the first sub-model and the second sub-model are two-dimensional network models based on the U-Net architecture.

[0018] Furthermore, the orthogonal label of the i-th earthquake volume is:

[0019]

[0020] Where: represents the i-th 3D seismic volume x i Orthogonal label labels, and y i [h,l,w] represent and y i Any voxel in, s is the predefined annotation step, y i represents full-resolution fault annotation, and [h, l, w] represents the coordinate points of the 3D seismic data volume.

[0021] Furthermore, the combining of the main model prediction value and the orthogonal annotation label to generate the auxiliary pseudo label, and performing a first post-processing operation on the main model prediction value to generate the auxiliary mask, includes:

[0022] The binarized result of the main model prediction value is combined with the orthogonal annotation label to generate the auxiliary pseudo label, which is expressed as follows:

[0023]

[0024] Where: represents auxiliary pseudo labels, Represents auxiliary pseudo labels For any voxel in represents the i-th earthquake volume x i Orthogonal labeling labels For any voxel in Represents the main model prediction value M1(x i ,θ M1 ) is the binarization result, s is the predefined annotation step size;

[0025] A main confidence cube is generated according to the prediction value of the main model, and the main confidence cube is used to generate a main mask factor according to the hard-soft threshold, which is expressed as:

[0026]

[0027] Where: F1(M1(x i ,θ M1 )) represents the main mask factor, F1(M1(x i ,θ M1 ))[h,l,w] represents the main mask factor F1(M1(x i ,θ M1 )) in any voxel, represents the confidence cube, β is a predefined parameter, Th is a soft threshold or a hard threshold, and [c, h, l, w] represents the probability that the predicted category of the coordinate point [h, l, w] on the 3D seismic data volume is c.

[0028] The auxiliary mask is generated based on the main mask factor, and the formula is expressed as:

[0029]

[0030] Where: is the auxiliary mask, and s is the predefined annotation step size.

[0031] Furthermore, the method also includes: when ACC is greater than the accuracy critical value, Th is a soft threshold; when ACC is less than or equal to the accuracy critical value, Th is a hard threshold, wherein ACC is the accuracy calculated based on the main model prediction value and the orthogonal annotation label.

[0032] Furthermore, the step of inputting the slice data of the three-dimensional seismic data volume into the auxiliary model includes:

[0033] Inputting two-dimensional slice data obtained by slicing the three-dimensional seismic data volume along the direction of the survey line into the first sub-model, and inputting two-dimensional slice data obtained by slicing the three-dimensional seismic data volume along the direction of the connecting survey line into the second sub-model;

[0034] Accordingly, the method of combining the auxiliary model prediction value with the orthogonal annotation label to generate the main pseudo label specifically includes:

[0035] The main pseudo label is generated by combining the binarization result of the prediction value of the first sub-model with the orthogonal annotation label, or the main pseudo label is generated by combining the binarization result of the prediction value of the second sub-model with the orthogonal annotation label.

[0036] Furthermore, performing a second post-processing operation on the auxiliary model prediction value to generate a primary mask includes:

[0037] The prediction value of the first sub-model is fused with the prediction value of the second sub-model to determine an auxiliary mask factor. The auxiliary mask factor represents the high-confidence prediction area where the first sub-model and the second sub-model are consistent. The public expression is:

[0038]

[0039] Where: F2(M2(x i ′ ,θ M2 )) represents the auxiliary mask factor, β is a predefined parameter, and They represent the binarization results of the main model prediction value, the binarization results of the first sub-model prediction value, and the binarization results of the second sub-model prediction value, respectively. and denote the confidence cube of the main model, the confidence cube of the first submodel, and the confidence cube of the second submodel, respectively; [h, l, w] denotes the coordinate points of the 3D seismic data volume; the set Ξ denotes the region where the first submodel and the second submodel are consistent; the set Ψ denotes the position where the second submodel and the main model are consistent; the set Υ and the set Λ denote the region where the second submodel has a higher confidence than the main model, respectively; ∩ denotes the intersection; and ∪ denotes the union;

[0040] The main mask is generated based on the auxiliary mask factor, and the formula is expressed as:

[0041]

[0042] Where: Main mask, F2(M2(x i ′ ,θ M2)) represents the auxiliary mask factor, and s is the predefined annotation step size.

[0043] Furthermore, the loss function used in mutual guidance training is:

[0044]

[0045] Where: and represents the loss function, and Represent the optimization parameters of the main model and the auxiliary model respectively, M1(x i ,θ M1 ) and M2(x i ′ ,θ M2 ) represent the prediction values ​​of the main model and the auxiliary model respectively, represents the primary pseudo label and primary mask, represents auxiliary pseudo labels and auxiliary masks, and Represent the total loss function of the main model and the total loss function of the auxiliary model, express The model parameter value that takes the minimum value.

[0046] Furthermore, the total loss function of the main model is expressed as follows:

[0047]

[0048] Where: is the cross entropy loss, Represents the main model prediction value and the main pseudo label The probability of matching, and Each represents a voxel; is the Dice loss, represents the confidence cube output by the main model for category c, represents the position of category c in the main pseudo label, Δ, Δ o , Δ p and Δ op It is an intermediate variable in the calculation process, V, I, C represent time, survey line and contact survey line respectively, and ∈ is a smoothing term.

[0049] Furthermore, the total loss function of the auxiliary model is expressed as follows:

[0050]

[0051] Where: is the loss of the first sub-model, is the loss of the second sub-model, is the predicted value of the first sub-model, is the predicted value of the second sub-model.

[0052] In a second aspect, the present invention further proposes a fault identification model training system, the system comprising:

[0053] A data acquisition module is used to acquire a seismic data set, which includes a three-dimensional seismic data volume and full-resolution fault annotations;

[0054] A prediction module is used to input the 3D seismic data volume into the main model and the slice data of the 3D seismic data volume into the auxiliary model to obtain the main model prediction value and the auxiliary model prediction value respectively;

[0055] a first processing module, configured to combine the primary model prediction value with the orthogonal annotation label to generate an auxiliary pseudo label, and perform a first post-processing operation on the primary model prediction value to generate an auxiliary mask;

[0056] a second processing module, configured to combine the auxiliary model prediction value with the orthogonal annotation label to generate a primary pseudo label, and perform a second post-processing operation on the auxiliary model prediction value to generate a primary mask;

[0057] The guided training module is used to iteratively perform mutual guided training on the seismic dataset, which includes training the auxiliary model based on the auxiliary pseudo-labels and auxiliary masks and training the main model based on the main pseudo-labels and main masks, until a trained main model is obtained to perform the fault identification task.

[0058] In a third aspect, the present invention further proposes a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the fault identification model training method described above is implemented.

[0059] The advantages of the present invention are:

[0060] (1) The present invention directly inputs the three-dimensional seismic data volume into the main model, and simultaneously inputs the two-dimensional slice data of the three-dimensional seismic data volume into the auxiliary model, and combines the predicted values ​​of the main model with the orthogonal annotation labels to generate auxiliary pseudo-labels for training the auxiliary model, and generates auxiliary masks for training the auxiliary model according to the predicted values ​​of the main model; at the same time, combines the predicted values ​​of the auxiliary model with the orthogonal annotation labels to generate main pseudo-labels for training the main model, and generates main masks for training the main model according to the predicted values ​​of the auxiliary model; then utilizes these main pseudo-labels, main masks, auxiliary pseudo-labels and auxiliary masks to perform relative training on the main model and the auxiliary model. Mutually guided training combines the directional fault features proposed by the auxiliary model with the orthogonal annotation labels to form main pseudo-labels for training the main model, which introduces more reliable, continuous and multi-dimensionally aligned fault information into the main model. At the same time, the predicted values ​​of the main model are combined with the orthogonal annotation labels to guide the training of the auxiliary model, which introduces additional multi-dimensional fault information into the auxiliary model, realizing collaborative training and mutual enhancement between the main model and the auxiliary model. The main model finally trained is used as the fault recognition model. This collaborative training mechanism significantly reduces the cost of manual annotation while maintaining high fault interpretation accuracy, and exhibits stronger generalization ability in fault interpretation of real-world data.

[0061] (2) The present invention combines orthogonal annotation labels with the predicted values ​​of the main model and the auxiliary model to generate pseudo labels for model training. It explores the application of orthogonal annotation labels as a novel sparse annotation method in seismic fault interpretation. Compared with traditional annotation strategies, it reduces redundant annotations, provides some complete two-dimensional slice labels, and utilizes the complementary information of slices in different directions.

[0062] (3) The present invention uses a two-dimensional auxiliary model to assist in training the three-dimensional main model, effectively utilizing the real annotated slices in the longitudinal and transverse survey lines, and enhancing multi-dimensional alignment, three-dimensional continuity and accuracy.

[0063] (4) The present invention introduces soft and hard threshold methods to generate auxiliary masks during the multi-network collaborative training process. The soft and hard thresholds are dynamically adjusted according to the accuracy of the model to generate reliable auxiliary masks, select high-confidence areas to the greatest extent, improve the quality of pseudo labels, and guide the training of auxiliary models.

[0064] (5) The present invention introduces a confidence-based primary mask selection strategy in the multi-network collaborative training process to ensure that more reliable voxels are always used to guide the training of the fault interpretation model.

[0065] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0066] Figure 1 Schematic diagrams of annotation labels obtained using different annotation strategies in one embodiment of the present invention, where (a) is subvoxel annotation, (b) is slice-by-slice annotation, (c) is sparse point annotation, and (d) is orthogonal annotation.

[0067] Figure 2 1 is a flow chart of a fault identification model training method proposed in one embodiment of the present invention;

[0068] Figure 3 1 is an example diagram of a 3D seismic data volume, a label, and an orthogonal annotation in one embodiment of the present invention, wherein (a) is a 3D seismic data volume, (b) is a label, and (c) is an orthogonal annotation;

[0069] Figure 4 This is a schematic diagram of the main model training principle in one embodiment of the present invention;

[0070] Figure 5 Schematic diagram of auxiliary model training principle in one embodiment of the present invention;

[0071] Figure 6 1 is a structural diagram of a fault identification model training system proposed in one embodiment of the present invention;

[0072] Figure 7 is a schematic diagram of fault identification results for a selected slice of synthetic seismic data in one embodiment of the present invention;

[0073] Figure 8 is a schematic diagram of ablation experiment results of a selected slice of synthetic seismic data in one embodiment of the present invention;

[0074] Figure 9 1 is a schematic diagram of the prediction results of the V-Net model trained using different methods on the F3 dataset in one embodiment of the present invention, wherein (a) is a time slice of the F3 seismic volume, (b) is the fault prediction result of the V-Net model trained using only 1 / 4 of the labeled synthetic data in the Fault-Seg-OA framework, (c) is the fault prediction result of the V-Net model trained using only 1 / 8 of the labeled synthetic data in the Fault-Seg-OA framework, (d) is the prediction result of the V-Net trained using all labeled synthetic data for full supervision, (e) shows the result of the V-Net model trained using the DeSCO method, and (f) shows the result of the V-Net model trained using the OABS method;

[0075] Figure 10: This is a schematic diagram of the prediction results of different methods on the Thebes dataset in one embodiment of the present invention, wherein (a) is a survey line slice from the Thebes seismic data, (b) is the true label, (c) is the prediction result of Fault-Seg-OA (8%) on the Thebes dataset, (d) is the prediction result of Fault-Seg-OA (4%) on the Thebes dataset, (e) is the prediction result of Fault-Seg-OA (2%) on the Thebes dataset, (f) is the prediction result of Fault-Seg-OA (1%) on the Thebes dataset, (g) is the prediction result of V-Net trained only using full supervision on the Thebes dataset, and (h) is the prediction result of FaultSeg3D on the Thebes dataset. DETAILED DESCRIPTION

[0076] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0077] like Figure 2 As shown, the first embodiment of the present invention proposes a fault identification model training method, which includes the following steps:

[0078] S10, obtaining a seismic data set, where the seismic data set includes a three-dimensional seismic data volume and full-resolution fault annotations;

[0079] S20, inputting the 3D seismic data volume into the main model and inputting the slice data of the 3D seismic data volume into the auxiliary model, to obtain the main model prediction value and the auxiliary model prediction value respectively;

[0080] It should be noted that the slice data is two-dimensional slice data of a three-dimensional seismic data volume.

[0081] S30, combining the main model prediction value with the orthogonal annotation label to generate an auxiliary pseudo label, and performing a first post-processing operation on the main model prediction value to generate an auxiliary mask;

[0082] S40, combining the auxiliary model prediction value with the orthogonal annotation label to generate a primary pseudo label, and performing a second post-processing operation on the auxiliary model prediction value to generate a primary mask;

[0083] S50, iteratively looping the mutual guided training of the auxiliary model guided by the auxiliary pseudo-labels and the auxiliary mask and the main model guided by the main pseudo-labels and the main mask on the seismic dataset until a trained main model is obtained for performing the fault identification task.

[0084] It should be noted that this embodiment repeatedly cycles through the mutual guidance training of the main and auxiliary models on the entire seismic data set until the maximum number of rounds is reached, and then saves the trained model parameters to obtain the trained main model as a fault identification model for performing downstream fault identification tasks. Since the directional fault characteristics proposed by the auxiliary model are combined with the orthogonal annotation labels to form the main pseudo-labels for training the main model during the training process, more reliable, continuous and multi-dimensionally aligned fault information is introduced into the main model. At the same time, the predicted values ​​of the main model are combined with the orthogonal annotation labels to guide the training of the auxiliary model, which introduces additional multi-dimensional fault information into the auxiliary model, realizing collaborative training and mutual enhancement between the main model and the auxiliary model. The main model finally trained is used as the fault identification model. This collaborative training mechanism can achieve more continuous and accurate fault segmentation while using significantly less labeled data, and exhibits stronger generalization ability in the interpretation of faults in real-world data.

[0085] It should be noted that orthogonal labeling is a novel sparse labeling method. The mainstream labeling strategy will vary according to the training paradigm of semi-supervised learning and weakly supervised learning. Semi-supervised learning aims to reduce the number of labeled samples. Its training set usually consists of a small amount of voxel-level labeled data and a large amount of unlabeled data. However, this labeling strategy still requires some sub-volumes with complete voxel-level 3D labels. Due to the extremely high similarity between adjacent slices in seismic data, this labeling method will lead to redundant labeling and overfitting problems in the labeled area. The weakly supervised learning used in this embodiment focuses on reducing the labeling quality by utilizing weak labels such as image-level labels, graffiti, points or partially labeled slices. However, if the traditional sparse point labeling method is used, it may cause weak labels to ignore the fine fault details in a single slice. Orthogonal labeling selects partial slices along two orthogonal directions for labeling, which can effectively reduce redundant labeling work while retaining some accurate 2D fault slice labels. Moreover, labels from different orthogonal directions provide complementary information, enhancing the 3D integrity of the fault label. Therefore, this embodiment combines the processing output of one model with orthogonal annotation as pseudo-labels for another model. By applying orthogonal annotation as a novel sparse annotation method in seismic fault interpretation, compared with traditional annotation strategies, it reduces redundant annotations, provides some complete two-dimensional slice labels, and utilizes the complementary information of slices in different directions.

[0086] As a further preferred technical solution, the three-dimensional seismic data volume xi and the full-resolution tomographic annotation y i They are all defined in three orthogonal dimensions, namely time V, survey line I and contact survey line C;

[0087] The slice data is the three-dimensional seismic data volume x i Two-dimensional slice data x cut along the direction of the survey line or the direction of the connecting survey line i ′ .

[0088] Specifically, given an earthquake dataset in, represents the i-th earthquake body, y i Indicates the corresponding full-resolution fault annotation, Indicates the dimension , both are defined in three orthogonal dimensions: V (time), I (line), and C (connection line). Figure 3 As shown, the earthquake dataset consists of N earthquake samples x i Orthogonal marking only selectively marks a portion of the slices in the direction of the survey line and the connecting survey line to form a two-dimensional slice data x i ′ , significantly reducing annotation redundancy while retaining key fault information.

[0089] Furthermore, the orthogonal label of the i-th earthquake volume is:

[0090]

[0091] Where: represents the i-th 3D seismic volume x i Orthogonal label labels, and y i [h,l,w] represent and y i Any voxel in, s is the predefined annotation step, y i It represents full-resolution fault annotation, which can be understood as [h, l, w] being the coordinate points of the 3D seismic data volume in [V, I, C].

[0092] As a further preferred technical solution, the main model M1 is a three-dimensional network model based on the V-Net architecture; the auxiliary model M2 includes a first sub-model and the second sub-model The first sub-model And the second sub-model Both are two-dimensional network models based on U-Net architecture; θ M1 ,θ M2are the trainable parameters of the main model M1 and the auxiliary model M2 respectively.

[0093] It should be noted that in this embodiment, the main model uses the V-Net architecture and the auxiliary model uses the U-Net architecture for illustration only. The main model and the auxiliary model may also use network models of other architectures, and this embodiment does not make any specific restrictions on this.

[0094] This embodiment uses a two-dimensional auxiliary model to assist in the training of a three-dimensional main model. The first and second sub-models in the auxiliary model effectively utilize real annotated slices in the longitudinal and transverse directions, respectively, enhancing multi-dimensional alignment, three-dimensional continuity, and accuracy. Furthermore, a collaborative training mechanism is employed, where the processed output of one model is combined with orthogonal annotations to serve as pseudo-labels for another model. The output of a 2D convolutional neural network (CNN) is used to assist in the training of the 3D main model, and the output of the 3D main model is used to train the 2D auxiliary model. This enables the final fault interpretation model to capture detailed features in seismic slices like a 2D CNN, while maintaining the 3D model's advantage in preserving the spatial continuity of the seismic volume.

[0095] As a further preferred technical solution, step S30: combining the main model prediction value with the orthogonal annotation label to generate an auxiliary pseudo label, and performing a first post-processing operation on the main model prediction value to generate an auxiliary mask, specifically includes the following steps:

[0096] S31, combining the binarization result of the main model prediction value with the orthogonal annotation label to generate the auxiliary pseudo label, which is expressed as:

[0097]

[0098] Where: represents auxiliary pseudo labels, yes For any voxel in represents the i-th earthquake volume x i Orthogonal labeling labels For any voxel in Represents the main model prediction value M1(x i ,θ M1 ) is the binarization result, s is the predefined annotation step size;

[0099] S32. Generate a main confidence cube according to the prediction value of the main model, and generate a main mask factor based on the main confidence cube according to the soft and hard thresholds, which is expressed as:

[0100]

[0101] Where: F1(M1(x i ,θ M1)) represents the main mask factor, represents the main confidence cube, β is a predefined parameter, Th is a soft threshold or a hard threshold, and [c, h, l, w] represents the probability that the predicted category of the coordinate point [h, l, w] on the 3D seismic data volume is c;

[0102] It should be noted that the input shape of the main model is h×l×w, and the output shape is c×h×l×w, where the dimension c is understood as the segmentation category.

[0103] S33, generating the auxiliary mask based on the main mask factor, specifically setting the orthogonally marked positions in the 3D seismic data volume to 1, and setting the other positions to be consistent with the main mask factor, to obtain the auxiliary mask, which is expressed as:

[0104]

[0105] Where: is the auxiliary mask, and s is the predefined annotation step size.

[0106] As a further preferred technical solution, when ACC is greater than the accuracy critical value ι, Th is a soft threshold; when ACC is less than or equal to the accuracy critical value ι, Th is a hard threshold, where ACC is the accuracy calculated based on the main model prediction value and the orthogonal annotation label.

[0107] It should be noted that in this embodiment, Th is calculated based on the predicted value of the main model and the orthogonal annotation. The soft threshold or hard threshold is determined by the accuracy ACC calculated between . When ACC exceeds the accuracy threshold, Th uses a soft threshold to retain more information; when ACC is less than or below the accuracy threshold, Th uses a hard threshold to effectively filter out erroneous information.

[0108] It should be noted that β is a predefined parameter that increases with the number of training rounds and is used to adjust the weight of pseudo labels in the loss calculation.

[0109] Furthermore, in this embodiment, the auxiliary mask is derived Finally, set the mask to 1 at the correct labeled position and set the mask to 0 at the wrong labeled position, and then combine the first sub-model and the second sub-model The output of Calculate the loss.

[0110] This embodiment introduces soft and hard threshold methods to generate auxiliary masks during the multi-network collaborative training process. The soft and hard thresholds are dynamically adjusted according to the accuracy of the model to generate reliable auxiliary masks. More reliable pseudo labels can be selected to guide the training of the auxiliary model.

[0111] As a further preferred technical solution, in step S20, inputting the slice data of the 3D seismic data volume into the auxiliary model includes:

[0112] Inputting slice data obtained by slicing the three-dimensional seismic data volume along the direction of the survey line into the first sub-model, and inputting slice data obtained by slicing the three-dimensional seismic data volume along the direction of the connecting survey line into the second sub-model;

[0113] Specifically, this embodiment converts the three-dimensional seismic data x i It is cut into I slices along the survey line direction and C slices along the connecting survey line direction. In the training data set used in this embodiment, I and C are both 128. The auxiliary model M2 consists of two 2D network models based on the U-Net architecture, namely the first sub-model and the second sub-model These two sub-models assist in the training of the main fault identification model M1. Specifically, the slice data along the survey line direction is input into the first sub-model Input the slice data along the connecting survey line into the second sub-model

[0114] Accordingly, in step S40, the auxiliary model prediction value is combined with the orthogonal annotation label to generate the main pseudo label, which specifically includes:

[0115] The main pseudo label is generated by combining the binarization result of the prediction value of the first sub-model with the orthogonal annotation label, or the main pseudo label is generated by combining the binarization result of the prediction value of the second sub-model with the orthogonal annotation label.

[0116] Specifically, the formula for the main pseudo-label used for main model training is as follows:

[0117]

[0118] Where: represents the primary pseudo label, for For any voxel in represents the i-th earthquake volume x i Orthogonal labeling labels Any voxel in ; Specifically or Represents the predicted value of the first sub-model The binarized result or the second sub-model prediction value The binarization result of .

[0119] It should be noted that the first sub-model (or the second sub-model ) is then binarized and labeled with the orthogonal Combined to generate the main pseudo label for training the main model M1 First sub-model or the second submodel Both can be used to generate pseudo labels because only the areas where both sub-models are consistent will guide the training of the main model M1 through subsequent mask operations.

[0120] As a further preferred technical solution, in step S40, the auxiliary model prediction value is subjected to a second post-processing operation to generate a main mask, which specifically includes the following steps:

[0121] S41: Fusing the prediction value of the first sub-model with the prediction value of the second sub-model to determine an auxiliary mask factor, where the auxiliary mask factor represents a high-confidence prediction region where the first sub-model and the second sub-model are consistent, and is expressed as:

[0122]

[0123] Where: F2(M2(x′ i ,θ M2 )) represents the auxiliary mask factor, β is a predefined parameter, and the parameter β is used to control the weight of the pseudo label in the loss calculation. and They represent the binarization results of the main model prediction value, the binarization results of the first sub-model prediction value, and the binarization results of the second sub-model prediction value, respectively. and denote the confidence cube of the main model, the confidence cube of the first submodel, and the confidence cube of the second submodel, respectively; [h, l, w] denotes the coordinate points of the 3D seismic data volume; the set Ξ denotes the region where the first submodel and the second submodel are consistent; the set Ψ denotes the position where the second submodel and the main model are consistent; the set Υ and the set Λ denote the region where the second submodel has a higher confidence than the main model, respectively; ∩ denotes the intersection; and ∪ denotes the union;

[0124] It should be noted that, because only the areas where both sub-models are consistent will guide the training of the main model M1 through subsequent mask operations, the set Ψ can also represent the locations where the first sub-model and the main model are consistent.

[0125] It should be noted that in order to train a multi-dimensional aligned and high-precision master model for fault identification, this embodiment designs a fusion operation F2(·) to generate a master mask by selecting the areas where the two sub-models produce consistent and high-confidence predictions. The mask is set to 1 at the correct labeled position and to 0 at the wrong labeled position, and then the loss of the main model is calculated based on the output of the main model.

[0126] It should be understood that and Represents the confidence cube, calculated in the same way as the main confidence cube above The calculation method is the same and will not be repeated here.

[0127] S42, generating the main mask based on the auxiliary mask factor, specifically setting the orthogonally marked positions in the 3D seismic data volume to 1, and setting the other positions to be consistent with the auxiliary mask factor, to obtain the main mask, which is expressed as:

[0128]

[0129] Where: Main mask, F2(M2(x i ′ ,θ M2 )) represents the auxiliary mask factor, and s is the predefined annotation step size.

[0130] It should be noted that this embodiment introduces a confidence-based main mask selection strategy in the multi-network collaborative training process to ensure that more reliable voxels are always used to guide the training of the fault interpretation model.

[0131] As a further preferred technical solution, Figures 4 and 5 As shown, in step S50, the main and auxiliary models are trained to guide each other. The training goal is to minimize the difference between the model prediction and the pseudo-label, thereby improving the fault interpretation performance with the least amount of annotation work. Therefore, the optimization process of learning the fault interpretation model using orthogonal annotated data can be expressed as:

[0132]

[0133] Where: and represents the loss function, and Represent the optimization parameters of the main model and the auxiliary model respectively, M1(x i ,θ M1 ) and M2(x i ′ ,θ M2 ) represent the prediction values ​​of the main model and the auxiliary model respectively, represents the primary pseudo label and primary mask, represents auxiliary pseudo labels and auxiliary masks, and Represent the total loss function of the main model and the total loss function of the auxiliary model, express The model parameter θ corresponding to the minimum value M1 ,θ M2 .

[0134] As a further preferred technical solution, the total loss function of the main model is expressed as follows:

[0135]

[0136] Where: is the cross entropy loss, Represents the main model prediction value and the main pseudo label The probability of matching, and Each represents a voxel, and The shape is V×I×C; Dice loss

[0137] lose, represents the confidence cube output by the main model for category c, Indicates the position of category c in the main pseudo label. For a single data sample, the output dimension of the main model M1 is 2×V×I×C, where the first channel indicates whether the voxel belongs to a fault; Δ, Δ o , Δ p and Δ op It is an intermediate variable in the calculation process, V, I, and C represent time, survey line, and contact survey line respectively; ∈ is a smoothing term, which is a small constant to prevent it from being divided by 0 and improve numerical stability.

[0138] It should be noted that, considering that the fault area usually occupies only a small part of the seismic image, this embodiment introduces the Dice loss To alleviate the problem of class imbalance.

[0139] As a further preferred technical solution, the total loss function of the auxiliary model is expressed as follows:

[0140]

[0141] Where: is the loss of the first sub-model, is the loss of the second sub-model, and represents the loss function, is the predicted value of the first sub-model, is the predicted value of the second sub-model.

[0142] It should be noted that the loss function The specific formula is similar to the loss function representation of the main model mentioned above:

[0143]

[0144] Where: Represents the predicted value of the first sub-model and the auxiliary pseudo-label The probability of matching, represents the confidence cube output by the first sub-model for category c, Represents the position of category c in the auxiliary pseudo-label.

[0145] Similarly, The calculation process is as follows:

[0146]

[0147] Where: Represents the second sub-model prediction value and auxiliary pseudo label The probability of matching, represents the confidence cube output by the second sub-model for category c, Represents the position of category c in the auxiliary pseudo-label.

[0148] Specifically, the input for training in this embodiment is a 3D seismic dataset with orthogonal fault annotations, and the goal is to obtain spatially consistent 3D fault prediction results. The entire process integrates a 3D volume model and a 2D model in two directions to improve the accuracy and robustness of the prediction under weak supervision. The training process begins with the initialization of key hyperparameters. Model parameters include the maximum number of training rounds, learning rate, batch size, etc. These parameters define the basic settings for model optimization. Next, a 3D seismic dataset with orthogonal fault annotations is loaded. Subsequently, a 3D model for volume data and a 2D model in two directions for slice analysis are constructed. The network parameters and optimizer are then initialized to prepare for the optimization process. In each training round, the model performs forward propagation to generate prediction results from the 3D and 2D branches. These prediction results are then binarized, and corresponding confidence cubes are obtained based on the model prediction values. After obtaining the consistency weights, the pseudo labels required for training and their corresponding masks are calculated according to the defined formula. The overall loss function used in the training process is calculated as follows:

[0149]

[0150] Once the loss value is obtained, backpropagation is performed to update the model parameters. This cycle is repeated on the entire dataset until the maximum number of epochs is reached. Finally, the trained model parameters are saved for use in downstream fault prediction tasks.

[0151] The overall training process is as follows:

[0152] Input: 3D orthogonal annotation dataset

[0153] Output: Fault prediction

[0154] Initialization parameters

[0155] Loading the dataset

[0156] Create 3D model M1

[0157] Creating a 2D Model and

[0158] Creating an optimizer

[0159] for inD a :

[0160] Output:

[0161] Binarization:

[0162] calculate:

[0163] Get the consistency weight β

[0164] calculate and

[0165] calculate

[0166] Calculate ACC for M1

[0167] If ACC>ι:

[0168] Set Th as the soft threshold

[0169] otherwise:

[0170] Set Th as hard threshold

[0171] calculate

[0172] calculate

[0173] calculate

[0174] calculate

[0175] Backpropagation

[0176] End this loop

[0177] End the outer loop

[0178] Save the model parameters.

[0179] In addition, if Figure 6 As shown, another embodiment of the present invention further provides a fault identification model training system, the system comprising:

[0180] A data acquisition module 10 is used to acquire a seismic data set, which includes a three-dimensional seismic data volume and full-resolution fault annotations;

[0181] Prediction module 20, for inputting the 3D seismic data volume into the main model and the slice data of the 3D seismic data volume into the auxiliary model, and obtaining the main model prediction value and the auxiliary model prediction value respectively;

[0182] a first processing module 30 for combining the primary model prediction value with the orthogonal annotation label to generate an auxiliary pseudo label, and performing a first post-processing operation on the primary model prediction value to generate an auxiliary mask;

[0183] A second processing module 40 is configured to combine the auxiliary model prediction value with the orthogonal annotation label to generate a primary pseudo label, and perform a second post-processing operation on the auxiliary model prediction value to generate a primary mask;

[0184] The guided training module 50 is used to iteratively perform mutual guided training on the seismic data set, which includes training the auxiliary model based on the auxiliary pseudo-labels and auxiliary masks and training the main model based on the main pseudo-labels and main masks, until a trained main model is obtained for performing the fault identification task.

[0185] As a further preferred technical solution, the first processing module 30 specifically includes:

[0186] an auxiliary pseudo-label generating unit, configured to combine the binarized result of the main model prediction value with the orthogonal annotation label to generate the auxiliary pseudo-label;

[0187] A main mask factor generating unit, configured to generate a main confidence cube according to the main model prediction value, and to generate a main mask factor from the main confidence cube according to soft and hard thresholds;

[0188] An auxiliary mask generating unit is configured to generate the auxiliary mask based on the main mask factor.

[0189] As a further preferred technical solution, when ACC is greater than the accuracy critical value, Th is a soft threshold; when ACC is less than or equal to the accuracy critical value, Th is a hard threshold, where ACC is the accuracy calculated based on the main model prediction value and the orthogonal annotation label.

[0190] As a further preferred technical solution, the auxiliary model includes a first sub-model and a second sub-model, and the prediction module 20 is specifically used to input the slice data cut from the three-dimensional seismic data body along the direction of the survey line into the first sub-model, and to input the slice data cut from the three-dimensional seismic data body along the direction of the connecting survey line into the second sub-model.

[0191] As a further preferred technical solution, the second processing module 40 specifically includes:

[0192] a primary pseudo-label generating unit, configured to generate a primary pseudo-label by combining the binarized result of the prediction value of the first sub-model with the orthogonal annotation label, or to generate a primary pseudo-label by combining the binarized result of the prediction value of the second sub-model with the orthogonal annotation label;

[0193] A main mask generating unit is used to fuse the prediction value of the first sub-model with the prediction value of the second sub-model to determine a target area, where the target area is a consistent high-confidence prediction area, and generate the main mask based on the target area.

[0194] As a further preferred technical solution, the loss function used by the guided training module 50 during mutual guided training is:

[0195]

[0196] Where: and represents the loss function, and Represent the optimization parameters of the main model and the auxiliary model respectively, M1(x i ,θ M1 ) and M2(x i ′ ,θ M2 ) represent the prediction values ​​of the main model and the auxiliary model respectively, represents the primary pseudo label and primary mask, represents auxiliary pseudo labels and auxiliary masks, and Represent the total loss function of the main model and the total loss function of the auxiliary model, express The model parameter θ corresponding to the minimum value M1 ,θ N2 .

[0197] In addition, another embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the fault identification model training method described in the first embodiment is implemented.

[0198] It should be noted that other embodiments or specific implementation methods of the fault identification model training system and computer-readable storage medium of the present invention can refer to the above-mentioned fault identification model training method embodiments, which will not be repeated here.

[0199] It should be noted that this embodiment proposes a fault identification model training framework based on orthogonal annotation. In order to make full use of accurate 2D slice labels in different directions, this method uses slices in the longitudinal and transverse directions to train two two-dimensional segmentation networks to extract directional fault features. Then, the areas where the two two-dimensional networks produce consistent and high-confidence outputs are selected and combined with the orthogonal annotations as pseudo-labels for training the 3D network. This introduces more reliable, continuous and multi-dimensionally aligned fault information into the 3D fault identification model. At the same time, the three-dimensional network prediction results after confidence-based filtering are combined with the orthogonal annotations to guide the 2D network. This introduces additional multi-dimensional fault information into the 2D network, realizing collaborative training and mutual enhancement between 2D and 3D network models.

[0200] To verify that the orthogonal annotation-based fault recognition model training framework can train high-performance fault recognition networks with minimal annotations, experiments were conducted on a synthetic dataset. The results show that this approach achieves superior performance compared to other methods. Furthermore, experiments on two real-world datasets demonstrate that models trained on synthetic data using the orthogonal annotation-based fault recognition model training framework generalize well to real data, and that models trained with a small amount of real annotations perform well on real-world fault recognition tasks.

[0201] The following is a detailed description of the experimental verification process:

[0202] (1) Synthetic dataset used in the experiment:

[0203] The synthetic dataset used in the experiment is a known dataset. In this synthetic seismic data generation pipeline, a one-dimensional horizontal reflectivity model is first constructed using random values. Fold structures are introduced by vertically shearing the reflectivity model. The lateral variation is controlled by a combination of two-dimensional Gaussian functions, and the vertical attenuation is achieved using a linear scaling function. To increase structural complexity, a planar shear with laterally uniform and vertically invariant displacements is applied. Subsequently, planar faults are added, with displacements following a Gaussian or linear distribution, whose parameters are randomly selected to vary the dip, strike, and displacement pattern. The fault and fold reflectivity models are then convolved with a Ricker wavelet with a randomly selected peak frequency to enhance seismic realism, followed by the addition of random noise. Finally, a total of 220 3D seismic volumes with fault interpretation labels are obtained. Each volume has a shape of 128×128×128, corresponding to the time, survey line, and tie line orientation. Of these, 200 volumes are used to train the fault detection model, and the remaining 20 are retained to test the model's fault interpretation performance.

[0204] (2) Experimental process:

[0205] The orthogonal annotation-based fault identification model training framework (hereinafter referred to as Fault-Seg-OA) proposed in this example was implemented using Python 3.8.20 and PyTorch 1.12.0. Model training was performed in parallel on four NVIDIA GTX 2080Ti GPUs, each with 11,264 MiB of memory. The number of data loading worker threads was set to 8. The Adam optimizer was used with an initial learning rate of 0.001. For the 3D primary model, the batch size was set to 1, and each input was a 128×128×128 seismic volume. For the 2D auxiliary model, the batch size was 128, and each batch consisted of 128 128×128 seismic data slices. When processing the 3D primary model output, a hard threshold of 0.9 and a soft threshold of 0.7 were applied, and the accuracy threshold ι was set to 0.98. The weighting parameter β of the pseudo-labels in the loss function was gradually increased from 0 to 0.1 during training.

[0206] To evaluate the model performance, a set of statistical indicators are used, including precision, recall, mean intersection over union (mIoU), Dice coefficient and F1 score. Fault detection is formulated as a binary segmentation task on seismic images, where fault areas are considered as positive classes and non-fault areas are considered as negative classes. A confusion matrix is ​​constructed accordingly. True positives (TP) represent correctly classified fault pixels; false negatives (FN) are fault pixels that are misclassified as non-faults; false positives (FP) are non-fault pixels that are mislabeled as faults; and true negatives (TN) are correctly identified non-fault pixels. The variable k+1 represents the number of classes. These statistical indicators are derived from the confusion matrix:

[0207]

[0208] (3) Comparison with other state-of-the-art methods:

[0209] This example compares the proposed Fault-Seg-OA framework with other methods. The comparison includes seven state-of-the-art methods: dense-sparse co-training (DeSCO), which trains medical image segmentation models based on orthogonal annotations; orthogonal annotation and barely supervised learning (OABS), which improves DeSCO and makes its design more suitable for fault identification; cross-consistency training (CCT), which uses prediction consistency for training; direction context-aware consistency (DCC), which aligns features of the same semantic identity in different contexts; self-supervised semi-supervised learning (S4L), which incorporates self-supervised strategies into a semi-supervised learning framework to enhance feature representation; Fault-Seg-ST and Fault-Seg-SST, which introduce perturbation-based self-training and pseudo-label stability-guided retraining, respectively.

[0210] In the experiment, 20 synthetic earthquake samples not used in model training were selected as the test set. Model performance was evaluated using precision, recall, Dice coefficient, and mean Intersection over Union (MIoU). Furthermore, the fault prediction results on selected slices of the test data were visualized to qualitatively assess the segmentation quality. The following quantitative and qualitative analysis of the results verifies the effectiveness of Fault-Seg-OA compared to fully supervised (SupOnly) and other advanced semi-supervised and weakly supervised methods.

[0211] 3-1) Quantitative comparison: The performance indicators of Fault-Seg-OA and other comparison methods are summarized in Table 1. Among them, the fully supervised baselines Fault-Seg-OA, DeSCO and OABS use 3D convolutional neural network V-Net for fault recognition training. CCT, S4L, DCC, Fault-Seg-ST and Fault-Seg-SST are based on 2D convolutional neural network U-Net. The "Label" column in Table 1 records the amount of labeled data used by each method. SupOnly uses all labels of 200 synthetic training samples (i.e., 200×128×128×128) and normalizes them to 1 unit. Other methods are represented by scores corresponding to their relative label usage.

[0212] Table 1 Experimental results of Fault-Seg-OA and state-of-the-art methods

[0213]

[0214]

[0215] Analysis of Table 1 shows that Fault-Seg-OA consistently outperforms or performs on par with all comparison methods when using only 1 / 4 of the labels. Even when the number of labels is reduced to 1 / 8 or 1 / 16, Fault-Seg-OA maintains competitive performance across all metrics. A significant performance drop only occurs when the number of labels is reduced to 1 / 32. These results demonstrate the effectiveness of the orthogonal annotation strategy and 3D-2D cross-supervision framework in training high-performance fault recognition models with limited annotations.

[0216] Compared to other methods, Fault-Seg-OA achieves metrics that are higher than or comparable to S4L and DCC under the same label conditions. Under the same label conditions, it is slightly inferior to Fault-Seg-ST, Fault-Seg-SST, and CCT in terms of recall, but outperforms them in terms of precision, Dice coefficient, mean Intersection over Union (mIoU), and F1 score. Compared to SupOnly, Fault-Seg-OA achieves higher precision, Dice coefficient, mIoU, and F1 score using only 1 / 4 of the labels. The number of labels used by DeSCO and OABS is not directly shown in the table because both methods require an additional step of knowledge integration before training. DeSCO relies on registration-based pseudo-label generation, which generates pseudo-labels from a small set of annotated data—a common technique in medical image segmentation. However, its direct application to earthquake fault recognition results in poor pseudo-label quality, resulting in poor performance across all metrics. OABS, a modified version of DeSCO for earthquake data, requires full supervision for pre-training. In comparison, Fault-Seg-OA achieves higher precision, Dice coefficient, mIoU, and F1 score using fewer labels.

[0217] The method proposed in this example achieved significantly higher precision than all other methods, while having a slightly lower recall, indicating fewer false positives but more false negatives. This may be attributed to the confidence threshold used in the pseudo-label generation process. In the model output, even if the prediction is correct, the confidence scores at the fault boundaries are often low. After applying the confidence threshold, these boundary voxels are often excluded from the pseudo-label, causing the model to incorrectly classify some actual fault areas as non-faults during training. This leads to a higher false negative rate.

[0218] 3-2) Qualitative comparison:

[0219] Figure 7The fault identification results for selected slices of synthetic seismic data are shown, demonstrating the fault interpretation results of V-Net models trained using different methods on a synthetic test set. The model trained using Fault-Seg-OA with only 1 / 4 of the labels produced predictions that were most consistent with the ground truth. When the number of labels was reduced to 1 / 8 and 1 / 16, there was only a slight decrease in continuity, and the fault locations remained clearly discernible. Only when the number of labels was further reduced to 1 / 32 did a significant loss of continuity occur, demonstrating that the method in this example can achieve high-quality fault interpretation with fewer labels. In contrast, despite using all available labels, SupOnly predicted fault areas that were significantly larger than the ground truth, resulting in a high false positive rate—particularly noticeable in the upper right corner of the survey line slice. DeSCO and OABS exhibited significant false positives and missed positives, and their predictions were less consistent than those of Fault-Seg-OA. These visualizations further confirm that Fault-Seg-OA has fewer false positives and higher precision, consistent with the analysis in the quantitative comparison.

[0220] It should be noted that Figure 7 middle, DeSCO requires image registration to generate pseudo labels; the total number of pseudo labels is 1 / 10. OABS (Optimal Adaptive Binarization Selection) requires a pre-training phase; the total number of labels used in the pre-training and main training phases is 51 / 640.

[0221] (4) Ablation experiment:

[0222] The main contributions of the method proposed in this embodiment can be summarized in three aspects. First, orthogonally annotated seismic data are used to train the fault interpretation model. Second, a hard-soft threshold mechanism is designed to process the output of the three-dimensional model and generate reliable pseudo-labels for the two-dimensional network. Third, when selecting consistent outputs from two two-dimensional networks, a confidence threshold is introduced to identify areas where the two-dimensional network predictions are more confident than the three-dimensional network, thereby generating a mask to guide the training of the three-dimensional model. In this embodiment, V-Net is adopted as the unified backbone network, and an ablation study is performed to evaluate the effectiveness of each proposed component in improving the fault interpretation performance.

[0223] 4-1) Validity of orthogonal annotation:

[0224] The experimental results are shown in Table 2 and Figure 8Ablation experiment results on selected slices of synthetic seismic data shown. When using 1 / 16 of the complete annotation, training with orthogonally annotated seismic slices along the survey line and tie line directions enables Fault-Seg-OA to achieve 90.38% precision, 79.51% recall, 83.40% Dice coefficient and 74.56% mIoU. Under the same annotation budget (1 / 16), we further train the model using only labels annotated along the tie line direction. Under this setting, all evaluation metrics drop significantly, with precision dropping to 88.96%, recall dropping to 72.97%, Dice coefficient dropping to 77.02% and mean intersection over union (mIoU) dropping to 68.22%. As Figure 8 As shown, the model trained using unidirectional annotation exhibits significant errors in fault localization and poor continuity in prediction results. Because the training labels are limited to the tie-line direction, the model performs poorly in the time and survey line directions. These findings clearly demonstrate that at the same annotation cost, orthogonal annotation can significantly improve fault identification performance compared to unidirectional annotation. Orthogonal annotation provides complementary information from multiple spatial perspectives, enabling the model to capture more comprehensive and robust fault characteristics. This multi-dimensional guidance proves to be very effective in our proposed Fault-Seg-OA framework.

[0225] Table 2 Ablation experiment results of Fault-Seg-OA

[0226]

[0227] 4-2) Effectiveness of hard-soft threshold:

[0228] As shown in Table 2, we evaluated the impact of replacing the hard-soft threshold with a fixed hard threshold when processing the output of the 3D master model. While using a fixed hard threshold reduced the number of false positives in the generated pseudo-labels, improving precision from 90.38% to 94.49%, it also resulted in a significant decrease in other performance metrics. Specifically, recall dropped from 79.51% to 72.10%, the Dice coefficient dropped from 83.40% to 78.41%, and the mean Intersection over Union (mIoU) dropped from 74.56% to 69.26%. Figure 8 We further demonstrate that removing the hard-soft threshold has limited impact on overall fault localization. However, it is clear that the predicted fault regions become narrower and more discontinuous, indicating an increase in false negatives. These results demonstrate that the hard-soft threshold plays a key role in filtering reliable pseudo-labels from the 3D network output. Its adaptive design provides a better balance between precision and recall, contributing to improved overall fault segmentation performance.

[0229] 4-3) Effectiveness of confidence threshold in fusion:

[0230] In the process of fusing the outputs of two 2D networks, it is necessary to select locations where both networks produce consistent predictions with higher confidence scores than the 3D network in order to generate reliable pseudo-labels to guide the training of the 3D master model. In this section, we study the impact of removing the confidence-based selection criterion. The experimental results are summarized in Tables 2 and Figure 8 The results are similar to those observed when removing the hard-soft threshold. Eliminating this step also reduces false positives in the pseudo-labels used to supervise the 3D master model, improving the precision from 90.38% to 94.89%. However, other metrics drop significantly, with the recall dropping from 79.51% to 71.21%, the Dice coefficient dropping from 83.40% to 77.56%, and the mean Intersection over Union (mIoU) dropping from 74.56% to 68.47%. Figure 8 As shown in Figure 3, removing this confidence-based filtering also results in a narrower predicted fault region and an increase in the number of false negatives. These observations confirm the effectiveness of the confidence threshold in the fusion strategy in filtering unreliable predictions in the 2D network, thereby improving the quality of pseudo-labels and ultimately enhancing the 3D fault interpretation performance.

[0231] The following describes the application on two real-world datasets:

[0232] (1) Application on F3 data

[0233] To demonstrate the robustness and generalization ability of the proposed method, we conducted experiments on the F3 block earthquake dataset from the Netherlands, which has a dimension of 128 (time) × 384 (survey lines) × 512 (connection survey lines) samples. All evaluated methods used the same neural network architecture (V-Net) and were trained on 200 synthetic earthquake samples, but with different training strategies. The results are shown in Figure 2. Figure 9 As shown, Figure 9 Prediction results of the V-Net model trained with different methods on the F3 dataset: (a) a time slice of F3 earthquake data, (b) Fault-Seg-OA (1 / 4 labeled samples), (c) Fault-Seg-OA (1 / 8 labeled samples), (d) full supervision only, (e) DeSCO, (f) OABS.

[0234] Figure 9 (a) in Figure 3 shows a time slice of the F3 seismic volume. Figure 9 (b) and Figure 9 (c) shows the fault prediction results of the V-Net model trained with only 1 / 4 and 1 / 8 of the labeled synthetic data using the proposed Fault-Seg-OA framework, respectively. Figure 9(d) shows the prediction results of V-Net trained with full supervision using all labeled synthetic data. Figure 9 (e) and (f) in Figure 3 show the results of the V-Net model trained using the DeSCO and OABS methods, respectively. Figure 9 In (b) and (d), Fault-Seg-OA achieves more accurate fault detection on real data using only 1 / 4 of the labeled synthetic data, showing stronger robustness than the fully supervised model. Figure 9 The comparison between (b), (c), (e), and (f) in Figure 3 further demonstrates that our approach outperforms DeSCO and OABS in both accuracy and generalization ability. Notably, Fault-Seg-OA successfully identifies several faults that were missed by DeSCO and OABS.

[0235] Figure 9 We show that Fault-Seg-OA accurately identifies all major faults in real earthquake data using only one-quarter of the labeled synthetic data. Although the predicted faults exhibit some discontinuities, this is primarily due to distribution differences between the synthetic and real datasets. Due to the lack of accurate ground-truth labels for the F3 data, further improvement is expected if a small amount of orthogonal annotations are provided during training. Overall, the experiments demonstrate that Fault-Seg-OA achieves strong performance and robustness on real earthquake datasets, even when trained solely on synthetic data.

[0236] (2) Application to Thebe Data

[0237] To further evaluate the direct generalization ability of Fault-Seg-OA on real-world seismic data, we conducted experiments on the Thebe dataset. This dataset is the largest publicly available earthquake fault dataset and was originally obtained from a seismic survey of the Thebe gas field on the Exmouth Plateau in the Carnarvon Basin on the Northwest Shelf of Australia. Its size is 1804 (survey lines) × 3174 (connection survey lines) × 1537 (time steps). We extracted a sub-volume of size 96×3174×1537 for training and testing. The V-Net fault segmentation model was trained using Fault-Seg-OA at orthogonal label ratios of 8%, 4%, 2% and 1%. The model was then evaluated on unseen blocks without using its labels during training. The fault interpretation results on the Thebe dataset are shown in Figure 2. Figure 10 As shown, a line slice with a size of 300 (time) × 512 (connection line) is displayed. Figure 10Prediction results of different methods on the Thebes dataset. (a) A line slice from the Thebes seismic data. (b) True label. (c) Fault-Seg-OA (8%). (d) Fault-Seg-OA (4%). (e) Fault-Seg-OA (2%). (f) Fault-Seg-OA (1%). (g) V-Net trained using only full supervision. (h) FaultSeg3D.

[0238] Results show that Fault-Seg-OA demonstrates stable, accurate, and consistent performance at label rates of 8%, 4%, and 2%. Fault identification performance only degrades significantly when the label ratio drops to 1%, demonstrating the method's strong generalization capabilities even in extremely low-label scenarios. In contrast, fully supervised models trained solely on synthetic data, such as V-Net, often produce false positives due to the domain gap between synthetic and real datasets. The fully supervised method FaultSeg3D, while able to capture the long fault structure on the right side of the slice, still suffers from many false detections.

[0239] In summary, experiments show that models trained solely on synthetic data experience significant performance degradation when applied to real earthquake data with widely varying distributions. In contrast, the method proposed in this example significantly improves performance, achieving accurate and consistent fault identification with excellent generalization capabilities, requiring only minimal orthogonal annotation of real data.

[0240] A large number of experiments conducted on synthetic datasets and real earthquake datasets show that on synthetic datasets, using 1 / 4 of the labeled data can achieve results comparable to fully supervised learning, and even when the labeled data is reduced to 1 / 8 or 1 / 16, a high segmentation accuracy can still be maintained. On real datasets (F3 data and Thebe data), the generalization ability of the model is demonstrated, and only a small amount of real data annotation is required to significantly improve the accuracy and continuity of fault identification. Therefore, this embodiment significantly reduces the annotation workload through orthogonal annotation and multi-network collaborative training, while improving the adaptability of the model to complex real scenes, and consistently outperforms full supervision and representative semi-supervised and weakly supervised methods, achieving excellent performance in terms of accuracy, Dice coefficient and mean intersection over union (mIoU). In addition, the proposed framework exhibits strong generalization ability on real complex geological structures. This method provides a new approach to the fault segmentation task by using a small amount of labeled real data for training.

[0241] It should be noted that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device), or in conjunction with such instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by an instruction execution system, apparatus, or device, or in conjunction with such instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wires (electronic device), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.

[0242] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0243] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0244] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0245] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A fault recognition model training method, characterized in that: include: Acquire a seismic dataset, which includes a 3D seismic data volume and full-resolution fault annotations; Inputting the 3D seismic data volume into the main model and inputting the slice data of the 3D seismic data volume into the auxiliary model, respectively obtaining the main model prediction value and the auxiliary model prediction value; Combining the primary model predictions with the orthogonal annotation labels to generate auxiliary pseudo-labels, and performing a first post-processing operation on the primary model predictions to generate auxiliary masks; Combining the auxiliary model predictions with the orthogonal annotation labels to generate primary pseudo labels, and performing a second post-processing operation on the auxiliary model predictions to generate a primary mask; An auxiliary model guided by auxiliary pseudo labels and auxiliary masks and a main model guided by main pseudo labels and main masks are trained iteratively on the seismic dataset until a trained main model is obtained for performing the fault identification task.

2. The fault identification model training method according to claim 1, characterized in that: The three-dimensional seismic data volume x i and the full-resolution tomographic annotation y i They are all defined in three orthogonal dimensions, namely time V, survey line I and contact line C; The slice data is the two-dimensional slice data obtained by slicing the three-dimensional seismic data along the direction of the survey line. Or two-dimensional slice data cut in the direction of the contact line 3. The fault identification model training method according to claim 1, characterized in that: The main model is a three-dimensional network model based on the V-Net architecture.

4. The fault identification model training method according to claim 1, wherein: The auxiliary model includes a first sub-model and a second sub-model, and the first sub-model and the second sub-model are both two-dimensional network models based on the U-Net architecture.

5. The fault identification model training method according to claim 1, wherein: The orthogonal label of the i-th earthquake volume is: Where: represents the i-th 3D seismic volume x i Orthogonal label labels, and y i [h,l,w] represent and y i Any voxel in, s is the predefined annotation step, y i represents full-resolution fault annotation, and [h, l, w] represents the coordinate points of the 3D seismic data volume.

6. The fault identification model training method according to claim 1, characterized in that: The step of combining the main model prediction value with the orthogonal annotation label to generate an auxiliary pseudo label, and performing a first post-processing operation on the main model prediction value to generate an auxiliary mask, includes: The binarized result of the main model prediction value is combined with the orthogonal annotation label to generate the auxiliary pseudo label, which is expressed as follows: Where: represents auxiliary pseudo labels, Represents auxiliary pseudo labels For any voxel in represents the i-th earthquake volume x i Orthogonal labeling labels For any voxel in Represents the main model prediction value M1(x i ,θ M1 ) is the binarization result, s is the predefined annotation step size; A main confidence cube is generated according to the prediction value of the main model, and the main confidence cube is used to generate a main mask factor according to the hard-soft threshold, which is expressed as: Where: F1(M1(x i ,θ M1 )) represents the main mask factor, F1(M1(x i ,θ M1 ))[h,l,w] represents the main mask factor F1(M1(x i ,θ M1 )) in any voxel, represents the confidence cube, M1(x i ,θ M1 ) represents the prediction value of the main model, β is a predefined parameter, Th is a soft threshold or a hard threshold, and [c, h, l, w] represents the probability that the predicted category of the coordinate point [h, l, w] on the 3D seismic data volume is c; The positions of the orthogonal annotations in the three-dimensional seismic data volume are set to 1, and the other positions are consistent with the main mask factor to obtain the auxiliary mask.

7. The fault identification model training method according to claim 6, characterized in that: The method further includes: when ACC is greater than the accuracy critical value, Th is a soft threshold; when ACC is less than or equal to the accuracy critical value, Th is a hard threshold, wherein ACC is the accuracy calculated based on the main model prediction value and the orthogonal annotation label.

8. The fault identification model training method according to claim 4, characterized in that: The step of inputting the slice data of the three-dimensional seismic data volume into the auxiliary model comprises: Inputting slice data obtained by slicing the three-dimensional seismic data volume along the direction of the survey line into the first sub-model, and inputting slice data obtained by slicing the three-dimensional seismic data volume along the direction of the connecting survey line into the second sub-model; Accordingly, the method of combining the auxiliary model prediction value with the orthogonal annotation label to generate the main pseudo label specifically includes: The main pseudo label is generated by combining the binarization result of the prediction value of the first sub-model with the orthogonal annotation label, or the main pseudo label is generated by combining the binarization result of the prediction value of the second sub-model with the orthogonal annotation label.

9. The fault identification model training method according to claim 8, characterized in that: The step of performing a second post-processing operation on the auxiliary model prediction value to generate a primary mask includes: The prediction value of the first sub-model is fused with the prediction value of the second sub-model to determine an auxiliary mask factor. The auxiliary mask factor represents the high-confidence prediction area where the first sub-model and the second sub-model are consistent. The public expression is: Where: F2(M2(x i ′ ,θ M2 )) represents the auxiliary mask factor, F2(M2(x i ′ ,θ M2 ))[h,l,w] represents the auxiliary mask factor F2(M2(x i ′ ,θ M2 )) for any voxel, β is a predefined parameter, and They represent the binarization results of the main model prediction value, the binarization results of the first sub-model prediction value, and the binarization results of the second sub-model prediction value, respectively. and denote the confidence cube of the main model, the confidence cube of the first submodel, and the confidence cube of the second submodel, respectively; [h, l, w] denotes the coordinate points of the 3D seismic data volume; the set Ξ denotes the region where the first submodel and the second submodel are consistent; the set Ψ denotes the position where the second submodel and the main model are consistent; the set Υ and the set Λ denote the region where the second submodel has a higher confidence than the main model, respectively; ∩ denotes the intersection; and ∪ denotes the union; The orthogonally marked positions in the three-dimensional seismic data volume are set to 1, and the other positions are consistent with the auxiliary mask factors to obtain the main mask.

10. The fault identification model training method according to claim 1, wherein: The loss function used for mutual guidance training is: Where: and represents the loss function, and Represent the optimization parameters of the main model and the auxiliary model respectively, M1(x i ,θ M1 ) and M2(x i ′ ,θ M2 ) represent the prediction values ​​of the main model and the auxiliary model respectively, represents the primary pseudo label and primary mask, represents auxiliary pseudo labels and auxiliary masks, and Represent the total loss function of the main model and the total loss function of the auxiliary model, express The model parameter value that takes the minimum value.

11. The fault identification model training method according to claim 10, characterized in that: The formula of the total loss function of the main model is expressed as: Where: is the total loss function of the main model, is the cross entropy loss, Represents the main model prediction value and the main pseudo label The probability of matching, and Each represents a voxel; is the Dice loss, represents the confidence cube output by the main model for category c, represents the position of category c in the main pseudo label, Δ, Δ o , Δ p and Δ op It is an intermediate variable in the calculation process, V, I, C represent time, survey line and contact survey line respectively, and ∈ is a smoothing term.

12. The fault identification model training method according to claim 10, wherein: The formula of the total loss function of the auxiliary model is expressed as: Where: is the total loss function of the auxiliary model, is the loss of the first sub-model, is the loss of the second sub-model, is the predicted value of the first sub-model, is the predicted value of the second sub-model.

13. A fault identification model training system, characterized in that: include: A data acquisition module is used to acquire a seismic data set, which includes a three-dimensional seismic data volume and full-resolution fault annotations; A prediction module is used to input the 3D seismic data volume into the main model and the slice data of the 3D seismic data volume into the auxiliary model to obtain the main model prediction value and the auxiliary model prediction value respectively; a first processing module, configured to combine the primary model prediction value with the orthogonal annotation label to generate an auxiliary pseudo label, and perform a first post-processing operation on the primary model prediction value to generate an auxiliary mask; a second processing module, configured to combine the auxiliary model prediction value with the orthogonal annotation label to generate a primary pseudo label, and perform a second post-processing operation on the auxiliary model prediction value to generate a primary mask; The guided training module is used to iteratively perform mutual guided training on the seismic dataset, which includes training the auxiliary model based on the auxiliary pseudo-labels and auxiliary masks and training the main model based on the main pseudo-labels and main masks, until a trained main model is obtained to perform the fault identification task.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the fault identification model training method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Semi-supervised 3D fault identification method based on sparse labeling

    CN119251613A