Orthogonal labeling-based fault recognition model training method, fault recognition method, electronic equipment and storage medium
By adopting the training method of fault recognition model based on orthogonal annotation in fault recognition technology, using segmentation network and semi-supervised learning framework to generate and optimize pseudo-labels, the dependence problem on high-quality labels in fault recognition is solved, and the accuracy and efficiency of recognition are improved.
Patent Information
- Application Number
- CN202510122474.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-26
AI Technical Summary
Current fault identification technology faces the problems of seismic data complexity and noise interference, fault morphology diversity, data quality and resolution limitations, and the lack of high-quality labels, which leads to the increasing difficulty of fault feature extraction and accurate identification.
The training method of the fault recognition model based on orthogonal annotation is adopted. By slicing the seismic data sample set on multiple orthogonal dimensions, a segmentation network is used to generate pseudo-labels, and through multiple rounds of pseudo-label generation and parameter updates, a semi-supervised learning framework of optimized loss function and context prototype perception learning technology is trained.
It significantly reduces the dependence on high-quality full labels, reduces the cost of manual labeling during fault recognition, improves the accuracy and efficiency of fault recognition, and is suitable for other segmentation tasks that are difficult to use in high-quality labeling.
Smart Images

Figure CN120044591A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of fault identification, and specifically to a training method for a fault identification model based on orthogonal annotation, a fault identification method, an electronic device, and a storage medium. Background Art
[0002] Fault identification is an important link in seismic data interpretation. By identifying faults, the movement history of strata, fracture characteristics, and the distribution of oil and gas reservoirs can be revealed, which is of great significance for oil and gas exploration, geological disaster assessment, etc. However, current fault identification technologies face many difficulties, including the complexity and noise interference of seismic data, the diversity of fault morphologies, the limitations of data quality and resolution, and the lack of high-quality labels, which increase the difficulty of fault feature extraction and accurate identification. Current fault identification methods are mainly divided into supervised learning methods, weakly supervised learning, and semi-supervised learning methods, etc.
[0003] Existing supervised learning-based fault identification methods rely on a large number of high-quality labels, resulting in high annotation costs. Based on weakly supervised learning and semi-supervised learning methods, although the demand for labels is reduced compared to supervised learning methods, a certain number of high-quality labels are still usually required. The introduction of extremely weakly supervised learning methods based on orthogonal annotation can solve some of the above technical problems, but existing pseudo-label generation strategies are difficult to generate high-quality fault pseudo-labels while being independent of true label pre-training. Summary of the Invention
[0004] In view of the above problems, the present invention provides a training method for a fault identification model based on orthogonal annotation, a fault identification method, an electronic device, and a storage medium that improve the accuracy and efficiency of fault identification.
[0005] According to a first aspect of the present invention, there is provided a training method for a fault identification model based on orthogonal annotation, including:
[0006] Performing data slicing on a seismic data sample set in multiple orthogonal dimensions to obtain a multi-dimensional slice sample set, where the multi-dimensional slice sample set includes a time dimension slice sample set, an inline dimension slice sample set, and a cross dimension slice sample set;
[0007] Generating pseudo-labels for the multi-dimensional slice sample set using a segmentation network, and updating the parameters of the segmentation network using the loss value during the pseudo-label generation process, and obtaining initial pseudo-labels for the multi-dimensional slice sample set by alternately performing multiple rounds of pseudo-label generation operations and parameter update operations;
[0008] The initial pseudo-labels of the multi-dimensional slice sample set are cross-optimized between different dimensions in multiple rounds using an optimized loss function, and by introducing the three-dimensional information of the seismic data volume during the cross-optimization process, pseudo-labels integrating multi-dimensional information are obtained;
[0009] Based on the semi-supervised learning framework of context prototype-aware learning technology, the fault identification model is trained in multiple rounds using the seismic data sample set with pseudo-labels and the unlabeled seismic data training set, and the training process of the fault identification model is supervised using the training loss function to obtain the trained fault identification model.
[0010] According to an embodiment of the present invention, the above-mentioned method of generating pseudo-labels for the multi-dimensional slice sample set using a segmentation network, and updating the parameters of the segmentation network using the loss value during the pseudo-label generation process, and obtaining the initial pseudo-labels of the multi-dimensional slice sample set by alternately performing multiple rounds of pseudo-label generation operations and parameter update operations includes:
[0011] Using the first segmentation network to perform downsampling and upsampling operations on the first time slice in the time dimension slice sample set successively to generate the first-round iterative pseudo-label of the first time slice;
[0012] Processing the first-round iterative pseudo-label of the first time slice and the corresponding true label of the first time slice through the pseudo-label generation loss function to calculate the first-round iterative pseudo-label loss value;
[0013] Through backpropagation operations, using the first-round iterative pseudo-label loss value to update the parameters of the first segmentation network to obtain the first segmentation network after the first round of iteration;
[0014] Iteratively perform pseudo-label generation operations, pseudo-label loss value calculation operations, and parameter update operations until the number of iterations meets the first preset iteration value to obtain the initial pseudo-label of the first time slice;
[0015] Using the first segmentation network after the previous round of iteration and the initial pseudo-label of the previous time slice to perform the same processing operations on each time slice in the time dimension slice sample set to obtain the initial pseudo-label of the time dimension slice sample set.
[0016] According to an embodiment of the present invention, the above-mentioned method of generating pseudo-labels for the multi-dimensional slice sample set using a segmentation network, and updating the parameters of the segmentation network using the loss value during the pseudo-label generation process, and obtaining the initial pseudo-labels of the multi-dimensional slice sample set by alternately performing multiple rounds of pseudo-label generation operations and parameter update operations further includes:
[0017] Using the second segmentation network to perform the same processing operations on the inline dimension slice sample set as on the time dimension slice sample set to obtain the initial pseudo-label of the inline dimension slice sample set;
[0018] The third segmentation network is used to perform the same processing operations on the cross-dimensional slice sample set as on the time-dimensional slice sample set, and an initial pseudo-label of the cross-dimensional slice sample set is obtained.
[0019] According to an embodiment of the present invention, the above-mentioned multi-round cross-optimization of the initial pseudo-labels of the multi-dimensional slice sample set using the optimized loss function, and by introducing the three-dimensional information of the seismic data volume during the cross-optimization process, obtaining the pseudo-labels that fuse multi-dimensional information includes:
[0020] Performing a mask generation operation on the pixels in the seismic data sample set that are lower than the preset pixel threshold to obtain a pixel mask, and pairwise combining the multi-dimensional slice sample set in terms of dimensions to obtain a cross-dimensional slice sample set;
[0021] Processing the pixel mask and the initial pseudo-labels of the cross-dimensional slice sample set corresponding to the pixel mask using the optimized loss function to obtain a cross-optimization loss value;
[0022] Using the cross-optimization loss value to update the parameters of the segmentation network in the cross-optimization stage during the generation process of the initial pseudo-labels;
[0023] Using the segmentation network obtained in the cross-optimization stage to regenerate the initial pseudo-labels in the multi-dimensional slice sample set;
[0024] Iteratively performing the calculation operation of the cross-optimization loss value, the operation of updating the parameters of the segmentation network in the cross-optimization stage, and the operation of regenerating the initial pseudo-labels until the cross-optimization loss value is lower than the preset threshold, and obtaining the pseudo-labels that fuse multi-dimensional information.
[0025] According to an embodiment of the present invention, the above-mentioned semi-supervised learning framework based on the context prototype perception learning technology uses the seismic data sample set with pseudo-labels and the seismic data training set of unlabeled data to train the fault recognition model in multiple rounds, and uses the training loss function to supervise the training process of the fault recognition model, and the trained fault recognition model obtained includes:
[0026] Constructing a semi-supervised learning framework based on the teacher neural network - student neural network, and regarding the student neural network as the fault recognition model;
[0027] Introducing the context prototype perception learning technology into the semi-supervised learning framework, and constructing a training loss function based on the class activation map;
[0028] Using the teacher neural network to process the seismic data training set of unlabeled data to obtain a first output result, and using the student neural network to process the seismic data sample set with pseudo-labels to obtain a second output result;
[0029] The training loss value is calculated by processing the first output result and the second output result through a training loss function. Based on the backpropagation mechanism, the training loss value is used to update the parameters of the student neural network.
[0030] Based on the student neural network with updated parameters, the parameters of the teacher neural network are updated through an exponential moving average operation.
[0031] The operations of calculating the training loss value, obtaining the output result, and updating the neural network parameters are iteratively performed until the preset training conditions are met, and a trained fault recognition model is obtained.
[0032] According to an embodiment of the present invention, introducing the context prototype perception learning technology into the semi-supervised learning framework and constructing a training loss function based on the class activation map includes:
[0033] The output result of the teacher neural network is thresholded to generate a binary mask, and the unlabeled seismic data training set is subjected to masked average pooling according to the binary mask to obtain instance prototypes.
[0034] The instance prototypes are stored in a dynamically updatable support library, and the instance prototypes are clustered using a preset clustering algorithm.
[0035] A candidate context prototype set is constructed using the clustered instance prototypes, and the cosine similarity between each candidate context prototype in the candidate context prototype set and the instance prototype is calculated.
[0036] Multiple candidate context prototypes most similar to the instance prototype are selected from the candidate context prototype set based on the cosine similarity, and a context-aware prototype set is constructed using the selected multiple candidate context prototypes.
[0037] According to an embodiment of the present invention, introducing the context prototype perception learning technology into the semi-supervised learning framework and constructing a training loss function based on the class activation map further includes:
[0038] The positive correlation weight between each context-aware prototype in the context-aware prototype set and the instance prototype is calculated based on a parameter-free identity mapping layer and an activation function.
[0039] Each context-aware prototype in the context-aware prototype set is weighted using the positive correlation weight to obtain a weighted context-aware prototype set.
[0040] The distribution gap between the instance prototype and each context-aware prototype in the context-aware prototype set is calculated to obtain an offset term representing the distribution gap.
[0041] Replace the instance prototype with the result after the operation of the instance prototype and the offset item to obtain a corrected instance prototype, and use the corrected instance prototype to correct the weighted context-aware prototype set to obtain a corrected context-aware prototype set;
[0042] Use the corrected context-aware prototype set and the unlabeled seismic data training set to obtain a class activation map, and use the class activation map to construct a training loss function.
[0043] The second aspect of the present invention provides a fault identification method, including:
[0044] Use the trained fault identification model to process the fault data to be identified to obtain a fault identification result, where the trained fault identification model is trained by the training method of the fault identification model based on orthogonal annotation as described above, and the fault data to be identified includes seismic fault images.
[0045] The third aspect of the present invention provides an electronic device, including: one or more processors; a memory for storing one or more computer programs, where the above one or more processors execute the above one or more computer programs to implement the steps of the above method.
[0046] The fourth aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the above computer program or instruction is executed by a processor, the steps of the above method are implemented.
[0047] The training method of the fault identification model based on orthogonal annotation provided by the present invention significantly reduces the dependence on high-quality full labels by introducing the orthogonal annotation method, reducing the manual annotation cost in the fault identification process; at the same time, a dedicated pseudo-label generation strategy is designed for the complexity of seismic data, which is more adaptable to the non-linear and high-dimensional characteristics of seismic data; in addition, the introduction of the semi-supervised learning framework enables the present invention to only generate pseudo-labels for a small amount of data and complete model training by combining a large amount of unlabeled data. At the same time, the context-aware learning technology is introduced into the semi-supervised learning framework, further enhancing the ability of the present invention to extract seismic data features. Therefore, the fault identification model trained by the training method of the fault identification model based on orthogonal annotation provided by the present invention has wide applicability, is not only applicable to fault identification, but also can be extended to other segmentation tasks with difficult high-quality annotation, and has high application value. Description of the Drawings
[0048] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other purposes, features and advantages of the present invention will be clearer. In the drawings:
[0049] Figure 1It is an application scenario diagram of a training method and a fault identification method for an orthogonal annotation-based fault identification model according to an embodiment of the present invention;
[0050] Figure 2 It is a flowchart of a training method for an orthogonal annotation-based fault identification model according to an embodiment of the present invention;
[0051] Figure 3 It is a schematic diagram of a pseudo-label generation process according to an embodiment of the present invention;
[0052] Figure 4 It is a schematic diagram of training a fault identification model using a semi-supervised learning framework according to an embodiment of the present invention;
[0053] Figure 5 It is a structural block diagram of a training device for an orthogonal annotation-based fault identification model according to an embodiment of the present invention;
[0054] Figure 6 It is a block diagram of an electronic device suitable for implementing a training method and a fault identification method for an orthogonal annotation-based fault identification model according to an embodiment of the present invention. Detailed implementation manners
[0055] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present invention.
[0056] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0057] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0058] In the case of using expressions such as "at least one of A, B, and C, etc.", generally, it should be interpreted according to the meaning that those skilled in the art usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0059] Existing fault identification includes supervised fault identification methods, weakly supervised and semi-supervised fault identification methods.
[0060] Supervised fault identification methods rely on high-quality labels manually annotated and learn the features of faults from data through deep learning models. Such methods have the advantages of high accuracy and stable models, but are highly dependent on labels. Seismic data has the characteristics of complexity, non-linearity, and high noise, and the process of obtaining real fault labels is very difficult, requiring experts to spend a lot of time and energy on annotation. At the same time, although synthetic data can make up for the lack of labels to a certain extent, there are distribution differences between synthetic data and real data, resulting in limited generalization ability.
[0061] Due to the above various technical problems existing in supervised fault identification methods, therefore, it is of great value to explore weakly supervised learning and semi-supervised learning fault identification techniques. Weakly supervised learning and semi-supervised learning methods can use some labeled data and a large amount of unlabeled data for training in the case of insufficient labels, which reduces the dependence on labels to a certain extent. However, these methods usually still require a certain number of high-quality labels as initial guidance, and it is often difficult to balance the quality of pseudo-labels and the robustness of the model.
[0062] In the medical image segmentation task, an orthogonal annotation method has been proposed and applied. This method only annotates on some orthogonally perpendicular slices, effectively reducing the need for full annotation, lightening the annotation burden, and achieving extremely weakly supervised learning. Its advantage is that it can use extremely few annotations to generate pseudo-labels to guide model training. However, this method cannot be directly applied to the field of seismic fault identification. The current pseudo-label generation strategies for orthogonal annotation mainly include two types of methods: image registration and neural networks based on pre-training. The former requires strict registration relationships between images, and seismic data is difficult to meet this condition due to the complexity of the acquisition method. The latter requires real labels to pre-train the neural network, and the quality of pseudo-labels is directly related to it, which still cannot completely get rid of the dependence on real labels.
[0063] To overcome at least one of the problems in the prior art, embodiments of the present invention provide a method for training a fault recognition model based on orthogonal annotation and a fault recognition method, which rely on a very small number of labeled slices to generate a large number of reliable pseudo-labels, and then use the pseudo-labels to guide the training of the fault recognition model according to the semi-supervised learning framework, so that the trained fault recognition model can improve the recognition efficiency and accuracy of fault data.
[0064] Figure 1 FIG. is an application scenario diagram of the method for training a fault recognition model based on orthogonal annotation and the fault recognition method according to an embodiment of the present invention.
[0065] As Figure 1 shown, the application scenario 100 according to this embodiment may include the field of fault recognition technology. The network 104 is used to provide a medium for communication links between the first terminal device 101, the second terminal device 102, the third terminal device 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or fiber optic cables, etc.
[0066] Users can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0067] The first terminal device 101, the second terminal device 102, and the third terminal device 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0068] The server 105 may be a server that provides various services, such as a background management server that supports the websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (only as an example). The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0069] It should be noted that the training method and fault identification method of the fault identification model based on orthogonal annotation provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the training device of the fault identification model based on orthogonal annotation provided by the embodiments of the present invention can generally be set in the server 105. The training method and fault identification method of the fault identification model based on orthogonal annotation provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Correspondingly, the training device of the fault identification model based on orthogonal annotation provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0070] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in
[0071] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 1 The following will be based on Figures 2 to 4 the described scenario and will describe in detail the training method and fault identification method of the fault identification model based on orthogonal annotation of the disclosed embodiments through
[0072] Figure 2 is a flowchart of the training method of the fault identification model based on orthogonal annotation according to the embodiments of the present invention.
[0073] As Figure 2 shown, the above-mentioned training method of the fault identification model based on orthogonal annotation includes operation S210 to operation S240.
[0074] In operation S210, the seismic data sample set is sliced in multiple orthogonal dimensions to obtain a multi-dimensional slice sample set, where the multi-dimensional slice sample set includes a time dimension slice sample set, an inline dimension slice sample set, and a cross dimension slice sample set.
[0075] Operation S210 is used to construct the data sample set used for model training, that is, to slice the seismic data sample set into several three-dimensional bodies of specific dimensions. In the present invention, the data dimensions are selected as (128, 128, 128), and these three dimensions respectively represent vertical (time dimension, the same below), inline (inline dimension, the same below), and crossline (cross dimension, the same below) directions. Among them, the th three-dimensional seismic body is denoted as , and the The slice is denoted as , the th slice in the inline direction is denoted as , the th slice in the crossline direction is denoted as . The label of the th 3D seismic volume is denoted as . One label slice is required in each direction to guide the generation of pseudo-labels, that is , and .
[0076] The above three dimensions: vertical, inline, and crossline are pairwise orthogonal dimensions. Inspired by the technology of the present invention, those skilled in the art can adopt other orthogonal multiple dimensions not limited to the above three dimensions, or through data preprocessing, convert non-orthogonal dimensions into orthogonal dimensions first.
[0077] In operation S220, a segmentation network is used to generate pseudo-labels for a multi-dimensional slice sample set, and the loss value during the pseudo-label generation process is used to update the parameters of the segmentation network. An initial pseudo-label for the multi-dimensional slice sample set is obtained by alternately performing multiple rounds of pseudo-label generation operations and parameter update operations.
[0078] Operation S220 is used for the initial generation of pseudo-labels. During the initial generation of pseudo-labels, a segmentation network is used, such as a U-Net network, and the U-Net network structure is adopted to predict pseudo-labels. U-Net is a convolutional neural network with a symmetric structure of an encoder and a decoder. The encoder extracts multi-level features through convolution and pooling, and the decoder gradually restores the spatial resolution through convolution and upsampling, while using skip connections to fuse the high-resolution features of the encoder to achieve end-to-end segmentation from the input to the output. Three U-Net networks are initialized for the expansion of the three direction labels of orthogonal annotation. The U-Net network for pseudo-label expansion in the vertical direction is denoted as , the U-Net network for pseudo-label expansion in the inline direction is denoted as , and the U-Net network for pseudo-label expansion in the crossline direction is denoted as .
[0079] The above operation S220 involves two processes, namely the pseudo-label generation process and the parameter update process of the U-Net network. The above two processes are alternately performed. After iteratively updating the parameters of the U-Net network to a certain extent, the finally updated U-Net is used to generate the initial pseudo-label.
[0080] Under the technical inspiration provided by the present invention, those skilled in the art can adopt other reasonable image segmentation networks according to actual needs.
[0081] In operation S230, the initial pseudo-labels of the multi-dimensional slice sample set are cross-optimized in multiple rounds between different dimensions by using the optimized loss function, and by introducing the three-dimensional information of the seismic data volume during the cross-optimization process, pseudo-labels integrating multi-dimensional information are obtained.
[0082] The above operation S230 mainly optimizes the calculation of the loss value and updates the parameters of the U-Net network again. Through cross-optimization, the integration of information between different dimensions is achieved, and the pseudo-labels integrating multi-dimensional information are used as the labels of the seismic data sample set.
[0083] In operation S240, based on the semi-supervised learning framework of the context prototype perception learning technology, the fault recognition model is trained in multiple rounds by using the seismic data sample set with pseudo-labels and the unlabeled seismic data training set, and the training process of the fault recognition model is supervised by using the training loss function to obtain the trained fault recognition model.
[0084] The above operations S210 to S240 schematically show the training method of the fault recognition model for seismic fault data recognition. Under the technical inspiration of the present invention, the above operations S210 to S240 can also be applied to the training of other fault image recognition models, such as the recognition of computed tomography images (CT images). That is, through the above operations S210 to S240, a model or method for other fault image recognition can be obtained.
[0085] The training method of the fault recognition model based on orthogonal annotation provided by the present invention significantly reduces the dependence on high-quality full labels by introducing the orthogonal annotation method, and reduces the manual annotation cost in the fault recognition process. At the same time, a dedicated pseudo-label generation strategy is designed for the complexity of seismic data, which is more adaptable to the non-linear and high-dimensional characteristics of seismic data. In addition, the introduction of the semi-supervised learning framework enables the present invention to only generate pseudo-labels for a small amount of data and complete model training by combining a large amount of unlabeled data. At the same time, the context-aware learning technology is introduced into the semi-supervised learning framework, which further enhances the ability of the present invention to extract seismic data features. Therefore, the fault recognition model trained by the training method of the fault recognition model based on orthogonal annotation provided by the present invention has wide applicability, is not only applicable to fault recognition, but also can be extended to other segmentation tasks with difficult high-quality annotation, and has high application value.
[0086] According to an embodiment of the present invention, the above method for generating pseudo-labels for a multi-dimensional slice sample set using a segmentation network, and updating the parameters of the segmentation network using the loss value during the pseudo-label generation process, to obtain the initial pseudo-labels of the multi-dimensional slice sample set by alternately performing multiple rounds of pseudo-label generation operations and parameter update operations includes: using a first segmentation network to perform downsampling and upsampling operations on the first time slice in the time-dimensional slice sample set in sequence to generate the first-round iterative pseudo-labels of the first time slice; processing the first-round iterative pseudo-labels of the first time slice and the true labels corresponding to the first time slice through a pseudo-label generation loss function to calculate the first-round iterative pseudo-label loss value; through backpropagation operations, using the first-round iterative pseudo-label loss value to update the parameters of the first segmentation network to obtain the first segmentation network after the first round of iteration; iteratively performing pseudo-label generation operations, pseudo-label loss value calculation operations, and parameter update operations until the number of iterations meets a first preset iteration value to obtain the initial pseudo-labels of the first time slice; using the first segmentation network after the previous round of iteration and the initial pseudo-labels of the previous time slice to perform the same processing operations on each time slice in the time-dimensional slice sample set to obtain the initial pseudo-labels of the time-dimensional slice sample set.
[0087] According to an embodiment of the present invention, the above method for generating pseudo-labels for a multi-dimensional slice sample set using a segmentation network, and updating the parameters of the segmentation network using the loss value during the pseudo-label generation process, to obtain the initial pseudo-labels of the multi-dimensional slice sample set by alternately performing multiple rounds of pseudo-label generation operations and parameter update operations further includes: using a second segmentation network to perform the same processing operations on the inline-dimensional slice sample set as on the time-dimensional slice sample set to obtain the initial pseudo-labels of the inline-dimensional slice sample set; using a third segmentation network to perform the same processing operations on the cross-dimensional slice sample set as on the time-dimensional slice sample set to obtain the initial pseudo-labels of the cross-dimensional slice sample set.
[0088] The above embodiments relate to the generation process of initial pseudo-labels for multiple dimensions (i.e., three orthogonal dimensions of vertical, inline, and crossline). The present invention further elaborates on the following specific implementation process and attached Figure 3 to further illustrate the above process in detail.
[0089] Figure 3 is a schematic diagram of the pseudo-label generation process according to an embodiment of the present invention.
[0090] The pseudo-label generation process includes the initial pseudo-label generation process (1) and the initial pseudo-label optimization process (2).
[0091] (1.1) Initial generation of pseudo labels: Generate the pseudo labels of the second slice based on the true labels of the first slice (the true labels of the seismic data sample set itself, and use orthogonal annotation to take out the true labels of the dimension for initial generation of pseudo labels): Earthquake data samples The pseudo-label generation process is as follows: Figure 3 As shown, take the vertical direction as an example. First, The first slice in the vertical direction Enter the U-net network , and forward propagation is performed. According to the network Output and the true label slice Calculating supervised loss The process is shown in formulas (1) to (3):
[0092] (1),
[0093] (2)
[0094] (3),
[0095] in, is the cross entropy loss function, is the dice loss function, No. The pixels are recorded as , No. The pixels are recorded as According to the supervision loss and cross loss Calculate the total loss function as shown in formula (4):
[0096] (4),
[0097] Among them, the cross loss in the first round of pseudo label generation process The value of is always 0. is a hyperparameter that does not participate in the update. Back propagation calculates the loss function The gradient of the U-net network is updated according to the gradient descent Repeat the forward propagation, back propagation, and parameter update operations several times, and record the final network output, i.e., the pseudo label. .
[0098] Repeat the entire process above, using the second slice As a network Input, according to the network Output and the previous pseudo-label slice Calculate the supervised loss . Network After several trainings, record the final output pseudo-label The process of generating pseudo-labels in the Inline and crossline directions is synchronized with the vertical direction and is exactly the same. Network Obtain pseudo-labels through training and , Network Obtain pseudo-labels through training and .
[0099] (1.2) Generation of remaining pseudo-labels: Simultaneously perform in three directions: vertical, Inline, and crossline, and repeat the steps in (1.1). Update the neural network parameters under the guidance of the previous slice pseudo-label, and then obtain new pseudo-labels. Finally, all the pseudo-label slices obtained in the vertical direction constitute , all the pseudo-label slices obtained in the Inline direction constitute , all the pseudo-label slices obtained in the crossline direction constitute .
[0100] According to the embodiments of the present invention, the above-mentioned multi-round cross-optimization of the initial pseudo-labels of the multi-dimensional slice sample set by using the optimized loss function, and by introducing the three-dimensional information of the seismic data volume during the cross-optimization process, the pseudo-labels integrating multi-dimensional information are obtained, including: performing a masking generation operation on the pixels in the seismic data sample set that are lower than the preset pixel threshold to obtain a pixel mask, and combining the multi-dimensional slice sample set pairwise in dimensions to obtain a cross-dimensional slice sample set; processing the pixel mask and the initial pseudo-labels of the cross-dimensional slice sample set corresponding to the pixel mask through the optimized loss function to obtain a cross-optimization loss value; using the cross-optimization loss value to update the parameters of the segmentation network in the cross-optimization stage during the generation process of the initial pseudo-labels; using the segmentation network obtained in the cross-optimization stage to re-generate the initial pseudo-labels in the multi-dimensional slice sample set; iteratively performing the calculation operation of the cross-optimization loss value, the operation of updating the parameters of the segmentation network in the cross-optimization stage, and the operation of re-generating the initial pseudo-labels until the cross-optimization loss value is lower than the preset threshold to obtain the pseudo-labels integrating multi-dimensional information.
[0101] The following further elaborates on the above optimization process of the initial pseudo-labels through specific implementation manners.
[0102] Optimization process of initial pseudo-labels (2): AsFigure 3 As shown, repeat all steps in the generation process (1) of the initial pseudo-labels. After generating new pseudo-label slices, update the recorded , and . During the calculation of the loss, it is necessary to select those pixels with uncertainty lower than the threshold and generate the corresponding mask . The cross loss is no longer 0, and its calculation is as shown in formulas (5) to (8):
[0103] (5),
[0104] (6),
[0105] (7),
[0106] (8),
[0107] where, is the single-pixel value of the mask of the two selected pseudo-label bodies, is 's single-pixel value, is 's single-pixel value, is 's single-pixel value. Repeat the process of generating pseudo-labels until the value of drops below the set threshold. The pseudo-labels obtained at this time introduce three-dimensional information of the seismic data volume compared with the pseudo-labels initially generated in the generation process (1) of the initial pseudo-labels. , and The pseudo-label bodies generated in these three directions gradually tend to be unified, and finally any one is selected as the finally obtained pseudo-label . Repeating step (1) of the initial pseudo-label generation process can obtain fault pseudo-labels for any number of samples.
[0108] According to an embodiment of the present invention, the above semi-supervised learning framework based on context prototype-aware learning technology trains a fault recognition model using a seismic data sample set with pseudo-labels and a seismic data training set of unlabeled data for multiple rounds, and uses a training loss function to supervise the training process of the fault recognition model. The trained fault recognition model obtained includes: constructing a semi-supervised learning framework based on a teacher neural network - student neural network, and regarding the student neural network as the fault recognition model; introducing the context prototype-aware learning technology into the semi-supervised learning framework, and constructing a training loss function based on a class activation map; using the teacher neural network to process the seismic data training set of unlabeled data to obtain a first output result, and using the student neural network to process the seismic data sample set with pseudo-labels to obtain a second output result; calculating a training loss value by processing the first output result and the second output result through the training loss function, and based on the backpropagation mechanism, using the training loss value to update the parameters of the student neural network; based on the student neural network with updated parameters, updating the parameters of the teacher neural network through an exponential moving average operation; iteratively performing the training loss value calculation operation, the output result acquisition operation, and the neural network parameter update operation until a preset training condition is met, to obtain a trained fault recognition model.
[0109] Figure 4 It is a schematic diagram of training a fault recognition model by the semi-supervised learning framework according to an embodiment of the present invention.
[0110] For the forward propagation of the fault recognition model, using the pseudo-label with fused multi-dimensional information as the label of the seismic data sample set, and denoting the above data as , and denoting the corresponding pseudo-label as . The data without labels is denoted as .
[0111] The semi-supervised learning framework in the present invention is based on the common teacher-student framework. The overall framework structure is as Figure 4 shown, where both the teacher network and the student network adopt a U-Net network structure constructed by three-dimensional convolutional layers. The teacher network is denoted as , and the student network is denoted as . Inputting the pseudo-labeled data into the student network to obtain the output . Inputting the unlabeled data into the teacher model and the student model respectively to obtain and . The calculation of the loss functions and is shown in Formulas (9) and (10):
[0112] (9),
[0113] (10),
[0114] Wherein, is the cross - entropy loss function, is the class activation map.
[0115] Obtaining the class activation map The process will be elaborated in the subsequent embodiments.
[0116] According to an embodiment of the present invention, introducing the context prototype - aware learning technology into the semi - supervised learning framework and constructing the training loss function based on the class activation map includes: thresholding the output result of the teacher neural network to generate a binary mask, and performing masked average pooling on the unlabeled seismic data training set according to the binary mask to obtain instance prototypes; storing the instance prototypes in a dynamically updatable support library, and performing clustering processing on the instance prototypes using a preset clustering algorithm; constructing a candidate context prototype set using the clustered instance prototypes, and calculating the cosine similarity between each candidate context prototype in the candidate context prototype set and the instance prototypes; selecting multiple candidate context prototypes most similar to the instance prototypes from the candidate context prototype set based on the cosine similarity, and constructing a context - aware prototype set using the selected multiple candidate context prototypes.
[0117] According to an embodiment of the present invention, introducing the context prototype - aware learning technology into the semi - supervised learning framework and constructing the training loss function based on the class activation map further includes: calculating the positive - correlation weight between each context - aware prototype in the context - aware prototype set and the instance prototypes based on a parameter - free identity mapping layer and an activation function; performing weighted calculation on each context - aware prototype in the context - aware prototype set using the positive - correlation weight to obtain a weighted context - aware prototype set; calculating the distribution gap between the instance prototypes and each context - aware prototype in the context - aware prototype set to obtain an offset term representing the distribution gap; replacing the instance prototypes with the result after the operation of the instance prototypes and the offset term to obtain corrected instance prototypes, and using the corrected instance prototypes to correct the weighted context - aware prototype set to obtain a corrected context - aware prototype set; obtaining the class activation map using the corrected context - aware prototype set and the unlabeled seismic data training set, and constructing the training loss function using the class activation map.
[0118] The following further elaborates in detail the process of training the fault recognition model using the above - mentioned semi - supervised learning framework provided by the present invention through specific embodiments and in combination with the attached Figure 4 drawings.
[0119] As Figure 4 shown, in terms of context prototype - aware learning, it includes masked average pooling, k - means clustering, outputting the class activation map, and calculating the loss function.
[0120] Mask average pooling: For the output of the teacher network perform thresholding to generate a binary mask . According to the unlabeled data perform mask average pooling to obtain , as shown in formula (11):
[0121] (11),
[0122] where is three-dimensional seismic data, and the indices corresponding to its three dimensions are respectively represented by , and . As the output of the teacher model, it has four dimensions. The first dimension is the recognition category, that is, fault and non-fault, and its index is represented by , n takes the value of 0 or 1, and the remaining three dimensions are consistent. The dimension of is the same as
[0123] k-means clustering: Store the instance prototype of each data in the support library, and this support library will be updated dynamically during training. Apply the k-means clustering algorithm to divide the prototypes of the same category into multiple clustering groups. Each clustering group represents a context feature pattern of this category. For the category , extract all instance prototypes from the support library to construct a candidate context prototype set .
[0124] Output class activation map: Calculate the similarity between each context prototype and the current instance prototype using the cosine similarity metric, as shown in formula (12):
[0125] (12).
[0126] Based on the calculated similarity, select the context prototype most similar to the current instance from the candidate set to form a context-aware prototype set, as shown in formula (13):
[0127] (13),
[0128] For each context prototype , calculate its positive correlation weight with the instance prototype according to the parameter-free identity mapping layer and the softmax function , as shown in formula (14):
[0129] (14),
[0130] where and are parameter - free identity mapping layers in feature transformation, is the scale factor for adjusting the weight . Weight the selected context prototypes, as shown in formula (15):
[0131] (15).
[0132] Calculate the distribution gap between the dense distribution center of the current instance feature and the context prototype. Introduce an offset term for the current instance prototype, as shown in formula (16):
[0133] (16),
[0134] Align the instance feature to the distribution center of the context feature, that is, use to replace the original .
[0135] Use the context - aware prototype set to calculate the enhanced class activation value for each pixel position of the image, as shown in formula (17):
[0136] (17),
[0137] where, is the number of context prototypes selected as the most relevant to the class . The class activation map containing both the whole fault and non - fault categories is denoted as .
[0138] Calculate the loss function: Combine the output Figure 4 of the teacher model shown in with the class activation map to calculate the loss function, as shown in formula (18):
[0139] (18).
[0140] Loss function weight sum: The loss function of the semi - supervised learning framework introducing context prototype - aware learning technology based on the teacher - student framework is , and weighted sum, calculated as described in formula (19):
[0141] (19).
[0142] Backpropagation: According to the loss function calculated in the loss function weights, gradient backpropagation is performed to update the network parameters in the framework. The gradient propagation is only carried out within the student network and the context prototype perception learning model, while the weights of the teacher network are updated by the exponential moving average of the student network weights. At the same time, during the entire training process, the context prototype set and instance features are dynamically updated through the support library, and the output of the teacher model is continuously optimized.
[0143] The second aspect of the present invention provides a fault identification method, including: using the trained fault identification model to process the fault data to be identified to obtain a fault identification result, where the trained fault identification model is trained as described in the above-mentioned training method of the fault identification model based on orthogonal annotation, and the fault data to be identified includes seismic fault images.
[0144] Using the training method of the fault identification model provided by the present invention, a fault identification model is obtained based on the weights of the student neural network for identifying fault data.
[0145] Those skilled in the art can, under the inspiration given by the present invention, identify other fault data, such as computed tomography images (CT images), etc.
[0146] Based on the above-mentioned training method of the fault identification model based on orthogonal annotation, the present invention also provides a training device for the fault identification model based on orthogonal annotation. The following will be combined with Figure 5 This device will be described in detail.
[0147] Figure 5 It is a structural block diagram of a training device for a fault identification model based on orthogonal annotation according to an embodiment of the present invention.
[0148] As Figure 5 shown, the above-mentioned training device 500 for the fault identification model based on orthogonal annotation includes an orthogonal slicing module 510, a pseudo-label generation module 520, a pseudo-label optimization module 530, and a semi-supervised training module 540.
[0149] The orthogonal slicing module 510 is used to slice the seismic data sample set in multiple orthogonal dimensions to obtain a multi-dimensional slice sample set, where the multi-dimensional slice sample set includes a time-dimensional slice sample set, an inline-dimensional slice sample set, and a cross-dimensional slice sample set; in one embodiment, the orthogonal slicing module 510 can be used to perform the operation S210 described above, which will not be elaborated here.
[0150] The pseudo-label generation module 520 is used to generate pseudo-labels for the multi-dimensional slice sample set by using the segmentation network, and update the parameters of the segmentation network by using the loss value during the pseudo-label generation process. The initial pseudo-labels of the multi-dimensional slice sample set are obtained by alternately performing multiple rounds of pseudo-label generation operations and parameter update operations. In one embodiment, the pseudo-label generation module 520 can be used to execute the operation S220 described above, which will not be elaborated here.
[0151] The pseudo-label optimization module 530 is used to perform cross-optimization among different dimensions of the initial pseudo-labels of the multi-dimensional slice sample set by using the optimization loss function, and obtain pseudo-labels that fuse multi-dimensional information by introducing the three-dimensional information of the seismic data volume during the cross-optimization process. In one embodiment, the pseudo-label optimization module 530 can be used to execute the operation S230 described above, which will not be elaborated here.
[0152] The semi-supervised training module 540 is used to perform multi-round training on the fault identification model by using the semi-supervised learning framework based on the context prototype perception learning technology, and use the training loss function to supervise the training process of the fault identification model to obtain the trained fault identification model. In one embodiment, the semi-supervised training module 540 can be used to execute the operation S240 described above, which will not be elaborated here.
[0153] According to the embodiments of the present invention, any multiple of the orthogonal slice module 510, the pseudo-label generation module 520, the pseudo-label optimization module 530, and the semi-supervised training module 540 can be combined and implemented in one module, or any one of them can be split into multiple modules. Or, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to the embodiments of the present invention, at least one of the orthogonal slice module 510, the pseudo-label generation module 520, the pseudo-label optimization module 530, and the semi-supervised training module 540 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or can be implemented by any other reasonable way of integrating or packaging circuits, etc., in hardware or firmware, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Or, at least one of the orthogonal slice module 510, the pseudo-label generation module 520, the pseudo-label optimization module 530, and the semi-supervised training module 540 can be at least partially implemented as a computer program module, and when the computer program module is run, it can execute the corresponding functions.
[0154] Figure 6 It is a block diagram of an electronic device suitable for implementing a training method and a fault identification method of an orthogonal annotation-based fault identification model according to an embodiment of the present invention.
[0155] As Figure 6 shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage section 608 into a random access memory (RAM) 603. The processor 601 can include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 601 can also include on-board memory for caching purposes. The processor 601 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.
[0156] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. The processor 601 performs various operations of the method flow according to an embodiment of the present invention by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the program can also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 can also perform various operations of the method flow according to an embodiment of the present invention by executing the programs stored in the one or more memories.
[0157] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the input / output (I / O) interface 605: an input section 606 including a keyboard, a mouse, etc.; an output section 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 608 including a hard disk, etc.; and a communication section 609 including a network interface card such as a LAN card, a modem, etc. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the input / output (I / O) interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read from it can be installed into the storage section 608 as needed.
[0158] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the methods according to the embodiments of the present invention are implemented.
[0159] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 602 and / or the RAM 603 described above and / or one or more memories other than the ROM 602 and the RAM 603.
[0160] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0161] Those skilled in the art can understand that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.
[0162] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.
Claims
1. A training method for a fault recognition model based on orthogonal annotation, characterized in that: The method comprises: Slicing the seismic data sample set in multiple orthogonal dimensions to obtain a multi-dimensional slice sample set, wherein the multi-dimensional slice sample set includes a time dimension slice sample set, an inline dimension slice sample set, and a cross dimension slice sample set; Generate pseudo labels for the multi-dimensional slice sample set using a segmentation network, and update parameters of the segmentation network using the loss value in the pseudo label generation process, and obtain initial pseudo labels for the multi-dimensional slice sample set by alternating multiple rounds of pseudo label generation operations and parameter update operations; The initial pseudo labels of the multi-dimensional slice sample set are cross-optimized in multiple rounds in different dimensions by using an optimized loss function, and the three-dimensional information of the seismic data volume is introduced in the cross-optimization process to obtain pseudo labels that integrate the multi-dimensional information; A semi-supervised learning framework based on contextual prototype-aware learning technology uses a seismic data sample set with the pseudo-labels and an unlabeled seismic data training set to perform multiple rounds of training on the fault recognition model, and uses a training loss function to supervise the training process of the fault recognition model to obtain a trained fault recognition model.
2. The method according to claim 1, characterized in that Generating pseudo labels of the multi-dimensional slice sample set by using a segmentation network, and updating parameters of the segmentation network by using the loss value in the pseudo label generation process, and obtaining initial pseudo labels of the multi-dimensional slice sample set by alternating multiple rounds of pseudo label generation operations and parameter update operations, including: Using a first segmentation network, successively performing downsampling operations and upsampling operations on a first time slice in the time dimension slice sample set to generate a first round of iterative pseudo labels for the first time slice; Processing the first-round iteration pseudo-label of the first time slice and the real label corresponding to the first time slice through the pseudo-label generation loss function, and calculating the first-round iteration pseudo-label loss value; Through a back-propagation operation, the first segmentation network is parameter updated using the pseudo-label loss value of the first iteration to obtain a first segmentation network after the first iteration; Iterate the pseudo-label generation operation, the pseudo-label loss value calculation operation, and the parameter update operation until the number of iterations meets a first preset iteration value, thereby obtaining an initial pseudo-label for the first time slice; The same processing operation as the first time slice is performed on each time slice in the time dimension slice sample set using the first segmentation network after the previous iteration and the initial pseudo label of the previous time slice to obtain the initial pseudo label of the time dimension slice sample set.
3. The method according to claim 2, characterized in that Also includes: Using a second segmentation network, the inline dimension slice sample set is subjected to the same processing operation as the time dimension slice sample set to obtain an initial pseudo label of the inline dimension slice sample set; The third segmentation network is used to perform the same processing operation on the cross-dimensional slice sample set as that on the time dimension slice sample set to obtain an initial pseudo label of the cross-dimensional slice sample set.
4. The method according to claim 1, characterized in that: The initial pseudo labels of the multi-dimensional slice sample set are cross-optimized in multiple rounds in different dimensions by using the optimized loss function, and the pseudo labels integrating the multi-dimensional information are obtained by introducing the three-dimensional information of the seismic data volume in the cross-optimization process, including: Performing a mask generation operation on pixels in the seismic data sample set that are lower than a preset pixel threshold to obtain a pixel mask, and performing two-by-two dimensional combination of the multi-dimensional slice sample sets to obtain a cross-dimensional slice sample set; Processing the pixel mask and the initial pseudo labels of the cross-dimensional slice sample set corresponding to the pixel mask through the optimization loss function to obtain a cross-optimization loss value; Using the cross-optimization loss value to update the parameters of the segmentation network obtained in the initial pseudo-label generation process in the cross-optimization phase; Regenerating the initial pseudo-labels for the multi-dimensional slice samples using the segmentation network obtained in the cross optimization stage; The calculation operation of the cross-optimization loss value, the update operation of the segmentation network parameters in the cross-optimization stage, and the regeneration operation of the initial pseudo-label are iteratively performed until the cross-optimization loss value is lower than a preset threshold, thereby obtaining the pseudo-label that integrates the multi-dimensional information.
5. The method according to claim 1, characterized in that: A semi-supervised learning framework based on contextual prototype-aware learning technology is used to perform multiple rounds of training on the fault recognition model using a seismic data sample set with the pseudo-labels and a seismic data training set with unlabeled data, and a training loss function is used to supervise the training process of the fault recognition model. The trained fault recognition model includes: Constructing the semi-supervised learning framework based on the teacher neural network-student neural network, and using the student neural network as the fault recognition model; Introducing the contextual prototype-aware learning technology into the semi-supervised learning framework, and constructing a training loss function based on a class activation map; Using the teacher neural network to process the unlabeled seismic data training set to obtain a first output result, and using the student neural network to process the seismic data sample set with the pseudo-label to obtain a second output result; Processing the first output result and the second output result by the training loss function to calculate a training loss value, and updating the parameters of the student neural network by using the training loss value based on a back-propagation mechanism; Based on the student neural network after parameter update, updating the parameters of the teacher neural network through an exponential average moving operation; The training loss value calculation operation, the output result acquisition operation, and the neural network parameter update operation are iterated until the preset training conditions are met to obtain a trained fault recognition model.
6. The method according to claim 5, characterized in that Introducing the context prototype-aware learning technology into the semi-supervised learning framework and constructing a training loss function based on a class activation map includes: Thresholding the output result of the teacher neural network to generate a binary mask, and performing mask average pooling on the unlabeled seismic data training set according to the binary mask to obtain an instance prototype; The instance prototypes are stored in a dynamically updateable support library, and a preset clustering algorithm is used to perform clustering processing on the instance prototypes; Constructing a candidate context prototype set using the instance prototype after clustering processing, and calculating the cosine similarity between each candidate context prototype in the candidate context prototype set and the instance prototype; Based on the cosine similarity, a plurality of the candidate context prototypes that are most similar to the instance prototype are selected from the candidate context prototype set, and a context-aware prototype set is constructed using the selected plurality of the candidate context prototypes.
7. The method according to claim 6, characterized in that Also includes: Calculate the positive correlation weight of each context-aware prototype in the context-aware prototype set and the instance prototype based on a parameter-free identity mapping layer and an activation function; Performing weighted calculation on each context-aware prototype in the context-aware prototype set by using the positive correlation weight to obtain a weighted context-aware prototype set; Calculating a distribution gap between the instance prototype and each context-aware prototype in the context-aware prototype set to obtain an offset term representing the distribution gap; Replacing the instance prototype with a result of an operation between the instance prototype and the offset term to obtain a modified instance prototype, and modifying the weighted context-aware prototype set with the modified instance prototype to obtain a modified context-aware prototype set; The class activation map is obtained by using the modified context-aware prototype set and the unlabeled seismic data training set, and the training loss function is constructed by using the class activation map.
8. A fault identification method, characterized in that: The method comprises: The fault data to be identified are processed using the trained fault identification model to obtain a fault identification result, wherein the trained fault identification model is trained using the training method described in any one of claims 1 to 7, and the fault data to be identified include seismic fault images.
9. An electronic device, comprising: one or more processors; a memory for storing one or more computer programs, It is characterized in that the one or more processors execute the one or more computer programs to implement the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Urban area tail gas pollution prediction method
CN110503139A
Lotus phenotype identification method and device based on pseudo tag algorithm and MobileNetV2 network
CN117953281A
Advanced prediction, interpretation and imaging method and system for drilling and blasting tunnel drill jumbo
CN118091739A
Road target identification method and system in severe weather scene, and medium
CN118196735A
Detection of High Incident Reflective Boundaries Using Near-Field Shear Waves
US20160363684A1