Method and device for generating a set of seismic first arrival training
Patent Information
- Application Number
- CN202110603979.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-31
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2041-05-31
AI Technical Summary
目前现有智能拾取技术的初至标签制作基本上依赖人工进行筛选,但在海量数据中人工筛选具有代表性的样本费时费力,且具有一定程度的偏向性和冗余性,人工筛选标签样本的质量和时效性制约了深度神经网络拾取方法的精度和效率
[0008]第四方面,本发明实施例还提供一种计算机可读存储介质,所述计算机可读存储介质存储有执行上述地震初至训练集生成方法的计算机程序。
Smart Images

Figure CN115481710B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oil and gas exploration and development technology, and in particular to a method and apparatus for generating a first-arrival training set of earthquakes. Background Technology
[0002] This section is intended to provide background or context for the embodiments of the invention set forth in the claims. The description herein is not an admission that it is prior art simply because it is included in this section.
[0003] First-arrival travel times of seismic data can be used to construct near-surface structural information, which helps improve the accuracy of static correction, seismic imaging, and reservoir inversion. First-arrival picking is one of the most fundamental and time-consuming tasks in seismic data processing. For the problem of first-arrival picking in complex regions with massive amounts of low signal-to-noise ratio data, conventional methods such as energy ratio methods, correlation methods, and fractal dimension methods are based on single-rule designs, resulting in poor noise resistance, low accuracy of automatic picking results, and reliance on manual editing, which can take several months. Intelligent picking methods, such as first-arrival picking based on fully convolutional neural networks, can directly learn the inherent characteristics of first-arrival waves from seismic waveform data, resulting in high accuracy and strong noise resistance. These methods require samples with known first-arrival times as learning models; the more types of labeled samples, the better the model learning effect. The quantity and quality of labeled samples determine the accuracy and efficiency of first-arrival picking. Currently, the initial label creation of existing intelligent picking technologies basically relies on manual screening. However, manually screening representative samples from massive amounts of data is time-consuming and laborious, and has a certain degree of bias and redundancy. The quality and timeliness of manually screened label samples restrict the accuracy and efficiency of deep neural network picking methods. Summary of the Invention
[0004] This invention provides a method and apparatus for generating an earthquake first arrival training set, which can reduce the bias and randomness of the first arrival training sample selection, obtain a better earthquake first arrival training set, and thus facilitate the improvement of the accuracy and efficiency of the deep neural network first arrival picking method.
[0005] In a first aspect, embodiments of the present invention provide a method for generating an earthquake first arrival training set, the method comprising: acquiring sample pool information and benchmark model information; generating prediction information using the sample pool information and the benchmark model information; calculating information entropy data and classification data of the prediction information; and determining an earthquake first arrival training set based on the information entropy data and the classification data.
[0006] Secondly, embodiments of the present invention also provide an earthquake first arrival training set generation device, the device comprising: an acquisition module for acquiring sample pool information and baseline model information; a generation module for generating prediction information using the sample pool information and the baseline model information; a calculation module for calculating the information entropy data and classification data of the prediction information; and a determination module for determining the earthquake first arrival training set based on the information entropy data and the classification data.
[0007] Thirdly, embodiments of the present invention also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described method for generating the earthquake first arrival training set.
[0008] Fourthly, embodiments of the present invention also provide a computer-readable storage medium storing a computer program for executing the above-described method for generating an earthquake first arrival training set.
[0009] The embodiments of this invention bring the following beneficial effects: This invention provides a method and apparatus for generating an earthquake first-arrival training set. The method includes: acquiring sample pool information and benchmark model information; generating prediction information using the sample pool information and benchmark model information; calculating the information entropy data and classification data of the prediction information; and determining the earthquake first-arrival training set based on the information entropy data and classification data. This invention can automatically filter more representative and valuable sample data from the sample pool information using information entropy data and classification data, thereby obtaining a higher-quality earthquake first-arrival training set, thus improving the efficiency and accuracy of intelligent data acquisition.
[0010] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention are realized and obtained in accordance with the structures particularly pointed out in the description, claims and drawings.
[0011] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0012] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0013] Figure 1This is a flowchart of the earthquake first arrival training set generation method provided in an embodiment of the present invention;
[0014] Figure 2 This is a schematic diagram of the U-GAN used for first-arrival segmentation provided in an embodiment of the present invention;
[0015] Figure 3 This is a schematic diagram of the deep active learning process framework provided in an embodiment of the present invention;
[0016] Figure 4 The validation set loss curves (left) for conventional learning and (right) for active learning are provided for embodiments of the present invention.
[0017] Figure 5 A schematic diagram of F1 scores for conventional learning and active learning provided in an embodiment of the present invention;
[0018] Figure 6 This is a structural block diagram of an earthquake first arrival training set generation device provided in an embodiment of the present invention;
[0019] Figure 7 A structural block diagram of a computer device provided in an embodiment of the present invention;
[0020] Figure 8 This is a structural block diagram of another earthquake first arrival training set generation device provided in an embodiment of the present invention;
[0021] Figure 9 This is a structural block diagram of another earthquake first arrival training set generation device provided in an embodiment of the present invention;
[0022] Figure 10 This is a block diagram of the determination module structure provided in an embodiment of the present invention. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] To further improve the accuracy and efficiency of intelligent first arrival picking technology, this invention provides a method and apparatus for generating an earthquake first arrival training set. This method uses deep active learning technology to automatically select the most representative and valuable earthquake data for first arrival labeling, which greatly enriches the diversity of label samples and reduces the number of labels. It does not require manual intervention, thereby avoiding the bias and randomness of manual sample selection to a certain extent, and achieving optimal picking accuracy and efficiency at the lowest cost.
[0025] To facilitate understanding of this embodiment, a method for generating an earthquake first arrival training set disclosed in this embodiment of the invention will first be described in detail.
[0026] This invention provides a method for generating an earthquake first arrival training set, see [link to relevant documentation]. Figure 1 The flowchart shown is a method for generating an earthquake first arrival training set. The method includes the following steps:
[0027] Step S102: Obtain sample pool information and benchmark model information.
[0028] In this embodiment of the invention, the sample pool information includes multiple sets of initial arrival data, each set of initial arrival data can be used to train a benchmark model so that the benchmark model can recognize the initial arrival information.
[0029] The baseline model can be an active learning model, which can be selected according to actual needs. This embodiment of the invention does not impose specific limitations on this.
[0030] Step S104: Generate prediction information using sample pool information and benchmark model information.
[0031] In this embodiment of the invention, a benchmark model is used to predict samples in the sample pool to obtain prediction information.
[0032] Step S106: Calculate the information entropy data and classification data of the prediction information.
[0033] In this embodiment of the invention, information entropy data is used to describe the degree of uncertainty of the predicted information, and classification data is used to describe the differential distribution of the predicted information. After obtaining the predicted information, its information entropy data and classification data are calculated respectively to obtain the evaluation of the predicted information.
[0034] Step S108: Determine the earthquake first arrival training set based on information entropy data and classification data.
[0035] In this embodiment of the invention, the sample information corresponding to the prediction information is filtered according to the information entropy data and classification data, thereby realizing the selection of the earthquake first arrival training set from the sample pool information.
[0036] This invention provides a method for generating an earthquake first-arrival training set. The method includes: acquiring sample pool information and baseline model information; generating prediction information using the sample pool information and baseline model information; calculating the information entropy data and classification data of the prediction information; and determining the earthquake first-arrival training set based on the information entropy data and classification data. This invention can automatically filter more representative and valuable sample data from the sample pool information using information entropy data and classification data, thereby obtaining a higher-quality earthquake first-arrival training set, thus improving the efficiency and accuracy of intelligent data acquisition.
[0037] In one embodiment, before obtaining the sample pool information and the benchmark model information, the following steps may also be performed:
[0038] Obtain initial model information and labeled sample data; use the labeled sample data to train the initial model information to obtain baseline model information.
[0039] In this embodiment of the invention, the initial model information is the initial form of the baseline model, and the parameters in the initial model are obtained through initialization. The labeled sample data are pre-labeled samples. For example, in order to minimize the workload of manually previewing and selecting data, a few shots can be randomly selected at intervals according to the ground firing sequence as the initial training set, and the first arrivals can be picked and labeled to obtain labeled sample data.
[0040] After obtaining labeled sample data, the initial model information is trained using the labeled sample data to obtain optimized model parameters, and then the training results are used as the baseline model information.
[0041] In one embodiment, before obtaining the sample pool information and the benchmark model information, the following steps may also be performed:
[0042] Obtain pre-trained model information and use it as baseline model information.
[0043] In this embodiment of the invention, pre-trained model information trained on other work areas can be directly used as the baseline model information for the current target data. This reduces the workload caused by sample labeling and training initial model information.
[0044] In one embodiment, the information entropy data for calculating prediction information can be performed by the following steps: generating probability map data based on the prediction information; and calculating information entropy data based on the probability map data.
[0045] In this embodiment of the invention, the prediction information is data output from the baseline model information. The prediction information is processed to obtain probability map data, and information entropy data is calculated after obtaining the probability map data.
[0046] In one embodiment, information entropy data is calculated from probability graph data using the following formula:
[0047] H(p)=-∑p i Log2(p i ), where H(p) represents the information entropy data, p i This represents the probability that a given piece of information in a probability graph belongs to category i.
[0048] In this embodiment of the invention, the probability of category i is determined in the probability graph data based on given information, and then the probability is calculated according to the formula H(p)=-∑p i Log2(p i Calculate information entropy data.
[0049] In one embodiment, the classification data for calculating the prediction information can be performed by the following steps: calculating the difference value of image semantic features between the prediction information and the sample set information to obtain the classification data.
[0050] In this embodiment of the invention, the training set information and the sample pool information have similar sample set information. For the same sample, the image semantic feature difference value is calculated based on the prediction information and the sample set information, thereby obtaining the descriptive information of the difference between the two, that is, obtaining the classification data.
[0051] In one embodiment, determining the earthquake first arrival training set based on information entropy data and classification data can be performed according to the following steps:
[0052] The sample pool information corresponding to the prediction information is sorted according to the information entropy data and classification data to obtain the sorting results; the earthquake first arrival training set is determined according to the sorting results.
[0053] In this embodiment of the invention, the specific sorting criteria can be selected according to the actual needs and the degree of emphasis on information entropy data and classification data, to obtain the sorting result.
[0054] In one embodiment, the sample pool information corresponding to the prediction information is sorted according to the information entropy data and classification data to obtain the sorting result, including:
[0055] The sample pool information corresponding to the prediction information is sorted according to the information entropy data to obtain the first sorting result; the target sample pool information is determined according to the first sorting result; the target sample pool information is sorted according to the classification data and the prediction information to obtain the sorting result.
[0056] In this embodiment of the invention, the sample pool information corresponding to the prediction information can be sorted first based on the information entropy data, or the sample pool information corresponding to the prediction information can be sorted first based on the classification data, or the two can be weighted and then the sample pool information can be sorted. In implementation, the settings can be configured according to actual needs.
[0057] In one embodiment, the method may also perform the following steps:
[0058] A deep neural network is trained based on the earthquake first arrival training set data; the trained deep neural network is then used to generate the first arrival picking results.
[0059] In this embodiment of the invention, after using the method to select a higher quality sample set of data, namely the earthquake first arrival training set data, the deep neural network is trained using the earthquake first arrival training set data, thereby improving the accuracy and efficiency of the deep neural network first arrival picking method.
[0060] The implementation process of this method will be described below with a specific example.
[0061] Based on the characteristics of the initial arrival picking task, this invention designs an active learning model, U-GAN, to automatically select the most valuable data as training samples. The sample selection process considers both uncertainty and variability, achieving optimal model performance with minimal annotation cost. U-GAN includes a generator and a discriminator. Valuable uncertain samples are selected based on the information entropy predicted by the generator, and the discriminator further classifies the uncertain samples before selecting those with high variability.
[0062] See Figure 2 The diagram illustrates the U-GAN process for first-arrival segmentation. The generator can be a fully convolutional neural network with identical input and output dimensions. Taking U-Net as an example, U-Net consists of a symmetrical encoder and decoder, along with skip connections connecting features of the same level across the two paths. The encoder extracts multi-scale features layer by layer through multiple convolutions and pooling operations. The decoder fuses multi-scale features layer by layer through multiple deconvolutions and convolutions to achieve the target output. Skip connections copy and stitch the high-resolution encoder features from each stage onto the decoder features, resulting in more accurate results during decoding and precise classification of each pixel. The generator U-Net structure is as follows: Figure 1 As shown, both the encoder and decoder consist of multiple stages. Each stage of the encoder includes two convolutional layers (3×3 kernels) and one max-pooling layer (stride 2). Each stage of the decoder includes one deconvolutional layer (stride 2) and two convolutional layers (3×3 kernels). In the final stage of the decoder, a convolutional layer (1×1 kernel) is used to transform the high-resolution feature map into a binary classification result. Each convolutional layer is followed by an activation function; the last convolutional layer uses the Sigmoid activation function, and the remaining convolutional layers use the ELU activation function.
[0063] The network structure of the discriminator is as follows: Figure 1As shown, it consists of two parts: a convolutional unit and a fully connected unit. The convolutional unit consists of multiple sets of convolutional layers and max pooling layers with the same structure. The convolutional kernel size is 3×3 and the pooling stride is 2. The fully connected unit consists of multiple fully connected layers. The last fully connected layer outputs the binary classification result. The last fully connected layer uses the Sigmoid activation function. The remaining fully connected layers and convolutional layers are followed by the ELU activation function.
[0064] The error of U-GAN consists of two parts: one is the error between the predicted segmentation result and the actual segmentation result of the generator U-net, and the other is the distribution difference between the training data and the unlabeled data of the discriminator. U-GAN uses cross-entropy as the objective function and uses the Adam optimizer to optimize the objective function. The training process of U-GAN follows the conventional GAN network, fixing one network and updating the parameters of another network, alternating and iterating to minimize the error of the other network.
[0065] During network training, the generator U-net takes seismic data as input and outputs predicted segmentation results. The discriminator receives two inputs: seismic data plus actual segmentation, or seismic data plus predicted segmentation. Its output is a binary classification value of 0 and 1, measuring whether the two inputs belong to the same distribution. During testing, the information entropy of the generator's predicted segmentation probability reflects the uncertainty differences in the test data. The discriminator's output classification value represents the difference in image semantic features between the test data and its segmentation results; different classification values reflect the differences in type between the test data and the training data. The information entropy calculated based on the probability distribution of the generator's output is defined as follows:
[0066] H(p)=-∑p i Log2(p i )
[0067] Where, p i H([0.5,0.5]) represents the probability that a given sample belongs to class i. For a binary classification problem, H([0.5,0.5]) = 1.0 and H([1.0,0.0]) = 0.0, meaning that samples with a predicted probability of 0.5 are samples that the model is likely to mispredict.
[0068] See Figure 3 The diagram illustrates the deep active learning workflow framework. The initial arrival picking process for U-GAN-based deep active learning is as follows:
[0069] (1) Constructing an initial training sample set and benchmark model: A high-precision benchmark model can accelerate the active learning process, so the benchmark model must be initially trained before starting active learning. For the earthquake first arrival picking problem, we can directly use the model trained on other work areas as the benchmark model for the current target data. If there is no pre-trained model, in order to minimize the workload of manually previewing and selecting data, we use a few shots at intervals according to the ground firing sequence as the initial training set, and perform first arrival labeling. The remaining unlabeled data is put into the sample pool.
[0070] (2) Model prediction: Predict all unlabeled samples in the sample pool, calculate the uncertainty (information entropy calculated using the probability distribution map output by the generator) and the difference distribution (classification size output by the discriminator).
[0071] (3) Sample selection and annotation: In each annotation suggestion stage, the k most uncertain and representative images are extracted from all unannotated sample sets. Since uncertainty is a more important indicator, firstly, k images are selected from the set based on their uncertainty level to form a candidate set Sc (subscript c represents candidate). Then, based on the difference information, the k most representative images are selected from the candidate set Sc to form an annotation set Sa (subscript a represents annotated). Sa is initially annotated and added to the training set.
[0072] (4) Model training: Train the active learning model using the current training set.
[0073] Repeat steps (2)-(4) above until the model performance no longer increases significantly or reaches the expected standard. The active learning process framework is as follows: Figure 3 As shown. Due to the differences in data quality across different work areas, K and k can be determined based on the actual situation, taking into account both the time cost of active learning iterations and the cost of sample labeling.
[0074] This invention introduces the concept of deep active learning on the basis of initial arrival picking in deep neural networks. The active learning method proactively selects or generates the most valuable samples through appropriate strategies, performs expert annotation, and supplements them to the existing training set, enabling the model to achieve strong generalization ability with low annotation costs. The core of active learning lies in the design of the sample sampling strategy. The most commonly used criteria for formulating sample sampling strategies in deep active learning are the uncertainty criterion and the difference criterion. Difference sampling selects the most representative samples, aiming to cover information about the overall data distribution, with each sample providing non-repetitive and non-redundant information, i.e., samples have a certain degree of difference. If each iteration queries a batch of samples, a difference metric is needed to ensure sample diversity and avoid data redundancy. Uncertainty sampling selects the samples that the model is most uncertain about; generally, the most uncertain samples are also those that the model is likely to mispredict. However, some samples with the highest uncertainty often have high similarity, so the uncertainty and representativeness of the samples must be considered comprehensively when selecting examples.
[0075] The data for a certain mountainous area contains 131 manually collected first-arrival labels. To ensure diversity in first-arrival types, 56 shots were manually selected as training data based on factors such as firing point location, signal-to-noise ratio, and first-arrival wavelength. This number 56 was not rigorously verified, as there may be some redundancy among different types of first-arrivals. A validation set of 75 shots was then selected every 10 shots based on firing point location, ensuring the validation set covers the entire work area as much as possible. Four shots were randomly selected from the 56 labeled samples as the initial training set. During the active learning iteration process, based on uncertainty and difference rules, 4, 8, and 16 shots were sequentially selected from the training pool and added to the training set. Simultaneously, samples with the same data were selected as the training set for regular learning (U-net model, random selection of training data).
[0076] The number of training samples in the five rounds of active learning and conventional learning were 4, 8, 16, 32 and 56 respectively, and the number of validation / test samples was 75. There was no sample selection in the initial training of the first round and the training of all samples in the last round. The two learning modes shared the same model and results.
[0077] Figure 4 The figures show the loss curves of the two learning modes on the validation set. Under the condition of the same number of samples, the loss curve of active learning decreases faster and the loss at the optimal point of the model is lower. This indicates that the samples selected by active learning can accelerate the convergence of the model and improve its performance.
[0078] Figure 5The F1 score curves (F1 score = 2 × picking accuracy / (picking accuracy + picking rate), F1 score comprehensively reflects the effect of first-arrival segmentation) are shown for two learning modes. In active learning, the model performance with 32 training samples is comparable to that with all 56 training samples, indicating that there is indeed a large redundancy in manually selected samples, or that the near offset data information in difficult samples contains data information from simple samples. In this example, active learning only requires half the amount of data to achieve the same optimal model performance, which halves the annotation cost and halves the work cycle. For the first-arrival picking problem of seismic data, in the early stage of learning, only a few samples need to be added for the model performance to improve rapidly, while in the later stage of learning, the number of samples increases exponentially, but the improvement in model performance is very slow. Therefore, when doing production projects, it is possible to balance picking accuracy and project cycle, and good picking results can be achieved by selecting a small amount of data for annotation based on active learning.
[0079] This invention automatically selects the most representative and valuable seismic data for initial arrival labeling without manual intervention, greatly enriching the diversity of labels and reducing the number of labels. Compared with conventional learning, active learning only requires half the number of labels to achieve the same or even higher model performance, greatly improving the efficiency and accuracy of intelligent picking.
[0080] This invention also provides an earthquake first-arrival training set generation device, as described in the following embodiments. Since the principle behind this device is similar to that of the earthquake first-arrival training set generation method, its implementation can be found in the implementation of the earthquake first-arrival training set generation method; repeated details will not be elaborated further. See also... Figure 6 The diagram shown illustrates the structural block diagram of the earthquake first arrival training set generation device. The device includes:
[0081] The acquisition module 61 is used to acquire sample pool information and benchmark model information; the generation module 62 is used to generate prediction information using sample pool information and benchmark model information; the calculation module 63 is used to calculate the information entropy data and classification data of the prediction information; and the determination module 64 is used to determine the earthquake first arrival training set based on the information entropy data and classification data.
[0082] In one embodiment, see Figure 8 The diagram shows another structural block diagram of an earthquake first arrival training set generation device. This device also includes a preprocessing module 65, which is used to: acquire initial model information and labeled sample data; and use the labeled sample data to train the initial model information to obtain baseline model information.
[0083] In one embodiment, the preprocessing module is further configured to: obtain pre-trained model information and use the pre-trained model information as baseline model information.
[0084] In one embodiment, the calculation module is configured to: generate probability map data based on prediction information; and calculate information entropy data based on the probability map data.
[0085] In one embodiment, the calculation module is used to calculate information entropy data based on the probability graph data using the following formula:
[0086] H(p)=-∑p i Log2(p i )
[0087] Where H(p) represents the information entropy data, p i This represents the probability that a given piece of information in a probability graph belongs to category i.
[0088] In one embodiment, the calculation module is used to: calculate the difference value of image semantic features between the prediction information and the sample set information to obtain classification data.
[0089] In one embodiment, see Figure 10 The block diagram of the determination module shown includes: a sorting unit 641, used to sort the sample pool information corresponding to the prediction information according to the information entropy data and classification data, and obtain the sorting result; and a training set unit 642, used to determine the earthquake first arrival training set according to the sorting result.
[0090] In one embodiment, the sorting unit is specifically configured to: sort the sample pool information corresponding to the prediction information according to the information entropy data to obtain a first sorting result; determine the target sample pool information according to the first sorting result; and sort the target sample pool information according to the classification data and the prediction information to obtain a sorting result.
[0091] In one embodiment, see Figure 9 The diagram shows another structural block diagram of an earthquake first arrival training set generation device. This device also includes a picking module 66, which is used to: train a deep neural network based on the earthquake first arrival training set data; and generate first arrival picking results using the trained deep neural network.
[0092] Based on the same inventive concept, this invention also provides an embodiment of an electronic device for implementing all or part of the above-described method for generating earthquake first arrival training sets. This electronic device specifically includes the following:
[0093] The device comprises a processor, memory, a communications interface, and a bus; wherein the processor, memory, and communications interface communicate with each other via the bus; the communications interface is used to realize information transmission between related devices; the electronic device can be a desktop computer, tablet computer, or mobile terminal, etc., and this embodiment is not limited to these. In this embodiment, the electronic device can be implemented with reference to the embodiments for implementing the above-mentioned earthquake first arrival training set generation method and the embodiments for implementing the above-mentioned earthquake first arrival training set generation device, the contents of which are incorporated herein by reference, and repeated details will not be described again.
[0094] Figure 7 This is a schematic diagram of the system composition structure of an electronic device provided in an embodiment of the present invention. Figure 7 As shown, the electronic device 70 may include a processor 701 and a memory 702; the memory 702 is coupled to the processor 701. It is worth noting that... Figure 7 This is an example; other types of structures can also be used to supplement or replace this structure to achieve telecommunications functions or other functions.
[0095] In one embodiment, the functionality implemented by the earthquake first arrival training set generation method can be integrated into the processor 701. The processor 701 can be configured to perform the following controls:
[0096] Acquire sample pool information and baseline model information; generate prediction information using sample pool information and baseline model information; calculate information entropy data and classification data of the prediction information; determine the earthquake first arrival training set based on information entropy data and classification data.
[0097] As can be seen from the above, the electronic device provided in the embodiments of the present invention can use information entropy data and classification data to automatically filter more representative and valuable sample data in the sample pool information, thereby obtaining a higher quality earthquake first arrival training set, so as to improve the efficiency of intelligent picking.
[0098] In another embodiment, the earthquake first arrival training set generation device can be configured separately from the processor 701. For example, the earthquake first arrival training set generation device can be configured as a chip connected to the processor 701, and the function of the earthquake first arrival training set generation method can be realized through the control of the processor.
[0099] like Figure 7 As shown, the electronic device 70 may further include: a communication module 703, an input unit 704, an audio processing unit 705, a display 706, and a power supply 707. It is worth noting that the electronic device 70 does not necessarily need to include these components. Figure 7All components shown; in addition, the electronic device 70 may also include Figure 7 For components not shown, please refer to existing technology.
[0100] like Figure 7 As shown, processor 701, sometimes also referred to as controller or operation control, may include a microprocessor or other processor device and / or logic device, which receives input and controls the operation of various components of electronic device 70.
[0101] The memory 702 may be, for example, one or more of a cache, flash memory, hard drive, removable media, volatile memory, non-volatile memory, or other suitable device. It may store the aforementioned failure-related information, and also store a program for executing that information. The processor 701 may execute the program stored in the memory 702 to perform information storage or processing, etc.
[0102] Input unit 704 provides input to processor 701. Input unit 704 may be, for example, a keypad or touch input device. Power supply 707 provides power to electronic device 70. Display 706 displays images and text. Display may be, for example, an LCD display, but is not limited to this.
[0103] The memory 702 can be a solid-state memory, such as a read-only memory (ROM), random access memory (RAM), a SIM card, etc. It can also be a memory that retains information even when power is off, can be selectively erased, and contains more data; examples of this type of memory are sometimes referred to as EPROMs. The memory 702 can also be some other type of device. The memory 702 includes a buffer memory 7021 (sometimes called a buffer). The memory 702 may include an application / function storage unit 7022 for storing application programs and function programs or processes for executing the operation of the electronic device 70 via the processor 701.
[0104] The memory 702 may also include a data storage unit 7023 for storing data, such as contacts, digital data, pictures, sounds, and / or any other data used by the electronic device. The driver storage unit 7024 of the memory 702 may include various drivers for the electronic device's communication functions and / or for performing other functions of the electronic device (such as messaging applications, address book applications, etc.).
[0105] The communication module 703 is a transmitter / receiver that transmits and receives signals via the antenna 708. The communication module (transmitter / receiver) 703 is coupled to the processor 701 to provide input signals and receive output signals, which is the same as in a conventional mobile communication terminal.
[0106] Based on different communication technologies, multiple communication modules 703 can be configured in the same electronic device, such as cellular network modules, Bluetooth modules, and / or wireless LAN modules. The communication module (transmitter / receiver) 703 is also coupled to a speaker 709 and a microphone 710 via an audio processing unit 705 to provide audio output via the speaker 709 and receive audio input from the microphone 710, thereby realizing typical telecommunications functions. The audio processing unit 705 may include any suitable buffer, decoder, amplifier, etc. Additionally, the audio processing unit 705 is also coupled to a processor 701, enabling on-device recording via the microphone 710 and on-device playback of stored audio via the speaker 709.
[0107] In embodiments of the present invention, a computer-readable storage medium is also provided for implementing all steps of the earthquake first arrival training set generation method in the above embodiments. The computer-readable storage medium stores a computer program, which, when executed by a processor, implements all steps of the XX method in the above embodiments. For example, when the processor executes the computer program, it implements the following steps:
[0108] Acquire sample pool information and baseline model information; generate prediction information using sample pool information and baseline model information; calculate information entropy data and classification data of the prediction information; determine the earthquake first arrival training set based on information entropy data and classification data.
[0109] As can be seen from the above, the computer-readable storage medium provided in the embodiments of the present invention can use information entropy data and classification data to automatically filter more representative and valuable sample data in the sample pool information, thereby obtaining a higher quality earthquake first arrival training set, so as to improve the efficiency of intelligent picking.
[0110] While this invention provides the method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-inventive labor. The order of steps listed in the embodiments is merely one possible execution order among many and does not represent the only possible execution order. In actual device or client product execution, the methods shown in the embodiments or drawings can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment).
[0111] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0112] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0113] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0114] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0115] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0116] In this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, without necessarily requiring or implying any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "upper," "lower," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the invention.
[0117] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to the internal connection between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0118] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention is not limited to any single aspect, nor to any single embodiment, nor to any combination and / or substitution of these aspects and / or embodiments. Each aspect and / or embodiment of the present invention can be used alone, or in combination with one or more other aspects and / or other embodiments.
[0119] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A method for generating an earthquake first arrival training set, characterized in that, include: Acquire sample pool information and benchmark model information; the sample pool information includes multiple sets of initial arrival data, each set of initial arrival data is used to train the benchmark model so that the benchmark model can recognize the initial arrival information; Prediction information is generated using the sample pool information and the baseline model information; Calculate the information entropy data and classification data of the predicted information; the information entropy data is used to describe the degree of uncertainty of the predicted information, and the classification data is used to describe the differential distribution of the predicted information; Based on the information entropy data and the classification data, the sample information corresponding to the prediction information is filtered, and the earthquake first arrival training set is selected from the sample pool information; The process involves filtering sample information corresponding to the prediction information based on the information entropy data and the classification data, and selecting an earthquake first arrival training set from the sample pool information. This includes: sorting the sample pool information corresponding to the prediction information according to the emphasis of the information entropy data and the classification data based on actual needs, and obtaining a sorting result; and determining the earthquake first arrival training set based on the sorting result. Based on actual needs, the sample pool information corresponding to the prediction information is sorted according to the emphasis of the information entropy data and the classification data, resulting in a sorting result, including: First, sort the sample pool information corresponding to the prediction information based on the information entropy data; second, sort the sample pool information corresponding to the prediction information based on the classification data; or third, sort the sample pool information after weighting the information entropy data and the classification data. Before obtaining the sample pool information and the baseline model information, the process also includes: obtaining initial model information and using the initial training set extracted at intervals according to the ground firing sequence, and picking up the initial arrival to label the labeled sample data; using the labeled sample data to train the initial model information to obtain the baseline model information; Before obtaining sample pool information and benchmark model information, the process also includes: obtaining pre-trained model information and using the pre-trained model information as benchmark model information.
2. The method according to claim 1, characterized in that, The information entropy data for calculating the predicted information includes: Generate probability map data based on the prediction information; Calculate the information entropy data based on the probability map data.
3. The method according to claim 2, characterized in that, This includes calculating information entropy data based on the probability map data using the following formula: in, Represents information entropy data, This indicates that given information in probabilistic graphical data belongs to a category. The probability of.
4. The method according to claim 1, characterized in that, The classification data used to calculate the predicted information includes: The difference in image semantic features between the predicted information and the sample set information is calculated to obtain classification data.
5. The method according to claim 1, characterized in that, The sample pool information corresponding to the predicted information is sorted based on the information entropy data and the classification data to obtain the sorting result, including: The sample pool information corresponding to the predicted information is sorted according to the information entropy data to obtain a first sorting result; The target sample pool information is determined based on the first sorting result; Based on the classification data and the prediction information, the target sample pool information is sorted to obtain the sorting result.
6. The method according to any one of claims 1-5, characterized in that, Also includes: A deep neural network is trained based on the earthquake first arrival training set data; The initial arrival picking results are generated using a trained deep neural network.
7. A device for generating an earthquake first arrival training set, characterized in that, include: The acquisition module is used to acquire sample pool information and benchmark model information. The sample pool information includes multiple sets of initial arrival data. Each set of initial arrival data is used to train the benchmark model so that the benchmark model can recognize the initial arrival information. The generation module is used to generate prediction information using the sample pool information and the benchmark model information; The calculation module is used to calculate the information entropy data and classification data of the predicted information; the information entropy data is used to describe the degree of uncertainty of the predicted information, and the classification data is used to describe the differential distribution of the predicted information. The determination module is used to filter the sample information corresponding to the prediction information based on the information entropy data and the classification data, and to select the earthquake first arrival training set from the sample pool information. The determining module includes: The sorting unit is used to sort the sample pool information corresponding to the prediction information according to the emphasis of the information entropy data and the classification data according to actual needs, and obtain the sorting result; Training set unit, used to determine the earthquake first arrival training set based on the sorting results; The sorting unit is specifically used to sort the sample pool information corresponding to the prediction information first based on the information entropy data, sort the sample pool information corresponding to the prediction information first based on the classification data, or sort the sample pool information after weighting the information entropy data and the classification data. The preprocessing module is used to: acquire initial model information and extract initial training set at intervals according to the ground firing sequence, and pick up the labeled sample data obtained by labeling the initial arrival; use the labeled sample data to train the initial model information to obtain baseline model information; The preprocessing module is also used for: Obtain the pre-trained model information and use it as the baseline model information.
8. The apparatus according to claim 7, characterized in that, The computing module is used for: Generate probability map data based on the prediction information; Calculate the information entropy data based on the probability map data.
9. The apparatus according to claim 8, characterized in that, The calculation module is used to calculate information entropy data based on the probability map data using the following formula: in, Represents information entropy data, This indicates that given information in probabilistic graphical data belongs to a category. The probability of.
10. The apparatus according to claim 7, characterized in that, The computing module is used for: The difference in image semantic features between the predicted information and the sample set information is calculated to obtain classification data.
11. The apparatus according to claim 7, characterized in that, The sorting unit is specifically used for: The sample pool information corresponding to the predicted information is sorted according to the information entropy data to obtain a first sorting result; The target sample pool information is determined based on the first sorting result; Based on the classification data and the prediction information, the target sample pool information is sorted to obtain the sorting result.
12. The apparatus according to any one of claims 7-11, characterized in that, It also includes a pickup module, used for: A deep neural network is trained based on the earthquake first arrival training set data; The initial arrival picking results are generated using a trained deep neural network.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the earthquake first arrival training set generation method according to any one of claims 1 to 6.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the earthquake first arrival training set generation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Segmentation model training method, OCT image segmentation method, device, equipment and medium
CN109829894A
Object-oriented classification method based on active learning
CN111259961A
Seismic data first arrival pickup method based on Unet++ convolutional neural network
CN111626355A