Method for generating training data, method for generating a model, and apparatus for generating training data

JP2026153012APending Publication Date: 2026-09-30PREFERRED NETWORKS INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2021170348
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2021-06-07
Filing Date
2021-10-18
Publication Date
2026-09-30

Smart Images

  • Figure 2026153012000001_ABST
    Figure 2026153012000001_ABST
Patent Text Reader

Abstract

To generate training data related to structure. [Solution] The learning data generation method according to the embodiment is a learning data generation method that is executed using at least one processor, and comprises: generating a model of a structure based on a plurality of features relating to the structure; generating simulated data that simulates observed values ​​relating to the structure by wave propagation simulation on the model of the structure; and generating learning data by associating the generated model of the structure with the simulated data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Embodiments of the present invention relate to a method for generating training data, a method for generating a model, and a device for generating training data. [Background technology]

[0002] Inferring the spatial distribution of physical properties in the subsurface from observational data such as seismic waveforms is equivalent to seismic inversion. Solving the seismic inversion problem requires numerous seismic wave propagation simulations, which are highly specialized in their application. In recent years, researchers have begun implementing deep learning methods to directly estimate physical properties such as subsurface velocity from seismic record data. For example, by inputting observational data acquired through seismic exploration into a trained deep neural network (DNN), the spatial distribution of physical properties in the subsurface can be inferred from the observational data. This approach reduces the time required to solve the seismic inversion problem.

[0003] However, when using models such as DNNs for subsurface exploration using seismic waveforms, for example, there is a problem in that it is difficult to obtain a large amount of data to use for model training, due to reasons such as the inability to grasp the actual subsurface structure. [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Ke Wang, Laura Bandura, Dimitri Bevc, Shuxing Cheng, Jim DiSiena, Adam Halpert, Konstantin Osypov, Bruce Power, Ellen Xu, Chevron Energy Technology Company End-to-End Deep Neural Network for Seismic Inversion SEG International Exposition and 89th Annual Meeting 10.1190 / segam2019-3216464.1 Page 4982-4986 https: / / library.seg.org / doi / 10.1190 / segam2019-3216464.1 [Overview of the Initiative] [Problems that the invention aims to solve]

[0005] The problem that the invention aims to solve is generating training data related to structure. [Means for solving the problem]

[0006] The training data generation method according to the embodiment is a training data generation method that is performed using at least one processor, and comprises: generating a model of a structure based on a plurality of features relating to the structure; generating simulated data that simulates observed values ​​relating to the structure by performing a wave propagation simulation on the model of the structure; and generating training data by associating the generated model of the structure with the simulated data. [Brief explanation of the drawing]

[0007] [Figure 1] Figure 1 is a block diagram showing an example of the hardware configuration of a learning system having a learning data generation device according to the embodiment. [Figure 2] Figure 2 shows an example of a functional block in a processor according to an embodiment. [Figure 3] FIG. 3 is a diagram illustrating an example of a generation region before generation of a subsurface structure model according to the embodiment. [Figure 4] FIG. 4 is a diagram illustrating an example of a generation region having a plurality of strata deposited therein according to the embodiment. [Figure 5] FIG. 5 is a diagram illustrating an example of a generation region having a plurality of folded strata according to the embodiment. [Figure 6] FIG. 6 is a diagram illustrating an example of a generation region having a fault generated in a plurality of strata according to the embodiment. [Figure 7] FIG. 7 is a diagram illustrating an example of a generation region obtained by eroding a stratum shallower than an unconformity surface according to the embodiment. [Figure 8] FIG. 8 is a diagram illustrating an example of a generation region obtained by re-depositing a plurality of strata in an eroded region according to the embodiment. [Figure 9] FIG. 9 is a diagram illustrating an example of a generation region obtained by folding a plurality of strata and then generating a plurality of faults according to the embodiment. [Figure 10] FIG. 10 is a diagram illustrating an example of a generation region having folding and faults generated in a plurality of strata according to the embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of a generation region into which rock salt has intruded according to the embodiment. [Figure 12] FIG. 12 is a diagram illustrating an example of an image of common shot gather data (shot image) according to the embodiment. [Figure 13] FIG. 13 is a flowchart illustrating an example of a procedure of learning data generation processing according to the embodiment. [Figure 14] FIG. 14 is a diagram illustrating an example of a P-wave velocity (Vp) model generated by a parameter generation system according to the embodiment. [Figure 15] FIG. 15 is a diagram illustrating an example of functional blocks in a processor mounted on a learning apparatus according to the embodiment. [Figure 16] FIG. 16 is a diagram illustrating an example of estimation of P-wave velocity (Vp) in a Marmousi2 geological structure model according to the embodiment. [Figure 17] Figure 17 shows an example of P-wave velocity (Vp) estimation in the 1994 Amoco static correction test dataset, relating to an embodiment. [Figure 18] Figure 18 shows an example of an application example of the embodiment, illustrating the overview of the generation of training data, the generation of a subsurface structure estimation model using the training data, and the subsurface structure estimation process using the generated subsurface structure estimation model. [Figure 19] Figure 19 is a flowchart illustrating an example of the application of the embodiment, showing the steps for generating training data, generating a subsurface structure estimation model using the training data, and the model generation estimation process. [Modes for carrying out the invention]

[0008] Embodiments relating to the training data generation method, model generation method, and training data generation apparatus will be described in detail below with reference to the drawings. The training data generation method is performed, for example, using at least one processor.

[0009] (Embodiment) Figure 1 is a block diagram showing an example of the hardware configuration of a learning system 1 having a learning data generation device 3 according to this embodiment. As shown in Figure 1, the learning system 1 includes a learning data generation device 3, a learning device 7 connected to the learning data generation device 3 via a communication network 5, an external device 9A connected to the learning data generation device 3 via the communication network 5, and an external device 9B connected via a device interface 39. The learning system 1 generates multiple training data sets using the learning data generation device 3. The learning system 1 uses the generated multiple training data sets to train the deep neural network to be trained and generate a trained model.

[0010] A trained model is a model that, for example, outputs sound waves, electromagnetic waves, or radiation to an object being observed, and estimates the structure of that object based on the reflected waves that propagate within the object. This structure is, for example, the internal structure of the object being observed. Objects being observed include underground structures, artificial structures such as pillars and bridges, clouds, and living organisms. Trained models can be applied to, for example, non-destructive testing, acoustic diagnostics of structures, echo detection, submarine sonar, and remote sensing. To make the explanation more specific, the object being observed will be described as an underground structure. In this case, the trained model will take seismic waves (elastic waves), electromagnetic waves, or radiation as input and output the underground structure of the object being observed. For example, when seismic waves are used as input, the trained model is used for seismic exploration. Also, for example, when electromagnetic waves (electromagnetic fields) are used as input, the trained model is used for electromagnetic exploration. To make it even more specific, the trained model will be described as being used for seismic exploration. In this case, multiple training data correspond to reflected waves that propagate within the object being observed due to seismic waves radiated to the underground object being observed.

[0011] The learning data generation device 3 includes a computer 30 and an external device 9B connected to the computer 30 via a device interface 39. The learning device 7 may also be connected to the computer 30 via the device interface 39. The computer 30, as an example, includes a processor 31, a main memory 33, an auxiliary memory 35, a network interface 37, and a device interface 39. The learning data generation device 3 may be implemented as a computer 30 in which the processor 31, main memory 33, auxiliary memory 35, network interface 37, and device interface 39 are connected via a bus 41. The computer 30 may also be mounted on the learning device 7.

[0012] The computer 30 shown in Figure 1 has one of each component, but it may have multiple identical components. Also, although Figure 1 shows one computer 30, the software may be installed on multiple computers, and each of these computers may execute the same or different parts of the software's processing. In this case, it may be a distributed computing configuration in which each computer communicates via a network interface 37 or the like to execute processing. In other words, the learning data generation device 3 in this embodiment may be configured as a system that realizes the various functions described later by having one or more computers execute instructions stored in one or more storage devices. Furthermore, information transmitted from a terminal may be processed by one or more computers located on the cloud, and the processing results may be transmitted to a terminal such as a display device (display unit) corresponding to an external device 9B. The display device can be realized, for example, by various displays.

[0013] The various calculations performed by the learning data generation device 3 in this embodiment may be executed in parallel using one or more processors, or using multiple computers connected via a network. Alternatively, the various calculations may be distributed to multiple processing cores within a processor and executed in parallel. Furthermore, some or all of the processing and means of this disclosure may be executed by at least one of a processor and a storage device located on a cloud that can communicate with computer 30 via a network. Thus, the various methods described later in this embodiment may take the form of parallel computing using one or more computers.

[0014] The processor 31 may be an electronic circuit (processing circuit, processing circuitry, CPU (Central Processing Unit), GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), or ASIC (Application Specific Integrated Circuit), etc.) including the control unit and arithmetic unit of the computer 30. Alternatively, the processor 31 may be a semiconductor device including a dedicated processing circuit. The processor 31 is not limited to an electronic circuit using electronic logic elements, but may also be realized by an optical circuit using optical logic elements. Furthermore, the processor 31 may include arithmetic functions based on quantum computing.

[0015] The processor 31 performs calculations based on data and software (programs) input from various devices within the computer 30, and can output calculation results and control signals to these devices. The processor 31 may also control the various components of the computer 30 by executing the computer 30's OS (Operating System) or applications.

[0016] The learning data generation device 3 in this embodiment may be implemented by one or more processors 31. Here, the processor 71 may refer to one or more electronic circuits arranged on one chip, or one or more electronic circuits arranged on two or more chips or two or more devices. When multiple electronic circuits are used, each electronic circuit may communicate by wire or wireless.

[0017] The main memory 33 is a storage device that stores instructions executed by the processor 31 and various data, and the information stored in the main memory 33 is read by the processor 31. The auxiliary storage device 35 is a storage device other than the main memory 33. These storage devices refer to any electronic component capable of storing electronic information, and may be semiconductor memory. The semiconductor memory may be either volatile memory or non-volatile memory. The storage device for storing various data used in the learning data generation device 3 in this embodiment may be implemented by the main memory 33 or the auxiliary storage device 35, or by the built-in memory of the processor 31. For example, the storage unit in this embodiment may be implemented by the main memory 33 or the auxiliary storage device 35.

[0018] Multiple processors may be connected to one memory device, or a single processor 31 may be connected to it. Multiple memory devices may be connected to one processor. In this embodiment, if the learning data generation device 3 consists of at least one memory device and multiple processors connected to this at least one memory device, it may include a configuration in which at least one of the multiple processors is connected to at least one memory device. This configuration may also be realized by memory devices and processors 31 included in multiple computers. Furthermore, it may include a configuration in which the memory device is integrated with the processor 31 (for example, a cache memory including an L1 cache and an L2 cache).

[0019] The network interface 37 is an interface for connecting to the communication network 5 wirelessly or via a wired connection. The network interface 37 can be any appropriate interface, such as one conforming to existing communication standards. Information may be exchanged between the learning device 7 and the external device 9A connected via the communication network 5 through the network interface 37. The communication network 5 may be a WAN (Wide Area Network), LAN (Local Area Network), PAN (Personal Area Network), or a combination thereof, as long as information is exchanged between the computer 30 and the external device 9A. An example of a WAN is the Internet; an example of a LAN is IEEE 802.11 or Ethernet®; and an example of a PAN is Bluetooth® or NFC (Near Field Communication).

[0020] The device interface 39 is an interface such as USB (Universal Serial Bus) that directly connects to an output device such as a display device, an input device (input section), and an external device 9B. The output device may also have a speaker or the like that outputs sound.

[0021] External device 9A is a device connected to computer 30 via a network. External device 9B is a device directly connected to computer 30.

[0022] External device 9A or external device 9B may, for example, be an input device. The input device may be, for example, a camera, microphone, motion capture device, various sensors, keyboard, mouse, or touch panel, and will provide the acquired information to the computer 30. Alternatively, external device 9A or external device 9B may be a personal computer, tablet terminal, or smartphone, or other device equipped with an input unit, memory, and processor.

[0023] Furthermore, external device 9A or external device 9B may, for example, be an output device (output unit). The output device may be a display device (display unit) such as an LCD (Liquid Crystal Display), CRT (Cathode Ray Tube), PDP (Plasma Display Panel), or organic EL (Electro Luminescence) panel, or it may be a speaker that outputs sound, etc. Also, external device 9A or external device 9B may be a device such as a personal computer, tablet terminal, or smartphone that has an output unit, memory, and processor.

[0024] Furthermore, external device 9A or external device 9B may be a storage device (memory). For example, external device 9A may be network storage, and external device 9B may be storage such as an HDD.

[0025] Furthermore, the external device 9A or external device 9B may be a device that has some of the functions of the components of the learning data generation device 3 in this embodiment. In other words, the computer 30 may transmit or receive some or all of the processing results of the external device 9A or external device 9B.

[0026] Figure 2 shows an example of a functional block in the processor 31. The processor 31 has the following functions that it implements: a setting unit 311, a determination unit 313, a model generation unit 315, a simulated data generation unit 317, and a learning data generation unit 319. The functions implemented by the setting unit 311, the determination unit 313, the model generation unit 315, the simulated data generation unit 317, and the learning data generation unit 319 are each stored as programs in, for example, the main memory 33 or the auxiliary memory 35. The processor 31 reads and executes the programs stored in the main memory 33 or the auxiliary memory 35 to implement the functions related to the setting unit 311, the determination unit 313, the model generation unit 315, the simulated data generation unit 317, and the learning data generation unit 319.

[0027] The setting unit 311 sets a range of values ​​for each of the multiple feature quantities related to the subsurface structure (hereinafter referred to as the feature range). For example, feature quantities are geological parameters. Geological structure elements described using geological parameters include deposition, folding, faulting, erosion, redeposition, refolding, refaulting, salt intrusion, the way strata are bent laterally, the way unconformities are incorporated, the size of the structure (width, depth) related to the modeling of the subsurface structure, the insertion of slow-moving layers into the surface, or the distribution of layer thickness. Geological parameters include, for example, the thickness of the deposited strata, the amplitude and wavelength of the folds of the strata (or the number of vibrations of the strata per unit length), the angle of the fault relative to the horizontal, the length of the fault, the curvature of the fault (the degree of curvature of the fault in the lateral (horizontal) direction), the depth-dependent amplification of fault displacement, P-wave velocity, the ratio of P-wave velocity to S-wave velocity, the distribution location of salt domes, the P-wave velocity of salt domes, the amount of erosion of the strata, the depth of the unconformity, and the depth dependence of rock velocity. Furthermore, the setting unit 311 can set things that do not occur simultaneously in a geological context. For example, the setting unit 311 can set whether it is a normal fault or a reverse fault in the fault characteristics.

[0028] Specifically, the setting unit 311 sets multiple feature ranges for multiple feature quantities based on default settings stored in the main memory 33 or auxiliary memory 35. The default settings are, for example, combinations of multiple feature ranges that cover all geological layer patterns. Note that the default settings are not limited to one, and multiple settings may be set according to multiple regions. In this case, the default settings are set in advance according to geological information corresponding to the region that includes the area to be investigated for subsurface structure (hereinafter referred to as the investigation area). Geological information refers to various data acquired in the past in the region, and includes, for example, logging data from wells drilled in the region, observational data on subsurface structure in the region, and the spatial distribution of feature quantities inferred in the region, at least one of these. The setting unit 311 also sets, as a default setting, the region (hereinafter referred to as the generation region) in which a subsurface structure model (hereinafter referred to as the subsurface structure model) is generated by the model generation unit 315. The generation region is, for example, a region schematically showing the horizontal length and depth.

[0029] The setting unit 311 displays the feature range set by the default settings on the display device. Specifically, the display device displays multiple ranges corresponding to multiple feature quantities and two indicators for changing the upper and lower limits of each of the multiple ranges. At this time, the setting unit 311 may change the feature range as appropriate based on user instructions via the input device. In addition, the setting unit 311 may display radio buttons to allow the user to decide between a normal fault and a reverse fault. At this time, the setting unit 311 sets between a normal fault and a reverse fault based on user instructions via the radio buttons.

[0030] Furthermore, the setting unit 311 may display on the display device the feature range set by the default settings and a model of the subsurface structure (hereinafter referred to as the verification model) generated by the model generation unit 315 using, for example, representative values ​​of the feature range. Specifically, the display device displays a plurality of ranges corresponding to a plurality of feature quantities, two indicators for each of the plurality of ranges, and the verification model. The setting unit 311 displays the verification model on the display device by changing the hue to a predetermined color that changes from blue to yellow along the depth direction, for example, according to the magnitude of the P-wave velocity of the geological layer in the verification model. When the feature range is adjusted by the user moving the two indicators, the setting unit 311 displays the verification model generated using the representative values ​​that have changed in accordance with the adjustment, along with the changed feature range, on the display device.

[0031] The determination unit 313 determines model generation parameters, which are information necessary for generating a subsurface structure model, based on the feature range and random numbers for multiple features. As model generation parameters, for example, the determination unit 313 determines the values ​​within the feature range based on the feature range and random numbers for multiple features. Specifically, once the feature range is set by the setting unit 311, the determination unit 313 determines an identifier (hereinafter referred to as the subsurface model ID) for the subsurface structure model (hereinafter referred to as the subsurface structure model) generated using the model generation parameters within the feature range. The determination unit 313 generates random numbers using the subsurface model ID as a seed for random numbers. The determination unit 313 determines model generation parameters within the set feature range based on the set feature range and the generated random numbers for multiple features. Specifically, the determination unit 313 sets a probability distribution within the feature range and determines the model generation parameters within the feature range by using random numbers on the set probability distribution. The probability distribution is, for example, a uniform distribution, but is not limited to this, and other probability distributions such as the Poisson distribution may be used. The probability distribution may be set, for example, by the setting unit 311.

[0032] The model generation unit 315 generates a structural model based on multiple structural features. Specifically, the model generation unit 315 generates a subsurface structure model using model generation parameters determined by the determination unit 313. The model generation unit 315 also generates a verification model of the subsurface structure using representative values ​​within a range specified by two indicators. Figures 3 to 11 show an example of the process of generating a subsurface structure model generated by the model generation unit 315. The generation region 10 in Figures 3 to 11 indicates a region where a two-dimensional subsurface structure model is generated, but the model generation unit 315 may also generate a three-dimensional subsurface structure model. The model generation unit 315 may also generate multiple structural models using different features.

[0033] Figure 3 shows an example of the generated region 10 before the generation of the subsurface structure model. As shown in Figure 3, the model generation unit 315 generates a generated region 10 filled with zeros as features. The legend 11 shown in Figure 3 indicates the P-wave velocity. The upper end of the generated region 10 shown in Figure 3 represents the ground surface.

[0034] (deposition) The model generation unit 315 deposits multiple strata in the generation region 10 using the determined model generation parameters. For example, the model generation unit 315 uses the thickness of each of the multiple strata and the P-wave velocity of each of the multiple strata to arrange multiple strata in the generation region 10 so as to deposit multiple strata in the generation region 10. Figure 4 shows an example of a generation region 10 in which multiple strata have been deposited. As shown in Figure 4, multiple strata with different P-wave velocities are deposited in the generation region 10.

[0035] (fold) The model generation unit 315 folds multiple strata in the generation region 10, where multiple strata are deposited, using the determined model generation parameters. For example, the model generation unit 315 folds multiple strata in the generation region 10 using the wavelength and amplitude of the fold. Figure 5 shows an example of a generation region 10 in which multiple strata have been folded. As shown in Figure 5, the amplitude of the fold can also be varied depending on the depth. For example, the amplitude of the fold increases in proportion to the shallowness. Also, as shown in Figure 5, the fold corresponds to pulling the multiple strata upward in the vertical direction. Therefore, the deepest part of the generation region 10 is filled with the strata that existed before folding.

[0036] (Fault) The model generation unit 315 generates faults in multiple layers of rock in a generation region 10 having multiple folded layers, using the determined model generation parameters. For example, the model generation unit 315 forms faults in multiple layers using the fault location, fault angle, fault displacement, and fault curvature. Figure 6 shows an example of a generation region 10 in which faults have been generated in multiple layers. Typically, normal faults or reverse faults consistently exist in a single region. Therefore, depending on user settings or default settings, the model generation unit 315 generates either normal faults or reverse faults, as shown in Figure 6.

[0037] (scraping) The model generation unit 315 performs erosion processing in the generation region 10, which has multiple strata where faults have formed, using the determined model generation parameters. For example, the model generation unit 315 removes all strata shallower than the depth of the erosion surface from the generation region 10. Specifically, the model generation unit 315 fills the area of ​​all strata shallower than the depth of the erosion surface (hereinafter referred to as the erosion region) with 0. Figure 7 shows an example of a generation region 10 in which erosion has been performed on all strata shallower than the depth of the erosion surface. At this time, the model generation unit 315 may form an unconformity surface in the region 13 directly above the erosion surface, as shown in Figure 7.

[0038] (redeposition) The model generation unit 315 places multiple strata in the eroded area 15 of the generation region 10 using the thickness of each of the multiple strata and the P-wave velocity of each of the multiple strata. Figure 8 shows an example of the generation region 10 in which multiple strata have been re-deposited in the eroded area 15. As shown in Figure 8, multiple strata with different P-wave velocities are deposited in the eroded area 15.

[0039] (re-fold) The model generation unit 315 folds multiple strata in the generation region 10, where multiple strata are deposited in the eroded region 15, using the wavelength and amplitude of the folds. Figure 9 shows an example of a generation region 10 in which multiple strata have been folded. Folding corresponds to pulling multiple strata vertically upward, as shown in Figure 9. Therefore, as shown in Figure 9, the deepest part of the generation region 10 is filled with strata that existed before folding, similar to Figure 5.

[0040] (Re-fault) The model generation unit 315 forms faults in the generation region 10, which has multiple folded strata, using the fault location, fault angle, fault displacement, and fault curvature. Figure 10 shows an example of a generation region 10 in which faults have been generated in multiple strata.

[0041] (Rock salt intrusion) The model generation unit 315 intrudes rock salt into the generation region 10 using the determined model generation parameters. For example, the model generation unit 315 intrudes and places rock salt into the generation region 10 using the distribution location of the rock salt domes and the P-wave velocity of the rock salt domes. Figure 11 shows an example of the generation region 10 into which rock salt 17 has been intruded.

[0042] The procedures for the above geological events (deposition, folding, faulting, erosion, redeposition, fault reactivation (re-folding, re-faulting), salt intrusion) executed by the model generation unit 315 follow the timeline of their occurrence. Therefore, the order of the above geological events can be changed as appropriate. Furthermore, the geological events are not limited to deposition, folding, faulting, erosion, redeposition, fault reactivation (re-folding, re-faulting), and salt intrusion; other events may also be executed.

[0043] The model generation unit 315 cuts out a predetermined range from the generation region 10 in which the rock salt 17 has penetrated. The predetermined range is, for example, in the generation region 10 shown in Figure 11, a depth of 0 to 5 km and a horizontal length of 0 to 25 km. The predetermined range may be set and changed as appropriate by the setting unit 311 under the instruction of the user via the input device. The model generation unit 315 generates a subsurface structure model by cutting out a range from the generation region 10. The model generation unit 315 generates multiple subsurface structure models according to the generation of random numbers. Since the random number seed corresponds to the subsurface model ID, the model generation unit 315 can regenerate the subsurface structure model according to the selection of the subsurface model ID.

[0044] The simulated data generation unit 317 generates simulated data that simulates observed values ​​for a structure by performing a wave propagation simulation on the structural model. The wave propagation simulation is, for example, a simulation of seismic waves (hereinafter referred to as seismic wave simulation). Specifically, the simulated data generation unit 317 performs a seismic wave propagation simulation on the underground structure model generated by the model generation unit 315. For seismic wave propagation simulation, known techniques that are commonly used for wave propagation simulations of elastic waves or acoustic waves can be used as appropriate. The seismic wave propagation simulation is realized, for example, by applying initial conditions and boundary conditions to the partial differential equations that constitute the equation of motion (wave equation) of an elastic body and solving them sequentially. As numerical calculation methods for these partial differential equations, for example, the finite difference method (FDM) or the finite element method (FEM) can be used. The simulated data generation unit 317, for example, inputs a subsurface structure model into an earthquake simulator and executes an earthquake wave propagation simulation. This allows the simulated data generation unit 317 to generate simulated data that simulates observed values ​​related to the subsurface structure of the subsurface structure model. The simulated data includes, for example, shot data (shot data of the subsurface structure) received by a seismometer after seismic waves generated by a virtually artificial earthquake (shot) propagate through the subsurface structure model. The simulated data generation unit 317 may also generate simulated data by executing wave propagation simulations for multiple structural models.

[0045] Figure 12 shows an example of a common shot gather data image (shot image) 19. In the shot image 19 shown in Figure 12, the vertical axis corresponds to the time from the time the shot was executed, and the horizontal axis of the shot image 19 indicates the horizontal position in the subsurface structure model used to generate the shot image 19.

[0046] The learning data generation unit 319 generates learning data by associating the generated model of the structure with the simulated data. Specifically, the learning data generation unit 319 associates multiple underground structure models generated according to random numbers with multiple simulated data sets generated according to the multiple underground structure models, using the input / output relationships in earthquake wave propagation simulations. The learning data generation unit 319 generates multiple learning data sets from the associated multiple underground structure models and multiple simulated data sets. The learning data generation unit 319 stores the multiple learning data sets in the main memory 33 or the auxiliary memory 35. The learning data generation unit 319 may also store the multiple learning data sets in an external device A as network storage. The learning data generation unit 319 may also generate learning data by associating multiple structure models with the simulated data for each of them.

[0047] The components of the learning data generation device 3 have been described above. The following describes the procedure for generating learning data by the learning data generation device 3 (hereinafter referred to as the learning data generation process). The procedure for the learning data generation process corresponds to the learning data generation method. The learning data generation method generates a structural model based on multiple feature quantities related to the structure, generates simulated data that simulates observed values ​​related to the structure through wave propagation simulation on the structural model, and generates learning data by associating the generated structural model with the simulated data. For example, the structure is a subsurface structure, and the multiple feature quantities include geological information. Geological information includes, for example, the way the strata are bent laterally, the way unconformities are incorporated, the size of the structure (width, depth) related to the modeling of the subsurface structure, the insertion of a slow-velocity layer into the surface, or one of the distribution of layer thickness. The learning data generation method determines multiple feature quantities based on the range of feature quantity values ​​and random numbers. Specifically, the structure is a subsurface structure, and the range of feature quantity values ​​is set based on geological information in the target area. The geological information includes at least one of the following: logging data in the target area, observational data on the subsurface structure in the target area, and the spatial distribution of inferred features in the target area.

[0048] For example, the structure is a subsurface structure, and the training data generation method generates a model of the structure by sequentially executing events related to at least one of the following based on multiple features, as shown in Figures 3 to 11: deposition, folding, faulting, erosion, redeposition, refolding, re-faulting, or salt intrusion. The training data generation method generates a model of the structure by executing events in the order of deposition, folding, and faulting. Alternatively, the training data generation method may generate a model of the structure by executing events in the order of faulting, erosion, redeposition, refolding, and re-faulting. Or, a model of the structure may be generated by executing events in the order of re-faulting and salt intrusion. Figure 13 is a flowchart showing an example of the training data generation process.

[0049] (Training data generation process) (Step S101) The setting unit 311 sets a feature range for each of the multiple feature quantities related to the subsurface structure. The upper and lower limits of the feature range are set by default. The feature range may also be set based on geological information, for example, by user instruction via an input device. The above functions performed by the setting unit 311 may also be set by other input devices such as the external device 9A.

[0050] (Step S102) The determination unit 313 determines the underground model ID. The determination unit 313 generates random numbers using the underground model ID as a random number seed. For multiple features, the determination unit 313 determines the model generation parameters within the feature range based on the feature range and the random numbers. The determined model generation parameters are stored in the main memory 33 or auxiliary memory 35.

[0051] (Step 103) The model generation unit 315 generates subsurface structure models using model generation parameters. Specifically, the model generation unit 315 generates multiple subsurface structure models according to multiple random numbers. The model generation unit 315 stores the multiple subsurface structure models in the main memory 33 or the auxiliary memory 35.

[0052] (Step S104) The simulated data generation unit 317 performs an earthquake wave propagation simulation on the underground structure model and generates simulated data corresponding to the said underground structure model. The simulated data generation unit 317 stores the simulated data generated according to the underground structure model in the main memory 33 or the auxiliary memory 35.

[0053] (Step S105) The learning data generation unit 319 generates multiple learning data by associating multiple underground structure models generated according to random numbers with multiple simulated data generated according to the multiple underground structure models. The learning data generation unit 319 stores the multiple learning data in the main memory 33, auxiliary memory 35, or an external device A as network storage. Through the above process, one learning data generation device 3 generates, for example, more than 500,000 learning data for one underground model ID and stores the generated learning data.

[0054] Below, we will describe a parametric velocity model generation system as an example of a training data generation device 3.

[0055] (Paraphrase-based velocity model generation system) To obtain a large training dataset, we propose a parametric velocity model generation system. This system is designed to prevent data leakage and enable proper evaluation of its generation performance, as it only utilizes geological information (knowledge) in generating subsurface structure models. Firstly, the system generates subsurface structure models, i.e., realistic and high-resolution velocity distributions. The velocity structure in the subsurface structure model is generated through a composite process corresponding to geological events including stratification, folding, faulting, intrusion, and erosion. Each process is modeled using geological parameters such as layer thickness, P-wave velocity, and fault dip angle. Figure 14 shows an example of a velocity model as an example of a subsurface structure model. Specifically, Figure 14 shows an example of a P-wave velocity (Vp) model generated by the parametric generation system. To create shot gathers (simulated data) corresponding to the velocity model, seismic wave propagation in the generated subsurface structure model is simulated.

[0056] This system incorporates multiple probability distributions to extract a large number of samples for geological parameters, generating large-scale training data that includes a wide range of diversity in subsurface structures. The decision unit 313 uses a uniform distribution with upper and lower limits for the geological parameters. The range of the geological parameters is set by the setting unit 311 based on geological insights, such as a rough estimate of the velocity structure. In specific cases, insights into the target subsurface are obtained using geological surveys, such as the results of analysis using conventional inverse problem methods or logging data from nearby areas. High-quality training data is generated by making reasonable assumptions about these geological parameters.

[0057] The above describes the training data generation process by the training data generation device 3. The training device 7 will now be described. The hardware configuration of the training device 7 is the same as that within the frame of dotted line 3 in Figure 1, so the explanation will be omitted. The training device 7 trains a deep neural network using multiple training data sets. This deep neural network is an example of a model that estimates information about structure (hereinafter referred to as the estimation model). The training device 7 generates an estimation model that estimates information about structure using the training data generated using the training data generation method described above. That is, the model generation method that generates the estimation model using the training data generated using the training data generation method described above is executed using at least one processor in the training device 7.

[0058] Figure 15 shows an example of a functional block in a processor 81 mounted on a learning device 7. The processor 81 has a pre-processing unit 811, a model setting unit 813, and a learning unit 815 as functions realized by the processor 81. The functions realized by the pre-processing unit 811, the model setting unit 813, and the learning unit 815 are each stored as programs in, for example, the main memory or auxiliary memory mounted on the learning device 7. The processor 81 reads and executes the programs stored in the main memory or auxiliary memory mounted on the learning device 7 to realize the functions related to the pre-processing unit 811, the model setting unit 813, and the learning unit 815.

[0059] The preprocessor 811 increases the number of simulated data points corresponding to a single underground structure model, depending on factors such as noise addition and various settings during data acquisition in seismic exploration. The noise added to the simulated data sets includes, for example, noise from sensors that are unable to receive vibrations among multiple vibration-receiving sensors in seismic exploration, noise caused by vehicles passing on roads in the survey area, and noise from exploratory wells installed in the survey area. Various settings include, for example, the positional relationship of vibration-receiving sensors relative to the seismic simulator and the method of generating shots by the seismic simulator. As a result, the preprocessor 811 augments the number of simulated data points corresponding to a single underground structure model in multiple training data sets. That is, the augmentation of simulated data by the preprocessor 811 results in multiple simulated data points corresponding to a single underground structure model (ground truth data). The preprocessor 811 stores the multiple training data points increased by the augmentation in the main memory or auxiliary memory of the learning device 7.

[0060] The model configuration unit 813 configures a pre-trained model that will be trained using multiple training datasets from among multiple training data. The pre-trained model is, for example, a deep neural network. First, the model configuration unit 813 divides the multiple training data into multiple training datasets and multiple validation datasets used to validate the trained model. Next, the model configuration unit 813 extracts from the multiple training datasets multiple model configuration datasets used to configure the pre-trained model and multiple model validation datasets used to validate the configured model. Hereinafter, the multiple datasets obtained after extracting the model configuration datasets and model validation datasets from the multiple training datasets will be referred to as the extracted datasets.

[0061] The model setting unit 813 sets up a pre-training model using multiple model setting datasets and multiple model validation datasets. For example, the model setting unit 813 applies Neural Architecture Search (NAS) using an automated hyperparameter optimization framework to the multiple model setting datasets and multiple model validation datasets. This allows the model setting unit 813 to set the structure of the pre-training model and its hyperparameters. Optuna® is used as the NAS, for example. However, the NAS is not limited to Optuna®; other architectures may be used. Furthermore, the model setting unit 813 may set a deep neural network model suitable for exploration and its hyperparameters based on user instructions via an input device.

[0062] For example, NAS automatically designs a neural network suitable for a given task. Once the search space for neural architectures is defined, NAS sequentially samples, trains, and evaluates candidate architectures using multiple model configuration datasets and multiple model validation datasets to find the optimal architecture within that search space. This architecture is defined by model hyperparameters, including the number of layers and channels. The model configuration unit 813 employs Optuna®, an automated hyperparameter optimization framework. Optuna®'s user-friendly interface enables simple and automated hyperparameter optimization via parallel processing, reducing computation time.

[0063] Specifically, the model configuration unit 813 defines the neural architecture's search space as an encoder-decoder model based on ResNet. ResNet is a deep learning model used in both the seismic inverse problem and computer vision. Encoder-decoder models that associate inputs and outputs with a common latent feature space are of particular interest. Typically, convolutional neural networks capture spatially local features. In an encoder-decoder model, the encoder and decoder are connected by a common latent feature space. Therefore, the encoder-decoder model can learn spatially global features that lack structure by breaking down the structure of the input simulated data (shot images). Learning spatially global features is useful for the seismic inverse problem.

[0064] The learning unit 815 trains the pre-training model set by the model setting unit 813 using the extracted dataset. For example, the learning unit 815 inputs each of the multiple simulated data points in the extracted dataset into the pre-training model. The learning unit 815 adjusts the weights of the pre-training model, for example, by stochastic gradient descent with backpropagation, to reduce the difference between the output of the pre-training model and the underground structure model corresponding to the simulated data input into the pre-training model. The learning unit 815 also validates the pre-training model with adjusted weights using a validation dataset. Through these steps, the learning unit 815 generates a trained model (estimated model). The learning unit 815 stores the generated trained model in the main memory or auxiliary memory of the learning device 7.

[0065] The above describes an example of generating a trained model using the learning device 7. Below, we will describe an example of the process for estimating the subsurface structure using the trained model (hereinafter referred to as the subsurface structure estimation process).

[0066] Prior to performing the subsurface structure estimation process, it is assumed that shot data related to the subsurface structure to be estimated has been acquired in advance. Furthermore, prior to performing the subsurface structure estimation process, data cleansing such as noise reduction and outlier removal may be performed on the raw data or shot data as appropriate.

[0067] The hardware configuration of the estimation device is the same as that within the dotted line 3 in Figure 1, so the explanation is omitted. The estimation device stores the trained model in its main memory or auxiliary memory. The estimation device may store multiple trained models generated according to the region. In this case, the estimation device selects a trained model according to the user's instructions via the input device. The estimation device estimates the underground structure of the target by inputting shot data into the trained model. In this way, the estimation device estimates the underground structure corresponding to the input shot data. The estimated underground structure data may be post-processed as appropriate. The estimation device stores the estimated underground structure data in its main memory or auxiliary memory. In this case, the estimation device may display the estimated underground structure data on a display unit (display) provided in the estimation device.

[0068] The above describes the estimation of underground structure data using the estimation device. Below, we will describe an example of experimental results using various processes related to this embodiment.

[0069] (An example of experimental results) The various processes in the embodiment are applied to the inverse problem of estimating the velocity structure of each strata in a subsurface cross-section from two-dimensional (2D) shot-gather images. The experiment comprises four steps: training data generation, training data partitioning, NAS, and evaluation of the optimal neural architecture.

[0070] The training data generation process created a training data dataset of 300,000 pairs, each containing a velocity model corresponding to the subsurface structure model and corresponding shot gathers (simulated data). The generation parameters (geological parameters) that control the geological properties in the velocity model were determined based on the geological insights of the Marmousi2 geological structure model (Martin, GS, Wiley, R., and Marfurt, KJ

[2006] Marmousi2: An elastic upgrade for Marmousi. The Leading Edge, 25(2): 156-166). Summary information about these geological properties was used to define the feature range, but the benchmark dataset itself used for benchmarking the trained model was not used. Seismic wave propagation was simulated using a supercomputer.

[0071] To prevent data leakage during NAS processing, the training data was split. The 300,000 training data points were split into a large training dataset of 240,000 samples and a large validation dataset of 60,000 samples. These large datasets were used to train the optimal neural architecture. Furthermore, 10,000 data points were sampled from the large training dataset, and this subset was split into a small model configuration dataset of 8,000 samples and a small model validation dataset of 2,000 samples. These smaller datasets were used to find the optimal neural architecture. This dataset splitting avoids non-standard access (data leakage) to the large validation dataset during the NAS step.

[0072] The model configuration unit 813 optimizes the ResNet-based encoder-decoder model by tuning hyperparameters such as the number of layers and channels. Ultimately, it obtains an optimal neural architecture (pre-trained model) with more than 100 hidden layers, which is much deeper than those used in previous studies.

[0073] The learning unit 815 trained an optimal neural architecture using a large training dataset. The generated trained model was evaluated using two standard benchmark datasets: the Marmousi2 geological structure model and the 1994 Amoco static correction test dataset. Figure 16 shows an example of P-wave velocity (Vp) estimation in the Marmousi2 geological structure model. Specifically, Figure 16 shows an example of a comparison between the results of an inverse problem applied to a conventional ResNet50-based model (He, K., Zhang, X., Ren, S., and Sun, J.

[2016] Deep residual learning for image recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 770-778.) and the results of an inverse problem applied to a trained model trained by the learning device 7. Figure 16(a) shows the ground truth. Figure 16(b) shows the estimation by the trained model generated according to the embodiment. Figure 16(c) shows the estimation using an encoder-decoder model based on ResNet50. Figure 16(d) shows the one-dimensional (1D) profile at 2km, corresponding to the red line in (a) through (c). Figure 16(e) shows the 1D profile at 11km, corresponding to the blue line.

[0074] The results of the inverse problem correspond to the results of the inverse problem of the Marmousi2 geological structure model. Due to the quality and quantity of the training dataset in this embodiment, the results of the inverse problem (c) using the conventional baseline ResNet50-based model roughly regenerated the velocity model (a). On the other hand, the results of the inverse problem (b) using the pre-trained model generated in this embodiment showed more understandable results compared to the conventional inverse problem results (c). For example, the results of the inverse problem (b) using the pre-trained model generated in this embodiment predicted the velocity of the salt layer (4 km deep in Figure 16 (d)) more accurately than the conventional inverse problem results (c), and obtained more detailed structures in the complex area around the fault (Figure 16 (e)).

[0075] Figure 17 shows the modeling results for the 1994Amoco static correction test dataset. Figure 17(a) shows the ground truth. Figure 17(b) shows the estimation by the trained model generated by this embodiment. The trained model generated by this embodiment estimated high resolution and good output despite the training data not containing information related to the 1994Amoco static correction test dataset.

[0076] According to the learning data generation device 3 of this embodiment (hereinafter referred to as "the learning data generation device 3"), as an example, a range of values ​​for each of a plurality of features relating to the subsurface structure is set, the values ​​of the features within the range are determined based on the range and random numbers for the plurality of features, a subsurface structure model is generated using the determined values, simulated data that simulates observed values ​​relating to the subsurface structure is generated by seismic wave propagation simulation on the subsurface structure model, and multiple learning data are generated by associating the plurality of subsurface structure models generated according to the random numbers and the plurality of simulated data generated according to the plurality of subsurface structure models with the input / output relationship in the seismic wave propagation simulation. According to the learning data generation device 3, the range is set based on geological information in the region including the area to be investigated for the subsurface structure. Here, the geological information in the learning data generation device 3 includes at least one of logging data in the region, observational data relating to the subsurface structure in the region, and the spatial distribution of inferred features in the region. Furthermore, according to the training data generation device 3, the features include the depth of the erosion surface in the subsurface structure model, the degree of fault curvature in the subsurface structure model, and the unconformity of the strata, which includes the unconformity surface formed by the flow of the salt layer. In addition, according to the training data generation device 3, normal faults or reverse faults are set in the features.

[0077] Based on these considerations, the learning data generation device 3 can, for example, set geological parameters using random numbers within a range of geological parameters set based on geological knowledge and insights, and then generate a large number of diverse subsurface structure models using the set geological parameters, following the geological process of subsurface structure generation. In other words, the learning data generation device 3 can generate a large number of realistic subsurface structure models using random numbers, referencing natural history, and generate a large amount of training data. The large amount of training data generated by the learning data generation device 3 contains a large number of high-quality, i.e., realistic subsurface structure models as training data, thus improving the generalization performance of the trained models.

[0078] Furthermore, according to the learning data generation device 3, as an example, it displays multiple ranges corresponding to multiple features and two indicators that change the upper and lower limits of each of the multiple ranges, and generates a subsurface structure verification model using the representative values ​​of the ranges specified by the two indicators, and displays multiple ranges corresponding to multiple features, the two indicators for each of the multiple ranges, and the verification model. As a result, with the learning data generation device 3, when inputting ranges that reflect geological knowledge etc. in the range of geological parameters, the user can easily grasp the changes in the subsurface structure model that accompany changes and adjustments to the ranges. This improves the user's operability in generating subsurface structure models and further improves the quality of learning data. This improves the generalization performance of the trained model.

[0079] (Examples of application) This application example involves using observational data obtained from observations of a structure, simulated data, and a loss function weighted according to the magnitude of the data values ​​in the simulated data. For example, a Generative Adversarial Network (GAN) is used to train an improver that improves the simulated data to more closely resemble the reality of the observed data based on the simulated data. The simulated data is then input into the trained improver to generate improved data, and training data is generated by associating the improved data with the structure model. This application example also describes training an estimation model for the structure (hereinafter referred to as the subsurface structure estimation model) using training data containing the improved data, and the subsurface structure estimation process using the subsurface structure estimation model.

[0080] Furthermore, in this application example, the learning data generation unit 319 is provided on the processor 81 mounted on the learning device 7. When the technical features of this application example are implemented in an estimation device, the processor in the estimation device comprises the learning data generation unit 319, the learning unit 815, and an estimation unit that estimates the structure of the observed data by inputting the observed data into a trained model.

[0081] Figure 18 shows an example of an overview of the process, including the generation of training data, the training of the subsurface structure estimation model TUSEM using the training data, and the subsurface structure estimation process using the trained subsurface structure estimation model TUSEM. The Sim shown in Figure 18 corresponds to the training data generation process in the embodiment. Specifically, the Sim shown in Figure 18 shows an overview of the process of generating multiple subsurface structure models and generating multiple simulated data SDs corresponding to the subsurface structure model USM using wave propagation simulation WPS. The processing details in the Sim shown in Figure 18 are the same as in the embodiment, so the explanation is omitted. In addition, Real in Figure 18 shows the collected observation data (e.g., shot data) OD used to perform the subsurface structure estimation process. The acquisition of observation data OD follows existing methods, so the explanation is omitted.

[0082] Figure 18 shows the S2R process, which involves learning a refiner (Refiner) RF using a generative adversarial network based on simulated data SD and observed data OD, and then converting the simulated data SD to improved data RD using the learned refiner TRF. The generative adversarial network in Figure 18 has a refiner (Refiner) RF and a discriminator (Discriminator) DCN to be trained. Furthermore, during the learning process of the refiner RF by the generative adversarial network, the output from the refiner RF during training corresponds to noisy data NAD, which is simulated data SD with realistic noise added to it.

[0083] In Figure 18, INV shows the learning process of the model LOM to be trained using the improved data RD and the subsurface structure model USM, and the subsurface structure estimation process in which the observed data OD is input to the trained subsurface structure estimation model TUSEM and the estimated subsurface structure UGS is output.

[0084] The learning data generation unit 319 executes a wave propagation simulation (WPS) on the underground structure model (USM). This causes the learning data generation unit 319 to generate simulated data (SD) corresponding to the underground structure model (USM), which is equivalent to the simulation results of the wave propagation simulation (WPS). The learning data generation unit 319 performs the above process on multiple underground structure models (USM) to generate multiple simulated data (SD) corresponding to multiple underground structure models (USM). Next, the learning data generation unit 319 adds randomly generated noise to each of the multiple simulated data (SD). This generates multiple simulated data with randomly added noise (hereinafter referred to as "noised simulated data"). The addition of noise is performed, for example, at the arrow NA shown in Figure 18. Note that the addition of random noise may be omitted as appropriate to shorten the processing time in this application example. The learning data generation unit 319 associates the multiple underground structure models (USM) and the multiple noise-added simulated data with the wave propagation simulation (WPS) via input / output and stores them in memory. Furthermore, the learning data generation unit 319 reads the observation data OD acquired by the existing acquisition device and the network before training from memory. The total number of observation data ODs may be less than, for example, the total number of noise-added simulated data, but it is desirable to have a large number of data, above a predetermined number, in order to improve the generalization performance of the training for the improvement device RF.

[0085] The learning data generation unit 319 applies observed data and noise-added simulated data (or simulated data if noise addition is not performed) to the network and alternately learns the improver RF and the discriminator DCN. The loss function used to learn the improver RF (hereinafter referred to as the improver loss function) is weighted according to the strength of the signal value in the simulated data SD, for example, in proportion to the magnitude of the signal value in the simulated data SD, in order to maintain that signal value. The loss function specific to this application example will be described below. Note that other network configurations and processes can be handled using existing technologies, so their explanation will be omitted.

[0086] Loss function of the improvement device L1represents an image corresponding to noise-added simulated data input to the network with respect to simulated data SD as I sim , and let an image corresponding to noise-added data NAD be I taint , then it is defined by, for example, the following equation (1).

[0087] Loss L1 =mean(abs(I sim )*abs(I sim -I taint ))···(1) The right-hand side of equation (1) is for image I sim and image I taint , this indicates that the absolute value of image I sim is multiplied for each pixel by the absolute value abs of the difference between, and the mean value mean across all pixels of the entire image is calculated. Multiplying abs(I sim -I taint ) by abs(I sim ) means that in the ordinary L1 loss mean(abs(I sim -I taint ), the absolute value image abs(I sim -I taint ) before calculating the average value is multiplied by the magnitude of the signal value in the simulated data SD as a weight abs(I sim ). Thereby, the learning data generation unit 319 repeatedly trains the refiner RF so that the refiner loss function Loss L1 represented by equation (1) becomes smaller. Accordingly, the refiner RF is trained to maintain the signal value in proportion to the magnitude of the signal value in the simulated data SD.

[0088] Note that the weight in the refiner loss function Loss L1 is not limited to the above abs(I sim ). For example, the weight in the refiner loss function Loss L1 may be any function as long as it is a broadly monotonic increasing function (broadly monotonic increasing function) of abs(I sim ). Specifically, the weight in the refiner loss function Loss L1 is a non-linear weight (abs(I sim ))^2, sqrt(abs(Isim )), or min(a, abs(I sim ) etc. are also acceptable.

[0089] The learning data generation unit 319 stores in memory an improved model RF (hereinafter referred to as the learned improved model TRF) that has been learned by learning a network using multiple observation data and multiple noise-added simulated data (multiple simulated data if noise addition is not performed). Since the learned improved model TRF is an improved model that reflects the observation data, it reflects the region in which the observation data was acquired, the collection conditions in the collection of the observation data (for example, observation equipment related to the collection of observation data, various characteristics of the operator when collecting observation data, characteristics of the survey company related to the collection of observation data, etc.), and is learned in a way that is specialized for the observation data, that is, it is automatically customized with respect to the observation data.

[0090] The learning data generation unit 319 generates noise-added simulated data by adding randomly generated noise to the simulated data. The learning data generation unit 319 reads the trained improver TRF from memory, inputs the noise-added simulated data or the simulated data, and generates improved data. The noise-added simulated data input to the trained improver TRF may be the noise-added simulated data used to train the improver RF. The learning data generation unit 319 generates learning data by associating the structural model with the improved data. Specifically, the learning data generation unit 319 generates learning data by associating the underground structure model USM with the improved data RD. The learning data generation unit 319 generates multiple learning data sets by repeating the process from generating the noise-added simulated data to generating the improved data RD for multiple simulated data sets. The learning data generation unit 319 stores the generated multiple learning data sets in memory.

[0091] The learning unit 815 learns the subsurface structure estimation model TUSEM by learning the target model LOM using multiple training data sets. Since known methods can be used appropriately for learning the model LOM using multiple training data sets, a detailed explanation is omitted. In learning the subsurface structure estimation model TUSEM, noise-enhanced simulated data is used, which is a simulation result (simulated data SD) with added realism from observed data OD. Therefore, the subsurface structure estimation model TUSEM learns a network (a pre-trained model) capable of inversion, which is the estimation of the subsurface structure, onto the actual observed data OD.

[0092] The estimation unit estimates the subsurface structure UGS by inputting the observation data OD into the subsurface structure estimation model TUSEM. The estimated subsurface structure is stored in memory. The estimated subsurface structure may also be displayed on the display.

[0093] To make the explanation more concrete, the observational data will be assumed to be shot data on the subsurface structure obtained in the area related to the estimation of the subsurface structure. Furthermore, the process of estimating the subsurface structure based on the acquisition of the shot data will be described. Figure 19 is a flowchart showing an example of the procedure for a series of processes (hereinafter referred to as the model generation estimation process) that includes the generation of training data, the generation of a subsurface structure estimation model using the training data, and the subsurface structure estimation process using the generated subsurface structure estimation model.

[0094] (Model generation and estimation process) (Step S191) The learning data generation unit 319 acquires multiple shot data ODs. For example, the learning data generation unit 319 acquires shot data ODs from an acquisition device, a server device storing shot data ODs, or a storage medium storing shot data ODs. The learning data generation unit 319 stores the acquired shot data ODs in memory.

[0095] (Step S192) The training data generation unit 319 reads multiple simulated data SDs from memory. The training data generation unit 319 adds random noise to each of the multiple simulated data SDs to generate noisy simulated data. The training data generation unit 319 stores the generated noisy simulated data in memory. Note that this step may be omitted in order to shorten the processing time in the model generation and estimation process.

[0096] (Step S193) The learning data generation unit 319 uses noise-added simulated data and multiple shot data to train the improver RF together with the discriminator DCN and generate a trained improver TRF. Specifically, the learning data generation unit 319 inputs the shot data to the improver RF to be trained. The learning data generation unit 319 outputs noise-added data NAD from the improver RF to be trained. The learning data generation unit 319 generates an image corresponding to the noise-added simulated data with respect to the simulated data SD. sim And, an image corresponding to the noise-added data NAD is I taint Based on this, the improvement device loss function Loss L1 The learning data generation unit 319 calculates the improvement loss function. L1 To reduce noise, for example, the improver RF is learned using backpropagation. In addition, the learning data generation unit 319 learns the classifier DCN. The learning data generation unit 319 learns the improver RF and the classifier DCN by repeating these processes according to each of the multiple noise-added simulated data and each of the multiple shot data. Once the learning process is complete, the learning data generation unit 319 generates a learned improver TRF. Note that the learning of the improver RF and the classifier DCN may be implemented by the learning unit 815 in the learning device.

[0097] (Step S194) The learning data generation unit 319 inputs each of the multiple noise-added simulated data sets into the trained improvement unit TRF to generate multiple improvement data sets. Specifically, the learning data generation unit 319 generates multiple noise-added simulated data sets by adding randomly generated noise to each of the multiple simulated data sets.

[0098] (Step S195) The learning data generation unit 319 generates multiple learning data sets by associating multiple underground structure models with multiple improvement data sets. The learning data generation unit 193 stores the multiple learning data sets in memory.

[0099] (Step S196) The learning unit 815 generates the subsurface structure estimation model TUSEM using multiple training data. That is, the learning unit 815 generates the subsurface structure estimation model TUSEM by learning the target model LOM across multiple training data. The learning unit 815 stores the subsurface structure estimation model TUSEM in memory.

[0100] (Step S197) The estimation unit inputs each of the multiple shot data into the TUSEM subsurface structure estimation model and estimates the UGS subsurface structure. The estimation unit stores the estimated subsurface structure in memory. The estimation unit may also display the estimated UGS subsurface structure on a display.

[0101] The learning data generation device 3 according to this embodiment uses observation data OD obtained by observation of the structure, simulated data SD, and a loss function Loss weighted according to the magnitude of the data values ​​in the simulated data SD. L1 Using a generative adversarial network, the improver RF is trained to generate improved data RD based on the simulated data SD, bringing the simulated data SD closer to the reality of the observed data OD. The simulated data SD is input to the trained improver TRF to generate the improved data RD, and the improved data RD is associated with the subsurface structure model USM to generate training data. For example, according to this training data generation device 3, the improver loss function Loss is proportional to the magnitude of the signal value in the simulated data SD. L1 The improver RF is trained using this method. As a result, the training data generation device 3 can generate a trained improver TRF that maintains the signal values ​​in the simulated data SD.

[0102] Furthermore, since the learning data generation device 3 uses the observed data OD to train the improvement device RF, variable factors such as noise due to the collection conditions of the observed data OD can also be learned during the training process of the improvement device RF. Therefore, the learning data generation device 3 can generate improved data RD that brings the simulated data SD closer to the reality of the observed data OD. In other words, the learning data generation device 3 can generate more realistic learning data by generating improved data that adds realism to the simulated data SD, which is the simulation result of the subsurface structure model USM.

[0103] Furthermore, this learning data generation device 3 generates noise-added simulated data based on simulated data SD, observed data OD, and the above-mentioned improvement loss function Loss. L1 The improver RF may be trained using the above. In this case, the training data generation device 3 can effectively train the improver RF and the discriminator DCN with respect to random noise in the observed data OD. As a result, the training data generation device 3 can train the improver RF to generate more realistic improved data.

[0104] Based on the above, this learning data generation device 3 can generate more realistic improvement data, thereby further improving the quality of the learning data.

[0105] This learning device uses the learning unit 815 to train the target model LOM using training data that contains improved data using the trained improver TRF. Thus, this learning device generates the subsurface structure estimation model TUSEM by training the target model LOM using highly realistic improved data and the subsurface structure model USM as ground truth data. Therefore, this learning device can train the subsurface structure estimation model TUSEM, which can output a more reliable subsurface structure UGS for the observed data OD.

[0106] Furthermore, this estimation device estimates the subsurface structure UGS by inputting the observational data OD used to train the improvement device RF into the subsurface structure estimation model TUSEM. Therefore, this estimation device can estimate a highly realistic subsurface structure UGS by using the subsurface structure estimation model TUSEM, which is generated using highly realistic training data.

[0107] Based on the above, the learning data generation device 3 according to this embodiment can generate learning data related to structure.

[0108] When the technical features of this embodiment are realized by the learning data generation method, the learning data generation method is a learning data generation method executed using at least one processor, which involves generating a structural model based on a plurality of structural feature quantities, generating simulated data that simulates observed values ​​related to the structure by performing a wave propagation simulation on the structural model, and generating learning data by associating the generated structural model with the simulated data. The processing procedure corresponding to the learning data generation method corresponds to the procedure for the learning data generation process, so its explanation is omitted. Furthermore, the effects of the learning data generation method are the same as in the embodiment, so its explanation is omitted. When the technical features of this embodiment are realized by the model generation method, the model generation method generates an estimation model that estimates information about the structure using the learning data generated using the learning data generation method described above. The processing procedure for the model generation method corresponds to the processing procedure in the learning unit 815, etc., so its explanation is omitted. Furthermore, the effects of the model generation method are the same as in the embodiment, so its explanation is omitted.

[0109] In the embodiments described above, some or all of the devices may be composed of hardware, or they may be composed of information processing by software (programs) executed by a CPU or GPU. If the information processing is composed of software, the software that realizes at least some of the functions of the devices in the embodiments described above may be stored on a non-temporary storage medium (non-temporary computer-readable medium) such as a flexible disk, CD-ROM (Compact Disc-Read Only Memory), or USB memory, and the information processing of the software may be executed by having the computer 30 read it. Alternatively, the software may be downloaded via a communication network 5. Furthermore, the information processing may be executed by hardware by implementing the software on a circuit such as an ASIC or FPGA.

[0110] The type of storage medium used to store the software is not limited. The storage medium is not limited to removable media such as magnetic disks or optical disks; it may also be a fixed storage medium such as a hard disk or memory. Furthermore, the storage medium may be located inside or outside the computer.

[0111] Where the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used in this specification (including the claims), it includes any of a, b, c, ab, ac, bc, or abc. It also includes multiple instances of any element, such as aa, abb, aabbcc, etc. Furthermore, it includes adding other elements other than the enumerated elements (a, b, and c), such as abcd which has d.

[0112] In this specification (including the claims), when expressions such as "data as input / based on / according to / in accordance with data" (including similar expressions) are used, unless otherwise specified, this includes cases where the data itself is used as input, or where the data has been processed in some way (e.g., data with added noise, normalized data, intermediate representations of the data, etc.) is used as input. Furthermore, when it is stated that some result is obtained "based on / according to / in accordance with data", this includes cases where the result is obtained based solely on the data in question, as well as cases where the result is also influenced by other data, factors, conditions, and / or states other than the data in question. Furthermore, when it is stated that "data is output", unless otherwise specified, this includes cases where the data itself is used as output, or where the data has been processed in some way (e.g., data with added noise, normalized data, intermediate representations of the data, etc.) is used as output.

[0113] In this specification (including the claims), the terms “connected” and “coupled” are intended to be non-restrictive terms that include any direct connection / coupling, indirect connection / coupling, electrical connection / coupling, communicative connection / coupling, operational connection / coupling, physical connection / coupling, etc. The terms should be interpreted as appropriate in the context in which they are used, but any form of connection / coupling that is not intentionally or naturally excluded should be interpreted non-restrictively as being included in the terms.

[0114] In this specification (including the claims), when the expression "A configured to B" is used, it may include that the physical structure of element A has a configuration capable of performing operation B, and that the permanent or temporary setting / configuration of element A is configured to actually perform operation B. For example, if element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and that it is configured to actually perform operation B by the setting of a permanent or temporary program (instruction). Furthermore, if element A is a dedicated processor or dedicated arithmetic circuit, it is sufficient that the circuit structure of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached.

[0115] Wherever terms meaning "comprising" or "possessing" (e.g., "comprising / including" and "having") are used in this specification (including the claims), they are intended to be open-ended terms, including cases where the subject matter of such terms is not the object of the term. Where the object of such terms meaning "comprising" or "possessing" is an expression that does not specify a quantity or suggests a singular number (an expression with the article "a" or "an"), such expression should be interpreted as not being limited to a specific number.

[0116] In this specification (including the claims), even if expressions such as "one or more" or "at least one" are used in one place, and expressions that do not specify a quantity or suggest a singularity (expressions using the articles a or an) are used in another place, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or suggest a singularity (expressions using the articles a or an) should be interpreted as not necessarily being limited to a specific number.

[0117] In this specification, if a particular configuration of an embodiment is described as yielding a specific advantage or result, it should be understood, unless otherwise stated, that the same advantage or result can also be obtained from one or more other embodiments having the same configuration. However, it should be understood that the presence or absence of such an advantage or result generally depends on various factors, conditions, and / or states, and that the configuration does not necessarily guarantee that the advantage or result can be obtained. The advantage or result can only be obtained from the configuration described in the embodiment when various factors, conditions, and / or states are met, and the advantage or result cannot necessarily be obtained in the invention claimed to define that configuration or a similar configuration.

[0118] In this specification (including the claims), when terms such as "maximize" are used, they include finding the global maximum value, finding an approximation of the global maximum value, finding the local maximum value, and finding an approximation of the local maximum value, and should be interpreted appropriately depending on the context in which the term is used. They also include finding approximations of these maximum values ​​probabilistically or heuristically. Similarly, when terms such as "minimize" are used, they include finding the global minimum value, finding an approximation of the global minimum value, finding the local minimum value, and finding an approximation of the local minimum value, and should be interpreted appropriately depending on the context in which the term is used. They also include finding approximations of these minimum values ​​probabilistically or heuristically. Similarly, when terms such as "optimize" are used, they include finding the global optimal value, finding an approximation of the global optimal value, finding the local optimal value, and finding an approximation of the local optimal value, and should be interpreted appropriately depending on the context in which the term is used. They also include finding approximations of these optimal values ​​probabilistically or heuristically.

[0119] In this specification (including the claims), when multiple hardware components perform a predetermined process, each component may cooperate to perform the predetermined process, or some components may perform all of the predetermined process. Alternatively, some components may perform part of the predetermined process, while other components perform the remainder. In this specification (including the claims), when expressions such as "one or more hardware components perform a first process, and the one or more hardware components perform a second process" are used, the hardware component performing the first process and the hardware component performing the second process may be the same or different. In other words, it is sufficient that the hardware component performing the first process and the hardware component performing the second process are included in the one or more hardware components. Hardware may include electronic circuits or devices containing electronic circuits.

[0120] In this specification (including the claims), when multiple memory devices store data, each of the multiple memory devices may store only a portion of the data or the entire data.

[0121] While embodiments of this disclosure have been described in detail above, this disclosure is not limited to the individual embodiments described above. Various additions, modifications, substitutions, and partial deletions are possible, provided that they do not depart from the conceptual idea and spirit of the present invention derived from the claims and their equivalents. For example, where numerical values ​​or mathematical formulas are used in the description in all of the embodiments described above, they are provided as examples only and are not limited thereto. Also, the order of operations in the embodiments is provided as examples only and is not limited thereto. [Explanation of Symbols]

[0122] 1. Learning System 3. Training Data Generation Device 5. Communication Network 7. Learning device 9A external device 9B External device 10 Generation area 13. The area directly above the scraped surface 15 Abrasion area 17 Rock salt 19 Shot Images 30 Computers 31 processors 33 Main memory 35 Auxiliary storage device 37 Network Interfaces 39 Device Interfaces 41 Bus 81 processors 311 Settings Section 313 Decision Section 315 Model Generation Unit 317 Simulated Data Generation Unit 319 Training Data Generation Unit 811 Pre-processing Unit 813 Model Setting Section 815 Learning Department

Claims

1. A method for generating training data that is performed using at least one processor, To generate a model of the structure based on multiple features related to the structure, By performing wave propagation simulations on the model of the aforementioned structure, simulated data is generated that simulates observed values ​​related to the aforementioned structure. The process involves associating the model of the aforementioned structure with the simulated data to generate training data, A method for generating training data, comprising the following features.

2. The aforementioned structure is an underground structure, The aforementioned wave propagation simulation is a simulation relating to seismic waves. The method for generating training data according to claim 1.

3. Based on the aforementioned multiple features, a model of the structure is generated by performing at least one event related to deposition, folding, faulting, erosion, redeposition, refolding, refaulting, or salt intrusion. The method for generating training data according to claim 2.

4. A model of the aforementioned structure is generated by executing events in the order of deposition, folding, and faulting. The method for generating training data according to claim 3.

5. A model of the aforementioned structure is generated by executing events in the following order: faulting, erosion, redeposition, refolding, and re-faulting. The method for generating training data according to claim 3 or claim 4.

6. A model of the aforementioned structure is generated by executing events in the order of re-faulting and salt intrusion. A method for generating training data according to any one of claims 3 to 5.

7. The aforementioned multiple features include geological information. A method for generating learning data according to any one of claims 2 to 6.

8. The aforementioned geological information includes one of the following: the way the strata are bent laterally, the way unconformities are incorporated, the size of the structures (width, depth) related to the modeling of the subsurface structure, the insertion of slow-velocity layers into the surface, or the distribution of layer thickness. The method for generating learning data according to claim 7.

9. Based on the range of values ​​of the aforementioned features and a random number, the plurality of features are determined. A method for generating learning data according to any one of claims 1 to 8.

10. The aforementioned structure is an underground structure, The range of values ​​for the aforementioned feature quantities is set based on geological information in the target area. A method for generating learning data according to any one of claims 1 to 9.

11. The geological information comprises at least one of the following: logging data in the target area, observational data concerning the subsurface structure in the target area, and the spatial distribution of the feature quantities inferred in the target area. The method for generating training data according to claim 10.

12. The aforementioned structure is an underground structure, The aforementioned simulated data includes shot data of the underground structure. A method for generating learning data according to any one of claims 1 to 11.

13. The simulated data is input to an improver that has been trained based on the observation data obtained from the observation of the structure, the simulated data, and a loss function weighted according to the magnitude of the data values ​​in the simulated data, in order to generate improved data. The model of the structure and the improvement data are associated to generate the training data. A method for generating learning data according to any one of claims 1 to 12.

14. An estimation model is generated that estimates information about structure using the training data generated using the training data generation method described in any one of claims 1 to 13. Model generation method.

15. A model generation unit generates a model of the structure based on multiple feature quantities relating to the structure, A simulated data generation unit generates simulated data that simulates observed values ​​related to the structure by performing a wave propagation simulation on a model of the structure, A learning data generation unit that associates the model of the structure with the simulated data and generates learning data, A learning data generation device equipped with the following features.

16. The system further includes a display unit that displays two indicators for changing the upper and lower limits of the aforementioned plurality of feature quantities. The model generation unit generates a model of the structure using representative values ​​within the range specified by the two indicators. The learning data generation device according to claim 15.