Synthetic data generation models

WO2025235927A3PCT designated stage Publication Date: 2025-12-11X DEVELOPMENT LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/028707
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-22
Filing Date
2025-05-09
Publication Date
2025-12-11

AI Technical Summary

Technical Problem

Generating realistic synthetic seismic images, especially for 3D seismic surveys, is challenging due to the complexity of the earth and limitations in seismic surveying processes, leading to inefficiencies in obtaining detailed and realistic training data for machine learning models.

Method used

A system using generative neural networks, specifically diffusion models with denoising neural networks, generates realistic synthetic seismic data items by processing conditioning inputs to produce high-detail and realistic seismic images, allowing for controlled data generation and improved training data for machine learning models.

Benefits of technology

The system efficiently generates high-realism seismic data, enabling better performance of machine learning models in seismic data processing tasks by providing larger and varied training datasets, improving generalization and reducing the need for manual data curation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025028707_11122025_PF_FP_ABST
    Figure US2025028707_11122025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for generating data items. One of the methods includes obtaining a conditioning input that characterizes a target seismic data item; and processing the conditioning input that characterizes the target seismic data item using a diffusion model that comprises a denoising neural network to generate the target seismic data item.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] SYNTHETIC DATA GENERATION MODELS

[0002] CROSS-REFERENCE TO RELATED APPLICATIONS

[0003] This application claims priority to U.S. Provisional Application No. 63 / 645,119, filed on May 9, 2024, and U.S. Provisional Application No. 63 / 792,853, filed on April 22, 2025. The disclosure of the prior application is considered part of and is incorporated by reference in the disclosure of this application.

[0004] BACKGROUND

[0005] This specification relates to generating synthetic data items using machine learning models.

[0006] For example, seismic data items can be generated as the output of a seismic survey in the real-world. A seismic survey uses seismic waves to create seismic data items (e.g., seismic images) of the earth through analysis of vibrations from those seismic waves. The seismic survey can identify subsurface discontinuities (e.g., faults), layering, probable rock structures, etc. The seismic survey is conducted by deploying an array of energy sources and an array of receivers in an area of interest. The array of energy sources can be, e.g., dynamite, a specialized air gun or a seismic vibrator. The energy that travels within the subsurface of the earth are the seismic waves, and the seismic waves are recorded at specific locations on the surface of the earth by the receivers (e.g., geophones or hydrophones).

[0007] A seismic survey simulation uses a synthetic seismic data generator that models the earth properties to generate synthetic seismic images that simulate the seismic images from real seismic surveys. Because the earth is complex and seismic surveying processes are limited in sampling in space and time, it is difficult to generate realistic synthetic seismic images using a synthetic seismic data generator, especially for a three-dimensional (3D) seismic survey.

[0008] SUMMARY

[0009] This specification describes a system implemented as computer programs on one or more computers that uses a generative neural network to generate realistic synthetic data items.

[0010] In general, one innovative aspect of the subject matter described in this specification can be embodied in methods performed by one or more computers that include the actions of obtaining a conditioning input that characterizes a target seismic data item; and processing the conditioning input that characterizes the target seismic data item using a diffusion model that comprises a denoising neural network to generate the target seismic data item. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0011] These and other implementations can each optionally include one or more of the following features.

[0012] In some implementations, the target seismic data item comprises a 2D seismic image.

[0013] In some implementations, the target seismic data item comprises a 3D seismic image.

[0014] In some implementations, the target seismic data item has a first size and wherein the denoising neural network has been trained on training data of a second size.

[0015] In some implementations, the first size and the second size are different sizes.

[0016] In some implementations, the conditioning input comprises a seismic data item that has a same modality as the target seismic data item.

[0017] In some implementations, obtaining the conditioning input comprises generating the seismic data item using a synthetic seismic data generator.

[0018] In some implementations, the method includes adding an amount of noise to the seismic data item.

[0019] In some implementations, the amount of noise is determined based on a parameter.

[0020] In some implementations, the conditioning input comprises any one or more of: a seismic image; one or more embeddings for a seismic image; log data; one or more embeddings for log data; well log data; one or more embeddings for well log data; metadata for a seismic survey; one or more embeddings for metadata for a seismic survey; audio data; one or more embeddings for audio data; text; or one or more embeddings for text.

[0021] In some implementations, processing the conditioning input that characterizes the target seismic data item using a diffusion model that comprises a denoising neural network to generate the target seismic data item comprises: at each of a plurality of reverse diffusion iterations, updating a representation of the target seismic data item by denoising the representation using the denoising neural network with the denoising neural network conditioned on the conditioning input. In some implementations, the diffusion model comprises an upsampling neural network configured to upsample a representation to generate the target seismic data item.

[0022] In some implementations, the target seismic data item is of a first data modality and wherein the upsampling neural network has been trained on training data of the first data modality.

[0023] In some implementations, the conditioning input comprises a text sequence.

[0024] In some implementations, the conditioning input includes an input seismic data item, and the text sequence specifies a modification to be applied to the conditioning input.

[0025] In some implementations, the modification requires identifying a particular class of objects or events in the input seismic data item.

[0026] In some implementations, the conditioning input is a first conditioning input, the method comprising: generating a second conditioning input that characterizes the input seismic data item; and processing the second conditioning input and the target seismic data item that characterizes the input seismic data item using the diffusion model to generate one or more reconstructions of the input seismic data item.

[0027] In some implementations, the method includes: determining whether to provide the target seismic data item as output based on a similarity between the one or more reconstructions and the input seismic data item.

[0028] In some implementations, the target seismic data item is one of a plurality of candidate target seismic data items.

[0029] In some implementations, the method includes filtering the plurality of candidate target seismic data items.

[0030] In some implementations, filtering the plurality of candidate target seismic data items comprises removing a subset of the plurality of candidate target seismic data items based on a respective embedding for each of the plurality of candidate target seismic data items.

[0031] In some implementations, the method includes performing a downstream task using the filtered data set.

[0032] In general, another innovative aspect of the subject matter described in this specification can be embodied in methods performed by one or more computers that include the actions of obtaining an initial data set, the initial data set comprising a plurality of synthetic seismic data items; generating a denoised data set by denoising at least a subset of the plurality of synthetic seismic data items from the initial data set; and performing a downstream task using the denoised data set.

[0033] Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0034] These and other implementations can each optionally include one or more of the following features.

[0035] In some implementations, the denoising is defined by a plurality of parameters, and wherein generating the denoised data set comprises: performing an optimization process to determine values for the parameters that optimize an objective; and identifying, as the denoised data set, a data set that has been generated by denoising the at least a subset of the plurality of seismic data items in accordance with the determined values of the parameters.

[0036] In general, another innovative aspect of the subject matter described in this specification can be embodied in methods performed by one or more computers that include the actions of receiving a first conditioning input of a first modality that characterizes a target seismic data item; processing the conditioning input of the first modality that characterizes the target seismic data item using an encoder neural network that is specific to the first modality to generate a first embedding of the first conditioning input; determining, based at least on the first embedding, a target embedding of the target seismic data item; and processing the target embedding of the target seismic data item using a decoder neural network to generate the target seismic data item. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0037] These and other implementations can each optionally include one or more of the following features.

[0038] In some implementations, the method includes receiving one or more additional conditioning inputs that are each of a respective additional modality and that each characterize the target seismic data item; and for each additional conditioning input, processing the additional conditioning input that characterizes the target seismic data item using a first encoder neural network that is specific to the respective additional modality of the additional conditioning input to generate a respective additional embedding of the additional conditioning input; wherein determining, based at least on the first embedding, a target embedding of the target seismic data item comprises determining the target embedding of the target seismic data item using the first embedding and each additional embedding.

[0039] In some implementations, the first modality is one of a seismic imaging modality; a log data modality; a well log data modality; a text modality; or one or more embeddings.

[0040] In some implementations, the respective additional modalities include one or more of: the seismic imaging modality; the log data modality; the well log data modality; the text modality; or the one or more embeddings.

[0041] In some implementations, determining, based at least on the first embedding, a target embedding of the target seismic data item comprises: using, as the target embedding of the target seismic data item, the first embedding.

[0042] In some implementations, determining the target embedding of the target seismic data item using the first embedding and each additional embedding comprises: combining the first embedding and each additional embedding to generate the target embedding.

[0043] In some implementations, the first encoder neural network and the respective additional encoder neural networks have been jointly trained.

[0044] In some implementations, the first encoder neural network and the respective additional encoder neural networks have been jointly trained through contrastive learning.

[0045] In some implementations, the first encoder neural network has been trained jointly with an encoder neural network that encodes seismic data items to generate embeddings of the seismic data items.

[0046] In some implementations, the decoder neural network has been trained on a seismic data item reconstruction objective after the joint training.

[0047] In some implementations, the decoder neural network has been trained on a seismic data item reconstruction objective during the joint training.

[0048] In some implementations, the method includes: processing the target embedding using an additional decoder neural network to generate additional data characterizing the target seismic data item.

[0049] In some implementations, the additional data comprises a text description or log data.

[0050] In general, another innovative aspect of the subject matter described in this specification can be embodied in methods performed by one or more computers that include the actions of obtaining an initial data set, the initial data set comprising a plurality of synthetic seismic data items; generating a fdtered data set, comprising removing a subset of the plurality of synthetic seismic data items from the initial data set; and performing a downstream task using the filtered data set. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0051] These and other implementations can each optionally include one or more of the following features.

[0052] In some implementations, performing a downstream task using the filtered data set comprises: training a neural network to perform the downstream task.

[0053] In some implementations, generating the filtered data set comprises: applying one or more filtering criteria to each of the plurality of synthetic seismic data items to determine whether to remove the data item from the initial data set.

[0054] In general, another innovative aspect of the subject matter described in this specification can be embodied in methods performed by one or more computers that include the actions of obtaining a conditioning input that characterizes a target data item; and processing the conditioning input that characterizes the target data item using a diffusion model that comprises a denoising neural network to generate the target data item. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0055] These and other implementations can each optionally include one or more of the following features.

[0056] In some implementations, the target data item is a waveform or an image.

[0057] In some implementations, the target data item is a seismic data item, a medical data item, or a materials science data item.

[0058] In some implementations, the target data item has a first size and wherein the denoising neural network has been trained on training data of a second size.

[0059] In some implementations, the first size and the second size are different sizes. In some implementations, the conditioning input comprises any one or more of: data representing an input data item, data representing text, data representing metadata, or one or more embeddings.

[0060] In some implementations, the input data item has a same modality as the target data item.

[0061] In some implementations, obtaining the conditioning input comprises generating the input data item using a synthetic data generator.

[0062] In some implementations, the synthetic data generator is configured to generate synthetic medical data items.

[0063] In some implementations, the synthetic data generator is configured to generate synthetic materials science data items.

[0064] In some implementations, wherein the synthetic data generator is configured to generate synthetic seismic data items.

[0065] In some implementations, the method includes adding an amount of noise to the input data item.

[0066] In some implementations, the amount of noise is determined based on a parameter.

[0067] In some implementations, processing the conditioning input that characterizes the target data item using a diffusion model that comprises a denoising neural network to generate the target data item comprises: at each of a plurality of reverse diffusion iterations, updating a representation of the target data item by denoising the representation using the denoising neural network with the denoising neural network conditioned on the conditioning input.

[0068] In some implementations, the diffusion model comprises an upsampling neural network configured to upsample a representation to generate the target data item.

[0069] In some implementations, the target data item is of a first data modality and wherein the upsampling neural network has been trained on training data of the first data modality.

[0070] In some implementations, the conditioning input comprises a text sequence.

[0071] In some implementations, the conditioning input includes an input data item, and the text sequence specifies a modification to be applied to the conditioning input.

[0072] In some implementations, the modification requires identifying a particular class of objects or events in the input data item. In some implementations, the conditioning input is a first conditioning input, the method comprising: generating a second conditioning input that characterizes the input data item; and processing the second conditioning input and the target data item that characterizes the input data item using the diffusion model to generate one or more reconstructions of the input data item.

[0073] In some implementations, the method includes: determining whether to provide the target data item as output based on a similarity between the one or more reconstructions and the input data item.

[0074] In some implementations, the target data item is one of a plurality of candidate target data items.

[0075] In some implementations, the method includes filtering the plurality of candidate target data items.

[0076] In some implementations, filtering the plurality of candidate target data items comprises removing a subset of the plurality of candidate target data items based on a respective embedding for each of the plurality of candidate target data items.

[0077] In some implementations, the method includes performing a downstream task using the filtered data set.

[0078] In general, another innovative aspect of the subject matter described in this specification can be embodied in methods performed by one or more computers that include the actions of obtaining an initial data set, the initial data set comprising a plurality of synthetic data items; generating a denoised data set by denoising at least a subset of the plurality of synthetic data items from the initial data set; and performing a downstream task using the denoised data set. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0079] These and other implementations can each optionally include one or more of the following features. In some implementations, the denoising is defined by a plurality of parameters, and wherein generating the denoised data set comprises: performing an optimization process to determine values for the parameters that optimize an objective; and identifying, as the denoised data set, a data set that has been generated by denoising the at least a subset of the plurality of data items in accordance with the determined values of the parameters.

[0080] In general, another innovative aspect of the subject matter described in this specification can be embodied in methods performed by one or more computers that include the actions of receiving a first conditioning input of a first modality that characterizes a target data item; processing the conditioning input of the first modality that characterizes the target data item using an encoder neural network that is specific to the first modality to generate a first embedding of the first conditioning input; determining, based at least on the first embedding, a target embedding of the target data item; and processing the target embedding of the target data item using a decoder neural network to generate the target data item. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0081] These and other implementations can each optionally include one or more of the following features.

[0082] In some implementations, the target data item is a seismic data item, a medical data item, or a materials science data item.

[0083] In some implementations, the method includes: receiving one or more additional conditioning inputs that are each of a respective additional modality and that each characterize the target data item; and for each additional conditioning input, processing the additional conditioning input that characterizes the target data item using a first encoder neural network that is specific to the respective additional modality of the additional conditioning input to generate a respective additional embedding of the additional conditioning input; wherein determining, based at least on the first embedding, a target embedding of the target data item comprises determining the target embedding of the target data item using the first embedding and each additional embedding. In some implementations, the first modality is one of: a seismic imaging modality; a medical imaging modality; a materials science imaging modality; a log data modality; a well log data modality; a text modality; or one or more embeddings.

[0084] In some implementations, the respective additional modalities include one or more of: the seismic imaging modality; the medical imaging modality; the materials science imaging modality; the log data modality; the well log data modality; the text modality; or the one or more embeddings.

[0085] In some implementations, determining, based at least on the first embedding, a target embedding of the target data item comprises: using, as the target embedding of the target data item, the first embedding.

[0086] In some implementations, determining the target embedding of the target data item using the first embedding and each additional embedding comprises: combining the first embedding and each additional embedding to generate the target embedding.

[0087] In some implementations, the first encoder neural network and the respective additional encoder neural networks have been jointly trained.

[0088] In some implementations, the first encoder neural network and the respective additional encoder neural networks have been jointly trained through contrastive learning.

[0089] In some implementations, the first encoder neural network has been trained jointly with an encoder neural network that encodes data items to generate embeddings of the data items.

[0090] In some implementations, the decoder neural network has been trained on a data item reconstruction objective after the joint training.

[0091] In some implementations, the decoder neural network has been trained on a data item reconstruction objective during the joint training.

[0092] In some implementations, the method includes: processing the target embedding using an additional decoder neural network to generate additional data characterizing the target data item.

[0093] In some implementations, the additional data comprises a text description or log data.

[0094] In general, another innovative aspect of the subject matter described in this specification can be embodied in methods performed by one or more computers that include the actions of obtaining an initial data set, the initial data set comprising a plurality of synthetic data items; generating a filtered data set, comprising removing a subset of the plurality of synthetic data items from the initial data set; and performing a downstream task using the filtered data set. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the methods.

[0095] These and other implementations can each optionally include one or more of the following features.

[0096] In some implementations, performing a downstream task using the filtered data set comprises: training a neural network to perform the downstream task.

[0097] In some implementations, generating the filtered data set comprises: applying one or more filtering criteria to each of the plurality of synthetic data items to determine whether to remove the data item from the initial data set.

[0098] Particular embodiments of the subject matter described in this specification can be implemented so as to realize one or more of the following advantages.

[0099] Machine learning models can be trained to perform machine learning tasks such as classification, segmentation, etc. Training a machine learning model to perform machine learning tasks requires a large amount of training data. In some domains, manually curating, e.g., labeling, a set of training data can be inefficient or infeasible. For example, obtaining a large number of data items, e.g., through sensing methods or techniques, can be inefficient. As another example, obtaining a large number of data items that exhibit features appropriate for the machine learning task can be infeasible, e.g., if the features are rare in the real-world.

[0100] As a particular example, machine learning models can be trained to perform seismic data item processing tasks, such as seismic characterization tasks, identifying geological features depicted in seismic data items, fault picking, facies classification, generating seismic data items, etc. Training a machine learning model to perform a seismic data item processing task requires a large amount of training data. Manually curating a set of training data can be inefficient or infeasible. For example, manually generating a set of training data can include performing seismic surveying to generate seismic data items, for which generating a large number can be infeasible. Manually generating a set of training data can also involve manually labeling the seismic data items, which can be inefficient.

[0101] The system described in this specification can efficiently generate training data for training machine learning models to perform machine learning tasks, such as seismic data item processing tasks. For example, the system can use one or more generative machine learning models to generate a synthetic seismic data item conditioned on at least a conditioning input. The system can train a machine learning model on the generated training data, resulting in better performance at inference compared to a machine learning model trained on existing data sets. For example, training the machine learning model on a larger number and greater variation of training examples allows the machine learning model to generalize better to previously unseen inputs at inference. The system thus enables large scale, automated generation of training examples for training a machine learning model to perform a machine learning task, e.g., a seismic data processing task.

[0102] Some conventional systems for generating synthetic data may generate data with few details or with a low level of realism. The system described in this specification addresses this issue by generating data items with a higher level of detail and / or realism that are closer to real-world data items. For example, for real-world seismic observations, the generated seismic data items can have a higher level of realism, e.g., in terms of texture, resolution, noise characteristics, and geological features. A machine learning model trained on more realistic training data can perform better when processing real -world data at inference. For example, the distribution of the more realistic training data can be closer to the distribution of real-world data than the distribution of synthetic data with a lower level of realism to the distribution of real-world data.

[0103] Some conventional methods for generating synthetic data do not allow for controlling specific properties of the generated data items. The system described in this specification addresses this issue by providing for control over properties of the generated data items. For example, the conditioning input can include data of different modalities, e.g., seismic data items, text data, log data, etc., on which the system is conditioned during generation of the seismic data items. For example, the conditioning input can characterize geological information, e.g., geological features, geographic information, source type, acquisition information, e.g., arrangement and movement of sources and receivers (e.g., nodal, cable-based, wide-azimuth), for the generated seismic data item. Thus, the system can provide for realistic and controlled data item generation. In some cases, the system can enable rapid scenario testing.

[0104] In some implementations, the conditioning input can include an initial synthetic data item. The initial synthetic data item can have been generated using a synthetic data generator that does not require a large amount of computing resources, but that does not generate data items that have a large number of details or a high level of realism. The system can generate multiple data items for each initial synthetic data item that each have a higher level of realism, allowing for the generation of multiple data items for each initial synthetic data item, and the generation of a large number of data items, without requiring the consumption of computing time and resources that would otherwise be required for generating detailed data items manually or through a physics-based technique. For example, generated seismic data items can retain the style depicted in the initial synthetic seismic data items, while mimicking more closely the geology of real-world images.

[0105] Furthermore, in some implementations, the system is configured to generate 3D data items. In example implementations, the system can generate 3D data items of large-scale, e.g., of a large size or resolution. For example, the system can include a diffusion model that denoises patches of an input 3D volume of a large size and reconstructs the denoised patches to generate the 3D data item of the large size. The system can thus generate realistic 3D data items of large sizes, allowing a machine learning model trained on more realistic training data of large sizes to perform better when processing real -world data of large sizes at inference.

[0106] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.

[0107] BRIEF DESCRIPTION OF THE DRAWINGS

[0108] FIG. l is a block diagram of an example seismic data generation system.

[0109] FIG. 2 is a flow diagram of an example process for generating a seismic data item.

[0110] FIG. 3 is a flow diagram of an example process for generating a denoised data set.

[0111] FIG. 4 is a flow diagram of another example process for generating a seismic data item. FIG. 5 is a flow diagram of an example process for generating a filtered data set.

[0112] FIGS. 6A-6B show example seismic data items.

[0113] FIG. 7 shows the performance of an example seismic data generation system.

[0114] DETAILED DESCRIPTION FIG. 1 shows an example seismic data generation system 100. The seismic data generation system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.

[0115] The seismic data generation system 100 can use one or more generative neural networks 110 to generate realistic seismic data items.

[0116] A seismic data item, as used in this specification, can be any appropriate data item that characterizes a region of the subsurface of the earth. For example, the seismic data item can be a one-dimensional waveform, two-dimensional seismic image generated through a seismic survey, or a three-dimensional seismic image generated through a seismic survey, a distributed acoustic sensing (DAS) image, data generated through other data acquisition techniques, and so on. For example, data generated through other data acquisition techniques can include structured well log or other log data. For example, seismic data items can record measurements such as gamma ray, resistivity, temperature, pressure, gravity, magnetic information, etc. Example seismic data items are described below with reference to FIGS. 6A- 6B.

[0117] In example implementations, a seismic data item can be a derivative product of another seismic data item. For example, the seismic data item can be a geological map or a fault model. For example, the seismic data item can be derived using at least one or more seismic images. As another example, the fault model can include a mask that identifies one or more geological features of a seismic image.

[0118] A more realistic seismic data item can have greater realism in terms of, for example, texture, resolution, noise characteristics, and / or geological features. As an example, the style of a data item can change over different portions corresponding to different depths.

[0119] In example implementations, e.g., for seismic data items that, in the real-world, are captured by sensors or require processing, the system can generate data items that are similar to real-world data items. For example, real -world data items can include features such as artifacts or missing data from sensors, the data sensing process, or from processing. For example, generating real-world log data can include obtaining data from one or more sensors, and processing the sensor data, e.g., by performing calibration, corrections, etc., resulting in noise or missing data. As another example, a seismic image can be processed by performing migration or other extended processing, resulting in noise or missing data. The system can thus generate data items that depict similar features to those depicted in real-world data items.

[0120] For example, the seismic data generation system 100 can generate a seismic data item 112 conditioned on a conditioning input 102. The conditioning input 102 characterizes the target seismic data item 112. For example, the conditioning input 102 can characterize geological information, e.g., geological features, geographic information, etc., source type, and / or acquisition information, e.g., arrangement and movement of sources and receivers (e.g., nodal, cable-based, wide-azimuth), for the target seismic data item 112.

[0121] The conditioning input can include data representing one or more inputs such as an input seismic data item, a text sequence, an audio input, or other structured data or metadata, e.g., for a synthetic survey. In example implementations, the data representing the input seismic data item, the text sequence, or the audio input can include the input seismic data item, the text sequence, the audio input, or the structured data or metadata.

[0122] In example implementations, the data representing the input seismic data item, the text sequence, the audio input, or the structured data or metadata can include one or more embeddings representing the input seismic data item, the text sequence, the audio input, or the structured data or metadata. For example, the system can generate the one or more embeddings using an appropriate encoder for each type of data.

[0123] In example implementations, the text sequence can describe the input seismic data item. For example, the text sequence can include local and global properties, e.g., general geology of the region, a general description of features such as salt domes, or a description of a salt dome depicted in the input seismic data item. In example implementations, the text sequence can describe one or more patches of the input seismic data item. For example, the text sequence can describe how a particular geological property is distributed in the input seismic data item.

[0124] In example implementations, the text sequence can be obtained from documents, e.g., geological reports, log data, velocity models, subsurface models.

[0125] In example implementations, the text sequence can be obtained from an audio input that represents speech. For example, the system can process the audio input using a machine learning model configured to perform speech-to-text to generate the text sequence.

[0126] In example implementations, the system can generate the text sequence by obtaining a text description and processing the text description using a language model neural network to generate a modified text description. For example, the modified text description can have more detail, be formatted in a particular manner appropriate for processing by the generative neural network 110, or can paraphrase the information in the text description. In example implementations, the system can obtain the text description from an audio input. For example, the system can process the audio input using a machine learning model configured to perform speech-to-text to generate the text description.

[0127] For example, the system can provide a seismic data item that identifies a particular depth range in a well log as sandy shale, and an instruction to generate a description of the strata. In example implementations, the instruction can include text that instructs the language model neural network to describe the strata as a geologist would do, taking into account features such as neighboring strata, provinces, etc. In example implementations, the instruction can include text that instructs the language model neural network to describe local and global properties.

[0128] In example implementations, the system can obtain multiple text sequences for the same input seismic data item. For example, the system can use the language model neural network to generate multiple different text sequences. The system can generate multiple seismic data items for the same input seismic data item and different text sequences. The system can thus generate a larger amount of training data compared to using a single text sequence.

[0129] In example implementations, the text sequence can specify a modification to be applied to the input seismic data item. For example, the modification can include application of the characterization of the target seismic data item according to the conditioning input. As an example, the modification can include the introduction of one or more geological features. For example, the text sequence can include “3 normal faults, non-intersecting” to guide the generation of the seismic data item 112 to be visually similar to the input seismic data item, but with the introduction of depicting faulting.

[0130] The input seismic data item can be, for example, a real-world seismic data item or a synthetic seismic data item.

[0131] In example implementations, the system can obtain the input seismic data item using a synthetic seismic data generator 150. The synthetic seismic data generator can be, for example, a simulation model, a mathematical model, a machine learning model, etc. that is configured to generate synthetic seismic data. The synthetic seismic data generated by the synthetic seismic data generator 150 can have a lower level of detail, realism, or both, than the seismic data generated by the system 100 and real-world seismic data. For example, a seismic image generated by the synthetic seismic data generator 150 may include fewer layers, fewer details in a particular strata, fewer imaging artifacts, less heterogeneity in structure, and / or less noise, than a seismic image generated by the system 100 and real-world seismic images.

[0132] The seismic data generation system 100 can generate the seismic data item 112 that is characterized by the conditioning input 102 using the one or more generative neural networks 110.

[0133] In example implementations, the generative neural network 110 can include a diffusion model that includes a denoising neural network. Generating seismic data items using a diffusion model is described in more detail below with reference to FIG. 2.

[0134] In example implementations, the generative neural network 110 can include multiple different modality-specific encoder neural networks that have been jointly trained, e.g., through self-supervised learning, e.g., contrastive learning. Generating seismic data items using multiple different modality-specific encoder neural networks is described in more detail below with reference to FIG. 4.

[0135] In example implementations, the system can optimize parameters for the generation of the seismic data item 112. Optimizing parameters for the generation of the seismic data item 112 is described below in further detail with reference to FIG. 3.

[0136] The system can generate a set of output training data for training a machine learning model to perform a seismic data item processing task by generating multiple seismic data items such as the seismic data item 112. The output training data for training the machine learning model to perform the seismic data item processing task can include multiple training examples that each include a training input and a target output appropriate for the seismic data item processing task.

[0137] In example implementations, one or more of the training inputs can include a seismic data item generated by the system 100. In these implementations, the target output can include one or more labels for the input seismic data item used to generate the seismic data item 112. As an example, the seismic data item processing task can be a feature identification task. Each training input can include a seismic data item generated by the system 100. The corresponding target output can include one or more labels for the input seismic data item that identify one or more geological features depicted in the input seismic data item, e.g., as a mask or as a class label.

[0138] In example implementations, the training input can include the one or more labels, and the target output can include the target seismic data item 112. As an example, the seismic data item processing task can be a seismic data item generation task. Each training input can include one or more labels for an input seismic data item that identify one or more geological features depicted in the input seismic data item, e g., as a mask or as a class label. The corresponding target output can include a seismic data item generated by the system 100 for the input seismic data item.

[0139] For each training example of the output training data that includes a seismic data item 112 generated from an input seismic data item, the system 100 can use the same labels as the input seismic data item, without having to obtain new labels for the generated seismic data item 112. The system can thus generate training examples that re-use existing target outputs for the input seismic data items, with training inputs that are more realistic than the input seismic data items.

[0140] In example implementations, the set of output training data can include one or more real -world data items. For example, the set of output training data can include a combination of real-world data items and synthetic data items generated by the system 100.

[0141] The one or more labels for the input seismic data item can specify geological features depicted in the input seismic data item, e.g., a fault, a geobody, salt, a mass transport event, a river channel, etc. In example implementations, the one or more labels can specify other features of the seismic data item, e.g., the type of seismic survey, survey parameters, noise conditions, and / or other information about the volume or subvolume depicted by the seismic data item, e.g., depth, a description for the seismic data item, etc. In example implementations, the one or more labels can include a mask, e.g., a segmentation mask, which identifies pixels of a 2D seismic data item, or voxels of a 3D seismic data item, that depict the geological feature.

[0142] In example implementations, the system can generate a set of output training data for a different task than the task of the existing target outputs. For example, the system can modify the one or more labels. As an example, the existing target outputs can be for a segmentation task, and the modified target outputs can be for a classification task. For example, the system can modify a label that includes a pixel-level or voxel-level mask that identifies pixels or voxels, respectively, depicting a geological feature, to be a label that identifies that the seismic data item, or a region of the seismic data item, depicts the geological feature. In example implementations, training a machine learning model to perform classification rather than segmentation requires a smaller amount of computing time and resources. In example implementations, performing inference using a machine learning model that performs classification rather than segmentation requires a smaller amount of computing time and resources compared to a machine learning model that performs segmentation.

[0143] In some implementations, the generative neural network 110 can be configured to generate different types of seismic data items and / or different styles of seismic data items. For example, the system 100 can specify the type and / or style of seismic data item to generate in the conditioning input 102. The generative neural network 110 can have been trained on a variety of types and / or styles of seismic data items.

[0144] In some implementations, the generative neural network 110 can be configured to generate a particular type and / or style of seismic data item. For example, the generative neural network 110 can have been trained on training data that includes a particular type and / or style of seismic data item.

[0145] Although this specification describes generating seismic data items, the techniques described in this specification can be used to generate any appropriate type of data item for any appropriate domain, e.g., ID data, e.g., waveforms, data of higher dimensions, e.g., 2D images, images of multiple dimensions, e g., 3D or 4D dense images, medical images, materials science images, or other sensor-based imagery, etc. The system can thus generate more realistic data items for any of a variety of domains, e.g., in label-sparse domains. In particular, the system can generate more realistic data items given a conditioning input that includes an input data item for a domain.

[0146] In example implementations, the input data item can be a synthetic input data item, e.g., generated using a synthetic data generator. In some cases the synthetic input data item can be an abstraction of a real-world data item. In some aspects, an abstraction of a real-world data item can have or include a sufficient amount of details for a manual annotator to determine one or more labels for the data item, but does not have or include a sufficient amount of details such that a machine learning model can learn to determine one or more labels for the data item. As an example, the system can generate data items that include medical images, e.g., 2D or 3D medical images. The conditioning input can include an input medical image. For example, the input medical image can be a 3D image depicting a human or animal organ (e.g., a kidney). In example implementations, the conditioning input can have been generated using a synthetic medical data generator, e.g., a 3D anatomy modeling program. For example, the system can perform the process described above and with reference to FIG. 2 to generate a more realistic 3D image depicting the kidney.

[0147] In example implementations, the conditioning input can characterize the target data item. For example, the conditioning input can characterize a type of pathology that could appear in the organ, imaging style, acquisition information, body types, and / or information about a patient. The generated 3D image can depict the organ (e.g., kidney) characterized by the conditioning input.

[0148] As another example, the system can generate data items that include materials science images, e.g., 2D or 3D materials science images. The conditioning input can include an input materials science image. For example, the input materials science image can be a 2D image depicting the structure of a material. In example implementations, the conditioning input can have been generated using a synthetic materials science data generator. In example implementations, the conditioning input can have been obtained using inspection cameras. The system can perform the process described above and with reference to FIG. 2 to generate a more realistic 2D image depicting the structure of the material.

[0149] In example implementations, the conditioning input can characterize the target data item. For example, the conditioning input can characterize a defect in the material. The generated 2D image can depict the structure of the material characterized by the conditioning input.

[0150] FIG. 2 is a flow diagram of an example process 200 for generating a seismic data item. For convenience, the process 200 will be described as being performed by a system of one or more computers located in one or more locations. For example, a seismic data generation system, e.g., the seismic data generation system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 200. In particular, the process 200 is an example process for generating a target seismic data item using a diffusion model. The target seismic data item can be, for example, a 2D seismic image, or a 3D seismic image.

[0151] The system obtains a conditioning input that characterizes a target seismic data item (step 202). The conditioning input can include data representing one or more inputs such as a seismic data item, e.g., a seismic image, log data, well log data. The one or more inputs can include audio data, structured data or metadata for a synthetic surveyor text. In example implementations, the data representing the one or more inputs can include one or more embeddings representing each of the one or more inputs. In example implementations, the system can generate the one or more embeddings for each of the one or more inputs, e.g., using an appropriate neural network for the input.

[0152] In example implementations, the conditioning input includes an initial seismic data item that has a same modality as the target seismic data item. In example implementations, the initial seismic data item can be a noisy version of the target seismic data item, e.g., can have at least some pixels that are sampled from a noise distribution. For example, the initial seismic data item can be a noisy version of a real-world seismic data item. In example implementations, the initial seismic data item can be a simplified version of the target seismic data item, e.g., depict fewer details than the target seismic data item.

[0153] In example implementations, the system can generate the initial seismic data item using a synthetic seismic data generator. In example implementations, the system can generate the initial seismic data item based on one or more parameters, as described below with reference to FIG. 3.

[0154] The system can add noise to the real-world seismic data item or the seismic data item generated using the synthetic seismic data generator to generate the initial seismic data item. In example implementations, the system can determine the amount of noise to add to the seismic data item based on a parameter. For example, the parameter can be a strength parameter. In example implementations, the strength parameter can have a default value. In example implementations, the system can determine the strength parameter based on an optimization process, as described below with reference to FIG. 3. In example implementations, the conditioning input includes a text sequence. The text sequence can specify a modification to be applied to the conditioning input, e.g., the seismic data item of the conditioning input.

[0155] As an example, the modification can include identifying a particular class of objects or events in the input seismic data item. The seismic data item generated by the diffusion model can characterize a segmented area in the input seismic data item that identifies the particular class of objects or events. For example, the seismic data item can include a mask of the same format as the input seismic data item that identifies regions in the input seismic data item that depict the particular class of objects or events. For example, the modification can include a prompt to identify geological features, e.g., identify all normal faults. In these examples, the diffusion model can generate a fault model that includes a mask for the normal faults depicted in the input seismic data item.

[0156] As another example, the modification can include modifying one or more properties of the input seismic data item. For example, the one or more properties can include the style of realism added during the process 200. The style can be defined by, for example, noise type, frequency content, specific artifacts, geographic region, etc. As particular examples, the modification can include "add typical Gulf of Mexico salt-related noise," "apply Australian land acquisition footprint," "simulate marine streamer feathering artifacts," or "match Otway Basin character."

[0157] In example implementations, the conditioning input can include embeddings for one or more of the types of data. An “embedding” as used in this specification is a sequence of one or more vectors of numeric values, e.g., floating point values or other values, each vector having a pre-determined dimensionality.

[0158] For example, the data representing text can include an embedding for the text. In example implementations, the system can generate the embedding for the text. For example, the system can process a given text sequence, e.g., using a text encoder neural network to generate the embedding for the text. The text encoder neural network can have any appropriate neural network architecture, e.g., a feedforward architecture, e.g., an encoder-only Transformer neural network, or a recurrent architecture, that allows the neural network to map the natural language sequence of text to the embedding of the text. As an example, the text encoder neural network can include a T5 text encoder, described in further detail in Raffel et al., Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer, arXiv preprint arXiv: 1910.10683 (2019).

[0159] As another example, the conditioning input can include latent representations of seismic data items. The latent representation of a seismic data item can include one or more embeddings for the seismic data item. For example, the conditioning input can include latent representations of log data, e.g., well log data, images, geological maps, fault models, etc. In example implementations, the system can generate the latent representations. For example, the system can process the data using an appropriate encoder for each type of seismic data item. Examples of suitable encoders are described below with reference to FIG. 4.

[0160] The conditioning input can include any appropriate combination of inputs that characterize the target seismic data item. As an example, the conditioning input can include data representing a seismic image and data representing a fault model. In example implementations, the conditioning input can also include data representing a text sequence that specifies a modification to add faults according to the fault model to the seismic image. The target seismic data item can include a modified version of the seismic image that depicts the faults from the fault model.

[0161] The system processes the conditioning input using a diffusion model to generate the target seismic data item (step 204). The diffusion model includes a denoising neural network.

[0162] The diffusion model can have any appropriate diffusion model architecture. For example, the diffusion model can have a U-Net based architecture or a diffusion Transformer based architecture. In example implementations, the diffusion model can be configured to generate 2D images, 3D images, or both.

[0163] For example, the diffusion model can include neural network components for performing diffusion. In example implementations, the diffusion model can use one or more positional embeddings, e.g., sinusoidal time embeddings.

[0164] In example implementations, the diffusion model can include one or more neural network layers. For example, the diffusion model can include convolutional layers, upsampling layers, and downsampling layers. As an example, for 3D images, the diffusion model can include one or more 3D convolutional layers, one or more 3D upsampling layers, and one or more 3D downsampling layers. In example implementations, the diffusion model can include one or more residual blocks, e.g., 2D or 3D residual blocks. In example implementations, the diffusion model can apply normalization, e.g., GroupNorm, and / or activation functions, e.g., SiLU.

[0165] In example implementations, the diffusion model can have a network depth, channel widths, and / or connections.

[0166] In example implementations, the diffusion model can be a pre-trained diffusion model. For example, the diffusion model can be pre-trained on a large number of images. In example implementations, the diffusion model can have been fine-tuned on training data for the seismic data item, e.g., 2D seismic data items.

[0167] In example implementations, the system can generate a 3D image using a diffusion model configured to generate 2D images, also referred to as a 2D diffusion model. For example, the system can process a volume using the 2D diffusion model. The diffusion model can be configured to iteratively update alternating orthogonal slices of the volume (e.g., inline, crossline, inline, and so on) for each iteration.

[0168] In example implementations, the diffusion model can be trained to generate 3D seismic data items. For example, the system or another training system can train the diffusion model on training data that includes real-world 3D seismic data items. The real-world 3D seismic data items can represent properties such as complex, domain-specific textures, noise patterns, frequency content variations, and structural or stratigraphic features. The diffusion model can thus learn to generate 3D seismic data items that have one or more of the properties of real- world 3D seismic data items.

[0169] To generate the target seismic data item, the system can perform step 206 at each of multiple reverse diffusion iterations. The system can update a representation of the target seismic data item (step 206). For example, the system can update a representation of the target seismic data item by denoising the representation using the denoising neural network with the denoising neural network conditioned on the conditioning input.

[0170] The output of the denoising neural network at the final time step defines the target seismic data item, i.e., is the target seismic data item or can be transformed into the target seismic data item.

[0171] At each iteration, the representation is the current version of the target seismic data item, and the output defines an estimate of the target seismic data item given the current version, e.g., is an estimate of the noise added to the target seismic data item to generate the current version or is the estimate of the target seismic data item given the current version.

[0172] The diffusion model then uses the output at the iteration to update the current version of the target seismic data item, e.g., using any appropriate diffusion model state transition rule, e.g., DDIM (further details of which can be found in J. Song et al., Denoising Diffusion Implicit Models, ICLR 2021, which is hereby incorporated by reference in its entirety), DDPM (further details of which can be found in J. Ho et al., Denoising Diffusion Probabilistic Models, NeurlPS, 2020, which is hereby incorporated by reference in its entirety), or another appropriate state transition rule.

[0173] For example, the output can be an estimate of the noise added to the target seismic data item to generate the current version of the target seismic data item. The diffusion model can determine a mean using the estimate of the noise and the current version of the target seismic data item. The diffusion model can update the current version of the target seismic data item according to a state transition rule that is based on the mean. For example, the diffusion model can sample the updated current version from a Gaussian distribution centered at the mean. In some examples, the diffusion model can combine the updated current version with additional noise.

[0174] As another example, the output can be an estimate of the noise added to the target seismic data item to generate the current version of the target seismic data item. The diffusion model can generate a prediction of the target seismic data item using the estimate of the noise. The diffusion model can update the current version of the target seismic data item according to a state transition rule that is based on a combination of at least the prediction of the target seismic data item and the estimate of the noise. In some examples, the state transition rule can be based on a combination of the prediction of the target seismic data item, the estimate of the noise, and additional noise. In some examples, the combination can be a weighted combination. In some examples, the additional noise can be weighted by a hyperparameter.

[0175] After the last iteration, the system can use the updated version of the target seismic data item as the final estimate of the target seismic data item, i.e., as the target seismic data item that is generated by the system.

[0176] In example implementations, the diffusion model can be a latent diffusion model. In these examples, the diffusion model can process the initial seismic data item using an encoder to generate a latent representation of the target seismic data item. At each iteration, the representation can be the current version of the latent representation of the target seismic data item and the output defines an estimate of the latent representation of the target seismic data item given the current version, e.g., is an estimate of the noise added to the latent representation of the target seismic data item to generate the current version or is the estimate of the latent representation of the target seismic data item given the current version. After the last iteration, the system can use the updated version of the latent representation of the target seismic data item as the final estimate of the latent representation of the target seismic data item. The system can process the final estimate of the latent representation of the target seismic data item using a decoder to generate the target seismic data item. In these examples, the encoder and decoder can be part of an autoencoder.

[0177] In these examples, at each iteration, updating the latent representation can include performing cross-attention over the latent representation and latent representations of inputs in the conditioning input, such as text, log data, etc.

[0178] In some implementations, the system can generate a target seismic data item of a different size than the seismic data item of the conditioning input. For example, the seismic data item of the conditioning input can be a 2D image of 64x64 pixels, and the target seismic data item can be a 2D image of 128x128 pixels.

[0179] In example implementations, the diffusion model can include an upsampling neural network. The upsampling neural network can be configured to upsample a representation to generate the target seismic data item of the different size. The upsampling neural network can have any architecture appropriate for upsampling a representation. In example implementations, the upsampling neural network can include a pretrained upsampling neural network, e.g., for upsampling images.

[0180] As an example, the upsampling neural network can be a second diffusion model. For example, the representation can be initialized to a noisy version of the target seismic data item, i.e., a version that has the same size as the target seismic data item but that includes at least some data elements that are sampled from a noise distribution. The system can generate the target seismic data item using the upsampling neural network over multiple reverse diffusion iterations, in a similar manner to the diffusion model described above. The upsampling neural network can be conditioned on the seismic data item generated by the diffusion model. In example implementations, the upsampling neural network can have been trained to generate seismic data items of a particular modality. For example, the target seismic data item can be of the particular modality, and the upsampling neural network can have been trained on training data of the particular modality. For example, the upsampling neural network can have been fine-tuned to generate seismic data items of the DAS modality.

[0181] The diffusion model can be configured to generate a seismic data item of any size, including sizes smaller or larger than one or more components of the diffusion model was trained to generate. As an example, the diffusion model can be configured to generate 3D images of larger volume compared to conventional diffusion models.

[0182] In example implementations, the system can generate a target seismic data item of a different size than the denoising neural network was trained to generate. For example, the target seismic data item can be a 3D volume of a first size, and the denoising neural network can have been trained on training data that includes 3D volumes of a second size. For example, the first size, e.g., 128x128x128 pixels, and the second size, e.g., 64x64x64 pixels, can be different sizes.

[0183] For example, the diffusion model can generate the seismic data item conditioned on at least a 3D volume of the target size. The 3D volume can be an input seismic data item, e.g., an input 3D seismic image, or a latent representation of the input seismic data item. The diffusion model can divide the 3D volume into overlapping 3D patches.

[0184] The diffusion model can denoise each 3D patch using the denoising neural network. In example implementations, the diffusion model can denoise multiple 3D patches in batches.

[0185] The system can combine, e.g., as a weighted reconstruction, the denoised 3D patches to generate the seismic data item. For example, the system can use a 3D Bartlett window taper to combine the denoised 3D patches in overlapping regions to reduce edge artifacts and provide for smooth transitions between patches. The system can thus generate realistic 3D seismic images of large sizes that preserves geological coherence.

[0186] In some implementations, the conditioning input is a first conditioning input. For example, the conditioning input can include a 3D seismic data item and a text prompt to identify all normal faults in the volume. The target seismic data item can include a mask depicting one or more normal faults in the 3D seismic data item. The system can generate a second conditioning input that characterizes the input seismic data item. For example, the second conditioning input can include a text prompt to recreate the original seismic data item for the mask of the target seismic data item.

[0187] The system can process the second conditioning input and the target seismic data item that characterizes the input seismic data item using the diffusion model to generate one or more reconstructions of the input seismic data item. In example implementations, the system can vary the random seed for the diffusion process to generate a distribution of reconstructions of the input seismic data item.

[0188] In example implementations, the system can determine whether to provide the target seismic data item as output based on a similarity between the one or more reconstructions and the input seismic data item. For example, the similarity can be an embedding similarity or a pixel-wise similarity. The system can thus ensure the quality of the target seismic data item.

[0189] In example implementations, the system can use the target seismic data item as a reinforcement learning input. For example, the system can train the diffusion model or another generative model to generate seismic data items similar to the target seismic data item. For example, the system can train the generative model to maximize a reward that is based on the similarity between a seismic data item generated by the generative model and the target seismic data item.

[0190] The target seismic data item for the first conditioning input can be referred to as the first target seismic data item. In example implementations, the system can generate multiple first target seismic data items, e.g., by varying the random seed.

[0191] In some of these examples, the second conditioning input can include the 3D seismic data item of the first conditioning prompt, and a text prompt to refine the mask of the first target seismic data item. The system can process the second conditioning input and the first target seismic data item to generate a refined target seismic data item. In example implementations, the system can generate multiple candidate refined target seismic data items.

[0192] In example implementations, the system can determine whether to provide one of the first target seismic data items as output based on a similarity between the candidate target seismic data items and the first target seismic data items. For example, the similarity can be an embedding similarity or a pixel-wise similarity. The system can determine to provide a first target seismic data item as output if the similarity between the first target seismic data item and one of the candidate target seismic data items for the first target seismic data item meets a threshold similarity. The system can thus ensure the quality of the first target seismic data item.

[0193] In some implementations, the generated target seismic data item is one of multiple candidate target seismic data items. In example implementations, the system can filter the multiple candidate target seismic data items. For example, the system can remove a subset of the candidate target seismic data items based on a respective embedding for each of the candidate target seismic data items. For example, the system can remove candidate target seismic data items for which the embedding distribution does not meet a threshold similarity with a real-world embedding distribution. The system can perform a downstream task using the filtered data set, e.g., training a machine learning model using the filtered data set.

[0194] FIG. 3 is a flow diagram of an example process for generating a denoised data set. For convenience, the process 300 will be described as being performed by a system of one or more computers located in one or more locations. For example, a seismic data generation system, e.g., the seismic data generation system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 300.

[0195] The system obtains an initial data set (step 302). The initial data set includes multiple synthetic seismic data items.

[0196] In example implementations, one or more of the synthetic seismic data items can have been generated by a synthetic seismic data generator. The synthetic seismic data items can have been generated by the synthetic seismic data generator according to different parameters. In example implementations, the system can use one or more of the parameters for a synthetic seismic data item as a label for the synthetic seismic data item. The initial data set can thus include synthetic data items that are diverse and have a large range in quality.

[0197] The system generates a denoised data set (step 304). For example, the system can denoise at least a subset of the multiple synthetic seismic data items from the initial data set.

[0198] In example implementations, the denoising is defined by multiple parameters. The system can determine one or more parameters for the denoising process based on one or more optimization processes. For example, the system can determine a starting diffusion timestep of the total timesteps of the diffusion process, a number of iterations for the scheduler, a noise schedule used during the diffusion sampling process, etc. For example, the system can perform an optimization process to determine values for the parameters that optimize an objective. The system can identify, as the denoised data set, a data set that has been generated by denoising at least a subset of the seismic data items in accordance with the determined values of the parameters.

[0199] In example implementations, the system can determine one or more properties of the initial synthetic seismic data item. The properties can include, for example, the quality, fidelity, and characteristics of the initial synthetic seismic data item. Characteristics can include, for example, stratigraphy (e.g., number and configuration of rock layers), existence of structural elements (e.g., dips, folding, and faulting), geobodies (e.g., channels or salt domes), the positions of structural elements and / or geobodies, intra-layer variability (e.g., showing how properties change within layers), layer boundary definitions (e.g., sharp or gradual), and an overall level of geologic "noise" or realism. In some cases, the properties of the initial synthetic seismic data item can affect the seismic data item. As an example, a more detailed initial synthetic seismic data item can result in a more realistic seismic data item generated by the diffusion model.

[0200] As another example, the system can transform the initial synthetic seismic data item. For example, the system can add noise to the initial synthetic seismic data item. The system can determine an amount of noise based on a strength parameter. The strength parameter can modulate the influence of the diffusion process relative to the initial synthetic seismic data item.

[0201] In example implementations, the strength parameter can have a value between 0 and 1. In example implementations, a lower strength value can indicate the introduction of less initial noise or limit the diffusion timestep range from which to select the starting timestep. Thus, the diffusion process makes smaller modifications to the initial synthetic seismic data item, and the output of the diffusion model is more similar to the initial synthetic seismic data item compared to a higher strength value.

[0202] In example implementations, a higher strength value can indicate the introduction of more initial noise, or a wider diffusion timestep range from which to select the starting timestep. Thus, the diffusion process has greater freedom to modify the initial synthetic seismic data item based on what the diffusion model learned during training. In example implementations, the seismic data item generated by the diffusion model can have additional realism or details not generated by the synthetic seismic data generator.

[0203] As an example, the objective can be based on a signal to noise ratio of geological features, e.g., a fault or horizon. For example, the system can maximize a signal to noise ratio. A higher signal to noise ratio can indicate a better quality seismic data item.

[0204] As another example, the objective can be based on a probability of geological feature identification by a machine learning model trained on the denoised data set.

[0205] As another example, the objective can be based on embedding space similarity between one or more denoised seismic data items and one or more other seismic data items, e.g., a real- world seismic data item. For example, the system can minimize a similarity measure between a cluster of real-world data and denoised data in the embedding space. As an example, the system can determine the similarity by determining a difference, e.g., distance in the embedding space, between an embedding for a denoised seismic data item and an embedding for a real-world seismic data item. For example, the system can determine the embedding using an encoder neural network. The system can thus reduce the amount of computing resources that would otherwise be required to determine the similarity between two seismic data items, e.g., in pixel space.

[0206] In example implementations, the real-world seismic data items can be broad, regional, or depict particular properties or geologic regimes. Thus the system can optimize the objective to generate data items to be closer to the real-world seismic data items, e.g., that are broad, regional, or depict particular properties or geologic regimes.

[0207] As another example, the objective can be based on a score based on Short-Time Average (STA) or Long-Time Average (LTA).

[0208] In example implementations, the system can perform the optimization process using one or more of a Gaussian process bandit, grid search, random search strategies, Bayesian optimization, genetic programming, or using a neural network as an optimizer.

[0209] The system performs a downstream task using the denoised data set (step 306). The downstream task, for example, can include training a particular machine learning model using the denoised data set, or training a neural network to perform the downstream task. For example, the system can generate training examples that each include a seismic data item of the denoised data set, and the corresponding labels of the initial seismic data item used to generate the seismic data item. The system can thus optimize different initial data sets for different processes, e.g., types of seismic data processing tasks, to improve the performance of the different processes.

[0210] FIG. 4 is a flow diagram of another example process 400 for generating a seismic data item. For convenience, the process 400 will be described as being performed by a system of one or more computers located in one or more locations. For example, a seismic data generation system, e.g., the seismic data generation system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 400.

[0211] In particular, the process 400 is an example process for generating a target seismic data item using modality-specific encoder neural networks. The target seismic data item can be, for example, a 2D seismic image, or a 3D seismic image.

[0212] The system receives a conditioning input (step 402). The first conditioning input characterizes the target seismic data item. The first conditioning input is of a first modality. The first modality can be one of a seismic imaging modality, a log data modality, a well log data modality, a text modality, or one or more embeddings. For example, the one or more embeddings can be embeddings representing the seismic imaging modality, the log data modality, the well log modality, the text modality, or other types of data.

[0213] In example implementations, the system receives one or more additional conditioning inputs that are each of a respective additional modality and that each characterize the target seismic data item. The respective additional modalities can include one or more of the seismic imaging modality, the log data modality, the log data modality, the text modality, or the one or more embeddings.

[0214] The system processes the conditioning input using an encoder neural network to generate a first embedding of the first conditioning input (step 404). For example, the encoder neural network can be a first encoder neural network that is specific to the first modality.

[0215] In example implementations, the first encoder neural network can have been trained jointly with an encoder neural network that encodes seismic data items to generate embeddings of the seismic data items. For example, the encoder neural network can be configured to encode seismic data items of one or more modalities, e.g., the encoder neural network can be one of the respective additional encoder neural networks described below.

[0216] In example implementations where the system receives additional conditioning inputs, the system can, for each additional conditioning input, process the additional conditioning input using a respective additional encoder neural network to generate a respective additional embedding of the additional conditioning input. Each respective additional encoder neural network is specific to the respective additional modality of the additional conditioning input.

[0217] In example implementations where the system receives additional conditioning inputs, the first encoder neural network and the respective additional encoder neural networks can have been jointly trained. For example, the first encoder neural network and the respective additional encoder neural networks can have been jointly trained through contrastive learning, e.g., using a contrastive loss such as the loss in Contrastive Language-Image Pre-training (CLIP).

[0218] In example implementations, each encoder neural network can have been trained individually, and then jointly. For example, each encoder neural network can be trained to encode a particular modality. The encoder neural networks can then be trained jointly, e.g., using a shared loss function such as a contrastive loss.

[0219] In example implementations, one or more of the encoder neural networks can have been trained to process geoscience data. For example, an encoder neural network for text can be pretrained to process text, and further trained, e.g., fine-tuned, on geoscience-related text.

[0220] In example implementations, one or more of the encoder neural networks can be configured to process 3D seismic data. The encoder neural network can be trained to optimize an objective function that preserves sharp boundaries. In example implementations, the objective function can be based on a coherency cube for the 3D seismic data. In example implementations, the objective function can be based on simulating the seismic data item on boundaries, e.g., convolving boundaries of the 3D seismic data with a wavelet.

[0221] The system determines, based on at least the first embedding, a target embedding of the target seismic data item (step 406).

[0222] In example implementations, the system can use the first embedding as the target embedding of the target seismic data item.

[0223] In example implementations where the system receives additional inputs, the system can determine the target embedding using the first embedding and each additional embedding. For example, the system can combine, e.g., average, the first embedding and each additional embedding to generate the target embedding.

[0224] The system processes the target embedding using a decoder neural network to generate the target seismic data item (step 408). In example implementations where the first encoder neural network has been jointly trained with an encoder neural network, the decoder neural network can have been trained on a seismic data item reconstruction objective after the joint training. For example, the seismic data item reconstruction objective can depend at least on a quality of the reconstruction of an output seismic data item for a given embedding of an input seismic data item.

[0225] In example implementations, the decoder neural network can have been trained on a seismic data item reconstruction objective during the joint training. For example, the seismic data item reconstruction objective can depend at least on a quality of the reconstruction of an output seismic data item for a given input seismic data item. In these implementations, during the joint training, the first encoder neural network and the encoder neural network can be trained using the reconstructive objective.

[0226] In some implementations, the system can further process the target embedding using an additional neural network to generate additional data characterizing the target seismic data item. The additional data can include, for example, a text description or log data. The system can use the additional data to generate additional training data for training a machine learning model to perform a seismic data item processing task. For example, the additional training data can be a set of output training data for a different task than the existing labels of the input seismic data items.

[0227] For example, the system can generate an additional training example that includes the additional data and the target seismic data item. In example implementations, the training input can include the additional data and the target output can include the target seismic data item. In example implementations, the training input can include the target seismic data item and the target output can include the additional data.

[0228] FIG. 5 is a flow diagram of an example process 500 for generating a filtered data set. For convenience, the process 500 will be described as being performed by a system of one or more computers located in one or more locations. For example, a seismic data generation system, e.g., the seismic data generation system 100 of FIG. 1, appropriately programmed in accordance with this specification, can perform the process 500.

[0229] The system obtains an initial data set (step 502). The initial data set includes multiple synthetic seismic data items. In example implementations, one or more of the synthetic seismic data items can have been generated by a synthetic seismic data generator. The synthetic seismic data items can have been generated by the synthetic seismic data generator according to different parameters. In example implementations, the system can use one or more of the parameters for a synthetic seismic data item as a label for the synthetic seismic data item. The initial data set can thus include synthetic data items that are diverse and have a large range in quality.

[0230] The system generates a fdtered data set (step 504). For example, the system can remove a subset of the multiple synthetic seismic data items from the initial data set.

[0231] As an example, the system can apply one or more filtering criteria to each of the synthetic seismic data items to determine whether to remove the data item from the initial data set. The filtering criteria can include, for example, heuristics such as unrealistic geology, frequency content, or classical geoscience techniques. For example, the system can remove a synthetic seismic data item in response to determining that the synthetic seismic data item does not satisfy one or more of the heuristics.

[0232] As another example, the filtering criteria can include embedding space similarity to a set of real data. For example, the system can generate embeddings for multiple synthetic seismic data items. The system can compare the embeddings with embeddings of real-world synthetic seismic data items. In example implementations, the system can compare the embeddings with embeddings of real-world synthetic seismic data items from a particular region, or embeddings of real-world synthetic seismic data items from a range of regions. For example, the system can remove synthetic seismic data items in response to determining that the embedding space similarity does not meet a threshold similarity.

[0233] As another example, the filtering criteria can include evaluation of an embedding space distribution. For example, the system can determine the similarity between the embedding space distribution and another embedding space distribution, e.g., a real-world embedding space distribution as described with reference to FIG. 7. For example, the system can remove synthetic seismic data items in response to determining that the embedding space distribution does not satisfy a threshold similarity with the other embedding space distribution. As another example, the system can determine the similarity between an embedding space distribution that represents one seismic data item, e.g., a distribution for the embeddings for the patches of a seismic data item, and an embedding space distribution for another seismic data item, e.g., a distribution for the embeddings for the patches of a real-world seismic data item. For example, the system can remove synthetic seismic data items in response to determining that the embedding space distribution does not satisfy a threshold similarity with the other embedding space distribution.

[0234] As another example, the filtering criteria can include the output from a trained machine learning model. For example, the system can determine whether a fault detector machine learning model detects a fault in a synthetic seismic data item that has a label that indicates the presence of faults. For example, the system can remove synthetic seismic data items in response to determining that the output from the machine learning model does not match the corresponding label for the synthetic seismic data item.

[0235] The system performs a downstream task using the filtered data set (step 506). The downstream task, for example, can include training a particular machine learning model using the filtered data set, or training a neural network to perform the downstream task. As another example, the downstream task can include processing the filtered data set to generate a larger data set, as described above with reference to FIGS. 1-4.

[0236] FIG. 6A shows example seismic data items. FIG. 6A shows 2D images that are slices of a 3D seismic data item, also referred to as a seismic cube.

[0237] For example, the row of images 610 can be slices of an input seismic 3D image. A system such as the system 100 of FIG. 1 can generate a 3D seismic image for the input seismic 3D image.

[0238] The row of images 620 can be slices of a generated 3D seismic data item from the input seismic 3D image according to a strength of 0.7. The row of images 630 can be slices of a generated 3D seismic data item for the input seismic 3D image according to a strength of 0.5. FIG. 6A shows that the slices for the higher strength value have more realism than the slices for the lower strength value. FIG. 6A also shows that the slices for the higher strength value are more different than the slices for the input, compared to the slices for the lower strength value.

[0239] FIG. 6B shows example seismic data items. FIG. 6B shows seismic data items 650, 660, and 670 that are 2D images.

[0240] The seismic data items 650, 660, and 670 are seismic 2D images generated using a system such as the system 100 of FIG. 1, for the input seismic 2D image 640. For example, the image 650 is generated according to a strength of 0.1, the image 660 is generated according to a strength of 0.2, and the image 670 is generated according to a strength of 0.3. FIG. 6B shows that as the strength value increases, the generated image has more realism, and is more different than the input 640. The system described in this specification can just generate seismic data items with variable amounts of realism.

[0241] FIG. 7 shows the performance of an example seismic data generation system. FIG. 7 shows the performance of the example seismic data generation system 100 of FIG. 1 in terms of an embedding density comparison.

[0242] For example, the graph 710 shows the t-SNE embedding density for synthetic seismic data items, e.g., that were generated by a synthetic seismic data generator. The graph 720 shows the t-SNE embedding density for seismic data items that were generated by the seismic data generation system 100. The graph 730 shows the t-SNE embedding density for real-world seismic data items. In example implementations, the system can determine the embedding densities according to the Kullback-Leibler (KL) divergence.

[0243] FIG. 7 shows that the graph 720 and the graph 730 are more similar than the graph 730 and the graph 710. Thus, the seismic data items generated by the seismic data generation system 100 are more similar to real -world seismic data items than the seismic data items generated by a synthetic seismic data generator.

[0244] This specification uses the term “configured” in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions. Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible non transitory storage medium for execution by, or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively, or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, which is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.

[0245] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0246] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a stand alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, sub programs, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.

[0247] In this specification, the term “database” is used broadly to refer to any collection of data: the data does not need to be structured in any particular way, or structured at all, and it can be stored on storage devices in one or more locations. Thus, for example, the index database can include multiple collections of data, each of which may be organized and accessed differently. Similarly, in this specification the term “engine” is used broadly to refer to a softwarebased system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.

[0248] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.

[0249] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e g., magnetic, magneto optical disks, or optical disks. However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.

[0250] Computer readable media suitable for storing computer program instructions and data include all forms of non volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices; magnetic disks, e.g., internal hard disks or removable disks; magneto optical disks; and CD ROM and DVD-ROM disks.

[0251] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.

[0252] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and computeintensive parts of machine learning training or production, i.e., inference, workloads.

[0253] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework, a Microsoft Cognitive Toolkit framework, an Apache Singa framework, or an Apache MXNet framework.

[0254] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication, e.g., a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.

[0255] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.

[0256] In addition to the embodiments of the attached claims and the embodiments described above, the following numbered embodiments are also innovative.

[0257] Embodiment l is a method performed by one or more computers, the method comprising: obtaining a conditioning input that characterizes a target seismic data item; and processing the conditioning input that characterizes the target seismic data item using a diffusion model that comprises a denoising neural network to generate the target seismic data item.

[0258] Embodiment 2 is the method of embodiment 1, wherein the target seismic data item comprises a 2D seismic image.

[0259] Embodiment 3 is the method of embodiment 1, wherein the target seismic data item comprises a 3D seismic image.

[0260] Embodiment 4 is the method of any one of embodiments 1-3, wherein the target seismic data item has a first size and wherein the denoising neural network has been trained on training data of a second size.

[0261] Embodiment 5 is the method of embodiment 4, wherein the first size and the second size are different sizes.

[0262] Embodiment 6 is the method of any one of embodiments 1-5, wherein the conditioning input comprises a seismic data item that has a same modality as the target seismic data item.

[0263] Embodiment 7 is the method of embodiment 6, wherein obtaining the conditioning input comprises generating the seismic data item using a synthetic seismic data generator.

[0264] Embodiment 8 is the method of embodiment 7, comprising adding an amount of noise to the seismic data item.

[0265] Embodiment 9 is the method of embodiment 8, wherein the amount of noise is determined based on a parameter.

[0266] Embodiment 10 is the method of any one of embodiments 1-9, wherein the conditioning input comprises any one or more of: a seismic image; one or more embeddings for a seismic image; log data; one or more embeddings for log data; well log data; one or more embeddings for well log data; metadata for a seismic survey; one or more embeddings for metadata for a seismic survey; audio data; one or more embeddings for audio data; text; or one or more embeddings for text.

[0267] Embodiment 11 is the method of any one of embodiments 1-10, wherein processing the conditioning input that characterizes the target seismic data item using a diffusion model that comprises a denoising neural network to generate the target seismic data item comprises: at each of a plurality of reverse diffusion iterations, updating a representation of the target seismic data item by denoising the representation using the denoising neural network with the denoising neural network conditioned on the conditioning input.

[0268] Embodiment 12 is the method of any one of embodiments 1-11, wherein the diffusion model comprises an upsampling neural network configured to upsample a representation to generate the target seismic data item.

[0269] Embodiment 13 is the method of embodiment 12, wherein the target seismic data item is of a first data modality and wherein the upsampling neural network has been trained on training data of the first data modality.

[0270] Embodiment 14 is the method of any one of embodiments 1-13, wherein the conditioning input comprises a text sequence.

[0271] Embodiment 15 is the method of embodiment 14, wherein the conditioning input includes an input seismic data item, and the text sequence specifies a modification to be applied to the conditioning input.

[0272] Embodiment 16 is the method of embodiment 15, wherein the modification requires identifying a particular class of objects or events in the input seismic data item.

[0273] Embodiment 17 is the method of any one of embodiments 15-16, wherein the conditioning input is a first conditioning input, the method comprising: generating a second conditioning input that characterizes the input seismic data item; and processing the second conditioning input and the target seismic data item that characterizes the input seismic data item using the diffusion model to generate one or more reconstructions of the input seismic data item. Embodiment 18 is the method of embodiment 17, comprising: determining whether to provide the target seismic data item as output based on a similarity between the one or more reconstructions and the input seismic data item.

[0274] Embodiment 19 is the method of any one of embodiments 1-18, wherein the target seismic data item is one of a plurality of candidate target seismic data items.

[0275] Embodiment 20 is the method of embodiment 19, comprising filtering the plurality of candidate target seismic data items.

[0276] Embodiment 21 is the method of embodiment 20, wherein filtering the plurality of candidate target seismic data items comprises removing a subset of the plurality of candidate target seismic data items based on a respective embedding for each of the plurality of candidate target seismic data items.

[0277] Embodiment 22 is the method of any one of embodiments 20-21, comprising performing a downstream task using the filtered data set.

[0278] Embodiment 23 is a system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of embodiments 1-22.

[0279] Embodiment 24 is one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of embodiments 1-22.

[0280] Embodiment 25 is a method performed by one or more computers, the method comprising: obtaining an initial data set, the initial data set comprising a plurality of synthetic seismic data items; generating a denoised data set by denoising at least a subset of the plurality of synthetic seismic data items from the initial data set; and performing a downstream task using the denoised data set.

[0281] Embodiment 26 is the method of embodiment 25, wherein the denoising is defined by a plurality of parameters, and wherein generating the denoised data set comprises: performing an optimization process to determine values for the parameters that optimize an objective; and identifying, as the denoised data set, a data set that has been generated by denoising the at least a subset of the plurality of seismic data items in accordance with the determined values of the parameters.

[0282] Embodiment 27 is a system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of either one of embodiments 25 or 26.

[0283] Embodiment 28 is one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of either one of embodiments 25 or 26.

[0284] Embodiment 29 is a method performed by one or more computers, the method comprising: receiving a first conditioning input of a first modality that characterizes a target seismic data item; processing the conditioning input of the first modality that characterizes the target seismic data item using an encoder neural network that is specific to the first modality to generate a first embedding of the first conditioning input; determining, based at least on the first embedding, a target embedding of the target seismic data item; and processing the target embedding of the target seismic data item using a decoder neural network to generate the target seismic data item.

[0285] Embodiment 30 is the method of embodiment 29, further comprising: receiving one or more additional conditioning inputs that are each of a respective additional modality and that each characterize the target seismic data item; and for each additional conditioning input, processing the additional conditioning input that characterizes the target seismic data item using a first encoder neural network that is specific to the respective additional modality of the additional conditioning input to generate a respective additional embedding of the additional conditioning input; wherein determining, based at least on the first embedding, a target embedding of the target seismic data item comprises determining the target embedding of the target seismic data item using the first embedding and each additional embedding.

[0286] Embodiment 31 is the method of any one of embodiments 29-30, wherein the first modality is one of: a seismic imaging modality; a log data modality; a well log data modality; a text modality; or one or more embeddings. Embodiment 32 is the method of embodiment 31 when dependent on embodiment 30, wherein the respective additional modalities include one or more of: the seismic imaging modality; the log data modality; the well log data modality; the text modality; or the one or more embeddings.

[0287] Embodiment 33 is the method of any one of embodiments 29 or 31, wherein determining, based at least on the first embedding, a target embedding of the target seismic data item comprises: using, as the target embedding of the target seismic data item, the first embedding.

[0288] Embodiment 34 is the method of any one of embodiments 29-33, when dependent on embodiment 30, wherein determining the target embedding of the target seismic data item using the first embedding and each additional embedding comprises: combining the first embedding and each additional embedding to generate the target embedding.

[0289] Embodiment 35 is the method of any one of embodiments 29-34, when dependent on embodiment 30, wherein the first encoder neural network and the respective additional encoder neural networks have been jointly trained.

[0290] Embodiment 36 is the method of embodiment 35, wherein the first encoder neural network and the respective additional encoder neural networks have been jointly trained through contrastive learning.

[0291] Embodiment 37 is the method of any one of embodiments 29-36, wherein the first encoder neural network has been trained jointly with an encoder neural network that encodes seismic data items to generate embeddings of the seismic data items.

[0292] Embodiment 38 is the method of embodiment 37, wherein the decoder neural network has been trained on a seismic data item reconstruction objective after the joint training.

[0293] Embodiment 39 is the method of embodiment 37, wherein the decoder neural network has been trained on a seismic data item reconstruction objective during the joint training.

[0294] Embodiment 40 is the method of any one of embodiments 29-39, further comprising: processing the target embedding using an additional decoder neural network to generate additional data characterizing the target seismic data item.

[0295] Embodiment 41 is the method of embodiment 40, wherein the additional data comprises a text description or log data. Embodiment 42 is a system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of embodiments 29-41.

[0296] Embodiment 43 is one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of embodiments 29-41.

[0297] Embodiment 44 is a method performed by one or more computers, the method comprising: obtaining an initial data set, the initial data set comprising a plurality of synthetic seismic data items; generating a fdtered data set, comprising removing a subset of the plurality of synthetic seismic data items from the initial data set; and performing a downstream task using the filtered data set.

[0298] Embodiment 45 is the method of embodiment 44, wherein performing a downstream task using the filtered data set comprises: training a neural network to perform the downstream task.

[0299] Embodiment 46 is the method of any one of embodiments 44-45, wherein generating the filtered data set comprises: applying one or more filtering criteria to each of the plurality of synthetic seismic data items to determine whether to remove the data item from the initial data set.

[0300] Embodiment 47 is a system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of embodiments 44-46.

[0301] Embodiment 48 is one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of embodiments 44-46. Embodiment 49 is a method performed by one or more computers, the method comprising: obtaining a conditioning input that characterizes a target data item; and processing the conditioning input that characterizes the target data item using a diffusion model that comprises a denoising neural network to generate the target data item.

[0302] Embodiment 50 is the method of embodiment 49, wherein the target data item is a waveform or an image.

[0303] Embodiment 51 is the method of any one of embodiments 49-50, wherein the target data item is a seismic data item, a medical data item, or a materials science data item.

[0304] Embodiment 52 is the method of any one of embodiments 49-51, wherein the target data item has a first size and wherein the denoising neural network has been trained on training data of a second size.

[0305] Embodiment 53 is the method of embodiment 52, wherein the first size and the second size are different sizes.

[0306] Embodiment 54 is the method of any one of embodiments 49-53, wherein the conditioning input comprises any one or more of: data representing an input data item, data representing text, data representing metadata, or one or more embeddings.

[0307] Embodiment 55 is the method of any one of embodiments 49-54, wherein the input data item has a same modality as the target data item.

[0308] Embodiment 56 is the method of embodiment 55, wherein obtaining the conditioning input comprises generating the input data item using a synthetic data generator.

[0309] Embodiment 57 is the method of embodiment 56, wherein the synthetic data generator is configured to generate synthetic medical data items.

[0310] Embodiment 58 is the method of embodiment 56, wherein the synthetic data generator is configured to generate synthetic materials science data items.

[0311] Embodiment 59 is the method of embodiment 56, wherein the synthetic data generator is configured to generate synthetic seismic data items.

[0312] Embodiment 60 is the method of any one of embodiments 55-59, comprising adding an amount of noise to the input data item.

[0313] Embodiment 61 is the method of embodiment 60, wherein the amount of noise is determined based on a parameter. Embodiment 62 is the method of any one of embodiments 49-61, wherein processing the conditioning input that characterizes the target data item using a diffusion model that comprises a denoising neural network to generate the target data item comprises: at each of a plurality of reverse diffusion iterations, updating a representation of the target data item by denoising the representation using the denoising neural network with the denoising neural network conditioned on the conditioning input.

[0314] Embodiment 63 is the method of embodiment 62, wherein the diffusion model comprises an upsampling neural network configured to upsample a representation to generate the target data item.

[0315] Embodiment 64 is the method of embodiment 63, wherein the target data item is of a first data modality and wherein the upsampling neural network has been trained on training data of the first data modality.

[0316] Embodiment 65 is the method of any one of embodiments 49-64, wherein the conditioning input comprises a text sequence.

[0317] Embodiment 66 is the method of embodiment 65, wherein the conditioning input includes an input data item, and the text sequence specifies a modification to be applied to the conditioning input.

[0318] Embodiment 67 is the method of embodiment 66, wherein the modification requires identifying a particular class of objects or events in the input data item.

[0319] Embodiment 68 is the method of any one of embodiments 66-67, wherein the conditioning input is a first conditioning input, the method comprising: generating a second conditioning input that characterizes the input data item; and processing the second conditioning input and the target data item that characterizes the input data item using the diffusion model to generate one or more reconstructions of the input data item.

[0320] Embodiment 69 is the method of embodiment 68, comprising: determining whether to provide the target data item as output based on a similarity between the one or more reconstructions and the input data item.

[0321] Embodiment 70 is the method of any one of embodiments 49-69, wherein the target data item is one of a plurality of candidate target data items.

[0322] Embodiment 71 is the method of embodiment 70, comprising filtering the plurality of candidate target data items. Embodiment 72 is the method of embodiment 71 , wherein filtering the plurality of candidate target data items comprises removing a subset of the plurality of candidate target data items based on a respective embedding for each of the plurality of candidate target data items.

[0323] Embodiment 73 is the method of any one of embodiments 71-72, comprising performing a downstream task using the filtered data set.

[0324] Embodiment 74 is a system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of embodiments 49-73.

[0325] Embodiment 75 is one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of embodiments 49-73.

[0326] Embodiment 76 is a method performed by one or more computers, the method comprising: obtaining an initial data set, the initial data set comprising a plurality of synthetic data items; generating a denoised data set by denoising at least a subset of the plurality of synthetic data items from the initial data set; and performing a downstream task using the denoised data set.

[0327] Embodiment 77 is the method of embodiment 76, wherein the denoising is defined by a plurality of parameters, and wherein generating the denoised data set comprises: performing an optimization process to determine values for the parameters that optimize an objective; and identifying, as the denoised data set, a data set that has been generated by denoising the at least a subset of the plurality of data items in accordance with the determined values of the parameters.

[0328] Embodiment 78 is a system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of either one of embodiments 76 or 77. Embodiment 79 is one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of either one of embodiments 76 or 77.

[0329] Embodiment 80 is a method performed by one or more computers, the method comprising: receiving a first conditioning input of a first modality that characterizes a target data item; processing the conditioning input of the first modality that characterizes the target data item using an encoder neural network that is specific to the first modality to generate a first embedding of the first conditioning input; determining, based at least on the first embedding, a target embedding of the target data item; and processing the target embedding of the target data item using a decoder neural network to generate the target data item.

[0330] Embodiment 81 is the method of embodiment 80, wherein the target data item is a seismic data item, a medical data item, or a materials science data item.

[0331] Embodiment 82 is the method of any one of embodiments 80-81, further comprising: receiving one or more additional conditioning inputs that are each of a respective additional modality and that each characterize the target data item; and for each additional conditioning input, processing the additional conditioning input that characterizes the target data item using a first encoder neural network that is specific to the respective additional modality of the additional conditioning input to generate a respective additional embedding of the additional conditioning input; wherein determining, based at least on the first embedding, a target embedding of the target data item comprises determining the target embedding of the target data item using the first embedding and each additional embedding.

[0332] Embodiment 83 is the method of any one of embodiments 80-82, wherein the first modality is one of: a seismic imaging modality; a medical imaging modality; a materials science imaging modality; a log data modality; a well log data modality; a text modality; or one or more embeddings.

[0333] Embodiment 84 is the method of embodiment 83, when dependent on embodiment 82, wherein the respective additional modalities include one or more of: the seismic imaging modality; the medical imaging modality; the materials science imaging modality; the log data modality; the well log data modality; the text modality; or the one or more embeddings. Embodiment 85 is the method of any one of embodiments embodiment 80 or 83, wherein determining, based at least on the first embedding, a target embedding of the target data item comprises: using, as the target embedding of the target data item, the first embedding.

[0334] Embodiment 86 is the method of any one of embodiments 80-85, when dependent on embodiment 82, wherein determining the target embedding of the target data item using the first embedding and each additional embedding comprises: combining the first embedding and each additional embedding to generate the target embedding.

[0335] Embodiment 87 is the method of any one of embodiments 80-86, when dependent on embodiment 82, wherein the first encoder neural network and the respective additional encoder neural networks have been jointly trained.

[0336] Embodiment 88 is the method of embodiment 87, wherein the first encoder neural network and the respective additional encoder neural networks have been jointly trained through contrastive learning.

[0337] Embodiment 89 is the method of any one of embodiments 80-88, wherein the first encoder neural network has been trained jointly with an encoder neural network that encodes data items to generate embeddings of the data items.

[0338] Embodiment 90 is the method of embodiment 89, wherein the decoder neural network has been trained on a data item reconstruction objective after the joint training.

[0339] Embodiment 91 is the method of embodiment 89, wherein the decoder neural network has been trained on a data item reconstruction objective during the joint training.

[0340] Embodiment 92 is the method of any one of embodiments 80-91, further comprising: processing the target embedding using an additional decoder neural network to generate additional data characterizing the target data item.

[0341] Embodiment 93 is the method of embodiment 92, wherein the additional data comprises a text description or log data.

[0342] Embodiment 94 is a system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of embodiments 80-93. Embodiment 95 is one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of embodiments 80-93.

[0343] Embodiment 96 is a method performed by one or more computers, the method comprising: obtaining an initial data set, the initial data set comprising a plurality of synthetic data items; generating a filtered data set, comprising removing a subset of the plurality of synthetic data items from the initial data set; and performing a downstream task using the filtered data set.

[0344] Embodiment 97 is the method of embodiment 96, wherein performing a downstream task using the filtered data set comprises: training a neural network to perform the downstream task.

[0345] Embodiment 98 is the method of any one of embodiments 96-97, wherein generating the filtered data set comprises: applying one or more filtering criteria to each of the plurality of synthetic data items to determine whether to remove the data item from the initial data set.

[0346] Embodiment 99 is a system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of embodiments 96-98.

[0347] Embodiment 100 is one or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of embodiments 96-98.

[0348] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0349] Similarly, while operations are illustrated in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0350] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims can be performed in a different order and still achieve desirable results. As one example, the processes correspond toed in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

[0351] What is claimed is:

Claims

CLAIMS1. A method performed by one or more computers, the method comprising: obtaining a conditioning input that characterizes a target seismic data item; and processing the conditioning input that characterizes the target seismic data item using a diffusion model that comprises a denoising neural network to generate the target seismic data item.

2. The method of claim 1, wherein the target seismic data item comprises a 2D seismic image.

3. The method of claim 1, wherein the target seismic data item comprises a 3D seismic image.

4. The method of any one of the preceding claims, wherein the target seismic data item has a first size and wherein the denoising neural network has been trained on training data of a second size.

5. The method of claim 4, wherein the first size and the second size are different sizes.

6. The method of any one of the preceding claims, wherein the conditioning input comprises a seismic data item that has a same modality as the target seismic data item.

7. The method of claim 6, wherein obtaining the conditioning input comprises generating the seismic data item using a synthetic seismic data generator.

8. The method of claim 7, comprising adding an amount of noise to the seismic data item.

9. The method of claim 8, wherein the amount of noise is determined based on a parameter.

10. The method of any one of the preceding claims, wherein the conditioning input comprises any one or more of: a seismic image; one or more embeddings for a seismic image; log data; one or more embeddings for log data; well log data; one or more embeddings for well log data; metadata for a seismic survey; one or more embeddings for metadata for a seismic survey; audio data; one or more embeddings for audio data; text; or one or more embeddings for text.

11. The method of any one of the preceding claims, wherein processing the conditioning input that characterizes the target seismic data item using a diffusion model that comprises a denoising neural network to generate the target seismic data item comprises: at each of a plurality of reverse diffusion iterations, updating a representation of the target seismic data item by denoising the representation using the denoising neural network with the denoising neural network conditioned on the conditioning input.

12. The method of any one of the preceding claims, wherein the diffusion model comprises an upsampling neural network configured to upsample a representation to generate the target seismic data item.

13. The method of claim 12, wherein the target seismic data item is of a first data modality and wherein the upsampling neural network has been trained on training data of the first data modality.

14. The method of any one of the preceding claims, wherein the conditioning input comprises a text sequence.

15. The method of claim 14, wherein the conditioning input includes an input seismic data item, and the text sequence specifies a modification to be applied to the conditioning input.

16. The method of claim 15, wherein the modification requires identifying a particular class of objects or events in the input seismic data item.

17. The method of either one of claims 15 or 16, wherein the conditioning input is a first conditioning input, the method comprising: generating a second conditioning input that characterizes the input seismic data item; and processing the second conditioning input and the target seismic data item that characterizes the input seismic data item using the diffusion model to generate one or more reconstructions of the input seismic data item.

18. The method of claim 17, comprising: determining whether to provide the target seismic data item as output based on a similarity between the one or more reconstructions and the input seismic data item.

19. The method of any one of the preceding claims, wherein the target seismic data item is one of a plurality of candidate target seismic data items.

20. The method of claim 19, comprising filtering the plurality of candidate target seismic data items.

21. The method of claim 20, wherein filtering the plurality of candidate target seismic data items comprises removing a subset of the plurality of candidate target seismic data items based on a respective embedding for each of the plurality of candidate target seismic data items.

22. The method of either one of claims 20 or 21, comprising performing a downstream task using the filtered data set.

23. A system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 1-22.

24. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of claims 1-22.

25. A method performed by one or more computers, the method comprising: obtaining an initial data set, the initial data set comprising a plurality of synthetic seismic data items; generating a denoised data set by denoising at least a subset of the plurality of synthetic seismic data items from the initial data set; and performing a downstream task using the denoised data set.

26. The method of claim 25, wherein the denoising is defined by a plurality of parameters, and wherein generating the denoised data set comprises: performing an optimization process to determine values for the parameters that optimize an objective; and identifying, as the denoised data set, a data set that has been generated by denoising the at least a subset of the plurality of seismic data items in accordance with the determined values of the parameters.

27. A system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of either one of claims 25 or 26.

28. One or more non -transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of either one of claims 25 or 26.

29. A method performed by one or more computers, the method comprising: receiving a first conditioning input of a first modality that characterizes a target seismic data item; processing the conditioning input of the first modality that characterizes the target seismic data item using an encoder neural network that is specific to the first modality to generate a first embedding of the first conditioning input; determining, based at least on the first embedding, a target embedding of the target seismic data item; and processing the target embedding of the target seismic data item using a decoder neural network to generate the target seismic data item.

30. The method of claim 29, further comprising: receiving one or more additional conditioning inputs that are each of a respective additional modality and that each characterize the target seismic data item; and for each additional conditioning input, processing the additional conditioning input that characterizes the target seismic data item using a first encoder neural network that is specific to the respective additional modality of the additional conditioning input to generate a respective additional embedding of the additional conditioning input; wherein determining, based at least on the first embedding, a target embedding of the target seismic data item comprises determining the target embedding of the target seismic data item using the first embedding and each additional embedding.

31. The method of either one of claim 29 or claim 30, wherein the first modality is one of: a seismic imaging modality; a log data modality; a well log data modality; a text modality; or one or more embeddings.

32. The method of claim 31, when dependent on claim 30, wherein the respective additional modalities include one or more of: the seismic imaging modality; the log data modality; the well log data modality; the text modality; or the one or more embeddings.

33. The method of claim 29 or claim 31, wherein determining, based at least on the first embedding, a target embedding of the target seismic data item comprises: using, as the target embedding of the target seismic data item, the first embedding.

34. The method of any one of claims 29-33, when dependent on claim 30, wherein determining the target embedding of the target seismic data item using the first embedding and each additional embedding comprises: combining the first embedding and each additional embedding to generate the target embedding.

35. The method of any one of claims 29-34, when dependent on claim 30, wherein the first encoder neural network and the respective additional encoder neural networks have been jointly trained.

36. The method of claim 35, wherein the first encoder neural network and the respective additional encoder neural networks have been jointly trained through contrastive learning.

37. The method of any one of claims 29-36, wherein the first encoder neural network has been trained jointly with an encoder neural network that encodes seismic data items to generate embeddings of the seismic data items.

38. The method of claim 37, wherein the decoder neural network has been trained on a seismic data item reconstruction objective after the joint training.

39. The method of claim 37, wherein the decoder neural network has been trained on a seismic data item reconstruction objective during the joint training.

40. The method of any one of claims 29-39, further comprising: processing the target embedding using an additional decoder neural network to generate additional data characterizing the target seismic data item.41 . The method of claim 40, wherein the additional data comprises a text description or log data.

42. A system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 29-41.

43. One or more non -transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of claims 29-41.

44. A method performed by one or more computers, the method comprising: obtaining an initial data set, the initial data set comprising a plurality of synthetic seismic data items; generating a filtered data set, comprising removing a subset of the plurality of synthetic seismic data items from the initial data set; and performing a downstream task using the filtered data set.

45. The method of claim 44, wherein performing a downstream task using the filtered data set comprises: training a neural network to perform the downstream task.

46. The method of either one of claim 44 or claim 45, wherein generating the filtered data set comprises: applying one or more filtering criteria to each of the plurality of synthetic seismic data items to determine whether to remove the data item from the initial data set.

47. A system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 44-46.

48. One or more non -transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of claims 44-46.

49. A method performed by one or more computers, the method comprising: obtaining a conditioning input that characterizes a target data item; and processing the conditioning input that characterizes the target data item using a diffusion model that comprises a denoising neural network to generate the target data item.

50. The method of claim 49, wherein the target data item is a waveform or an image.

51. The method of any one of claims 49-50, wherein the target data item is a seismic data item, a medical data item, or a materials science data item.

52. The method of any one of claims 49-51, wherein the target data item has a first size and wherein the denoising neural network has been trained on training data of a second size.

53. The method of claim 52, wherein the first size and the second size are different sizes.

54. The method of any one of claims 49-53, wherein the conditioning input comprises any one or more of: data representing an input data item, data representing text, data representing metadata, or one or more embeddings.

55. The method of any one of claims 49-54, wherein the input data item has a same modality as the target data item.

56. The method of claim 55, wherein obtaining the conditioning input comprises generating the input data item using a synthetic data generator.

57. The method of claim 56, wherein the synthetic data generator is configured to generate synthetic medical data items.

58. The method of claim 56, wherein the synthetic data generator is configured to generate synthetic materials science data items.

59. The method of claim 56, wherein the synthetic data generator is configured to generate synthetic seismic data items.

60. The method of any one of claims 55-59, comprising adding an amount of noise to the input data item.

61. The method of claim 60, wherein the amount of noise is determined based on a parameter.

62. The method of any one of claims 49-61, wherein processing the conditioning input that characterizes the target data item using a diffusion model that comprises a denoising neural network to generate the target data item comprises: at each of a plurality of reverse diffusion iterations, updating a representation of the target data item by denoising the representation using the denoising neural network with the denoising neural network conditioned on the conditioning input.

63. The method of claim 62, wherein the diffusion model comprises an upsampling neural network configured to upsample a representation to generate the target data item.

64. The method of claim 63, wherein the target data item is of a first data modality and wherein the upsampling neural network has been trained on training data of the first data modality.

65. The method of any one of claims 49-64, wherein the conditioning input comprises a text sequence.

66. The method of claim 65, wherein the conditioning input includes an input data item, and the text sequence specifies a modification to be applied to the conditioning input.

67. The method of claim 66, wherein the modification requires identifying a particular class of objects or events in the input data item.

68. The method of either one of claims 66 or 67, wherein the conditioning input is a first conditioning input, the method comprising: generating a second conditioning input that characterizes the input data item; and processing the second conditioning input and the target data item that characterizes the input data item using the diffusion model to generate one or more reconstructions of the input data item.

69. The method of claim 68, comprising: determining whether to provide the target data item as output based on a similarity between the one or more reconstructions and the input data item.

70. The method of any one of claims 49-69, wherein the target data item is one of a plurality of candidate target data items.

71. The method of claim 70, comprising filtering the plurality of candidate target data items.

72. The method of claim 71, wherein filtering the plurality of candidate target data items comprises removing a subset of the plurality of candidate target data items based on a respective embedding for each of the plurality of candidate target data items.

73. The method of either one of claims 71 or 72, comprising performing a downstream task using the fdtered data set.

74. A system, comprising: one or more computers; and one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations of the respective method of any one of claims 49-73.

75. One or more non -transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations of the respective method of any one of claims 49-73.