Generating data items using a spreading neural network

JP2026530429APending Publication Date: 2026-09-08ジーディーエム·ホールディング·エルエルシー
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026512125
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-21
Filing Date
2024-08-21
Publication Date
2026-09-08

Smart Images

  • Figure 2026530429000001_ABST
    Figure 2026530429000001_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus (including computer programs encoded on a computer storage medium) for generating data items using a spreading neural network or other generative neural network are disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Cross-reference of related applications This application claims priority to U.S. Provisional Patent Application No. 63 / 533,906, filed on 21 August 2023, the entirety of which disclosures are incorporated herein by reference.

[0002] This specification relates to generating conditioned outputs based on conditioned inputs using neural networks.

[0003] A neural network is a machine learning model that uses one or more layers of nonlinear units to predict an output for a given input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to one or more other layers in the network, i.e., one or more other hidden layers, the output layer, or both. Each layer of the network generates an output from a given input according to the current values ​​of its respective set of parameters. [Overview of the project]

[0004] This specification describes a system implemented as a computer program on one or more computers in one or more locations that generates conditioned output data items based on conditioned inputs.

[0005] Generally, a conditional input characterizes one or more desired properties of a data item, i.e., one or more properties that the final data item generated by the system should have.

[0006] More specifically, the system can generate output data items using a spreading neural network.

[0007] Certain embodiments of the subject matter described herein can be implemented to achieve one or more of the following advantages:

[0008] Compared to conventional techniques that use a spread neural network to generate data items, the described system can modify one or more of the following to improve the quality of output data items generated by the spread neural network: (i) training the spread neural network, (ii) inputs to the spread neural network, or (iii) how the spread neural network is used to generate output data items after training.

[0009] As an example, the system can be configured such that, in each update iteration, the spread neural network processes the spread input for the update iteration, which includes the noisy data item for the update iteration, and generates a denoising output that, unlike other techniques, defines an estimate of the residual error between the analytical estimate of the noise component of the noisy data item and the true noise component of the noisy data item. In other words, other techniques generally use spread neural networks that generate other types of denoising outputs, such as a denoising output that is an estimate of the noise component, or a denoising output that is an estimate of the ground truth data item.

[0010] By generating a denoising output that defines an estimate of the residual error between the analytical estimate of the noise component of a noisy data item and the true noise component of the noisy data item, the quality of the output data item can be improved, especially when high guidance weights for classifier-free guidance are used as part of the generation process. For example, using high guidance weights may result in the generated data item conforming more strongly to the provided conditioned input, which is advantageous, but it may also degrade the overall quality of the generated data item. For example, if the data item is an image, using high guidance weights may result in a highly saturated image. By using the denoising output described above, the system can generate data items that conform to the conditioned input while maintaining overall quality, for example, by reducing the saturation of the generated image to a realistic level.

[0011] As another example, a spread neural network can be configured to receive context embeddings that can represent either multiple different context data items or a single context data item. Context data items are generally of the same type as output data items and are used to guide the generation process. By allowing the user to condition the generation process based on a variable number of context data items, the system can use a spread neural network to generate output data items that accurately reflect the context provided by the user, for example, by matching a specific property of a single data item or by reflecting aggregated properties aggregated across multiple data items.

[0012] As yet another example, the system may use boundary conditioning. When boundary conditioning is used, the system conditions the diffuse neural network based on a lower or upper bound of input scalar values ​​that represent the values ​​of a particular property of the generated data item. That is, rather than requiring the diffuse neural network to generate output data items that have the exact values ​​of a particular property represented by the input scalar values, the system can give the diffuse neural network the flexibility to generate any suitable data item that has the values ​​of a particular property that are appropriately bounded by the scalar values.

[0013] As yet another example, a system can use normalization guidance when generating output data items using a spreading neural network. In particular, when a system uses normalization guidance, at each update iteration, the system can calculate the difference between the conditionally denoised output for the update iteration and the unconditionally denoised output or negative denoised output. The system can then determine the overall denoised output based on the direction rather than the magnitude of the difference. As mentioned above, using high guidance weights in conventional classifier-free guidance allows for good fit with the conditioned input and more consistent data items, but it can lead to a decrease in the quality of the data items. For example, when generating images with a high guidance scale using classifier-free guidance, the generation process may produce images with extremely high color saturation and images that are either very bright or very dark.

[0014] By using normalization guidance, for example, such saturation and the undesirable effects associated with it can be mitigated while maintaining the benefits of a high guidance scale, particularly regarding text-image alignment, as well as image quality and consistency.

[0015] Details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] [Figure 1] FIG. 1 is a diagram of an exemplary data generation system. [Figure 2] FIG. 2 is a flow diagram of an example process for generating a final data item using an analytic estimate. [Figure 3] FIG. 3 is a flow diagram of an example process for conditioning a diffusion neural network based on a variable number of context data items. [Figure 4] FIG. 4 is a flow diagram of an example process for training a diffusion neural network to be effectively conditioned based on multiple context data items. [Figure 5] FIG. 5 is a flow diagram of an example process for conditioning a diffusion neural network based on property value bounds. [Figure 6] FIG. 6 is a flow diagram of an example process for training a generative neural network based on property value bounds. [Figure 7] FIG. 7 is a flow diagram of an example process for determining a final denoised output for a given update iteration using normalized guidance.

[0017] Like reference numbers and symbols in the various drawings indicate like elements. DESCRIPTION OF EMBODIMENTS

[0018] This specification describes a system implemented as a computer program on one or more computers at one or more locations that generates conditioned output data items based on conditioning inputs.

[0019] Generally, a conditional input characterizes one or more desired properties of a data item, i.e., one or more properties that the final data item generated by the system should have.

[0020] The system can be configured to generate one of several conditioned output data items based on one of several conditioned inputs.

[0021] For example, the system can be configured to generate audio data, such as audio waveforms or audio spectrograms, such as Mel spectrograms or spectrograms with different frequency scales.

[0022] In this example, the conditional input may be the text or text features that the audio should represent; that is, the system functions as a text-to-speech machine learning model that converts the text or text features into audio data for the spoken text utterances.

[0023] As another example, a conditional input can identify a desired speaker in the audio, that is, cause the system to generate audio data representing speech by the desired speaker.

[0024] As another example, a conditional input could characterize properties of a song or other musical work, such as lyrics or genre, and the system would then generate a musical work having the properties characterized by the conditional input.

[0025] As another example, a conditional input could specify that audio data be classified into one class from a set of possible classes, thereby causing the system to generate audio data belonging to that class. For example, a class could represent a type of instrument or other audio output device, i.e., the system could generate audio output by the corresponding class or animal type, i.e., the system could generate audio representing noise produced by the corresponding animal.

[0026] As another specific example, the data item may be an image, thereby allowing the system to perform conditional image generation by generating luminance values ​​for the pixels of the image.

[0027] In this particular example, the conditional input may be a sequence of text, and the output data item may be an image describing the text; that is, the conditional input may be a caption for the output image.

[0028] As yet another specific example, the conditional input may be an object detection input specifying one or more bounding boxes and, optionally, each type of object to be drawn in each bounding box.

[0029] As yet another specific example, a conditional input can specify one object class from among several object classes to which the objects depicted in the output image should belong.

[0030] As another example, a conditional input can specify one or more images.

[0031] For example, a conditional input could specify an image of a first resolution, and the output data item could include an image of a higher second resolution.

[0032] For example, a conditional input could specify an image, and the output data item could include denoised, enhanced, stylized, or otherwise edited versions of that image.

[0033] As yet another specific example, the conditional input may specify images containing a target entity for detection, such as a tumor, and the output data items may include images that do not contain the target entity, for example, to facilitate the detection of the target entity by comparing the images.

[0034] As yet another specific example, the conditional input may be a segmentation that assigns each of several pixels in the output image to one of a set of categories, for example, a segmentation that assigns each pixel to one of each category.

[0035] As yet another example, the conditional input may be a different type of structured input, such as a mesh or graph that specifies the properties of the resulting image.

[0036] More generally, a conditional input can include one or more different types of inputs of one or more different modalities, such as text only, one or more images only, or both text and one or more images.

[0037] As yet another example, the output data item may be a video.

[0038] As a specific example, the conditional input may contain text, and the output data item may be a video described by text.

[0039] As yet another specific example, the conditional input may include one or more images, and the output data item may be a video that complements one or more images, for example, a video that starts with one or more images.

[0040] More generally, the task of generating output data items may be any task that outputs continuous data conditioned on a conditioned input. For example, the output may be the output of different sensors, such as lidar point clouds, radar point clouds, or electrocardiogram readings, and the conditioned input may represent the type of data to be measured by the sensor. If discrete output is desired, it can be obtained, for example, by thresholding the output generated by a spreading neural network.

[0041] In some application examples, output data items can be used in control tasks to control the actions of a mechanical agent that operates in a real-world environment to perform a mechanical task. For example, output data items can be processed by an agent's policy neural network to select one or more actions to be performed by the agent as part of the task. The agent can then perform one or more actions. The output data items (e.g., images) can characterize the state of the real-world environment that is expected to be captured by the agent performing one or more actions.

[0042] In any of the above examples, the output data items generated using the spreading neural network are either output data items in the output space, such that the value of the output data item is the value of an appropriate type of data item, such as an image pixel value or an audio signal amplitude value, or output data items in the latent space, such that the value of the output data item is the value of the latent representation of the output data item in the output space.

[0043] If output data items are generated in latent space, the system can generate the final output data items in output space by processing the output data items in latent space using a decoder neural network, for example, a decoder neural network pre-trained in an autoencoder framework. During training, the system can encode the target data items in output space and generate the target output of the spread neural network in latent space using an encoder neural network, for example, an encoder neural network pre-trained in conjunction with the decoder in an autoencoder framework.

[0044] Figure 1 shows an exemplary data generation system 100. The data generation system 100 is an example of a system implemented as a computer program on one or more computers in one or more locations where the systems, components, and technologies described below are implemented.

[0045] The system 100 acquires a conditional input 102 and uses the conditional input 102 to generate an output (final) data item 112 having one or more desired properties characterized by the conditional input 102.

[0046] In particular, to generate output data items 112, system 100 uses a spread neural network 110.

[0047] More specifically, system 100 uses a spread neural network 110 to perform despreading processing over multiple update iterations to generate output data items 112.

[0048] The spread neural network 110 may be any suitable spread neural network trained by, for example, system 100 or another training system, to process the spread input for the update iteration containing the current (time of the update iteration) data item in any given update iteration, and to generate a denoised output for the update iteration.

[0049] In some embodiments, the denoised output is an estimate of the noise component of the current data item, i.e., the noise that needs to be combined, for example, added to or subtracted from, the final data item, i.e., the output data item 112 generated by the system 100, in order to generate the current data item.

[0050] In some other embodiments, the denoising output is an estimate of the final data item given the current data item, i.e., an estimate of the data item obtained as a result of removing the noise component from the current data item.

[0051] In some other embodiments, the denoising output defines the predicted residual between the true noise component of the current data item and an analytical estimate of the noise component, i.e., an estimate analytically calculated from the current data item. For example, the estimate of the noise component can be obtained using a predetermined mathematical function with arguments corresponding to the elements of the (current) data item.

[0052] This type of noise-reduced output will be explained in more detail below.

[0053] For example, System 100 or another training system can train a spread neural network 110 on a set of training data items using denoising score matching to generate a denoising output.

[0054] The purpose of denoising score matching is to measure the error between (i) a denoising output generated by processing an input that includes noisy data items generated by adding sampled noise to training data items, and (ii) a target denoising output generated from training data items, sampled noise, or both, e.g., mean square error, L1 error, L2 error, or different types of errors.

[0055] For example, if the denoising output defines the predicted residual between the true noise component of the current data item and the analytical estimate of the noise component, the target denoising output can define the actual residual between the sampled noise and the analytical estimate of the noise component.

[0056] As another example, if the denoising output is an estimate of the noise component of the current data item, the target denoising output may be sampled noise.

[0057] As another example, if the denoised output is an estimate of the target data item, the target denoised output may be the target data item itself.

[0058] The spread neural network 110 may have any suitable architecture that enables the neural network to map a spread input, which includes input data items having the same number of dimensions as the output data items 112, to a denoised output that also has the same number of dimensions as the output data items 112.

[0059] For example, if the output data item is an audio signal or an image, the spreading neural network 110 may be a convolutional neural network, such as a U-Net or other architecture that maps one input of a given number of dimensions to an output of the same number of dimensions.

[0060] As another example, the spread neural network 110 may be a transformer neural network that processes the spread input through a set of self-attention layers to produce a denoised output.

[0061] The neural network 110 can be conditioned based on the conditioning input 102 in any of the following ways.

[0062] For example, system 100 can use an encoder neural network to generate one or more embeddings representing the conditioned input 102, and the spread neural network 110 can include one or more cross-attention layers, each performing cross-attention to one or more embeddings.

[0063] As used herein, an embedding is an ordered set of numbers, such as a vector of floating-point values ​​or other types of values.

[0064] For example, if the conditional input is text, the system can use a text encoder neural network, such as a transformer neural network, to generate a fixed or variable number of text embeddings that represent the conditional input.

[0065] If the conditioned input is an image, the system can use an image encoder neural network, such as a convolutional neural network or a vision transformer neural network, to generate a set of embeddings that represent the image.

[0066] If the conditioned input is audio, the system can generate one or more embeddings that encode the audio, for example, using an audio encoder neural network, such as an audio encoder neural network trained in conjunction with a decoder neural network as part of a neural audio codec.

[0067] If the conditional input is a scalar value, the system can, for example, use an embedding matrix to map the scalar value or a one-hot representation of a scalar value to the embedding.

[0068] In some cases, the conditional input 102 includes two or more of several different types of inputs, such as text, images, boundary values, or context embeddings.

[0069] In some of these cases, the system 100 generates one or more initial embeddings for each of the different types of inputs using an appropriate encoder neural network as described above, then processes the initial embeddings for all of the different types of inputs using a transformer encoder neural network, updating each initial embedding to generate a set of final embeddings. One or more cross-attention layers in the spread neural network 110 can then perform cross-attention on the set of final embeddings.

[0070] Apart from these cases, different cross-attention layers within the spreading neural network 110 can perform cross-attention on embeddings of different types of conditioned inputs.

[0071] In yet another case, the system 100 can concatenate initial embeddings of different types of inputs along the sequence dimension, and then one or more cross-attention layers can perform cross-attention on the concatenated set of final embeddings.

[0072] As another example, the diffuse neural network 110 may include one or more other types of neural network layers conditioned on one or more embeddings. Examples of such layers include feature-wise linear modulation (FiLM) layers and layers with conditional gate activation functions.

[0073] The spreading input in any given update iteration can also include data that defines the noise level of the iteration. Generally, each update iteration has a corresponding time step t, and the noise level of the iteration depends on the time step. For example, the noise level may be a decreasing function of the time step t. Examples of such functions include linear functions, cosine functions, and sigmoid functions. In these cases, data that identifies the noise level, time step, or both can be embedded using a suitable neural network, such as a multilayer perceptron (MLP), and can be used to condition the spreading neural network 110 with respect to the conditioning input as described above.

[0074] In each update iteration, the system 100 updates the current data item at the time of the update iteration using the denoising output 110 generated by the spreading neural network.

[0075] For example, the system can use the denoising output to determine an initial estimate of the final data item, and then apply an appropriate diffusion sampler to the initial estimate to update the current data item.

[0076] As another example, the system can adjust the denoised output using classifier-free guidance or negative guidance, use the adjusted denoised output to determine the initial estimate of the final data item, and then apply an appropriate diffusion sampler to the initial estimate to update the current data item. Classifier-free guidance is described, for example, Ho and Salimans, arXiv:2207.12598.

[0077] The system can update data items with estimates using an appropriate diffusion sampler, such as a DDPM (Denoising Diffusion Probability Model) sampler, a DDIM (Denoising Diffusion Implicit Model) sampler, or other suitable sampler, and generate updated current data items. DDPMs are described, for example, in Ho et al. arXiv:2006:11239.

[0078] If the denoised output is a prediction of the data item, the system can directly use the denoised output (or adjusted denoised output) as the estimate.

[0079] If the denoised output is a prediction of the noise component, the system can determine an initial estimate from the current data item, the denoised output, and the noise level for the current update iteration.

[0080] Optionally, after the last iteration, the system may refrain from using the diffusion sampler and instead use the initial estimate as the updated current data item.

[0081] After the last update iteration, system 100 outputs the current data item as the final output data item 112.

[0082] For example, system 100 can provide data item 112 to present or play back to the user on the user's computer, or it can store data item 112 for later use.

[0083] In some embodiments, the spread neural network 110 is one of a sequence of spread neural networks used by the system 100 to generate the final data item, such as a hierarchy or cascade of spread neural networks. For example, each spread neural network in the sequence can receive an output data item generated by a preceding spread neural network in the sequence as input and generate an output data item with improved resolution, such as improved spatial resolution, improved temporal resolution, or both, compared to the preceding spread neural network in the sequence. In these embodiments, all neural networks in the sequence can receive the conditioned input 102, or only a suitable subset of the spread neural networks in the sequence, such as only one or more of the earliest spread neural networks in the sequence, can receive the conditioned input 102.

[0084] Generally, system 100 can improve the quality of output data items generated by the spread neural network 110 by changing one or more of the following: training the spread neural network 110, inputting data to the spread neural network 110, or how the spread neural network 110 is used to generate output data items after training.

[0085] As an example, as described above, the system 100 can be configured such that, in each update iteration, it processes a spread input for the update iteration containing the current noisy data item for the update iteration, and generates a denoising output that defines an estimate of the residual error between the analytical estimate of the noise component of the noisy data item and the true noise component of the noisy data item.

[0086] This will be explained in more detail below, with reference to Figure 2.

[0087] As another example, the spreading neural network 110 can be configured to receive context embeddings, which can represent either multiple different context data items or a single context data item, as a conditional input 102 or as part of the conditional input 102. Context data items are generally of the same type as output data items and are used to guide the generation process.

[0088] This will be explained in more detail below with reference to Figures 3 and 4.

[0089] As yet another example, system 100 can employ boundary conditioning, in which case, in addition to the conditioning input 102, system 100 conditions the spread neural network 110 based on a lower or upper bound of input scalar values ​​representing the values ​​of a particular property of the generated data item. That is, rather than requiring the spread neural network 110 to generate output data items having the exact values ​​of a particular property represented by the input scalar values, system 100 can give the spread neural network 110 the flexibility to generate any suitable data item having the values ​​of a particular property bounded by the scalar values.

[0090] Boundary conditioning will be explained below with reference to Figures 5 and 6.

[0091] As yet another example, system 100 can use normalization guidance when generating output data items using a spreading neural network 100. In particular, when the system uses normalization guidance, in each update iteration, the system can calculate the difference between the conditionally denoised output for the update iteration and the unconditionally denoised output or a negative denoised output, and then determine the overall denoised output, for example, based on the direction rather than the magnitude of the difference. The unconditionally denoised output may be a denoised output that is not conditioned on any conditional input, while the negative denoised output may be a denoised output conditioned on a negative conditional input that indicates one or more properties that the denoised data item should not have.

[0092] Normalization guidance is explained below, with reference to Figure 7.

[0093] Figure 2 is a flowchart of an exemplary process 200 for generating final data items using a spreading neural network. For convenience, the process 200 is described as being performed by a system of one or more computers located in one or more locations. For example, a data generation system appropriately programmed according to this specification, e.g., the data generation system 100 shown in Figure 1, can perform the process 200.

[0094] The system receives the conditional input (step 202).

[0095] The system initializes the data items (step 204).

[0096] Generally, an initialized data item has the same number of dimensions as the final data item, but it has noisy values. That is, an initialized data item has the same number of elements as the final data item.

[0097] For example, the system can initialize a data item, or generate a first instance of the data item, by sampling the values ​​of each element within the data item from a corresponding noise distribution, such as a Gaussian distribution or a different noise distribution. That is, the output data item will contain multiple elements, the initial data item will contain the same number of elements, and the values ​​of each element will be sampled from a corresponding noise distribution.

[0098] Next, the system generates the final output data item by updating the data item in each of the multiple update iterations. In other words, the final output data item is the data item after the last iteration of the multiple update iterations.

[0099] In some cases, the number of iterations is fixed. In other cases, the system or other systems may adjust the number of iterations based on latency requirements for generating the final output data item; that is, they may select the number of iterations so that the final output data item is generated in a way that satisfies the latency requirements. In yet another case, the system or other systems may adjust the number of iterations based on computational resource consumption requirements for generating the final output data item; that is, they may select the number of iterations so that the final output data item is generated in a way that satisfies the requirements. For example, the requirement may be the maximum number of floating-point operations (FLOPS) performed as part of the generation of the final output data item.

[0100] As described above, the system performs a despreading process across update iterations by updating the current data item in each iteration. Each update iteration corresponds to a different point in time within a time interval, such as the interval between 0 and 1, or another appropriate time interval. A point in time is also called a time step t or time index t. For example, update iterations may be evenly spaced across the time interval, i.e., at regular intervals within the interval, or they may be spaced within the time interval according to a different scheme.

[0101] In particular, when an analytical estimation value is used, in each update iteration, the system performs steps 206 to 212 to update the data item.

[0102] The system determines an analytical estimation value of the noise component of the current data item (step 206).

[0103] For the first update iteration, the current data item is an initial noisy data item. For each subsequent update iteration, the current data item is the data item that has been updated in a preceding update iteration.

[0104] As described above, the noise component of the current data item is the noise added to the final data item to generate the current data item.

[0105] For example, in an iteration where the time index t, that is, the time point corresponding to the update iteration ("time step") is t, the current data item x t is x t = α t x0 + σ t ε, where ε is the noise component, x0 is the final data item, α t and σ t are weights for the iteration at time index t. For example, α t is [Math.]] and σ t may be a value between 0 and 1, inclusive. As another example, α t may be equal to 1, and σ t may be a value greater than 0. In general, one or both of α t , σ t can be determined according to a fixed schedule over the time index t, for example, a linear schedule, a quadratic schedule, a cosine schedule, etc.

[0106] To generate analytical estimates, the system can apply scaling factors to the current data item.

[0107] Generally, the scaling factor is a time-dependent scaling factor that depends on the update iteration; that is, it differs for different update iterations.

[0108] For example, the analytical estimate is k t x t It can be equal to, where k t This is the scaling factor.

[0109] Generally, the scaling factor is the estimated standard deviation s of the current data item in the update iteration t. t Because it depends on [something], it will be different for different update iterations.

[0110] For example, k t teeth,

number

number

[0111] For example, the system uses σ as the standard deviation of the training data items used to train the spreading neural network. data It can be approximated as follows: For example, σ dataThis can be equal to the standard deviation across all elements of the set of training data items used to train the spreading neural network, for example, all of the training data items used for training, or a subset of the training data items, for example, a randomly selected subset or another appropriate subset of the training data items.

[0112] As another example, the system is σ data The system can receive this as input from the user. As a specific example, the system can receive σ data It is possible to submit values ​​that allow users to modify specific properties of the generated data items. For example, if the data item is an image, σ data By changing the value of σ, the user can change the saturation of the generated image. As a specific example, data By setting the value to a lower value, for example, a value lower than the standard deviation of the training data items used to train the spreading neural network, the user can reduce the saturation of the image.

[0113] The system uses a spread neural network to process a first spread input for the update iteration, which includes the current data item and a representation of the conditioned input, and generates a first denoising output for the update iteration (step 208).

[0114] For example, before the first update iteration, and as described above, the system may process the conditioned input using one or more embedded neural networks to generate one or more embeddings of the conditioned input.

[0115] A first spreading input for any given update iteration may include one or more embeddings of the conditioning input.

[0116] The first diffusion input may also include one or more of the following: data identifying an update iteration (e.g., data identifying the corresponding time point for the update iteration), data characterizing one or more context data items to be used as context during data item generation, and scalar values ​​for one or more properties of the generated data item. These types of inputs are described in more detail below.

[0117] As described above, the first denoising output defines the residual error given the first denoising input, i.e., the predicted difference between the noise component of the current data item and the analytical estimate of the noise component. That is, the first denoising output is ε-k t x t Define the predicted value.

[0118] For example, the first noise-reduced output is the standard residual r t It may also be a predicted value of r, where r t teeth,

number

number

[0119] Optionally, i.e., when classifier-free guidance is used, the system can also process one or more additional spreading inputs for update iterations to generate a corresponding additional denoising output for each additional spreading input for each update iteration (step 210).

[0120] Each additional diffusion input includes the current data item at the time of the update iteration, but also includes a different conditional input.

[0121] For example, one of the additional diffusion inputs may be an unconditional diffusion input that includes a representation of a conditioned input specified to indicate that a data item should be generated unconditionally (i.e., without being conditioned on other conditioned inputs). For example, the representation of the conditioned input specified to indicate that a data item should be generated unconditionally may be a predetermined fixed embedding, such as an embedding containing all zeros.

[0122] As another example, one of the additional diffusion inputs may be a negative diffusion input that contains a representation of a negative conditioning input indicating a property that the generated data item should not have.

[0123] In other words, the system can also receive a negative conditioning input that indicates properties that the generated data item should not have, and the negative diffusion input can include a representation of the negative conditioning input, for example, one or more embeddings generated from the negative conditioning input.

[0124] Each additional denoising output defines the residual error given the corresponding additional denoising input, i.e., the predicted difference between the noise component of the current data item and the analytical estimate of the noise component.

[0125] The system determines the final denoised output for the update iteration from the first denoised output and any additional denoised outputs generated (step 212).

[0126] If no additional denoising output is generated, the system can set the final denoising output to be equal to the first denoising output.

[0127] If one or more additional denoising outputs are generated, the system can combine the first denoising output and the final denoising output according to guidance weights w for the update iteration. The guidance weights can be used to adjust the relative contributions of the first denoising output and the additional denoising outputs to the final denoising output.

[0128] For example, the system can set the final denoised output to be equal to (1+w) × first denoised output - w × additional denoised output, or, if there are multiple additional denoised outputs, to be equal to the sum of the additional denoised outputs (where × represents the multiplication operator). In other words, the final denoised output can be determined from the difference between the first denoised output scaled by (1+w) and the sum of one or more additional denoised outputs scaled by w.

[0129] As another example, the system can combine the denoised output using normalization guidance. Normalization guidance is described in more detail below with reference to Figure 7.

[0130] Next, the system updates the current data item using the final denoised output (step 214).

[0131] For example, the system can calculate the final estimate of the noise component from the final denoised output and analytical estimate, and then update the current data item using the final estimate of the noise component.

[0132] For example, the system calculates the final estimate of the noise component.

number

number

number

[0133] Next, the system can calculate the initial estimate of the final data item using the final estimate of the noise component, as follows:

number

[0134] In the final update iteration, the system can use the initial estimate as the updated data item.

[0135] For each update iteration other than the last update iteration, the system can apply an appropriate diffusion sampler to the initial estimate to generate updated data items. An example of a diffusion sampler is described above with reference to Figure 1.

[0136] The quality of the output data items can be improved, particularly when high guidance weights for classifier-free guidance are used as part of the generation process, by generating a denoising output that defines an estimate of the residual error between the analytical estimate of the noise component of a noisy data item and the true noise component of the noisy data item. For example, using high guidance weights may result in the generated data items conforming more strongly to the provided conditioned input, which is advantageous, but it may also degrade the overall quality of the generated image. For example, if the data item is an image, using high guidance weights may result in a highly saturated image. By using the denoising output described above, the system can generate data items that conform to the conditioned input while maintaining overall quality, for example, by reducing the saturation of the generated image to a realistic level.

[0137] If analytical estimates are not used, i.e., if the denoising output is an estimate of the noise component of the current data item or an initial estimate of the final data item, the system can calculate the initial estimate of the final data item without incorporating analytical estimates.

[0138] If the denoised output is an estimate of the noise component, that is, the system uses the initial estimate.

number

number

[0139] If the denoised output is an estimate of the final data item, the system can directly use the final denoised output as the initial estimate.

[0140] As described above, before generating data items using a spread neural network, the system trains the spread neural network, for example, for denoising and score matching purposes.

[0141] In particular, to train a spreading neural network for score matching purposes, the system may include (i) data items from a set of training data items, (ii) one or more corresponding conditioned inputs for the data items, (iii) a time step t for training, for example, uniformly randomly from time intervals or according to different distributions over time intervals, and (iv) sampling noise ε from a noise distribution.

[0142] Next, the system combines the target data item x0 with the sampled noise according to the sampled time step, for example, noisy data item x t to xt =α t x0+σ t By setting it to ε, the noisy data item x t It can generate [this].

[0143] Next, the system can use a spreading neural network to process inputs including noisy data items, data specifying time steps, and conditional inputs (multiple) to generate a denoised output.

[0144] Next, the system can train a spreading neural network by calculating the error between the denoised output and the target denoised output, using the error to, for example, determine the gradient of the error, and then using the gradient to update the parameters of the spreading neural network by applying an optimization tool to (at least) the gradient.

[0145] As a specific example, the purpose of denoising score matching can be to measure (i) the denoising output and (ii) the error between the target denoising output, such as mean square error, L1 error, L2 error, or different types of errors.

[0146] For example, if the denoising output defines the predicted residual between the true noise component of the current data item and the analytical estimate of the noise component, the target denoising output can define the actual residual between the sampled noise and the analytical estimate of the noise component.

[0147] As another example, if the denoised output is an estimate of the true noise component of the current data item, the target denoised output may be sampled noise.

[0148] As another example, if the denoised output is an estimate of the target data item, the target denoised output may be the target data item itself.

[0149] Figure 3 is a flowchart of an exemplary process 300 for conditioning a spreading neural network based on a variable number of contextual data items. For convenience, the process 300 is described as being performed by one or more computer systems located in one or more locations. For example, a data generation system appropriately programmed according to this specification, e.g., the data generation system 100 shown in Figure 1, can perform the process 300.

[0150] The system retrieves one or more context data items (step 302).

[0151] For example, a user can provide the system with context data items, or can identify one or more data items maintained by the system for use as context data items.

[0152] Contextual data items are generally of the same type as the final data items generated by the system, and they provide context for the system-generated data items. For example, if the generated data items are images, each contextual data item may be an image. If the generated data items are audio, each contextual data item may be an audio sample.

[0153] The system processes each context data item using an embedded neural network to generate its respective embedding (step 304).

[0154] As described above, an embedding is an ordered set of numbers. For example, an embedding may be a vector of numbers, such as floating-point values ​​or other vectors of numbers. Another example is that an embedding may be a set of vectors of numbers, such as a sequence or a two-dimensional grid.

[0155] The embedded neural network may be any suitable type of neural network capable of mapping corresponding types of inputs to the embedding. An example of an embedded neural network architecture is described above with reference to Figure 1.

[0156] Generally, embedded neural networks can be trained in conjunction with spread neural networks, or they can be pre-trained before training the spread neural network, for example, for self-supervised representation learning purposes. Examples of representation learning purposes include contrast learning and masked generative modeling purposes.

[0157] If multiple context data items exist, the system combines the embeddings of the context data items, for example, by averaging them, to generate a combined embedding.

[0158] In some embodiments, the system can receive user input assigning a weight to each of several context data items. For example, each weight may represent the degree to which each context data item should influence the content of the generated data item. In these embodiments, the system can calculate a weighted sum of the embeddings of the context data items according to the respective weights of the corresponding context data items.

[0159] Optionally, the system can then multiply the generated embedding by a non-negative scalar value to control the strength of the context embedding's influence on the generation of the output data items. The scalar value can be, for example, predetermined or received as user input.

[0160] The system processes the generated embeddings and optional additional conditioned inputs using a spreading neural network to produce output data items by performing a despreading process over multiple update iterations, as described above with reference to Figures 1 and 2 (step 306).

[0161] If multiple context data items exist, the system handles combined embedding. If a single context data item exists, the system handles embedding of that single context data item.

[0162] In other words, the system uses embeddings to provide additional context to the spread neural network, guiding it during the generation of output data items.

[0163] Therefore, the system allows the spread neural network to be conditioned on a variable number of user-specified context data items, meaning that the embeddings processed as input by the spread neural network may represent a single data item or be aggregated embeddings representing multiple different data items.

[0164] In particular, the system can generate output data items over multiple update iterations using a spread neural network, as described above with reference to Figures 1 and 2. As described above, in any given update iteration, the spread neural network can be conditioned on the generated embeddings in any of the following ways, for example, by including one or more cross-attention layers that perform cross-attention on the generated embeddings.

[0165] In some embodiments, the system modifies the training of the spreading neural network to improve its performance in conditioning with a variable number of contextual data items. An example of such a training process is described in more detail below with reference to Figure 4.

[0166] Figure 4 is a flowchart of an exemplary process 400 for training a spreading neural network to condition the spreading neural network based on a variable number of contextual data items. For convenience, process 400 is described as being performed by one or more computer systems located in one or more locations. For example, a data generation system appropriately programmed according to this specification, e.g., the data generation system 100 shown in Figure 1, can perform process 400.

[0167] The system maintains the embeddings of each of the multiple context data items and maintains the data to cluster the embeddings into multiple clusters (step 402). These context data items are also called “training” context data items and may be the same context data items referred to above (see Figure 3), or they may be a different set of context data items used solely for training the neural network.

[0168] For example, embeddings can be clustered using the k-means clustering method.

[0169] As another example, embeddings can also be clustered based on the similarity between corresponding data items. Specifically, in the case of images, different clusters may represent images of different objects; that is, each image within a cluster may contain an image of the object corresponding to that cluster.

[0170] In the case of audio, different clusters can represent different genres of music or music by different artists.

[0171] More generally, each cluster may correspond to different values ​​of a property and may include embedded data items that have the corresponding property values.

[0172] Embeddings are generated, for example, by pre-training an embedding neural network and then processing context data items using the aforementioned embedding neural network.

[0173] The system takes an input specifying a target data item from multiple context data items, and optionally takes a conditional input that characterizes the target data item (step 404). For example, the system may receive an input that randomly samples a target data item from multiple context data items, or an input that specifies a target data item to be used to train a neural network.

[0174] The system selects a context embedding for training a spreading neural network (step 406).

[0175] In particular, the system randomly selects either (i) an embedding of a target context data item or (ii) the centroid of the cluster to which the embedding of the target context data item belongs. That is, the system selects an embedding of a target context data item with probability p and a centroid of a cluster with probability 1-p. The value of probability p can be pre-set or received by the system as input from the user.

[0176] The system trains a spreading neural network using the selected context embeddings and target data items (step 408).

[0177] In other words, the system uses selected context embeddings as context to train a spread neural network on target data items, with the objective of denoising and score matching.

[0178] In particular, to train a spreading neural network for score matching purposes, the system can sample training time steps t, for example, uniformly and randomly from time intervals, or according to different distributions over time intervals, and sample noise ε from a noise distribution.

[0179] Next, the system combines the target data item x0 with the sampled noise according to the sampled time step, for example, noisy data item x t to x t =α t x0+σ t By setting it to ε, the noisy data item x t It can generate [this].

[0180] Next, the system processes the input, including the noisy data items, selected context embeddings (and optionally additional conditional inputs), to generate a denoised output.

[0181] Next, the system can train a spreading neural network by calculating the error between the denoised output and the target denoised output, using the error to, for example, determine the gradient of the error, and then using the gradient to update the parameters of the spreading neural network by applying an optimization tool to (at least) the gradient.

[0182] As a specific example, the purpose of denoising score matching can be to measure (i) the denoising output and (ii) the error between the target denoising output, such as mean square error, L1 error, L2 error, or different types of errors.

[0183] For example, if the denoising output defines the predicted residual between the true noise component of the current data item and the analytical estimate of the noise component, the target denoising output can define the actual residual between the sampled noise and the analytical estimate of the noise component.

[0184] As another example, if the denoised output is an estimate of the true noise component of the current data item, the target denoised output may be sampled noise.

[0185] As another example, if the denoised output is an estimate of the target data item, the target denoised output may be the target data item itself.

[0186] By repeatedly running process 400 for different context data items, the system can train a spreading neural network to effectively generate conditioned data items based on a variable number of context data items. That is, since the selected context embedding is randomly chosen to either (i) represent an embedding of a target context data item or (ii) represent the centroid of the cluster to which the target context data item belongs, the spreading neural network learns to effectively incorporate context from both single context items and collections of multiple context items by providing both types of inputs during training.

[0187] In other words, as a result of this training and using images as examples, the diffuse neural network learns to accurately denoise images conditioned on (i) image embeddings, or (ii) embeddings representing the cluster to which the image belongs. Therefore, after training, the diffuse neural network can be used to generate high-quality new images conditioned on embeddings from a single contextual image, or on embeddings aggregated from embeddings from multiple different contextual images.

[0188] In some cases, the system can run Process 400 after pre-training a spreading neural network with the aim of incorporating either no context data items or only embeddings that represent a single context data item.

[0189] Figure 5 is a flowchart of an exemplary process 500 for conditioning a generative neural network based on property value boundaries. For convenience, process 500 is described as being performed by one or more computer systems located in one or more locations. For example, a data generation system appropriately programmed according to this specification, e.g., the data generation system 100 shown in Figure 1, can perform process 500.

[0190] For example, the generative neural network may be the diffusion neural network described above, or a different type of generative neural network, such as an autoregressive generative neural network, such as a transformer-based generative neural network, or a neural network that first generates a latent representation of a data item and then decodes that latent representation to generate a data item. For example, the latent representation can represent a data item as a set of discrete latent vectors or as a set of continuous latent vectors.

[0191] For each of one or more properties of the output data item, the system receives an input scalar value from a range of scalar values, each representing a different value of the property (step 502).

[0192] One or more properties of the output data item may be any suitable properties of the type of data item that the spreading neural network is configured to generate (for example, properties that the user wants to change).

[0193] Generally, properties within a set of one or more properties are predetermined, but scalar values ​​may differ for each output data item. If a property does not have a scalar value corresponding to a given output data item, i.e., if the system does not receive an input scalar value for a given output data item, the system can set the scalar value to a predetermined value.

[0194] Examples of such properties include the realism, brightness, quality (determined by appropriate quality metrics depending on the data item type), and consistency of the generated data items.

[0195] Input scalar values ​​may be received, for example, as user input or input from another system. In some cases, the user can construct desired values ​​for the properties of each generated data item, while in other cases, the system can receive high-level contextual input from the user, such as natural language input or other input, and map that contextual input to the respective values ​​of the data items, for example, by applying a set of heuristics or using a trained model.

[0196] For each of one or more properties of the output data item, the system determines whether the input scalar value is an upper or lower limit of the desired value for the property of the output data item (step 504).

[0197] For example, if the received input value is part of a “positive” prompt, i.e., part of a positive conditioning input, the system can determine that each input scalar value is a lower bound for the desired value of the corresponding property. That is, since a given value is part of a “positive” prompt representing a desired property of an output data item, the system can determine that the prompt is satisfied if the value of the corresponding property of the output data item is greater than or equal to the received input value of that property.

[0198] If the received input value is part of a "negative" prompt, i.e., part of a negative conditioning input used for classifier-free guidance, the system can determine that the input scalar value is an upper limit of the desired value for the corresponding property. That is, since a given value is part of a "negative" prompt representing an undesirable property of an output data item, the system can determine that the prompt is satisfied if the value of the corresponding property of the output data item is at least as low as the received input value for that property.

[0199] As another example, the system may receive an input scalar value along with an instruction indicating whether the input scalar value is an upper or lower bound for the corresponding property value.

[0200] For each of one or more properties of the output data item, the system determines an indicator value that indicates whether the input scalar value is an upper or lower limit of the desired value for the property of the output data item (step 506). That is, for each property, the indicator value is equal to a first value, such as 1, if the input scalar value is the lower limit, and equal to a second value, such as 0 or -1, if the input scalar value is the upper limit.

[0201] The system uses a generative neural network to process an input that includes the input scalar value of one or more properties and an indicator value for each of those properties (i.e., a value indicating whether the input scalar value is an upper or lower bound) in order to generate an output data item for each of those properties (step 508).

[0202] For example, if the neural network is a spreading neural network, the system can run process 200 with the property values ​​as some or all of the conditional inputs to generate output data items.

[0203] By representing input values ​​not as exact targets but as ranges of property values, spread neural networks are given the flexibility to generate high-quality data items that meet the requirements. That is, generating data items with properties whose values ​​perfectly match the target value is a difficult task, but generating data items with property values ​​that are appropriately constrained by the target value is easier to achieve while respecting the constraints imposed by the rest of the conditioned input.

[0204] In some embodiments, the system modifies the training of the generative neural network to improve its performance in boundary conditioning. An example of such a training process is described in more detail below with reference to Figure 6.

[0205] Figure 6 is a flowchart of an exemplary process 600 for training a generative neural network based on property value boundaries. For convenience, the process 600 is described as being performed by one or more computer systems located in one or more locations. For example, a data generation system appropriately programmed according to this specification, e.g., the data generation system 100 shown in Figure 1, can perform the process 600.

[0206] The system retrieves the target data item and, for one or more properties of the target data item, retrieves the ground truth value of that property (step 602). The ground truth value of a given property is a scalar value that represents the actual value of the given property of the target data item.

[0207] For each property, the system samples a scalar value of the property uniformly from the range of possible values ​​of the property, for example, from another suitable distribution across the entire range of possible values ​​of the property (step 604).

[0208] For each property, the system processes the model input, which includes (i) a sampled value of the property and (ii) an indicator value indicating whether the sampled value of the property is greater than or less than the ground truth value of the property, using a generative neural network to generate a model output for the model input (step 606).

[0209] For example, the indicator value is 1 if the sampled value is greater than the target value, and 0 or -1 if the sampled value is less than the target value.

[0210] As another example, the indicator value is 0 or -1 if the sampled value is greater than the target value, and 1 if the sampled value is less than the target value.

[0211] If the sampled value is equal to the target value, the system can set the indicator value to 1, 0, or -1.

[0212] If the generative neural network is, for example, a spread neural network as described above, the model input may also include a noisy version of the target data item, and the model output may be one of the denoised outputs described above.

[0213] The system trains a generative neural network based on a loss function that measures the quality of the model output relative to the target model output for the target data item (step 608).

[0214] For example, if a generative neural network generates data items autoregressively or directly in a single forward pass, the target model output may be the target data items.

[0215] As another example, if a generative neural network generates latent representations of output data items, for example, if it generates tokens that represent the output data items, then the target model output may also be a latent representation of the target data items.

[0216] For example, if the generative neural network is a diffusion neural network as described above, the loss function may be for the purpose of matching the denoising score between the denoising output and the target denoising output, as described above.

[0217] The system can repeatedly run process 600 on different target data items to train a generative neural network.

[0218] Therefore, by repeatedly running process 600, the system has available ground truth values ​​for the training data items, but trains the generative neural network to generate data items through boundary conditioning rather than training the generative neural network to accurately replicate the ground truth values.

[0219] Figure 7 is a flowchart of an exemplary process 700 for determining the final denoising output in a given update iteration using normalization guidance. For convenience, the process 700 is described as being performed by a system of one or more computers located in one or more locations. For example, a data generation system appropriately programmed according to this specification, e.g., the data generation system 100 shown in Figure 1, can perform the process 700.

[0220] The system generates a first denoising output for the update iteration and an additional denoising output for the update iteration, as described above (step 702).

[0221] For example, the first denoising output may be a conditionally denoising output, and the additional denoising output may be an unconditionally denoising output or a negative denoising output.

[0222] As described above, each denoising output may be, for example, the predicted values ​​of each noise component of the noisy ("current") data item for a given update iteration, the predicted values ​​of each difference between the noise component and the analytical estimate of the noise component, or the predicted values ​​of each final data item.

[0223] The system calculates the difference between the value of the corresponding element in the first denoised output and the value of the corresponding element in the additional denoised output for each element of the final denoised output (step 704).

[0224] The system determines, for each element of the final denoised output, the sign of the difference calculated for that element, i.e., positive, negative, or none (if the difference is 0) (step 706).

[0225] For each element of the final denoised output, the system determines the value of the element in the final denoised output based on the sign of the difference calculated for that element and the guidance weights for a given update iteration (step 708).

[0226] For example, the system can set the value to 0 if the difference is zero and unsigned, set the value to be equal to the guidance weight if the sign is positive, and set the value to be equal to the negative value of the guidance weight if the sign is negative.

[0227] Therefore, the system determines the element values ​​of the final denoised output using only the direction, and not the magnitude, of the difference between the first denoised output and the additional denoised output.

[0228] The above explanation describes how to use guidance weights w when generating data items across multiple update iterations.

[0229] Generally, the guidance weight w for each update iteration is a scalar value. Specifically, the guidance weight may be a positive scalar value.

[0230] In some embodiments, the system uses the same guidance weight (also called the guidance scale) for each update iteration.

[0231] In some other embodiments, the system sets guidance weights for each update iteration as a function of time t corresponding to the update iteration.

[0232] The function may be any suitable function that maps t to guidance weights.

[0233] As a concrete example, the function can assign different guidance weights to different time bins. That is, the function maps each of several intervals of t to a corresponding guidance weight, and to determine the guidance weight for any given t, the system identifies the interval to which t belongs and sets the guidance weight for t to be equal to the guidance weight to which the function maps t. For example, if t is [0, 0.25], the function can set the weight to 10, and if t is [0.25, 0.5], the weight can be set to 25.

[0234] Generally, classifier-free guidance can have different effects depending on the diffusion time t. For example, when generating images, a high t may affect coarse, large structures within the image, while a low t may enhance fine details and edges. Therefore, the optimal guidance weights (in terms of the quality of the output data items) are not uniform across the time step. By setting the guidance weights as a function of t, the system can control this and improve the quality of the output data items.

[0235] If the system uses classifier-free guidance (e.g., conventional classifier-free guidance or normalized classifier-free guidance), the system can run periodically and unconditionally during training.

[0236] In this specification, the term “configured” is used in relation to systems and computer program components. A system of one or more computers being configured to perform a particular operation or action means that, while in operation, software, firmware, hardware, or a combination thereof is installed on the system that causes the system to perform that operation or action. A system of one or more computer programs being configured to perform a particular operation or action means that, when executed by a data processing device, the program contains instructions that cause the device to perform that operation or action.

[0237] The subject matter and functional embodiments described herein can be implemented in digital electronic circuits, tangibly embodied computer software or firmware, or computer hardware, including the structures disclosed herein and their structural equivalents, or one or more combinations thereof. Embodiments of the subject matter described herein may be implemented as one or more computer programs, for example, one or more modules of computer program instructions, encoded on tangible, non-temporary storage media, which are executed by or control the operation of a data processing device. The computer storage media may be a machine-readable storage device, a machine-readable storage board, a random-access memory device or a serial-access memory device, or one or more combinations thereof. Alternatively or additionally, the program instructions may be encoded into artificially generated propagating signals, for example, mechanically generated electrical signals, optical signals or electromagnetic signals, which are generated to encode information for transmission to a receiving device suitable for execution by a data processing device.

[0238] The term "data processing device" refers to data processing hardware and encompasses all types of devices, machines, and equipment for processing data, including, for example, programmable processors, computers, or multiple processors or computers. A device may also be, or further include, specialized logic circuits such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits). Optionally, in addition to hardware, a device may include code that constitutes the execution environment for computer programs, such as processor firmware, protocol stacks, database management systems, operating systems, or one or more combinations thereof.

[0239] Computer programs, which may be called or described as programs, software, software applications, apps, modules, software modules, scripts, or code, can be written in any form of programming language, including compiled or interpreted languages, declarative or procedural languages, and can be deployed in any form, including as standalone programs or as modules, components, subroutines, or other units suitable for use in a computing environment. A program may, but is not required to, correspond to a file in a file system. A program may be stored in a single file dedicated to a program of interest, or in a set of collaborative files, such as files storing one or more modules, subprograms, or parts of code, in part with other programs or data, for example, in a document in a markup language. A computer program may be deployed to run on one computer, or it may be located in one place or distributed across multiple locations and interconnected by a data communication network to run on multiple computers.

[0240] In this specification, the term “database” is used broadly to refer to any collection of data. The data does not need to be structured in any particular way, or not structured at all, and can be stored in one or more storage locations. Therefore, for example, an index database may contain multiple collections of data, each of which may be organized and accessed in a different way.

[0241] Similarly, in this specification, the term “engine” is used broadly to refer to a software-based system, subsystem, or process programmed to perform one or more specific functions. Typically, an engine is implemented as one or more software modules or components installed on one or more computers in one or more locations. In some cases, one or more computers are dedicated to a particular engine, while in other cases, multiple engines may be installed and run on the same one or more computers.

[0242] The processes and logic flows described herein may be performed by one or more programmable computers executing one or more computer programs to act on input data and produce outputs. Alternatively, the processes and logic flows may be performed by special-purpose logic circuits, such as FPGAs or ASICs, or by a combination of special-purpose logic circuits and one or more programmed computers.

[0243] A computer suitable for running computer programs can be based on a general-purpose or dedicated microprocessor, or both, or any other type of central processing unit. Generally, the central processing unit receives instructions and data from read-only memory, random-access memory, or both. The basic components of a computer are the central processing unit for executing instructions and one or more memory devices for storing instructions and data. The central processing unit and memory may be complemented by or incorporated into special-purpose logic circuits. Generally, a computer also includes one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operablely connected to them to receive data from them, transmit data to them, or both. However, a computer is not required to have such devices. Furthermore, a computer may be incorporated into other devices, such as mobile phones, personal digital assistants (PDAs), mobile audio or video players, game consoles, Global Positioning System (GPS) receivers, or portable storage devices (such as Universal Serial Bus (USB) flash drives) (these are just a few examples).

[0244] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks or removable disks, magneto-optical disks, and CD-ROM and DVD-ROM disks.

[0245] To provide user interaction, embodiments of the subject matter described herein can be implemented in a computer having a display device for displaying information to the user, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, and a keyboard and pointing device, such as a mouse or trackball, to which the user can provide input to the computer. Other types of devices can also be used to interact with the user. For example, the feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or haptic feedback, and input from the user may be received in any form, including acoustic, voice, or haptic input. Furthermore, the computer can interact with the user by sending and receiving documents to and from the user's device, for example, by sending a web page to a web browser on the user's device in response to a request received from a web browser. The computer may also interact with the user by sending text messages or other forms of messages to a personal device (for example, a smartphone running a messaging application) and then receiving a response message from the user.

[0246] Data processing devices for implementing machine learning models may also include dedicated hardware accelerator units for handling common, computationally intensive parts of machine learning training or production workloads, such as inference.

[0247] Machine learning models can be implemented and deployed using machine learning frameworks, such as the TensorFlow framework or the Jax framework.

[0248] Embodiments of the subject matter described herein may be implemented in a computing system that includes, for example, a data server as a backend component, or in a computing system that includes a middleware component, for example, an application server, or in a computing system that includes a client computer having a frontend component, for example, a graphical user interface, a web browser, or an application that enables a user to interact with embodiments of the subject matter described herein, or in a computing system that includes one or more such backend, middleware, or frontend components in any combination. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and, for example, the Internet.

[0249] A computing system can include clients and servers. Clients and servers are generally remote from each other and typically interact via a communication network. The client-server relationship arises from computer programs running on each computer and having a client-server relationship with each other. In some embodiments, the server sends data, such as an HTML page, to a user device for the purpose of displaying data to a user interacting with a device acting as a client and receiving user input from that user. Data generated on the user device, such as the results of user interactions, can be received by the server from the device.

[0250] While this specification includes details of many specific embodiments, these should not be construed as limiting the scope of any invention or claimable content, but rather as descriptions of features that may be specific to a particular embodiment of a particular invention. Certain features described herein as separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described as a single embodiment may also be implemented in multiple embodiments, individually or in any preferred secondary combination. Furthermore, features may be described above as functioning in a particular combination, and even if initially claimed as such, one or more features from the claimed combination may be removed from the combination, and the claimed combination may cover subcombinations or variations of subcombinations.

[0251] Similarly, while operations are shown in the drawings and described in a specific order in the claims, this should not be understood as requiring that such operations be performed in a specific or sequential order shown, or that all shown operations be performed, in order to obtain the desired results. In certain situations, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the program components and systems described can generally be integrated into a single software product or packaged into multiple software products.

[0252] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions described in the claims may be performed in a different order, and this may still yield desirable results. As an example, the process shown in the accompanying drawings does not necessarily require the actions to be performed in the specific order or sequence shown to obtain the desired results. In some cases, multitasking and parallel processing may be advantageous.

[0253] The nature of this disclosure may be as described in the following clauses.

[0254] Article 1. Initializing data items, Obtaining a conditional input that characterizes one or more desired properties of the aforementioned data item, The process includes generating a final data item having one or more desired properties, wherein the generation is performed in each of a plurality of update iterations. Identifying the current data item at the time of the update iteration, From the aforementioned current data item, determine the analytical estimate of the noise component of the aforementioned current data item, To generate a first denoising output for the update iteration, a spreading neural network is used to process a first spreading input for the update iteration, which includes the current data item and a representation of the conditional input, wherein the first denoising output defines a predicted value of the residual error between the noise component of the current data item and the analytical estimate of the noise component. At least the first denoising output is used to generate the final denoising output for the update iteration, A method performed by one or more computers, comprising updating the current data item using the final denoising output and the analytical estimate of the noise component.

[0255] Article 2. Determining an analytically estimated value of the noise component of the current data item from the current data item comprises: determining a scaling factor for the update iteration, wherein the scaling factor is a time-dependent scaling factor that depends on the update iteration; and applying the scaling factor to the current data item to generate the analytically estimated value. The method according to Clause 1, comprising the foregoing steps.

[0256] Clause 3. Determining the scaling factor comprises determining the scaling factor based on an estimated value of a standard deviation of a distribution of final data items that can be generated by the diffusion neural network. The method according to Clause 2, comprising the foregoing step.

[0257] Clause 4. The method according to Clause 3, further comprising receiving the estimated value of the standard deviation as an input from a user.

[0258] Clause 5. The method according to Clause 3, further comprising setting the estimated value of the standard deviation to a standard deviation of training items used for training the diffusion neural network.

[0259] Clause 6. The method according to any one of Clauses 1 to 5, wherein the first denoising output is a predicted value of a standardized residual error between the noise component of the current data item and the analytically estimated value of the noise component.

[0260] Clause 7. Updating the current data item using the final denoising output comprises calculating a final estimated value of the noise component from the final denoising output and the analytically estimated value; and updating the current data item using the final estimated value of the noise component. The method according to any one of Clauses 1 to 6, comprising the foregoing steps.

[0261] Article 8. Generating the final denoised output for the update iteration from at least the first denoised output is: The spreading neural network processes one or more additional spreading inputs for the update iteration, each including a representation of a conditioned input different from the current data item, in order to generate each additional denoising output for the update iteration for each additional spreading input, The method according to any one of the claims 1 to 7, comprising combining the first denoising output for the update iteration with the one or more additional denoising outputs for the update iteration in order to generate the final denoising output for the update iteration.

[0262] Article 9. To generate a final denoising output for the update iteration, the first denoising output for the update iteration and the one or more additional denoising outputs are combined. The method according to clause 8, comprising combining the first denoising output for the update iteration and the one or more additional denoising outputs for the update iteration based on guidance weights for the update iteration in order to generate the final denoising output for the update iteration.

[0263] Article 10. To generate the final denoising output for the update iteration, the first denoising output for the update iteration and the one or more additional denoising outputs for the update iteration are combined based on the guidance weights for the update iteration. The method according to Clause 9, comprising using normalization guidance to combine the first denoised output for the update iteration and the one or more additional denoised outputs for the update iteration based on guidance weights for the update iteration in order to generate the final denoised output for the update iteration.

[0264] Article 11. The method according to any one of the clauses 1 to 10, further comprising setting the final data item as the current data item after it has been updated in the last update iteration of the plurality of update iterations.

[0265] Article 12. The data item or target data item includes image, video, or audio data, as described in any one of the provisions of Clauses 1 to 11.

[0266] Article 13. Initializing data items, Obtaining a conditional input that characterizes one or more desired properties of the aforementioned data item, The process includes generating a final data item having one or more desired properties, wherein the generation is performed in each of a plurality of update iterations. Obtaining guidance weights for the aforementioned update iteration, Identifying the current data item at the time of the update iteration, To generate a first denoising output for the update iteration, a spreading neural network is used to process the first spreading input for the update iteration, which includes the current data item and a representation of the conditional input. To generate an additional denoising output for the update iteration, the spread neural network is used to process an additional spread input for the update iteration, including the current data item. The process involves generating a final denoised output for the update iteration from the first denoised output and the additional denoised output, For each element of the final noise reduction output, the difference between the value of the corresponding element in the first noise reduction output and the value of the corresponding element in the additional noise reduction output is determined. For each element of the final noise reduction output, the sign of the difference between the value of the corresponding element in the first noise reduction output and the value of the corresponding element in the additional noise reduction output is determined. for each element of said final denoised output, determining a value of the element in said final denoised output based on a sign of a difference of said element and a guidance weight for said update iteration; said generating comprising the above; updating said current data item using said final denoised output; a method executed by one or more computers, the method comprising the above.

[0267] Clause 14. said first denoised output defines a prediction value of a residual error between a noise component of said current data item and an analytic estimate of said noise component given said first diffusion input, the method of Clause 13, wherein said additional denoised output defines a prediction value of said residual error between said noise component of said current data item and said analytic estimate of said noise component given said additional diffusion input.

[0268] Clause 15. said first denoised output defines a prediction value of a noise component of said current data item given said first diffusion input, the method of Clause 13, wherein said additional denoised output defines a prediction value of a noise component of said current data item given said additional diffusion input.

[0269] Clause 16. determining a value of the element in said final denoised output based on the sign of the difference of said element and the guidance weight for said update iteration comprises: setting the value of said element equal to said guidance weight when said sign is positive; setting the value of said element equal to a negative value of said guidance weight when said sign is negative; the method according to any one of Clauses 13 to 15, comprising the above.

[0270] Clause 17. the method according to any one of Clauses 13 to 16, wherein said additional diffusion input for said update iteration comprises said current data item and a representation of a negative conditioning input.

[0271] Article 18. The method according to any one of the clauses 13 to 16, wherein the additional diffusion input for the update iteration includes the current data item and a representation of a conditional input designated to indicate unconditional generation.

[0272] Article 19. The guidance weights are the same in each update iteration, as described in any one of clauses 13 to 18.

[0273] Article 20. Obtaining the aforementioned guidance weights means The method according to any one of the clauses 13 to 18, comprising applying a function to the time step corresponding to the update iteration in order to determine guidance weights for the update iteration.

[0274] Article 21. The method according to clause 20, wherein the time step corresponding to the update iteration is selected from a time window, and the function maps each of the multiple intervals in the time window to a corresponding guidance weight.

[0275] Article 22. The data item or target data item includes image, video, or audio data, as described in any one of the provisions of Clauses 1 to 21.

[0276] Article 23. The output data item receives the respective input scalar values ​​for each of one or more properties, For each of the one or more properties, it is determined whether the input scalar value of each property is the upper or lower limit of the desired value of the property in the output data item. For each of the one or more properties, an indicator value is generated that indicates whether the respective input scalar value of the property is the upper or lower limit of the desired value of the property in the output data item. A method performed by one or more computers, comprising using a generative neural network to process an input for each of the one or more properties, which includes the respective input scalar values ​​of the property and the respective indicator values ​​of the property, in order to generate the output data items.

[0277] Article 24. The generative neural network is a diffuse neural network, The method according to Clause 23, wherein generating the output data item is to update the current data item in each of a plurality of update iterations by using the spread neural network to process the spread input for the update iteration, which includes the current data item and, for each of the one or more properties, the respective input scalar values ​​of the property and the respective indicator values ​​of the property, in order to generate a denoising output for the update iteration.

[0278] Article 25. The method according to clause 24, wherein the denoising output defines a predicted value of the residual error between the noise component of the current data item and the analytical estimate of the noise component given the spreading input.

[0279] Article 26. Generating the aforementioned output data items means The process includes determining, for each of the one or more properties, one or more embeddings representing the respective input scalar values ​​of the property and the respective indicator values ​​of the property. The method according to Clause 24 or Clause 25, wherein the spreading input for the update iteration includes the current data item for generating a denoising output for the new iteration using the spreading neural network, and for each of the one or more properties, the respective input scalar value of the property and the respective indicator value of the property.

[0280] Article 27. Determining whether each of the input scalar values ​​of the aforementioned property is an upper or lower limit of the desired value of the property of the output data item is: The method according to any one of the clauses 23 to 26, comprising receiving an input indicating whether the respective input scalar values ​​of the property are upper or lower limits of a desired value for the property of the output data item.

[0281] Article 28. Determining whether each of the input scalar values ​​of the aforementioned property is an upper or lower limit of the desired value of the aforementioned property of the output data item is: The method of any one of the clauses 23 to 26, which includes determining that, if each of the aforementioned input scalar values ​​is specified as part of a positive prompt, the respective input scalar values ​​of the property are the lower limit of the desired values ​​of the property of the output data item.

[0282] Article 29. Determining whether each of the input scalar values ​​of the aforementioned property is an upper or lower limit of the desired value of the aforementioned property of the output data item is: The method of any one of the clauses 23 to 26, which includes determining that each of the aforementioned input scalar values ​​of the property is an upper limit of the desired value of the property of the output data item, if each of the aforementioned input scalar values ​​is specified as part of a negative prompt.

[0283] Article 30. A method for training a generative neural network, which is performed by one or more computers, Obtaining target data items, For each of one or more properties of the target data item, obtain the ground truth value of the property. For each of the one or more properties, the value of the property is sampled, For each property, in order to generate a model output for a model input, a generative neural network is used to process a model input that includes (i) the sampling value of the property and (ii) an indicator value indicating whether the sampling value of the property is greater than or less than the ground truth value of the property. A method comprising training the generative neural network based on a loss function that measures the quality of the model output relative to the target model output of the target data item.

[0284] Article 31. The generative neural network is a diffuse neural network, The model input further includes noisy data items generated from the target data items, The aforementioned model output is a noise-reduced output. The method according to clause 30, wherein the target model output of the target data item is the target denoising output of the target data item.

[0285] Article 32. The method according to clause 31, wherein the denoising output defines a predicted value of the residual error between the noise component of the noisy data item and the analytical estimate of the noise component given the model input.

[0286] Article 33. After the aforementioned training, The output data item receives the respective input scalar values ​​for each of one or more properties, For each of the one or more properties, it is determined whether the respective input scalar value of the property is the upper or lower limit of the desired value of the property in the output data item. For each of the one or more properties, generate an indicator value that indicates whether the respective input scalar value of the property is the upper or lower limit of the desired value of the property in the output data item. The method according to any one of the claims 30 to 32, further comprising using the generative neural network to process an input for each of the one or more properties, which includes the respective input scalar values ​​of the property and the respective indicator values ​​of the property, in order to generate the output data items.

[0287] Article 34. The data item or target data item includes image, video, or audio data, as described in any one of the provisions of Clauses 1 to 33.

[0288] Article 35. To generate a first output data item, it receives a first input that identifies multiple context data items, Obtaining the respective embeddings of each of the aforementioned multiple context data items, To generate a combined embedding, the respective embeddings of each of the multiple context data items are combined, A method performed by one or more computers, comprising processing an input including the connected embedding using a spreading neural network to generate the first output data item.

[0289] Article 36. To generate a second output data item, receive a second input that identifies a specific data item, Obtaining the embedding of the aforementioned specific context data item, The method of Clause 35, further comprising processing the input, including the embedding of the particular context data item, using the spreading neural network to generate a second output data item.

[0290] Article 37. Obtaining the embedding of the aforementioned specific context data item is: The method according to Clause 36, comprising processing the specific context data item using an embedding neural network to generate the embedding of the specific context data item.

[0291] Article 38. To generate a combined embedding, combine the respective embeddings of each of the multiple context data items, The method described in any one of the provisions 35 to 37, which includes averaging each of the aforementioned embeddings.

[0292] Article 39. To generate the combined embedding, the respective embeddings of each of the plurality of context data items are combined. The system receives user input specifying the respective weights of each of the aforementioned context data items, The method according to any one of the clauses 35 to 37, comprising calculating a weighted sum of the respective embeddings of the context data items based on the respective weights of the corresponding context data items.

[0293] Article 40. Obtaining the respective embeddings of each of the aforementioned multiple context data items is: The method according to any one of the clauses 35 to 39, comprising processing each of the plurality of specific context data items using an embedding neural network to generate each of the aforementioned embeddings.

[0294] Article 41. The aforementioned spreading neural network is Maintaining the respective embeddings of each of the multiple training context data items, Maintaining data that clusters each of the embeddings of the aforementioned multiple training context data items into multiple clusters, Obtaining input to specify a target data item from the aforementioned multiple context data items, Selecting a context embedding for training the spreading neural network, the selection including (i) selecting the embedding of the target context data item, or (ii) selecting the centroid of the cluster to which the embedding of the target context data item belongs. The method according to any one of the clauses 35 to 40, which is trained by performing an operation that includes training the spreading neural network using the selected context embedding and the target data item.

[0295] Article 42. (i) Selecting the embedding of the target context data item, or (ii) Selecting the centroid of the cluster to which the embedding of the target context data item belongs, With probability p, (i) select the embedding of the target context data item, The method according to Clause 41, comprising (ii) selecting the centroid of the cluster to which the embedding of the target context data item belongs, with probability 1-p.

[0296] Article 43. Using a spread neural network to process the input, including the connected embeddings, in order to generate the first output data item, The method according to any one of the claims 35 to 42, comprising updating the current data item in each of a plurality of update iterations by processing the first spreading input for the update iteration, which includes the current data item and the combined embedding, using the spreading neural network in each update iteration to generate a first denoising output for the update iteration.

[0297] Article 44. The method according to clause 43, wherein the first denoising output defines a predicted value of the residual error between the noise component of the current data item and the analytical estimate of the noise component given the first spreading input.

[0298] Article 45. The method of Clause 43 or Clause 44, wherein updating the current data item in each of the plurality of update iterations further comprises, in each update iteration, processing an additional spread input for the update iteration, including the current data item and not the combined embedding, using the spread neural network to generate an additional denoising output for the update iteration; combining the first spread output for the update iteration and the additional denoising output based on guidance weights for the update iteration to generate the denoising output; and updating the current data item using the denoising output.

[0299] Article 46. The data item or target data item includes image, video, or audio data, as described in any one of the provisions of Clauses 1 to 45.

[0300] Article 47. One or more computers, A system comprising one or more storage devices communicably connected to one or more computers, wherein one or more storage devices store instructions that, when executed by one or more computers, cause one or more computers to perform the operations described in any one of the clauses 1 to 46.

[0301] Article 48. One or more non-temporary computer storage media, which, when executed by one or more computers, store instructions causing the one or more computers to perform the operations of each of the methods described in any one of the clauses 1 to 46.

Claims

1. Receiving a first input that identifies multiple context data items in order to generate a first output data item, Obtaining the respective embeddings of each of the aforementioned multiple context data items, To generate a combined embedding, the respective embeddings of each of the multiple context data items are combined, A method performed by one or more computers, comprising processing the input, including the connected embedding, using a spreading neural network to generate the first output data item.

2. To generate a second output data item, a second input is received that identifies a specific data item, Obtaining the embedding of the aforementioned specific context data item, The method according to claim 1, further comprising processing the input, including the embedding of the specific context data item, using the spreading neural network to generate the second output data item.

3. Obtaining the embedding of the aforementioned specific context data item is: The method according to claim 2, comprising processing the specific context data item using an embedding neural network to generate the embedding of the specific context data item.

4. To generate a combined embedding, combine the respective embeddings of each of the multiple context data items, The method according to any one of claims 1 to 3, comprising averaging each of the aforementioned embeddings.

5. To generate a combined embedding, combine the respective embeddings of each of the multiple context data items, The system receives user input specifying the weight of each of the aforementioned context data items, The method according to any one of claims 1 to 3, comprising calculating a weighted sum of the respective embeddings of the context data items based on the respective weights of the corresponding context data items.

6. Obtaining the respective embeddings of each of the aforementioned multiple context data items is: The method according to any one of claims 1 to 5, comprising processing each of the plurality of specific context data items using an embedding neural network to generate each of the aforementioned embeddings.

7. The aforementioned spreading neural network is Maintaining the respective embeddings of each of the multiple training context data items, Maintaining data that clusters each of the embeddings of the aforementioned multiple training context data items into multiple clusters, Obtaining input to specify a target data item from the aforementioned multiple context data items, Selecting a context embedding for training the spreading neural network, the selection including (i) selecting the embedding of the target context data item, or (ii) selecting the centroid of the cluster to which the embedding of the target context data item belongs. The method according to any one of claims 1 to 6, which is trained by performing an operation including training the spreading neural network with the selected context embedding and the target data item.

8. (i) Selecting the embedding of the target context data item, or (ii) selecting the centroid of the cluster to which the embedding of the target context data item belongs, With probability p, (i) select the embedding of the target context data item, The method according to claim 7, comprising (ii) selecting the centroid of the cluster to which the embedding of the target context data item belongs, with probability 1-p.

9. Using a spread neural network to process the input, including the connected embeddings, in order to generate the first output data item, The method according to any one of claims 1 to 8, comprising updating the current data item in each of a plurality of update iterations by processing a first spreading input for the update iteration, including the current data item and the combined embedding, using the spreading neural network in each update iteration to generate a first denoising output for the update iteration.

10. The method according to claim 9, wherein the first denoising output defines a predicted value of the residual error between the noise component of the current data item and the analytical estimate of the noise component given the first spreading input.

11. The method according to claim 9 or 10, wherein updating the current data item in each of the plurality of update iterations further comprises, in each update iteration, processing an additional spreading input for the update iteration, including the current data item and not the combined embedding, using the spreading neural network to generate an additional denoising output for the update iteration; combining the first spreading output for the update iteration and the additional denoising output based on guidance weights for the update iteration to generate a denoising output; and updating the current data item using the denoising output.

12. The method according to any one of claims 1 to 11, wherein the data item or target data item includes image, video, or audio data.

13. One or more computers, A system comprising one or more storage devices communicably connected to one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform the operation according to any one of claims 1 to 12.

14. One or more non-temporary computer storage media, which, when executed by one or more computers, store instructions causing one or more computers to perform the operations of each method described in any one of claims 1 to 12.