Contrastive language-image foundational models as detectors of generative model generated images

By employing a trained embedding neural network and classifier to analyze media assets, the method effectively identifies whether media were generated by generative models and determines the specific model involved, addressing the limitations of existing systems.

WO2025122163A1PCT designated stage expired Publication Date: 2025-06-12GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2023/083212
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-08
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing systems cannot accurately determine whether media, such as images, are originally generated or produced by generative models, and if so, which specific model generated the media.

Method used

The method involves using a trained embedding neural network to process input media assets, generating embeddings, and then using a classifier to produce scores for multiple categories, including real and generative model categories, to determine the likelihood and source of media generation.

Benefits of technology

This approach allows for accurate detection of media generated by generative models and identification of the specific model used, with a relatively low error rate, enabling users to make informed decisions about media incorporation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2023083212_12062025_PF_FP_ABST
    Figure US2023083212_12062025_PF_FP_ABST
Patent Text Reader

Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for detecting images generated by artificial intelligence using trained neural networks. In one aspect, a method comprises receiving a input media asset, processing the input media asset using a trained embedding neural network to generate an embedding, processing the embedding using a classifier to generate a respective score for each of the multiple categories, where two or more of the multiple categories each correspond to a different generative model of multiple generative models, and determining whether the input media asset was generated by one of the multiple generative models based on the respective scores.
Need to check novelty before this filing date? Find Prior Art

Description

CONTRASTIVE LANGUAGE-IMAGE FOUNDATIONAL MODELS ASDETECTORS OF GENERATIVE MODEL GENERATED IMAGESBACKGROUND

[0001] This specification relates to processing data using machine learning models.

[0002] Machine learning models receive an input and generate an output, e.g., a predicted output, based on the received input. Some machine learning models are parametric models and generate the output based on the received input and on values of the parameters of the model.

[0003] Some machine learning models are deep models that employ multiple layers of models to generate an output for a received input. For example, a deep neural network is a deep machine learning model that includes an output layer and one or more hidden layers that each apply a non-linear transformation to a received input to generate an output.SUMMARY

[0004] This specification describes a method for detecting inputs generated by artificial intelligence using trained neural networks.

[0005] According to a first aspect, there is a method by one or more data processing apparatus that includes receiving a input media asset, processing the input media asset using a trained embedding neural network to generate an embedding, processing the embedding using a classifier to generate a respective score for each of a set of multiple categories, where two or more of the multiple categories each correspond to a different generative models of multiple generative models, and determining whether the input media asset was generated by one of the multiple generative models based on the respective scores.

[0006] In some implementations, the input media asset is of a first modality’, and the trained embedding neural network is trained with another embedding neural network that processes an input of a second modality on a contrastive learning loss function using training examples that each include a training input of the first modality and a training input of a second modality.

[0007] In some implementations, the first modality is audio or images, and the second modality is text.

[0008] In some implementations, the input media asset is an image.

[0009] In some implementations, a first category' corresponds to a real image category' that represents images that are not generated by one of the multiple generative models, and determining whether the input media asset was generated by one of the multiple generative models based on the respective scores includes determining that the input media asset was notgenerated by one of the multiple generative models based on the respective score for each category.

[0010] In some implementations, the method further includes, prior to processing the input media asset using the classifier, processing the embedding using a second classifier to generate a respective score for each of another set of multiple categories, where the second set of multiple categories include a first category and a second category’, where the first category corresponds to the input media asset being real, and the second category corresponds to the input media asset being a synthetic input generated by a generative model.

[0011] In some implementations, the classifier is trained by, for each of the multiple generative models, generating multiple outputs using the generative model using multiple prompts, generating a respective embedding of each of the multiple outputs, and training the classifier on training data that includes multiple training examples that each include a respective embedding of a respective generated output and a respective label that identifies a generative model that generated the respective generated output.

[0012] In some implementations, the classifier is trained on a loss function that measures, for each of the multiple training examples, an error between an output generated by the classifier by processing the embeddings in the training example and the respective label in the training example.

[0013] The subject matter described in this specification can be implemented in particular embodiments so as to realize one or more of the following advantages.

[0014] In conventional systems, a user can obtain media, such images or audio, to incorporate into a digital domain of the user or for some other purpose. However, the user may not be aware whether the obtained media is original or has been generated by a neural network, such as a generative model, and it may be beneficial for the user be aware whether the media is synthetic (e.g., generated by a generative model), and if so, which generative model generated the media.

[0015] Some existing techniques address this issue by merely detecting whether the media was generated by a neural network. For example, an existing system can detect that an image was generated using a generative model, but the existing system cannot determine whether the image was specifically generated by a neural network that violates certain policies.

[0016] In contrast, this specification describes techniques that allow the system to determine the probability that media, such as an image, was generated by a generative model and which generative model generated the media. In particular, the system can receive an input and process the input using a trained embedding neural network to generate an embedding. The system can then process the embedding using a classifier to generate a score for each ofmultiple categories and determine the probability that the media was generated by a generative model and which generative model generated the media.

[0017] In some examples, the categories include that the input is “real’’ and multiple categories that each correspond to a different generative model. The system then determines the probability that the input was generated by any of the generative models based on the score for each category. For example, the system can process an embedding of an input using the classifier, and the classifier can generate a respective score for whether the input is real, whether the input was generated by a particular neural network of the multiple neural networks, or both.

[0018] Therefore, the system leverages the classifier to determine which generative model generated the media with high accuracy and relatively low error rate, and the system can provide the information of the determination to the user. Based on the information, the user can more accurately identify whether to incorporate the media into the digital domain.

[0019] The details of one or more embodiments of the subject matter of this specification are set forth in the accompanying drawings and the description below. Other features, aspects, and advantages of the subject matter will become apparent from the description, the drawings, and the claims.BRIEF DESCRIPTION OF THE DRAWINGS|00020| FIG. 1 is a block diagram of an example system.

[0021] FIG. 2 is a block diagram of an example training system.

[0022] FIG. 3 is a block diagram of an example classification system.

[0023] FIG. 4 is a flow diagram of an example process for detecting inputs generated by a neural network using trained neural networks.

[0024] Like reference numbers and designations in the various drawings indicate like elements.DETAILED DESCRIPTION

[0025] FIG. 1 shows an example system 100. The system 100 is an example of a system implemented as computer programs on one or more computers in one or more locations in which the systems, components, and techniques described below are implemented.

[0026] The system 100 is configured to detect whether an input is synthetic or real, where a synthetic input is generated by a generative model (e.g., a generative neural network), while a real input is measured or captured by a sensor. For example, real inputs can include imagescaptured by a camera, videos captured by a camera, audio captured by a recording device, or audio captured by a microphone. The synthetic inputs can include images generated by a generative neural network, video generated by a generative neural network, or audio generated by a generative neural network. The real inputs and the synthetic inputs do not include text.

[0027] The system 100 includes an input classification system 102, a user device 104, and training data 106. The user device 104 can be a computer, and the user device 104 can provide the input media asset 114 to the input classification system 102. The input media asset 114 can be audio or an image.

[0028] The input classification system 102 includes an embedding neural network 110, a training system 108, and a classifier 112.

[0029] The input classification system 102 is configured to process the input media asset 114 using the embedding neural network 110 and the classifier 112 to generate a classification 120 that identifies whether the input media asset 114 was generated by a generative model and, if so, which generative model generated the input media asset 114.

[0030] In particular, the system 102 uses the embedding neural network 110 to generate an embedding 116 by processing the input media asset 114. The embedding 1 16 is an ordered collection of numeric values (e.g., a vector or matrix of floating point or other numeric values that represents the input media asset 114. The embedding neural network 110 can include one or more convolutional neural network layers followed by one or more fully-connected layers. The embedding neural network 110 has been pre-trained to receive the input media asset 114 and to process the input media asset 1 14 to generate the embedding 1 1 .

[0031] The system 102 then processes the embedding 116 using the classifier 112 to generate the classification 120 based on category scores 118, as described in further detail below with reference to FIG. 3. The classifier 112 can be configured to perform any of a variety of classification tasks. As used in this specification, a classification task is any task that requires the classifier 112 to generate the classifier scores 118 that includes a respective score for each of a set of multiple categories and to then select one or more of the categories as a “classification” for the network input using the respective scores.

[0032] The classifier 112 can have any appropriate architecture that allows the classifier to map an input to an output that includes a respective score for each of the categories.

[0033] As one example, when the inputs are images, the classifier 112 can be a convolutional neural network, e.g., a neural network having a ResNet architecture, an Inception architecture, an EfficientNet architecture, a multi-layer perceptron (MLP) architecture, and so on. or a Transformer neural network, e.g., a vision Transformer.

[0034] In some examples, the category scores 118 correspond to multiple categories, which include a category corresponding to a real input (e.g., an input not generated by a neural network). If the system determines that the input is a real input based on the category scores 118, the system can output the classification 120 to the user device 104.

[0035] In some examples, the first set of categories can include a real input category' and multiple categories each corresponding to a generative model, and the system can determine whether the input is a real input, or which particular generative model generated the input based on the category scores 118. The system can then output the classification 120 based on the category scores 118, as described in further detail with reference to FIG. 3.

[0036] In some examples, prior to processing the embedding 116 using the classifier 112, the system can process the embedding 116 using a second classifier. The second classifier can generate category scores for a second set of categories that include a category for the input being real and a category' for the input being synthetic (e.g., generated by a generative model). In this case, in response to using the second classifier to determine that the input is synthetic based on the second set of category scores, the system can use the classifier 112 to determine which of the generative models generated the input based on the category scores 118 each corresponding to a generative model, and the system can output the classification 120 based on the second determination, as described in further detail below with reference to FIG. 3. In the case that the system determines the input is real, the system refrains from using the classifier 112 to determine which of the generative models generated the input.

[0037] Based on the classification 120, a user of the user device 104 can determine whether to incorporate the input into a respective digital domain, such as a website. For example, if the classification 120 indicates that the image w as generated by an “unsafe” generative model (e.g., a model that does not have rights to the generated image), the user may refrain from including the image in their website. In another example, if the classification 120 indicates that the image is real or if the image was generated by a “safe” generative model, the user can safely incorporate the image into their website.

[0038] Prior to using the classifier, the input classification system 102 trains the classifier 112 by processing training examples 122 using the pre-trained embedding neural network 110 and the training system 108. In particular, the training system 108 can process the training examples 122 generated by the embedding neural network 110 to train the classifier 112, as described in further detail below- with reference to FIG. 2. The embedding neural network 110 is pre-trained on a contrastive loss using training inputs 124.

[0039] In general, the system can determine a probability that the input was generated by a particular generative model with a relatively low error rate, i.e., can accurately distinguish between images generated by different generative models. While specific error rates will depend on the amount of training data available and the number of generative models in the set, two examples of accuracies achieved by the system are given in Table 1 below. In particular, Table 1 shows the accuracy of the model in distinguishing images generated by different versions of the Stable Diffusion (SD) generative model, with one row showing the accuracy in distinguishing between SD_1.5 and SD 1.5 - CFG 2.5 and the other row showing the accuracy in distinguishing between SD 1.5 and SD 1.5 - CFG_2.5-p, where CFG represents classifier-free guidance.Table 1

[0040] The training inputs 124 include multiple sets of inputs, where at least some of the inputs correspond to distinct generative models. The accuracy and the error rate depend on a number of the neural networks (e.g., generative models) that are represented by the training inputs, the overall size of the set of training inputs 124, and the distribution of the training inputs 124. In particular, the training inputs 124 can include multiple sets of images that correspond to multiple generative models (e.g., the three generative models of Table 1), where the size of the sets of images is configured by the training system 108. For example, in the first row of Table 1, Set 1 corresponds to a set of images generated by the generative model SD_1.5, and Set 2 corresponds to a set of images generated by the generative model SD_1.5 - CFG_2.5-p. The accuracy percentage is based on how accurately the system determines the probability that each image from each set was generated by generative model SD_1.5 or the generative model SD 1.5 - CFG 2.5-P.

[0041] In another instance, for each of the generative models, the training inputs 124 generated by the model can have a respective distribution of the images. For example. Set 1 and Set 2 of Table 1 can include the same distribution of different images (e.g., 100 images of a fish, 100 images of a dog, and 100 images of a cat).

[0042] In some examples, the training inputs 124 include real images (e.g., original images not generated by a neural network) and images generated by multiple generative models. Forexample, if the one or more categories include a ‘‘real’' category, the training inputs 124 include one or more real images.

[0043] FIG. 2 is a block diagram of an example training system, e.g., the training system 108 described with reference to FIG. 1.

[0044] The training system 108 can train the classifier 112 to generate category7scores 118 in order for the system 102 to detect whether the input was generated by a generative model.

[0045] The training system 108 trains the classifier 112 to generate training category scores 204 on a loss function using the training examples 122.

[0046] The system can use the embedding neural network 110 to generate the training examples 122 by processing the training inputs 124 extracted from the training data 106. The system can extract the training inputs 124 from the training data 106 and provide the training inputs 124 to the embedding neural network 1 10. The training inputs 124 include multiple real images and multiple images generated by multiple generative models 202

[0047] In particular, each of the generative models 202 generate one or more training inputs 124. The generative models 202 can be diffusion models that generate synthetic data items conditioned on prompts, such as hierarchical diffusion models, latent diffusion models, and denoising diffusion models. In this case, the same prompts are used to prompt each generative model to generate respective synthetic data items. In some other examples, the generative models 202 can be conditional generative models (e.g., conditional generative adversarial networks (cGANs).

[0048] For each of the generative models 202, the system can provide respective prompts for generating the inputs, and the generated model inputs have a certain distribution that is equal among the multiple generative models 202. For example, the system can provide a prompt to generate a certain amount of images of a fish, a clown, and drums to each of the generative models 202. Based on the prompts, the generative model 1 202-A can generate 1000 images of a fish, 1000 images of a clown, and 1000 images of drums, and the other generative models 202-N can have a same distribution. In some examples, the real images have the same distribution. That is, for each generative model 202, the training inputs 124 include approximately the same number of images for each prompt in a predetermined set of prompts.

[0049] Additionally, each of the inputs can correspond to a label that indicates whether the input is real or generated by a generated by a generative model 202, and, if generated by a generated by a generative model 202, which of the generative models 202 generated the image. For example, an image of the training inputs 124 provided to the embedding neural network 110 can include a label that the generative model 1 202-A. generated the image

[0050] The embedding neural network 110 can process the training inputs 124 to generate a training example 122 for each of the inputs. The training example 122 includes a respective embedding for the input and the label for the input based on whether the image was real, generated by a generative model, and, optionally, which generative model generated the image.

[0051] In particular, the embedding neural network 110 can be trained with another embedding neural network that processes inputs of a particular modality, such as an image or audio, and inputs of another modality, such as text, on a symmetric cross entropy loss function. For example, the inputs to the other embedding neural network can include an image and a prompt with text corresponding to the image. In another example, the inputs to the other embedding neural network can include audio and a prompt with text corresponding to the audio.

[0052] The training system 108 can then process the training examples 122 to train the classifier 1 12 to generate training category scores 206. For each training example 122, the classifier 112 processes the embedding of the training example, and the classifier 112 204 generates the training category' scores 206, where the training category' scores 206 include a respective score for a real data category, a category of whether a generative model generated the input, and. optionally, a category for each of the generative models. The training system 108 trains the classifier 112 on a loss function that measures an error between the training category' scores 206 and the label (e.g., the ground truth label) of the training example 122. For example, the loss function can be a cross entropy loss function.|00053| FIG. 3 is a block diagram of an example classification system e.g.. the input classification system 102 described with reference to FIG. 1. The input classification system 102 can use the embedding neural network 110 and the classifier 112 to generate a classification 120 for the input media asset 114.

[0054] The input classification system 102 can obtain an input media asset 114, such as an image. The input classification system 102 can process the input media asset 114 using the embedding neural network 110 to generate an embedding 116. The embedding 116 represents one or more features of the image in an embedding space.

[0055] The input classification system 102 can then process the embedding 116 using the trained classifier 112 to generate the category scores 118. The category scores 118 include a respective score for each category', and the input classification system 102 can generate the classification 120 based on a respective score being above a certain threshold, a respective score being relatively higher than the other respective scores by a certain amount, or both.

[0056] In some examples, the category scores 118 can include multiple categories each corresponding to a generative model and a category corresponding to the input being real. Thesystem can determine a classification 120 for the input by comparing each of the category scores. For example, if the category score of a first model is higher than a category score of the other models and the category score of the image being real, the system can determine that the classification 120 for the image indicates that the first model generated the image. That is, in this case, the system determines a classification from a first set of categories that include the image being real and multiple categories for each of the generative neural networks.

[0057] In some other examples, the input classification system 102 can include a second classifier 302. In this case, the second classifier 302 can generate category scores for the second set of categories that include the input being real and the input being synthetic. The system can use the second classifier 302 to process the embedding 116 prior to using the classifier 112 to generate category scores for whether the input is real or the input is generated by a generative model. In response to the second classifier 302 determining that a category score of the input being synthetic is higher than the input being real or that the category score of the input being synthetic is above a certain threshold, the system can provide the embedding 116 to the classifier 112.

[0058] The classifier 112 can then generate the category scores 1 18 by processing the embedding 116. The category scores 118 include a category for each of the multiple generative models. For example, the second classifier 302 can determine that the input media asset 114 is synthetic, and classifier 112 can determine which of the generative models generated the input based on the input being synthetic and based on the category scores 118. In this case, if the second classifier 302 determines that the input is real, the system refrains from processing the embedding 116 using the classifier 112, and the system outputs the classification 120 of the input being real. That is, in this case, the system first determines a classification from a first set of categories that include the input, such as an image, being real and the image being synthetic, and, based on the system determining that the image is synthetic, the system determines a classification from a second set of categories that include the multiple generative neural networks.

[0059] In some examples, none of the category scores 118 may be above a certain threshold or relatively higher than the other categories, and the system can output a classification 120 indicating that the classification is inconclusive, or that the classification is that the input is real.

[0060] FIG. 4 is a flow diagram of an example process for detecting inputs generated by a neural networks using trained neural networks. For convenience, the process 400 will be described as being performed by a system. For example, a system, e.g., the input classificationsystem 102 of FIG. 1, appropriately configured in accordance with this specification, can perform the process 400.

[0061] The system receives a test input (402). For example, the system can receive an image from a user device of a user.

[0062] The system processes the input media asset using the trained embedding neural network to generate an embedding (404). In particular, the trained embedding neural network has been trained jointly with another embedding neural network that processes an input of a modality different from a modality of the test input on a contrastive learning loss function using training examples. For example, the input of the different modality can be a text prompt, and the training examples can include images and corresponding text prompts. In some examples, the training examples can include audio and corresponding text prompts.

[0063] The system processes the embedding using a classifier to generate a respective score for each of multiple categories (406). In particular, at least two of the categories correspond to a different generative of multiple generative models. In some examples, one of the categories corresponds to a real image category that represents that the image was not generated by any of the generative models.

[0064] In some examples, before processing the input using the classifier, the system can process the embedding using a second classifier to generate a score for each of a second set of categories. The second set of categories include a first category corresponding to the input being real and a second category corresponding to the input being synthetic (e.g.. generated by a generative model). In this case, the system only processes the input using the first classifier and determines which generative model generated the input based on the category scores if the system determines that the second category score is higher than the first category score.

[0065] The classifier is trained on multiple outputs of each generative model using multiple corresponding prompts, where the training examples include a respective embedding for each output and a label that indicates whether the output is real, synthetic, and optionally, which generative model generated the output. In particular, the classifier is trained on a loss function that measures an error between the classification by the classifier and the label in the training example for each training example.

[0066] The system then determines whether the input media asset was generated by any one of the multiple generative models based on the respective scores (408). In some examples, the system can determine that the input media asset was not generated by any of the generative models based on the score corresponding to the real input category. In some other examples,the system determines that the input media asset was generated by one of the generative models based on the category score for that model being higher than the other category scores.

[0067] In some examples, none of the category scores 1 18 may be above a certain threshold or relatively higher than the other categories, and the system can output a classification 120 indicating that the classification is inconclusive, or that the classification is that the input is real.

[0068] This specification uses the term “configured’7in connection with systems and computer program components. For a system of one or more computers to be configured to perform particular operations or actions means that the system has installed on it software, firmware, hardware, or a combination of them that in operation cause the system to perform the operations or actions. For one or more computer programs to be configured to perform particular operations or actions means that the one or more programs include instructions that, when executed by data processing apparatus, cause the apparatus to perform the operations or actions.

[0069] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly-embodied computer software or firmware, in computer hardware, including the structures disclosed in this specification and their structural equivalents, or in combinations of one or more of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e.. one or more modules of computer program instructions encoded on a tangible storage medium, which may be non-transitory. for execution by. or to control the operation of, data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially-generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to suitable receiver apparatus for execution by a data processing apparatus.

[0070] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of apparatus, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit). The apparatus can optionally include, in addition to hardware, code that creates an execution environment for computer programs, e.g., code that constitutes processor firmware, a protocolstack, a database management system, an operating system, or a combination of one or more of them.

[0071] A computer program, which may also be referred to or described as a program, software, a software application, an app, a module, a software module, a script, or code, can be written in any form of programming language, including compiled or interpreted languages, or declarative or procedural languages; and it can be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A program may, but need not, correspond to a file in a file system. A program can be stored in a portion of a file that holds other programs or data, e.g., one or more scripts stored in a markup language document, in a single file dedicated to the program in question, or in multiple coordinated files, e.g., files that store one or more modules, subprograms, or portions of code. A computer program can be deployed to be executed on one computer or on multiple computers that are located at one site or distributed across multiple sites and interconnected by a data communication network.

[0072] In this specification the term ’ engine" is used broadly to refer to a software-based system, subsystem, or process that is programmed to perform one or more specific functions. Generally, an engine will be implemented as one or more software modules or components, installed on one or more computers in one or more locations. In some cases, one or more computers will be dedicated to a particular engine; in other cases, multiple engines can be installed and running on the same computer or computers.

[0073] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logic flows can also be performed by special purpose logic circuitry, e.g., an FPGA or an ASIC, or by a combination of special purpose logic circuitry and one or more programmed computers.

[0074] Computers suitable for the execution of a computer program can be based on general or special purpose microprocessors or both, or any other kind of central processing unit. Generally, a central processing unit will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a central processing unit for performing or executing instructions and one or more memory devices for storing instructions and data. The central processing unit and the memory can be supplemented by, or incorporated in, special purpose logic circuitry. Generally, a computer will also include, or be operatively coupled to receive data from or transfer data to, or both, one or more mass storage devices for storing data, e.g., magnetic, magneto-optical disks, or optical disks.However, a computer need not have such devices. Moreover, a computer can be embedded in another device, e.g., a mobile telephone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a Global Positioning System (GPS) receiver, or a portable storage device, e.g., a universal serial bus (USB) flash drive, to name just a few.

[0075] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media and memory devices, including by way of example semiconductor memory devices, e.g.. EPROM. EEPROM, and flash memory devices; magnetic disks, e g., internal hard disks or removable disks; magneto-optical disks; and CD- ROM and DVD-ROM disks.

[0076] To provide for interaction with a user, embodiments of the subject matter described in this specification can be implemented on a computer having a display device, e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user and a keyboard and a pointing device, e.g., a mouse or a trackball, by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback, e.g., visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including acoustic, speech, or tactile input. In addition, a computer can interact with a user by sending documents to and receiving documents from a device that is used by the user; for example, by sending web pages to a web browser on a user’s device in response to requests received from the web browser. Also, a computer can interact with a user by sending text messages or other forms of message to a personal device, e.g., a smartphone that is running a messaging application, and receiving responsive messages from the user in return.

[0077] Data processing apparatus for implementing machine learning models can also include, for example, special-purpose hardware accelerator units for processing common and computeintensive parts of machine learning training or production, i.e., inference, workloads.

[0078] Machine learning models can be implemented and deployed using a machine learning framework, e.g., a TensorFlow framework.

[0079] Embodiments of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface, a web browser, or an app through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-endcomponents. The components of the system can be interconnected by any form or medium of digital data communication, e.g.. a communication network. Examples of communication networks include a local area network (LAN) and a wide area network (WAN), e.g., the Internet.

[0080] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data, e.g., an HTML page, to a user device, e.g., for purposes of displaying data to and receiving user input from a user interacting with the device, which acts as a client. Data generated at the user device, e.g., a result of the user interaction, can be received at the server from the device.

[0081] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any invention or on the scope of what may be claimed, but rather as descriptions of features that may be specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially be claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.

[0082] Similarly, while operations are depicted in the drawings and recited in the claims in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0083] Particular embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the actions recited in the claims canbe performed in a different order and still achieve desirable results. As one example, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve desirable results. In some cases, multitasking and parallel processing may be advantageous.

[0084] What is claimed is:

Claims

CLAIMS1. A computer-implemented method, comprising: receiving an input media asset; processing the input media asset using a trained embedding neural network to generate an embedding; processing the embedding using a classifier to generate a respective score for each of a plurality of categories, wherein two or more of the plurality of categories each correspond to a different generative model of a plurality of generative models; and determining that the input media asset was generated by one of the plurality of generative models based on the respective scores.

2. The computer-implemented method of claim 1, wherein the input media asset is of a first modality, and wherein the trained embedding neural network is trained with another embedding neural network that processes an input of a second modality on a contrastive learning loss function using training examples that each include a training input of the first modality and a training input of a second modality.

3. The computer-implemented method of claim 2, wherein the first modality is audio or images, and wherein the second modality is text.

4. The computer-implemented method of claim 1, wherein the input media asset is an image.

5. The computer-implemented method of claim 1, wherein a first category corresponds to a real image category that represents images that are not generated by one of the plurality of generative models, and wherein determining whether the input media asset was generated by any one of the plurality of generative models based on the respective scores further comprises: determining that the input media asset was not generated by one of the plurality of generative models based at least in part on the respective score for the first category7.

6. The computer-implemented method of claim 1, further comprising: prior to processing the input media asset using the classifier: processing the embedding using a second classifier to generate a respectivescore for each of a second plurality' of categories, wherein the second plurality of categories comprise a first category and a second category, wherein the first category corresponds to the input media asset being real, and the second category corresponds to the input media asset being a synthetic input generated by a generative model.

7. The computer-implemented method of claim 6, the method further comprising: only performing the processing, processing, and determining in response to determining that the input media asset was generated by a generative model of a plurality of the generative models using the respective scores for the second plurality' of categories.

8. The computer implemented method of claim 1, wherein the classifier is trained by: for each of the plurality' of generative models, generating a plurality7of outputs using the generative model using a plurality' of prompts; generating a respective embedding of each of the plurality of outputs; and training the classifier on training data that comprises a plurality of training examples that each include a respective embedding of a respective generated output and a respective label that identifies a generative model that generated the respective generated output.

9. The computer implemented method of claim 8, wherein the classifier is trained on a loss function that measures, for each of the plurality7of training examples, an error between an output generated by the classifier by processing the embeddings in the training example and the respective label in the training example.

10. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform the operations of the respective method of any one of claims 1- 9.

11. One or more computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising the operations of the respective system of any one of claims 1-9.