Detecting synthetic visual media
By employing machine learning to separate and analyze foreground and background portions of visual media in non-spatial domains, the method effectively detects synthetic media, addressing inefficiencies in existing detection methods and enhancing reliability and scalability.
Patent Information
- Application Number
- PCT/EP2025/059652
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-17
- Filing Date
- 2025-04-08
- Publication Date
- 2026-01-22
AI Technical Summary
Existing methods for detecting synthetic visual media, such as deepfakes, are inefficient and unreliable, particularly when background objects are absent or when genuine images are manipulated, and they often require human expertise, making them costly and time-consuming.
A method utilizing machine learning models to identify and separate foreground and background portions of visual media, analyzing these portions in non-spatial domains like frequency or spectral domains to determine if the media is synthetic or genuine, using trained ML models to provide accurate classification.
Enables automated, scalable, and reliable detection of synthetic visual media, reducing false positives and costs associated with manual inspection, and improving accuracy by leveraging differences in frequency and spatial characteristics.
Smart Images

Figure EP2025059652_22012026_PF_FP_ABST
Abstract
Description
[0001] - 1 - Detecting Synthetic Visual Media Field of Invention 5 The present invention relates to a method for detecting synthetic images and video, and systems and computer programs for carrying out such methods. Background 10 There has been a recent proliferation in artificial intelligence (AI) tools for the generation and manipulation of images. This enables the production of an increasing volume of convincing synthetic visual media. 15 Identity verification often requires the submission of images of identity documentation (for instance a passport) and a current photo or video of the subject in question. Malicious actors may generate synthetic visual media to perpetrate fraud or gain access to sensitive information. Synthetic visual media may be manually doctored, may be fully AI generated, or may be AI-manipulated images that are based on an original or 20 genuine image. For example, “deepfake” images or videos, where visual media has been digitally manipulated to replace one person's likeness with that of another, can be used to perpetrate identify theft or gain unauthorised access to services. Of course, the malicious use of synthetic visual media is not limited to identity 25 verification. Synthetic visual media can be used to: facilitate false advertising, create images of non-existent products in use; enable the spread of disinformation, using a deepfake video impersonating a trusted source, or within social manipulation, presenting images of a sham lifestyle. 30 Synthetic visual media can be identified by inspection of a skilled individual. However, this is a costly and time-consuming process that requires a great deal of skill, knowledge and training. Additionally, there may be minute differences imperceivable to a human, which may give an indication that a piece of visual media is synthetic. Additionally, as AI tools have been developed for the creation of visual media, AI and machine learning 35 (ML) methods have been developed for their detection. However, standard ML methods 16216834.JAC.HSS P258779WO00 - 2 - may not have any additional understanding of the expected formats of visual media. Therefore, such existing techniques cannot achieve a high rate of accuracy and reliability. Korean patent application KR 20230070660 A discusses extracting information on 5 the boundary between foreground and background objects. However, this process is only applicable when background objects are present and often images used for identity verification require a plane background devoid of any background objects. Korean patent application KR 102523372 B1 discusses a method where the corneal 10 specular backscatter of eyes in images is analysed to determine if an image is synthetic or genuine. However, this method is only applicable to images that contain a face and may not identify a manipulated image containing a genuine face on an artificially generated background. 15 US patent application US 2022121868 A1 discusses separating audio from visual media in order to facilitate deepfake detection. However, this system requires the presence of audio and visual data to properly function. Chinese patent application CN 115100722 A discusses identifying coordinates of a 20 face (for example the leftmost end of the left eyebrow and the rightmost end of the right eyebrow) to better equip a classifier in identifying a face. However, this requires the image to contain the full face of a subject. Therefore, there is a need in a system and method that overcomes these problems. 25 The present system and method identify synthetic images and video, and can do so at scale or in an automated and more reliable fashion. 30 Entirely computer-generated visual media, such as images and video may be described as synthetic visual media. Typically, synthetic visual media is produced using generative artificial intelligence (AI) or similar technology. For example, the computer or AI system is provided with a prompt or other description (e.g., textural) of requested visual 35 media and in response, provides a response in the form of a computer file or stream of 16216834.JAC.HSS P258779WO00 - 3 - data forming the visual media. In contrast, genuine visual media is typically created using an optical system, such as a still or video camera that captures objects in the real world. Although a photograph or scan of a printed copy of an AI generated image is created using an optical system, this can still be synthetic media. Additionally, genuine visual media 5 include paintings or human-made digital images and / or videos. Computer processing may still be carried out on genuine visual media (e.g., to sharpen, re-colour, adjust contrast, remove artifacts, adjust the content, etc.) but the origin of genuine visual media is a real subject. The method and system provide an output indicating whether unknown visual media is synthetic or genuine. This output may trigger particular responses or actions 10 (e.g., access to a restricted area, the provision of an electronic service, issue of a verification certificate, etc.) or may form the basis of a report. Once visual media is provided (e.g., a digital image in the form of a file or data stream), a feature within the visual media is identified. This could be a face, in the case of 15 a portrait or a solid or complete or partially complete object. First and second portions of the visual media are also identified. The first portion contains the feature, and the second portion does not contain the feature. This may be achieved by detecting a boundary of the feature within the visual media. In an example implementation, the first portion may be a foreground of the visual media and the second portion may be a background of the visual 20 media. The portions may be contiguous portions of visual media or may comprise more than one separated areas (e.g., groups of pixels). The background, in particular, may comprise several separate parts of the image, video or visual media. These separate portions are provided to a machine learning (ML) model that has been trained (e.g., using many first and second portions of different items of visual media). The trained ML model 25 provides data indicating whether the provided visual media is synthetic or genuine. Separating out different portions of the visual media in this way provides a more efficient and effective way to determine the synthetic or genuine nature of visual media. The present method and system may take advantage of differences exhibited in a non-spatial domain (such as the frequency domain) of objects, features, foregrounds, and backgrounds 30 of synthetic visual media with genuine visual media. A spatial domain may be a domain where the representation of the image is in its original form, for example pixels and their positions (including RGB information). A non-spatial domain may be one in which pixels are not considered directly (such as the frequency domain). There are many other examples of spatial and non-spatial domains which may be known to the skilled person. 35 16216834.JAC.HSS P258779WO00 - 4 - In accordance with a first aspect, there is provided a method for detecting synthetic visual media, the method comprising the steps of: receiving visual media, wherein the visual media is synthetic visual media or genuine visual media; identifying at least one feature within the visual media; identifying a first portion of the visual media and second 5 portion of the visual media, the first portion of the visual media containing the at least one feature and the second portion of the visual media not including the at least one feature; providing the first portion and the second portion to a trained machine learning (ML) model; and the trained ML model providing data indicating the visual media to be synthetic and / or genuine. 10 This method utilises a machine learning (ML) or artificial intelligence (AI) model along with the identification of two portions of the visual media. This enables the method to more readily identify synthetic visual media compared to other available methods of detection. The synthetic visual media may be fully AI-generated, AI-manipulated genuine 15 images, or manually created fake visual media (or any combination of these). The applicability of the method to AI-manipulated images is of value in the identification of deep fakes and other fraudulent or dishonest uses of synthetic media, as discussed above. While ML models are discussed, it is readily apparent to the skilled person that AI models could be used interchangeably. The method may also be used to determine a quality or 20 realism of synthetic visual media. Optionally, the visual media may be an image and / or a video. The method is equally applicable to any type of visual media. The visual media may include a plurality of still images and / or video. The visual media (i.e., non-synthetic visual media) may be 25 captured by a camera or may be manually created (e.g., by a human using a computer, drawing or painting tools). This includes but is not limited to: digital or analogue photography, oil paint on canvas, watercolour paintings, graphite drawings, and / or digital art. Other visual media may be used and considered. Synthetic visual media may be any form of visual media generated by no or a minimum of human input (e.g., a human 30 provided a textual or aural prompt to an AI media generator). Optionally, the data indicating the visual media to be synthetic and / or genuine may be provided as a probability value. This enables appropriate action to be taken. For example, a particular action (e.g., refusal of an application based on the submitted visual 35 media) may be taken based on the probability value. 16216834.JAC.HSS P258779WO00 - 5 - Optionally, the method may further comprise the steps of: comparing the data indicating the visual media to be synthetic and / or genuine to a threshold; when the data is above the threshold then providing an output stating that the visual media was synthetically 5 generated; and when the indication is below the threshold then providing an output stating that the visual media is genuine. For example, the visual media may be considered as genuine or non-synthetic below a particular probability value threshold (providing a probability that the visual media is synthetic). 10 Therefore, a human readable output may be provided that is used to facilitate the user’s decision-making process. Optionally, the method may further comprise the step of extracting the visual media from an application. The visual media may be taken or extracted from other sources. This 15 may include screen capture techniques or image extraction processes. Optionally, the application may be an identity verification request. Other types of applications may be used. 20 Optionally, the method may further comprise the step of: when the provided data indicates the visual media is synthetic flagging the visual media as anomalous and / or flagging the application as anomalous. Other actions may be taken based on the outcome of the method. 25 Advantageously, the method is applicable to media present in applications (e.g., electronic and paper-based). This is due to the potentially high volumes of applications and the associated risks for successful dishonest or fraudulent applications. High volumes may hinder manual detection of synthetic content and incentivise the perpetration of such fraud. 30 Additionally, the flagging of an application as anomalous may facilitate the detection of fraud and help to reduce costs and improve results and reliability of an application process. False positives may also be reduced as well as reducing the incidence or suspicious applications being missed. This could be an application for a financial institution, or an application used to verify users for all kinds of services that could require 16216834.JAC.HSS P258779WO00 - 6 - identity verification (e.g. phone applications, betting, dating, taxi, rental, sport events, etc). Other applications may be reviewed using this method. Optionally, the method may further comprise the step of before being provided to 5 the trained ML model, generating embeddings for the first portion of the visual media and the second portion of the visual media. An embedding may be a lower dimensional vector into which can be translated high-dimensional vectors. Embeddings may also be described as numerical representations of real-world objects and that can capture inherent properties and relationships between real-world data. Embeddings may be vectors, for example. 10 Many kinds of transformation can be used to detect generated image artifacts. For example, a ML model may be trained to detect discrepancies between main and secondary image features identifiable in non-spatial domains or after other types of transformation. 15 Optionally, the step of generating embeddings may further comprise the step of converting the first portion of the visual media and the second portion of the visual media into a non-spatial domain. Visual Features Transformations may be used to convert the portions into the non-spatial domain. Other methods may be used. 20 Optionally, the step of generating embeddings may further comprise encoding features of the first portion of the visual media and the second portion of the visual media in a non-spatial domain. Visual Features Transformations may be used to encode the features into the non-spatial domain. Other methods may be used. 25 Optionally, the non-spatial domain may be a frequency domain. The Visual Features Transformation used may be a frequency domain transform such as Fourier transform, Hartly transform, and / or Wavelet transform. There are many algorithms for performing a transform, for example a Fourier transform may be performed with Discrete FT (DFT), Fast FT (FFT), Gabor Transform, Short-Time FT (STFT), Discrete Cosine 30 Transform (DCT), Discrete Sine Transform (DST), and Fractional FT (FrFT). Other methods may be used. This enables capturing frequency domain information. Optionally, the non-spatial domain may be a spectral domain. The Visual Features Transformation used may be a spectral transform method to analyse frequency 16216834.JAC.HSS P258779WO00 - 7 - components such as Graph Fourier Transform and / or Kernel PCA. Other methods may be used. This enables capturing frequency component information. Optionally, the non-spatial domain may be a multi-resolution domain. The Visual 5 Features Transformation used may be a multi-resolution transform such as Wavelet transform, Curvelet transform, Shearlet transform, and / or Contourlet transform. Other methods may be used. This enables capturing signals at different scales simultaneously. Optionally, the non-spatial domain may be a geometric domain. The Visual 10 Features Transformation used may be a geometric transform such as Radon transform, Hough transform, and / or Log-Polar transform. Other methods may be used. This enables capturing shape and spatial relationships. Optionally, the non-spatial domain may be a statistical and dimensionality reduction 15 domain. The Visual Features Transformation used may be a statistical and dimensionality reduction transform such as principal Component Analysis (PCA), Karhunen-Loève PCA, Kernel PCA, and Nonlinear Manifold Learning (t-SNE, UMAP). Other methods may be used. This enables capturing the largest variation in the data. 20 Optionally, the non-spatial domain may be an adaptive signal decomposition domain. The Visual Features Transformation used may be an adaptive signal decomposition technique such as Empirical Mode Decomposition (EMD) and Hilbert-Huang Transform (HHT). Other methods may be used. This enables capturing nonstationary signals. 25 Optionally, the non-spatial domain may be a fractal domain. The Visual Features Transformation used may be a fractal transform. Other methods may be used. This enables capturing fractal information. 30 Optionally, the non-spatial domain may be a nonlinear domain. The Visual Features Transformation used may be a Nonlinear Manifold Learning technique. Other methods may be used. This enables capturing non-linear information. Optionally, the non-spatial domain may be a basis domain. The Visual Features 35 Transformation used may be an Orthogonal and Basis Function Transform such as 16216834.JAC.HSS P258779WO00 - 8 - Hadamard Transform, Walsh Transform, Discrete Cosine Transform (DCT), and / or Discrete Sine Transform (DCT, DST). Other methods may be used. This enables capturing basis information. 5 Optionally, the step of converting the first portion of the visual media and the second portion of the visual media into the frequency domain may use Fast Fourier Convolution and / or the method may comprise the step of combining the embeddings before they are provided to the trained ML model. The embeddings may optionally be combined by concatenation. Other algorithms or functions may be used for this 10 conversion. Optionally, the step of encoding features of the first portion of the visual media and the second portion of the visual media in the frequency domain may use Fast Fourier Convolution and / or the method may comprise the step of combining the embeddings 15 before they are provided to the trained ML model. The encodings may optionally be combined by concatenation. Other algorithms or functions may be used for this conversion. Optionally, the step of encoding features of the first portion of the visual media and the second portion of the visual media may be done using a neural network with 3D 20 convolutions to encode volumetric data across at least three dimensions (height, width, and depth). Optionally, when the visual media is a video, a neural network with 3D convolutions may be enabled to encode spatiotemporal features of the first portion of the visual media and the second portion of the visual media to represent both visual and temporal aspects (motion, flow, frame differences). Optionally, the 3D convolution may be a 3D Fourier 25 Convolution. The generation of embeddings may allow greater accuracy. Bespoke embeddings may provide information in a format more readily exploitable by the ML model. In the case of visual media, the use of the frequency domain is of particular relevance as it can capture 30 information related to images in a way that better reflects the manner in which the media was created. The Fast Fourier Convolution may be considered a particularly effective tool for embedding visual media into the frequency domain. Combining embeddings reduces the computational load and facilitates the implementation of ML methods to reduce computing costs associated with the operation of a method to detect synthetic visual 16216834.JAC.HSS P258779WO00 - 9 - media. In particular, concatenation has a relatively low computational load with little or without information loss and as such it has been identified as a preferred tool for this task. In other embodiments, the combining of embeddings may be done via a simple 5 operation of addition of their values. Use of the sum of the embeddings even further reduces the computational load, but may result in the loss of information, which may lead to a less reliable classification. Depending on the resources and the priorities this may or may not be a viable alternative. 10 Optionally, when the visual media is a video or set of images, the first and second portion of the visual media may be different frames or images from the set of images. As AI-generated videos are typically generated frame-by-frame a disconnect between different frames may be evident when compared to genuine videos. The first and second portion of the video may be different frames of the video. The different frames may be consecutive 15 frames or sequential, but not strictly consecutive, frames belonging to the same scene. In genuine videos consecutive frames tend to have the smallest differences between them. Differences between characteristics of frames may be particularly evident when considering a non-spatial domain. 20 Optionally, it may be desirable to determine that the frames do not alter significantly or too much between scenes. The determination may be based on a determination the frames contain the same or similar objects. A determination that the frames are sufficiently similar may be based on a measure of similarity between frames (e.g., a similarity score generated using a pixel by pixel assessment or difference score) or exceeding a threshold 25 percentage difference (e.g., based on pixel intensity) between frames. This can improve the accuracy of determining whether the video is genuine or synthetic by reducing the likelihood of a genuine, albeit rapidly changing video, being determined to be synthetic. Optionally, when the visual media is a video or a set of images, the method may 30 further comprise the step of aggregating frames or images from the set of images. Optionally, the aggregation of the frames or images may be done by their averaging or blending to create a composite media. Optionally, the aggregation can be done via optical flow visualization resulting in a flow map or a heat map representing movement or transformation of objects in the media across different frames or images. Optionally, the 35 aggregation can be done via stacking the frames or images. In particular, this enables 16216834.JAC.HSS P258779WO00 - 10 - capturing information on how frames change over time or the images change across the set of images. Optionally, the ML model may be a combination of at least one or more ML models 5 comprising: at least one encoder model mapping the first and the second portions to corresponding embeddings; and a classifier model mapping the first portion embedding and the second portion embedding to an indication that the visual media is synthetic. Optionally, the method may further comprise at least two encoder models, wherein: 10 at least one encoder model may be configured to map the first portion to the first portion embedding; and at least one encoder model is configured to map the second portion to the second portion embedding. Additionally, the at least one encoder model may be a ML model configured to generate embeddings in the frequency domain and / or the at least one encoder model may be a Fast Fourier Convolution Network. 15 The division of the ML model allows bespoke ML models to be used for different purposes. This enables significantly improved accuracy as the encoder model and classifier model can have wholly different architectures more suited to their separate tasks. The use of two encoder models bespoke to the first portion embedding and second portion 20 embedding can enable greatly improved accuracy with comparable training times. Additionally, ML models configured to generate embeddings in the frequency domain and Fast Fourier Convolution Networks may be particularly suitable choices for the encoder model. 25 Optionally, the method may further comprise the step of providing one or more descriptor of the visual media to a trained ML classifier of the ML model. The descriptor may be a textural or other descriptor. 30 Optionally, the one or more descriptor may be an output of a second trained ML model. In particular, such output could be a vector embedding. Additionally, the method may or may not further comprise the step of combining the visual media descriptor with the first and the second portion embeddings. The visual media descriptor and the first and the second portion embeddings may be combined by concatenation, for example. 35 16216834.JAC.HSS P258779WO00 - 11 - Providing a descriptor of the whole image can enhance the performance of the ML model. The descriptor provides information relating to the whole of the visual media (e.g., image) at a macroscopic level. This information may be related to the composition of the visual media, overall brightness / darkness of the visual media, or even the context of the 5 visual media. This may be provided by the output of a trained ML model and / or combined with the first and second portion embeddings. The combination may simply be performed by concatenation, although the skilled person would appreciate that other methods are available. A descriptor may be long text, a vector of numbers, or any other information storage system capable of being interpreted by a trained ML classifier model. 10 Optionally, the method may further comprise the steps of: before the trained ML model provides data indicating the visual media to be synthetic and / or genuine, identifying one or more additional features in the visual media; for each of the one or more additional features: identifying a first portion of the visual media and second portion of the visual 15 media, the first portion of the visual media containing the one or more additional features and the second portion of the visual media may not include the one or more additional features; and providing the first portion and the second portion to the trained ML model. Optionally, the method may further comprise the step of selecting a subset of the 20 identified features to be separated into the first portion and the second portion. This method may be especially suited when identifying partially AI generated images where only certain features are AI generated. For example, when detecting a deep- fake, this can separate the visual media into a portion containing the face and a portion not 25 containing the face. A subset of features may be identified to further focus and improve this method. Preferably, the first portion may correspond to a foreground of the visual media and the second portion of the visual media may correspond to a background of the visual 30 media. The foreground may be a subject of an image or video (e.g., a person or face). The background may be a portion of a landscape, a room, a backdrop or other features. It will be readily apparent that this is of particular use in the detection of partially generated AI images wherein the foreground is genuine, and the background is AI 35 generated (or vice versa). 16216834.JAC.HSS P258779WO00 - 12 - Optionally, the first portion may correspond to one or more subject features of the visual media and the second portion of the visual media may correspond to the features other than the subject features of the visual media. This further improves the method for 5 deep-fake detection. We note that the term subject refers to the main focus of the visual media and not simply to the largest feature (or the largest feature may be detected). For example, visual media may contain only a face and a background. The face may be deemed the subject or main feature regardless of the exact proportion of the visual media containing the face. To provide an example of these concepts, in Joseph Mallord William 10 Turner’s “The Fighting Temeraire” the paddle-wheel steam tug may be deemed the subject or main feature, with the eponymous ship occupying a larger portion of canvas but fading into the background. In accordance with a second aspect, there is provided a method of training any 15 previous ML classification model to detect synthetic visual media, the method comprising the steps of: providing a plurality of genuine visual media; providing a plurality of synthetic visual media; for each visual media in the plurality of synthetic visual media and genuine visual media: identifying at least one feature within the visual media; identifying a first portion of the visual media and second portion of the visual media, the first portion of the 20 visual media containing the at least one feature and the second portion of the visual media not including the at least one feature; providing to an ML model the first and second portions of the visual media; obtaining data from the ML model indicating that the visual media is a synthetic and / or that the visual media is genuine; and based on the data obtained from the AI model and data stating that the visual media is synthetic or genuine, 25 updating the AI model. Training data may be provided or generated data. It will be apparent to the skilled person that the described ML classification models require bespoke training methods, which are described herein. Additionally, the embodiments may require yet further bespoke training methods which may be apparent to 30 the skilled person. Additionally, it will be apparent to the skilled person that the choice of training data will impact performance of the described ML models in different scenarios. Training data may include, but is not limited to: photographs or videos of people and / or real-world objects; scans or photographs of paintings; fully or partially AI-generated art; and AI-manipulated visual media. 35 16216834.JAC.HSS P258779WO00 - 13 - Optionally, the data obtained from the ML model indicating the visual media to be synthetic and / or genuine may comprise a probability value. Optionally, the first portion and the second portion of the visual media may be 5 transformed into the frequency domain to form embeddings. This training method enables the advantageous benefits of capturing frequency domain information. Optionally, the ML model may be a combination of at least one or more ML models comprising: at least one encoder model mapping the first and second portions to 10 corresponding embeddings; and a classifier model mapping the first portion embedding and the second portion embedding to an indication on if the visual media is synthetic. Optionally each of the one or more ML models may be updated based on the data obtained from the ML model and data stating that the visual media is synthetic or genuine. 15 This enables the training of a ML classifier comprising several different ML models. Optionally, for each visual media in the plurality of synthetic visual media and genuine visual media, one or more descriptor of the visual media from a trained ML classifier may be provided to the ML model. This enables the training of a ML classifier, 20 which includes a descriptor of the whole image to also be provided and thus achieve the advantageous benefits discussed above. Genuine images, for example with blurry (bokeh, tilt-and-shift) or monotonous (a white wall) backgrounds, may be more likely to be confused with AI-generated images 25 because of contrast in the detalization between the foreground and the background. Optionally, potentially confusing genuine images such as those with blurry or monotonous backgrounds, may be included in the plurality of genuine visual media to train the ML model to better differentiate between the true positives and false positives. This can reduce the incidents of wrong classification. 30 In accordance with a third aspect, there is provided a method for segmenting visual media, the method comprising the steps of: receiving a visual media; generating a descriptor of the image; detecting at least one feature in the visual media; and generating a first portion and second portion of the visual media, the first portion of the visual media 16216834.JAC.HSS P258779WO00 - 14 - containing the at least one feature. The use of a descriptor may allow the generation of the first and second portions to more accurately relate to a key feature in the visual media. This method may be used to aid in the identification of main and secondary objects 5 or a background for use in methods for detecting synthetic visual media. However, other applications are possible. For example, this method may be used in photography and videography to detect objects or a background within a frame. This can enable additional processing steps to enhance the visual media. Additionally, this method may be used in the editing or analysis of visual media. Enabling tools for object tracking, object removal, object 10 highlighting, optical focussing, or digital enhancement. Additionally, the background may be tracked, cropped, removed, enhanced, or painted out. Optionally, the method for segmenting visual media may further include the step of generating an embedding of the visual media in a non-spatial domain. A non-spatial 15 domain can capture AI artifacts not otherwise apparent. Any of the non-spatial domains discussed herein may be used. Other non-spatial domains may also be used. Optionally, the step of generating an embedding and generating a descriptor may be performed by the same model. This can enable the generating and embedding to be 20 more efficient. This model may be a BLIP-2 model. Optionally, the step of detecting at least one feature in the visual media may be based on the at least one descriptor. The use of a descriptor can improve the accuracy of detecting a feature. This may be performed by using DINO model. 25 Optionally, the step of detecting at least one feature in the visual media may further comprise creating at least one bounding box indicating the location of the at least one feature. Optionally, the method may further comprise a step of reducing the number of bounding boxes. The use of a bounding box can ensure features are fully captured and 30 avoid redundant features. A non-maximum suppression (NMS) technique may be used to reduce the number of bounding boxes. Optionally, the method for segmenting visual media may further include a step of generating at least one mask corresponding to the at least one feature detected in the 35 visual media and the step of generating a first and second portion of the visual media is 16216834.JAC.HSS P258779WO00 - 15 - based on the at least one mask. The use of a mask may enable the feature to be accurately captured. Optionally, the step of generating a first portion and second portion of the visual 5 media may be based on at least one mask. Optionally, the step of creating at least one mask may be based on at least one bounding box. Optionally, the at least one mask may accurately delineate the at least one feature. Segment anything model (SAM) from Meta AI may be used to create the at least one mask. 10 In accordance with a fourth aspect, there is provided a non-transitory computer- readable medium storing instructions that, when read by an apparatus, cause the apparatus to: on receipt of visual media, wherein the visual media is synthetic visual media or genuine visual media; identifying at least one feature within the visual media; identifying a first portion of the visual media and second portion of the visual media, the first portion of 15 the visual media containing the at least one feature and the second portion of the visual media not including the at least one feature; providing the first portion and the second portion to a trained machine learning (ML) model; and the trained ML model providing data indicating the visual media to be synthetic and / or genuine. 20 In accordance with a further aspect, there is provided a non-transitory computer- readable medium storing instructions that, when read by an apparatus, cause the apparatus to take the steps of any previous method. In accordance with a further aspect, there is provided apparatus comprising one or 25 more processors; and memory storing computer-executable instructions that, when executed by the processor, cause the apparatus to carry out the steps of any previous method. The methods described above may be implemented as a computer program 30 comprising program instructions to operate a computer. The computer program may be stored on a computer-readable medium, including a non-transitory computer-readable medium. The computer system may include a processor or processors (e.g., local, virtual or 35 cloud-based) such as a Central Processing Unit (CPU), and / or a single or a collection of 16216834.JAC.HSS P258779WO00 - 16 - Graphics Processing Units (GPUs). The processor may execute logic in the form of a software program. The computer system may include a memory including volatile and non- volatile storage medium. A computer-readable medium (CRM) may be included to store the logic or program instructions. For example, embodiments may include a non-transitory 5 computer-readable medium (CRM) storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform the disclosed methods. Non-transitory CRM may refer to a CRM that stores data for short periods or in the presence of power such as a memory device or Random Access Memory (RAM). For example, a non-transitory computer-readable medium may include 10 storage components, such as, a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optic disk, and / or a solid state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cartridge, and / or a magnetic tape. The different parts of the system may be connected using a network (e.g. wireless networks and wired networks). The computer system may include one or more interfaces. The computer system may contain 15 a suitable operating system such as UNIX, Windows (RTM) or Linux, for example. It should be noted that any feature described above may be used with any particular aspect or embodiment of the invention. 20 Brief Description of Figures The present invention may be put into practice in a number of ways and embodiments will now be described by way of example only and with reference to the accompanying drawings, in which: 25 FIG.1 shows a flowchart of a method for detecting synthetic images using a machine learning (ML) model; FIG.2 shows a schematic diagram of a computer system used to implement the method of Figure 1, given by way of example only; FIG.3 shows a flowchart of a method for training the ML model of the method of 30 Figure 1; FIG.4 shows a schematic diagram illustrating in more detail the method of Figure 1; and FIG.5 shows tables providing example results from the methods of Figures 1 and 3. 16216834.JAC.HSS P258779WO00 - 17 - It should be noted that the figures are illustrated for simplicity and are not necessarily drawn to scale. Like features are provided with the same reference numerals. 5 Detailed Description of Embodiments of the Invention Whilst the quality of AI-generated (synthetic images) images has increased, it has become more difficult to identify which images are genuine and which are synthetic. Studies have shown that detectors trained on images created using generative adversarial 10 network (GAN) models are not efficient at identifying images created by diffusion models. Figure 1 illustrates at a high level, a method 10 for detecting synthetic visual media when presented with visual media of unknown origin. At step 20, visual media is received. This may be a digital file or as a data stream. The visual media may be an image, a 15 plurality of images or a video (with or without audio). At step 30, at least one feature is identified within the visual media. This may be a subject of the visual media or one of a plurality of subjects. The feature may be identified using any suitable technique, such as boundary identification, contrast mapping, etc. The feature may be a most prominent feature and / or largest in the visual media, for example. However, it is not always 20 necessary to detect the largest feature as the identified feature. At step 40, a first portion of the visual media (e.g., image) is identified. This first portion contains the feature identified in step 30. At step 50, a second portion of the visual media is identified. The second portion does not contain the feature (or plurality of 25 features) identified in step 30. The first and second portions may be identified using a segmentation model or any suitable technique, for example. At step 60, the first and second portions of the visual media are provided to a machine learning (ML) model that has been previously trained. The training of the ML 30 model is described below and within Annex 1. The ML model provides an indication and / or data used to indicate that the visual media is synthetic or genuine at step 70. The data may take different forms. For example, the data may be a percentage indicating the confidence level that the visual media is 35 synthetic or genuine. A consumer of this information may compare the data to one or more 16216834.JAC.HSS P258779WO00 - 18 - thresholds to determine a binary outcome (e.g., synthetic / genuine). However, the data may be used in different ways and may trigger a further action based on its value. The method 10 may be implemented using a computer system in different forms, including a neural network, for example. 5 As shown in Figure 2, the computer system 100 includes a number of components including communication interfaces 120, system circuitry 130, input / output (I / O) circuitry 140, display circuitry and interfaces 150, and a datastore 170. The system circuitry 120 can include one or more processors or CPUs 180 and memory 190. The system circuitry 10 130 may include any combination of hardware, software, firmware, and / or other circuitry. The system circuitry 130 may be implemented, with one or more systems on a chip (SoC), application specific integrated circuits (ASIC), microprocessors, and / or analogue and digital circuits. 15 The display circuitry may provide one or more graphical user interfaces (GUIs) 160 and the I / O interface circuitry 140 may include touch sensitive or non-touch displays, sound, voice or other recognition inputs, buttons, switches, speakers, sounders, and other user interface elements. The I / O interface circuitry 140 may include microphones, cameras, headset and microphone input / output connectors, Universal Serial Bus (USB) 20 connectors, and SD or other memory card sockets. The I / O interface circuitry 140 may further include data media interfaces (e.g., a CD-ROM or DVD drive) and other bus and display interfaces. The memory 190 may include volatile (RAM) or non-volatile memory (e.g., ROM or 25 Flash memory). The memory may store the operating system 192 of the computer system 100, applications or software 194, dynamic data 196, and / or static data 198. The datastore or data source 170 may include one or more databases 172, 174 and / or a file store or file system, for example. 30 The method and system may be implemented in hardware, software, or a combination of hardware and software. The method and system may be implemented either as a server comprising a single computer system or as a distributed network of servers connected across a network. Any kind of computer system or other electronic apparatus may be adapted to carry out the described methods. 35 16216834.JAC.HSS P258779WO00 - 19 - The computer system 100 may also be used to implement the method 300 illustrated schematically in Figure 3. This is the method 300 for training the ML model described with reference to Figure 1. 5 A plurality of genuine visual media is provided to the ML model at step 310. Furthermore, a plurality of synthetic visual media is provided to the ML model at step 320. This may be achieved as a single step of combined synthetic or genuine visual media. The visual media of both types may be obtained from any suitable source, including standard training data sets. The ML model is provided with an indication as to the type of visual 10 media when they are provided. The plurality of genuine and synthetic media may be provided in any order. The method 300 loops through each visual media (e.g., image file) until they are all processed. The loop starts with step 330 and ends with step 380. At step 330, at least one 15 feature is identified within the current item of visual media. As with the method described with reference to Figure 1, this feature may be a subject of the visual media or one of a plurality of subjects. The feature may be identified using any suitable technique, such as boundary identification, contrast mapping, etc. The feature may be a most prominent feature in the visual media, for example. It is not necessary to process all visual media in 20 the training data. The loop may complete based on a suitable criterion (e.g., a particular number of files processed or when parameters of the model change below a threshold for additional items of visual media). At step 340, a first portion of the visual media (e.g., image) is identified. This first 25 portion contains the feature identified in step 330. At step 350, a second portion of the visual media is identified. The second portion does not contain the feature identified in step 30. The first and second portions may be identified using a segmentation model, for example. 30 At step 360, the first and second portions of the visual media are provided to a machine learning (ML) model. At step 370, the ML model is provided with data indicating whether or not the current visual media is synthetic or genuine. The ML model is updated at step 380, with the loop repeating until all items of visual media are processed (or a suitable number have been processed). The resultant trained ML model is used in the 35 method 10, described with reference to Figure 1. 16216834.JAC.HSS P258779WO00 - 20 - Figure 4 shows a schematic diagram providing additional details of the method. The method 400 of Figure 4 provides an example implementation, although variations may be made. The method 400 uses a combined neural network architecture, incorporating an 5 attention mechanism, spectral analysis, and multi-modal encoders as components. The multi-modal encoders extract common characteristic features from the image or other visual media. The multi-modal encoders possess an expressive ability in combining text and image modalities. 10 Spectral analysis is used with segmentation to take advantage of the concept that synthetic image generation procedures are based on the creation of both a central object or feature and a background of a scene. This leads to a frequency characteristic ratio that differs between genuine or real visual media and generated or synthetic visual media (especially images). 15 A spectral encoder may be incorporated as part of a combined architecture with central object segmentation. An attention mechanism is also present in the architecture to account for locality properties or a plurality of artifact characteristic of synthetic images. This may include accumulation of artifacts in the vicinity of contours of a central object or 20 feature (e.g., a person or face). Detecting AI-generated images of any kind is an important facility that may be used for identity verification, fraud detection, forensics, fake news, fake accounts detection, and plagiarism detection. The system and method may be used with different types of visual 25 media including still images and video. Figure 4 shows schematically a method 400 for carrying out an image analysis algorithm as well as an example neural network architecture. The high-level neural network consists of one or more lower-level ML models. The lower-level ML models may include 30 pre-trained freezed models, such as BLIP-2 and DINO, as well as bespoke or novel trained unfreezed models, such as the encoders OME and WME and the classifier (clf) that generates the clf value that is the final output that denotes the probability of the input image being AI-generated. 35 The following definitions are used in this example implementation: 16216834.JAC.HSS P258779WO00 - 21 - BLIP-2 model – Bootstrapping Language-Image Pre-training, is an AI model that can perform various multi-modal tasks like visual question answering, image-text retrieval (image-text matching) and image captioning; DINO – a grounded version of DINO, which is a self-supervised system by 5 Facebook AI that is able to learn representations from unlabelled data; NMS – non-maximum suppression technique, which is a post-processing technique used in object detection to eliminate duplicate detections and select the most relevant detected objects. This helps reduce false positives and the computational complexity of a detection algorithm; 10 Segment Anything Model (SAM) – which is a model designed to generate high- quality object masks based on various input prompts; OM is a subset of image parts belonging to the input image foreground, feature, or object (contains at least one image). OM is a first portion of the visual media containing at least one feature; 15 WM is a subset of image parts belonging to the input image background (contains at least one image). WM is a second portion of the visual media that does not contain the at least one feature; OME and WME are neural networks designed to generate embeddings for the OM and WM subset images respectively, preferably on the basis of Fast Fourier Convolution 20 (FFC) operations. The OME encoder may be trained on foreground images (or images of features or objects, such as faces) and the WME encoder may be trained on background images (e.g., images without close up objects, such as landscapes). In other example implementations, one single encoder may be used to generate embeddings for both the foreground and background images or portions of the same images. However, two 25 encoders optimised for the foreground (OM) and background (WM) portions of images operate more efficiently and effectively. As shown in Figure 4, the method 400 consists of several steps. A preparatory stage I (steps 1 to 4 described below) obtains full image embedding eblip(using the BLIP-2 30 model), which generates foreground and background elements from the image. These may be described as a first portion and a second portion of the image, respectively. As shown in Figure 4, this results in image subsets or portions OM and WM). An analysis stage II (steps 5 and 6) provides an output or data indicating a probability that the image is synthetic or AI-generated. 35 16216834.JAC.HSS P258779WO00 - 22 - The steps in stage I provide the OM and WM subsets but these can be obtained in different ways. Furthermore, different ML models may be used to generate the full input image embedding. The lower-level ML models used for stage II are preferably custom made but can be standard. 5 The following describes the steps of Figure 4 in more detail: 1. Image captions and embedding are extracted using the BLIP-2 model. The BLIP-2 model generates accurate textual descriptions of the main objects or features in an image. Other models may be used to generate a descriptor or textural 10 description of at least one feature or object in the visual media. For example, the textural description may be “a person”, “a man”, “a woman”, “a middle-aged person”, “an animal”, “a building”, etc. Furthermore, this step generates an embedding for the image (eblip). Preferably, this is a 1D vector embedding. In different example implementations other models can be used to generate the embedding. 15 2. One or more objects or features are detected in the image or other visual media, grounded by image captions or descriptors from the first stage. In this example implementation, DINO takes both the input visual media (e.g., image) and its textual description or descriptor and uses a suitable advanced deep learning 20 architecture to produce bounding boxes for the main objects or features in the image. By incorporating the textual prompt, the model effectively learns to associate visual features with semantic descriptions. To handle cases of overlapping bounding boxes and reduce the number of redundant detections, a non-maximum suppression (NMS) technique may be used. NMS 25 selects the most confident bounding boxes while suppressing overlapping ones. This process ensures that only the most relevant and accurate bounding boxes are retained to improve the object or feature detection results. As a result of this stage, a set of bounding boxes is obtained that indicate the location of one or more main objects or features in the input image. The method 400 can 30 also identify portions of the visual media that do not contain the one or more features. 3. A segmentation model uses the available or identified bounding boxes to provide regions or portions that contain the one or more object or features in the image. Segment anything model (SAM) from Meta AI may be used to process the input 35 bounding boxes and the visual media (image) to generate a set of masks. These masks 16216834.JAC.HSS P258779WO00 - 23 - represent the segmented regions corresponding to the main objects in the image. Each mask accurately delineates the shape and boundaries of a specific object or feature, enabling further analysis and classification tasks. 5 4. Regions or portions of the visual media containing the objects or features, and the background portion or region are created. This may be described as a masking stage. From the obtained masks, two subsets (or portions) of the image are formed. OM (object mask) and WM (background mask). 10 The OM subset or portion of the visual media may be obtained by element-wise multiplication between the input image and each mask. The WM subset or portion of the visual media may be obtained by subtracting the corresponding elements of the OM subset from the input or original visual media (e.g., image). These two subsets (e.g., of pixels), OM and WM, are created to facilitate separate analysis of the regions inside and outside 15 the main objects or features in the input image. Each subset includes at least one image or image portion. 5. The regions are analysed in the frequency domain. A Fourier transform (FT) is used to convert images into the frequency domain. Two encoders are used to perform 20 this operation on the two portions (OME and WME). These two encoders consist of neural networks that use the FT as their underlying operation. The encoders use fast Fourier convolution (FFC) internally. The FFC operator utilises fundamental concepts of Fourier analysis to enable efficient non-local receptive fields and multi-scale feature fusion. By utilising spectral analysis, FFC achieves global context capture and spatial information 25 fusion. Though we note that other kinds of transformation can be used to detect generated image artifacts. A ML model can be trained to detect discrepancies between main and secondary image features identifiable in non-spatial domains or after other types of transformation. 30 OME produces a set of embeddings for the current image in the OM subset. WME produces a set of embeddings for the current image in the WM subset. Both subsets of vectors are averaged to produce embeddings eOM, and eWM. The embeddings represent image features that were obtained from the frequency domain, e.g., via FFT. 16216834.JAC.HSS P258779WO00 - 24 - Preferably, both encoders have identical architecture (described below), but the OME encoder may be trained on foreground images (or portions) and the WME encoder may be trained on background images or portions. Therefore, the underlying ML models will typically comprise different weights. 5 In other example implementations one single encoder can be used to generate embeddings for both the foreground and background images. However, two dedicated encoders customised for the foreground and background images improve the method. 10 In other example implementations this stage can take different forms. For example, the method may first transform the images into a Fourier spectrum representation and then generate embeddings for the transformed images using one of more conventional convolutional neural network (CNN) models. However, test results show that use of FFC encoders yield better results. 15 6. During this stage, the embeddings from the first and fifth stages are combined and fed into a classifier. This step involves concatenating three embeddings: eblip, eOM, and eWM, to create a unified feature vector eufthat describes all the characteristics of the images. The unified feature vector eufcontains comprehensive information about the 20 image, including the textual description from the BLIP-2 model, the segmented main objects on a black background from the OME, and the image without the main object(s) or feature(s) from the WME. Once this combined feature embedding is formed, the image is classified based on this vector. This uses one linear layer as the head of the model. The classifier generates a clf value from 0 to 1 that denotes a probability that the image is AI- 25 generated or otherwise synthetic.0 means a 0% chance or confidence level and 1 means a 100% probability. An example threshold at this inference stage may be set as 0.5 but may be different, dependent on scenario and risk profile required by the decision. The following provides additional details of the Fourier transform encoders. As 30 discussed previously, the FFC encoder may be used. One of the most common models in deep learning (DL) uses ResNet blocks (see ResNet (Deep Residual Learning for Image Recognition, Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun). In the present example implementation, FFC ResNet blocks are used to create a model, which may be based on the ResNet-158 variant (there are many options with the similar architecture, but 35 having different numbers of layers). The classifier is preferably a one linear layer. 16216834.JAC.HSS P258779WO00 - 25 - Mathematically, the FFC process can be described as follows: Given an input feature map X with dimension H × W × C: 1. Split X into local (Xl) and global (Xg) parts along a channel dimension. 5 2. Update the local branch, Yl →l, through regular spatial convolution on Xl and Xg. 3. Update the global branch, Yg →g, through Fourier transforms, pointwise convolutions on the Fourier spectrum, and inverse Fourier transforms on Xg and regular convolutions on Xl. 4. Combine the local and global branches to obtain the final output Y. 10 5. Concatenate the outputs from both branches to obtain the final output Y. More detail on this mathematical process may be found in Annex 1. The model may be trained using one or more datasets. One dataset of images, 15 some of which are AI-generated may be formed by dividing or splitting into: ● 12000 labelled images as training set ● 6000 images as test set (other numbers of images in the dataset may be used). 20 A second example dataset may be a combination of real-world or genuine images from COCO2017 and AI-generated images from Stable Diffusion Wordnet Dataset (SD- WD), split into a combination of 200000 images for training and 40000 images for testing. Baselines: ResNet-50 and EfficientNetb4 in two variants, may be trained from 25 scratch and pre-trained on ImageNet. Tables 1 and 2 in Figure 5 show example results using these datasets. Table 1 shows results including a comparison between baseline results and results obtained according to an example implementation of the method and system. Table 2 shows example results of receiver operating characteristic curve – area under the curve (ROC-AUC) metrics on synthetic / genuine visual media. CLIP ResNet-50 30 was used and pre-trained on the LAION dataset. Embeddings from CLIP were also merged with embeddings from BLIP-2 and used in the classifier. Each model may be trained for approximately 30 epochs using early stopping, for example. Other training periods may be used. It has been determined that a batch size of 35 32 provided an optimum performance. To optimize training, the AdamW optimiser with a 16216834.JAC.HSS P258779WO00 - 26 - learning rate of 0.005 may be used. The CosineAnnealingLR scheduler may be used to dynamically adjust the learning rate throughout training. To increase the diversity and robustness of the models, various data augmentation 5 techniques may be applied. These include HorizontalFlip, which flips input images horizontally; RGBShift, which randomly shifts the values of the red, green, and blue channels to introduce colour variations. Other variations may be used to increase the size of training datasets. 10 Other variations may be generated using for example: A RandomBrightnessContrast process, which randomly adjust the brightness and contrast of the images; Mixup, which combines pairs of training examples by linearly interpolating their features and labels; and Cutmix, which replaces a portion of one image with a portion from another image to encourage the model to learn from different parts of the input space. 15 These augmentations aim to improve diversity and generalization. The models may use a cluster with multiple NVIDIA T416GB GPU workers, for example. Other processor configurations may be used. As used throughout, including in the claims, unless the context indicates otherwise, 20 singular forms of the terms herein are to be construed as including the plural form and vice versa. For instance, unless the context indicates otherwise, a singular reference herein including in the claims, such as "a" or "an" (such as an ion multipole device) means "one or more" (for instance, one or more ion multipole device). Throughout the description and claims of this disclosure, the words "comprise", "including", "having" and "contain" and 25 variations of the words, for example "comprising" and "comprises" or similar, mean "including but not limited to", and are not intended to (and do not) exclude other components. Also, the use of “or” is inclusive, such that the phrase “A or B” is true when “A” is true, “B is true”, or both “A” and “B” are true. 30 The use of any and all examples, or exemplary language ("for instance", "such as", "for example" and like language) provided herein, is intended merely to better illustrate the disclosure and does not indicate a limitation on the scope of the disclosure unless otherwise claimed. No language in the specification should be construed as indicating any non-claimed element as essential to the practice of the disclosure. 35 16216834.JAC.HSS P258779WO00 - 27 - The terms “first” and “second” may be reversed without changing the scope of the disclosure. That is, an element termed a “first” element may instead be termed a “second” element and an element termed a “second” element may instead be considered a “first” element. 5 Any steps described in this specification may be performed in any order or simultaneously unless stated or the context requires otherwise. Moreover, where a step is described as being performed after a step, this does not preclude intervening steps being performed. 10 It is also to be understood that, for any given component or embodiment described throughout, any of the possible candidates or alternatives listed for that component may generally be used individually or in combination with one another, unless implicitly or explicitly understood or stated otherwise. It will be understood that any list of such 15 candidates or alternatives is merely illustrative, not limiting, unless implicitly or explicitly understood or stated otherwise. Unless otherwise described, all technical and scientific terms used throughout have a meaning as is commonly understood by one of ordinary skill in the art to which the 20 various embodiments described herein belongs. As will be appreciated by the skilled person, details of the above embodiment may be varied without departing from the scope of the present invention, as defined by the appended claims. For example, different training data and ML models may be used. 25 Many combinations, modifications, or alterations to the features of the above embodiments will be readily apparent to the skilled person and are intended to form part of the invention. Any of the features described specifically relating to one embodiment or example may be used in any other embodiment by making the appropriate changes. 30 The following numbered clauses provide further example implementations: 1. A method for detecting synthetic visual media, the method comprising the steps of: receiving visual media, wherein the visual media is synthetic visual media or 35 genuine visual media; 16216834.JAC.HSS P258779WO00 - 28 - identifying at least one feature within the visual media; identifying a first portion of the visual media and second portion of the visual media, the first portion of the visual media containing the at least one feature and the second portion of the visual media not including the at least one feature; 5 providing the first portion and the second portion to a trained machine learning (ML) model; and the trained ML model providing data indicating the visual media to be synthetic and / or genuine. 10 2. The method of clause 1, wherein the visual media is an image or a video. 3. The method of clause 1 or clause 2, wherein the data indicating the visual media to be synthetic and / or genuine comprises a probability value. 15 4. The method according to any previous clause further comprising the steps of: comparing the data indicating the visual media to be synthetic and / or genuine to a threshold; when the data is above the threshold then providing an output stating that the visual media was synthetically generated; and 20 when the indication is below the threshold then providing an output stating that the visual media is genuine. 5. The method according to any previous clause further comprising the step of extracting the visual media from an application. 25 6. The method of clause 5 wherein the application is an identity verification request. 7. The method of clause 5 or clause 6 further comprising the step of when the provided data indicates the visual media is synthetic flagging the visual media as 30 anomalous and / or flagging the application as anomalous. 8. The method according to any previous clause further comprising the step of before being provided to the trained ML model generating embeddings for the first portion of the visual media and the second portion of the visual media. 35 16216834.JAC.HSS P258779WO00 - 29 - 9. The method of clause 8, wherein the step of generating embeddings further comprises the step of converting the first portion of the visual media and the second portion of the visual media into the frequency domain. 5 10. The method of clause 9, wherein the step of converting the first portion of the visual media and the second portion of the visual media into the frequency domain uses Fast Fourier Convolution. 11. The method of clause 8 or clause 9 further comprising the step of combining the 10 embeddings before they are provided to the trained ML model. 12. The method of clause 11, wherein the embeddings are combined by concatenation. 13. The method according to any previous clause, wherein the ML model is a 15 combination of at least one or more ML models comprising: at least one encoder model mapping the first and the second portions to corresponding embeddings; and a classifier model mapping the first portion embedding and the second portion embedding to an indication that the visual media is synthetic. 20 16. The method of clause 15 further comprising at least two encoder models, wherein: at least one encoder model is configured to map the first portion to the first portion embedding; and at least one encoder model is configured to map the second portion to the second 25 portion embedding. 17. The method of clause 15 or clause 16, wherein the at least one encoder model is a ML model configured to generate embeddings in the frequency domain. 30 18. The method according to any of clauses 15 to 17, wherein the at least one encoder model is a Fast Fourier Convolution Network. 19. The method according to any previous clause further comprising the step of providing one or more descriptor of the visual media to a trained ML classifier of the ML 35 model. 16216834.JAC.HSS P258779WO00 - 30 - 20. The method of clause 19, wherein the one or more descriptor is an output of a second trained ML model. 5 21. The method of clause 19 or clause 20 further comprising the step of combining the visual media descriptor with the first and the second portion embeddings. 22. The method of clause 21, wherein the visual media descriptor and the first and the second portion embeddings are combined by concatenation. 10 23. The method according to any previous clause further comprising the steps of: before the trained ML model provides data indicating the visual media to be synthetic and / or genuine identifying one or more additional features in the visual media; for each of the one or more additional features: 15 identifying a first portion of the visual media and second portion of the visual media, the first portion of the visual media containing the one or more additional features and the second portion of the visual media not including the one or more additional features; and providing the first portion and the second portion to the trained ML model. 20 24. The method of clause 23 further comprising selecting a subset of the identified features to be separated into the first portion and the second portion. 25. The method according to any previous clause, wherein the first portion is a foreground of the visual media and second portion of the visual media is a background of 25 the visual media. 26. The method according to any previous clause, wherein the first portion corresponds to one or more subject features of the visual media and second portion of the visual media corresponds to the features other than the subject features of the visual media. 30 27. A method for training a machine learning (ML) model to detect synthetic visual media, the method comprising the steps of: providing a plurality of genuine visual media; providing a plurality of synthetic visual media; 16216834.JAC.HSS P258779WO00 - 31 - for each visual media in the plurality of synthetic visual media and genuine visual media: identifying at least one feature within the visual media; identifying a first portion of the visual media and second portion of the visual 5 media, the first portion of the visual media containing the at least one feature and the second portion of the visual media not including the at least one feature ; providing to an ML model the first and second portions of the visual media; obtaining data from the ML model indicating that the visual media is a synthetic and / or that the visual media is genuine; and 10 based on the data obtained from the ML model and data stating that the visual media is synthetic or genuine, updating the ML model. 28. The method of clause 27, wherein the data obtained from the ML model indicating the visual media to be synthetic and / or genuine comprises a probability value. 15 29. The method of clause 27 or clause 28, wherein for each visual media in the plurality of synthetic visual media and genuine visual media before being provided to the ML model transforming the first portion and the second portion of the visual media into the frequency domain to form embeddings. 20 30. The method according to any of clauses 27 to 29, wherein the ML model is a combination of at least one or more ML models comprising: at least one encoder model mapping the first and second portions to corresponding embeddings; and 25 a classifier model mapping the first portion embedding and the second portion embedding to an indication on if the visual media is synthetic. 31. The method of clause 30, wherein each of the one or more ML models may be updated based on the data obtained from the ML model and data stating that the visual 30 media is synthetic or genuine. 32. The method according to any of clauses 27 to 31, wherein for each visual media in the plurality of synthetic visual media and genuine visual media, one or more descriptor of the visual media from a trained ML classifier is provided to the ML model. 35 16216834.JAC.HSS P258779WO00 - 32 - 33. A non-transitory computer-readable medium storing instructions that, when read by an apparatus, cause the apparatus to: on receipt of visual media, wherein the visual media is synthetic visual media or genuine visual media; 5 identifying at least one feature within the visual media; identifying a first portion of the visual media and second portion of the visual media, the first portion of the visual media containing the at least one feature and the second portion of the visual media not including the at least one feature; providing the first portion and the second portion to a trained machine learning (ML) 10 model; and the trained ML model providing data indicating the visual media to be synthetic and / or genuine. 16216834.JAC.HSS P258779WO00 - 33 - ANNEX 1 Merging Attention and Spectral Analysis for Synthetic Images Recognition This annex provides additional technical details of the ML model training and use, 5 including particular mathematical and computing techniques that may be used within the described system and method. 16216834.JAC.HSS 00 01 02 05603 Merging Attention and Spectral Analysis for Synthetic Images Recognition 05704 05805 05906 06007 06108 09 Paper ID **** 10 11 06512 Abstract simple idea that the objects and background of a generated 06613 image differ in the frequency domain. 06714 Recent advancements in image generation have led to a Our approach, which utilizes spectral analysis, shows 06815 wide range of applications. Unfortunately, there is also a promising results across several datasets, including the one 06916 worrying trend of malicious use of generated or synthetic used in the recent competition. 07017 images, including the creation of deep fakes. 07118 This paper proposes an approach to recognizing syn072
[0002] 2. Related Work 19 thetic images using spectral encoders and self-attention. 07320 The approach demonstrates promising results on several Progress has been made in detecting generated images, 07421 datasets. with much of the effort focused on Generative Adversarial 075*79 Networks (GANs) [27, 25, 9, 31], Recently, [4] studied the 07623 forensics traces left by diffusion models and examined how 07724 1. Introduction detectors, developed for GAN-generated images, performed 078?5 on images generated with diffusion models
[0026] , 07926 In recent years, image generation technologies have The recent work focusing specifically on text-to-image 08027 made significant progress. These mechanisms can create models such as Stable Diffusion
[0021] was DE-FAKE
[0022] , 08128 photorealistic synthetic images that are free of visible ar- which leveraged Contrastive Language-Image Pre-Training 08229 tifacts. Various methods have been developed, including (CLIP) [6] and proposed an approach for identifying syn- 08330 GANs [3, 10], Stable Diffusion
[0021] , DALL-E
[0019] , Latent thetic images generated with text-to-image models. The au- 08431 Diffusion
[0021] , and more. thors not only offer an approach to recognizing generated 08532 While image generation has many creative applications, images, but also a method for determining the generating 08633 it can also be used in harmful ways, such as creating mis- model based on the establishment of regularities in the gen- 08734 information. One recent trend is the creation of deep fakes erated artifacts.
[0012] studied the artifacts generated by dif- 08835 using fully generated face images. The most common types fusion models and GANs in the domain of deep fake detec- 08936 of electronic fraud involve falsifying photo IDs and replac- tion. 09037 ing faces during verification procedures. The trend in in- Extensive studies [5] showed that detectors trained only 09138 creasing fraud with synthetic face images has been reported on GAN images perform poorly on images generated with 09239 in the past few years1. other models, such as diffusion models. 09340 Recognition methods for artificially generated images Research has also been conducted in the area of fre- 09441 lag behind the development of generation methods. This quency domain analysis for synthetic image detection.
[0030] 09542 has resulted in a situation where the balance between gen- proposed using the frequency spectrum instead of image 09643 erative tools and artificial image recognition methods has pixels as input for classifier training. [8] demonstrated 09744 shifted in favor of the former, especially with the current that artifacts caused by upsampling operations found in all 09845 development of image generation methods using diffusion current GAN architectures can be identified using the fre- 09946 models.2quency representation to detect deep fake images. 10047 We believe that it is essential not only to develop and 10148 share detectors for artificial content but also to explore vari- 3. Proposed Approach 10249 ous ways to tackle this problem. In this paper, we explore a 10350 We propose a combined neural network architecture for 104511https: / / www.finextra.com / blogposting / 23223 / why-deepfake-fraud- recognizing synthetic i 105 losses-should-scare-financial-institutions mages. Our architecture incorporates
[0003] 2https: / / www wired com / story / deepfakes-not-very-good-nor-tools- the attention mechanism, spectral analysis, and multi-modal 10653 detect / encoders as components. 107 162
[0004] 163
[0005] 164
[0006] 165
[0007] 168
[0008] 169
[0009] 170
[0010] 171
[0011] 172
[0012] 173
[0013] 174
[0014] 175
[0015] 176 177
[0016] Figure 1. The scheme of the approach. We segment the image and then analyse its regions in frequency domain. 178
[0017] 179
[0018] 180
[0019] The multi-modal encoders are utilized to extract com- and image embedding, etup. The BLIP-2 model is designed mon characteristic features from the image. They possess to generate accurate textual descriptions of the main objects 182 great expressive ability in combining text and image modal- in an image. It produces the textual description by leverag- ities. Spectral analysis is used with segmentation to address ing its internal representation and contextual understanding 184 the idea that image generation procedures are based on the of the image content. In addition to the textual description, 185 creation of both the central object and the general back- we extract the image embedding ebupfrom the image en- 186 ground of the scene. This leads to a frequency characteristic coder of the BLIP-2 model. The image encoder embeds the 187 ratio that differs from those of actual images. image into a lower-dimensional feature representation that 188
[0020] To address this, we incorporated a spectral encoder as captures its global characteristics. This image embedding 189 part of our combined architecture in conjunction with cen- serves as a compact representation of the image and can be 190 tral object segmentation. Finally, we employed the attention used as a global feature for further analysis. 191 mechanism in our architecture to account for the locality 192 properties of many artifacts characteristic of synthetic im- captioribup, e-blip = BLIP2(I~) 193 ages, such as their accumulation in the vicinity of the con- l object. 3.2. Object Detect 194 tours of the centra ion Stage
[0021] 195
[0022] The approach consists of the following stages: In this stage, we use the Grounded DINO
[0016] model 196 to perform zero-shot object detection for all the main ob- 197
[0023] 1. Image captions and embedding are extracted. jects in an input image I, utilizing the textual description 198
[0024] 2. Object are detected in the image, grounded by image captionbHp obtained in the previous stage. The Grounded 199 captions from the first stage. DINO model is specifically designed for this task by lever- 200 aging textual descriptions as prompts. It takes both the in-
[0025] 3. A segmentation model uses available boxes to provide put image I and the textual description captionbupand uses 202 regions for objects in the image. advanced deep learning architectures to produce bounding 203 boxes for the main objects in the image. By incorporating 204
[0026] 4. Regions for objects and background are created. the textual prompt, the model effectively learns to associate 205
[0027] 5. The regions are analysed in frequency domain. visual features with semantic descriptions. 206
[0028] To handle cases of overlapping bounding boxes and re- 207
[0029] 6. Finally, the embeddings from the first and fifth stages duce the number of redundant detections, we employ the are combined and fed into a classifier. non-maximum suppression (NMS) technique. NMS selects 209 the most confident bounding boxes while suppressing over- 210
[0030] The stages are described in detail below. lapping ones. This process ensures that only the most rel- 211
[0031] 3.1. Captioning Stage evant and accurate bounding boxes are retained, thus im- proving the overall quality of the object detection results. 213
[0032] During the first stage, we feed the image I into the BLIP- As a result of this stage, we obtain a set of bounding boxes 2 model
[0014] to obtain its textual description, captioribup, boxesdino that indicate the location of the main objects in 215 the input image I. image’s height, and M represents its width. F(u, v) rep- resents the complex-valued frequency component at spatial frequency coordinates (u, v). boxesraw= NMS(DINO(I , captioribHp)) We have developed two new image encoders based on the Fourier Transform for our new solution: the OME and
[0033] 3.3. Segmentation Stage WME. These encoders consist of neural networks that use
[0034] To obtain segmented masks of the main objects, denoted the Fourier transform as their underlying operation. The as maskssam, in an image, we utilize the Segment Any- details are presented in 3.7. thing Model (SAM)
[0013] . SAM is a powerful model de- The OME takes each image from the OM subset as in- signed specifically to generate high-quality object masks put and produces a set of embeddings, denoted as e’OM, based on various input prompts. where i ranges from 1 to the size of the OM subset. To
[0035] SAM processes the input bounding boxes boxesdinoand obtain a unified embedding representing images with seg- the image I to generate a set of masks masks, ,am. These mented main objects on a black background, all eiOMem- masks represent the segmented regions corresponding to the beddings are averaged, resulting in a single vector CQM - main objects in the image. Each mask accurately delineates This process can be written according to the following for- the shape and boundaries of a specific object, enabling fur- mulas: ther analysis and classification tasks.
[0036] Thus, the first three stages aim to perform zero-shot pre- e’OM= OME(OM,}, processing for improved image classification. x |OAf| maskssam= SAM(I, boxesdino)€OM =~\1OM\1* Z=1e’°M
[0037] 3.4. Masking Stage Similarly, the WME takes each image from the WM
[0038] From the obtained masks maskssamfor the image I, subset as input. The encoder produces a set of embeddings, two subsets of images are formed: OM and WM. The denoted as e’WM, where i ranges from 1 to the size of the OM subset is obtained by element-wise multiplication be- WM subset. To obtain a unified embedding that represents tween the original image I and each mask masks’eamfrom the images with all the elements of the input image, except maskssam: for the main objects replaced with a black background, we average all elWMembeddings. This results in a single vec- tor eWM- This process can be written using the following
[0039] OMi = I * masks^arn, formulas: where OMtrepresents the / ’-th image from the OM sub- set. The WM subset is obtained by subtracting the cor- responding elements of the OM subset from the original image I :
[0040] WMi = I - OMt, where WMi represents the i-th image from the WM sub- 3.6. Final Stage set. These two subsets, OM and WM, are created to facil- The final step in our new approach involves concatenat- itate separate analysis of the regions inside and outside the ing three embeddings: Rbiip, COM, and ewM, to create a main objects in the image I. unified feature vector euf that describes all the characteris- tics of the images:
[0041] 3.5. Fourier Transform Stage
[0042] The Fourier Transform (FT) can be used to convert im-eu f = Concatenate(ei,iip, COM , RWM) ages into the frequency domain using the following for- The unified feature vector euf contains comprehensive mula: information about the image, including the textual descrip-
[0043] JV-l M-l tion from the BLIP-2 model, the segmented main objects on
[0044] F(u, v) = ^2 a black background from the OME, and the image with- a / =0 y=0 out the main objects from the WME. By combining these embeddings, we create a rich representation that captures
[0045] Here, I represents an input image, (a?, y) represents the diverse aspects of the image, which can be used for further pixel’s position on the input image, N represents the input analysis and classification.
[0046] 3 A visual representation of the approach can be seen in 5. Concatenate the outputs from both branches to obtain 378 Figure 1. the final output Y. 379
[0047] 380
[0048] 3.7. Fourier Transform Encoders In summary, FFC efficiently integrates spatial and spec- posed encoders use fast Fourier convolution in- tral information by leveragin 382
[0049] The pro g the principles of Fourier anal- ternally. The fast Fourier convolution (FFC) operator lever- ysis. By utilizing spectral decomposition, efficient Fourier convolutions, and cross-scale fusion, FFC enables non 384 ages the fundamental concepts of Fourier analysis to enable -local recepti 385 efficient non-local receptive fields and multi-scale feature ve fields and multi-scale feature fusion in an efficient and unified manner, significan 386 fusion. By utilizing spectral analysis, FFC achieves global tly benefiting various tasks in spatial information fusion through the computer vision and 387 context capture and signal processing. following key ideas:
[0050] 4. Experimen 389
[0051] Spectral Decomposition: The Fourier transform de- ts
[0052] 390 composes a signal into its frequency components, enabling 4.1. Datasets 391 operations in the Fourier domain to affect the entire original 392 signal globally. Inspired by this property, FFC uses Fourier The first dataset available for evaluating a proposed set of w u transforms to allow non-local receptive fields. techniques was provided by the organizers of the recent Al
[0053] 394
[0054] Efficient Convolution in Fourier Domain: The spectral or Not [1] competition. The full dataset consists of around
[0055] 395 convolution theorem states that convolution in the spatial 31,000 images, some of which were generated by Al. We
[0056] 396 domain equals pointwise multiplication in the Fourier do- used only the labelled training data. We split the data into a
[0057] 397 main. By leveraging this theorem, FFC can perform convo- 3 to 1 ratio, resulting in 12,000 labeled images as the train-
[0058] 398 lutions with large spatial kernels more efficiently as point- ing set and 6,000 images as the test set.
[0059] 399 wise multiplications on the Fourier spectrum. This leads to The second dataset we used is a combination of the
[0060] 400 significant computational savings. COCO2017
[0015] dataset for real-world images and the Sta-
[0061] Fourier Units (FU): FFC introduces Fourier Units that ble Diffusion Wordnet Dataset (SD-WD) [23, 2] for syn-
[0062] 402 apply FFT, modify the Fourier spectrum, and perform in- thetic images. We obtained around 200,000 samples for the verse FFT. These units enable efficient processing of large training set and about 40,000 images for the test set.
[0063] 404 receptive fields by operating on the entire image. We also combined all the datasets above and used the
[0064] 405
[0065] Multi-scale Feature Fusion: To capture multi-scale in- resulting dataset as the third separate benchmark.
[0066] 406 formation, FFC employs both local and global branches. 4.2. Baselines 407 The Local FU operates on patches to capture local features, while the Global FU uses the Fourier transform for non- As baselines, we used ResNet-50[l 1] and EfficientNet- 409 local processing, capturing global context via spectral con- b4
[0024] in two variants, trained from scratch and pretrained 410 volutions. on ImageNet [7]. 411
[0067] Cross-Scale Fusion: The local and global branches are As shown in
[0022] , Contrastive Language-Image Pre- combined to perform cross-scale fusion. This process in- Training models perform well on synthetic data detection. 413 volves separate local and global branches, and the outputs In our experiments, we used CLIP ResNet-50, pretrained are then concatenated to obtain the final result. The local on the LAION dataset [6]. Embeddings from CLIP were branch captures local features, while the global branch pro- also merged with embeddings from BLIP-2
[0014] and used 416 vides global context. in the classifer. 417
[0068] Mathematically, the FFC process can be described as fol- input feature map X of size H x W x C: 4.3. Training De 418 lows. Given an tails
[0069] 419
[0070] We trained each model for approximately 30 epochs us- 420
[0071] 1. Split X into local i'A7 ) and global Xy) parts along the ing early stopping. Through experimentation, we deter- 421 channel dimension. mined that a batch size of 32 provided the best performance. To optimize training, we employed the AdamW
[0018] opti- 423
[0072] 2. Update the local branch Yi^i through regular spatial mizer with a learning rate of 0.005. The CosineAnneal- 424 convolution on X / . ingLR
[0017] scheduler dynamically adjusted the learning rate 425
[0073] 3. Update the global branch Ya^gthrough Fourier trans- throughout training. 426 forms, pointwise convolutions on the Fourier spec- To increase the diversity and robustness of our models, 427 trum, and inverse Fourier transforms on Xg. we applied various data augmentation techniques. These 428 included HorizontalFlip, which flips input images horizon- 429
[0074] 4. Perform cross-scale fusion to combine the local and tally; RGBShift, which randomly shifts the values of the global branches. red, green, and blue channels to introduce color variations; 431 432
[0075] ‘KM
[0076] 434
[0077] 435
[0078] 436
[0079] 43'7
[0080] 438
[0081] 439
[0082] 440
[0083] 441
[0084] 442
[0085] 443
[0086] Table 1. Results for baselines and a proposed approach.
[0087] Table 2. ROC-AUC metric on Al or Not dataset.
[0088] RandomBrightnessContrast, which randomly adjusts the when FFC is used. brightness and contrast of the images; Mixup
[0029] , which We also experimented with using the Fourier transform combines pairs of training examples by linearly interpolat- as a preprocessing step
[0030] for images before they are fed ing their features and labels; and Cutmix
[0028] , which re- into OME or WME. However, these experiments did not places a portion of one image with a portion from another yield satisfactory results. We believe this means that in- image to encourage the model to learn from different parts corporating the Fourier transform within the encoder archi- of the input space. These augmentations aimed to improve tecture is more effective than using the Fourier transform diversity and generalization. as image preprocessing in combination with classical com-
[0089] We trained our models using a cluster with multiple puter vision models based on CNNs. NVIDIA T4 16GB GPU workers.
[0090] 5. Results & Conclusion
[0091] 4.4. Ablation Study
[0092] In this paper, we presented the approach for construct-
[0093] For the ablation study, we focused on experiments to find ing and training deep neural networks to recognize syn- the best combination of OME and WME features. The re- thetic images. Our proposed techniques have enabled us to sults are reported in Table 2. The concatenation of the FFC- develop a solution that outperforms counterparts based on ResNet-128 architecture for OME and WME that achieved widely-spread approaches to image recognition for fraud- the highest efficiency was used in the final version of the ulent images on several datasets. The proposed approach proposed architecture. It is worth noting that the GFNet combines common image features obtained using multi- architecture
[0020] did not achieve the same high results as modal encoders, specially designed Fourier encoders, seg-
[0025] Sheng-Yu Wang, Oliver Wang, Richard Zhang, Andrew Owens, and Alexei A Efros. Cnn-generated images are surprisingly easy to spot... for now. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pages 8695-8704, 2020.
[0094]
[0026] Ling Yang, Zhilong Zhang, Yang Song, Shenda Hong, Run- sheng Xu, Yue Zhao, Yingxia Shao, Wentao Zhang, Bin Cui, and Ming-Hsuan Yang. Diffusion models: A comprehensive survey of methods and applications, 2022.
[0095]
[0027] Ning Yu, Larry S Davis, and Mario Fritz. Attributing fake images to gans: Learning and analyzing gan fingerprints. In Proceedings of the IEEE / CVF international conference on computer vision, pages 7556-7566, 2019.
[0096]
[0028] S. Yun, D. Han, S. Chun, S. Oh, Y. Yoo, and I. Choe. Cut- mix: Regularization strategy to train strong classifiers with localizable features. In 2019 IEEE / CVF International Conference on Computer Vision (ICCV), pages 6022-6031, Los Alamitos, CA, USA, nov 2019. IEEE Computer Society.
[0097]
[0029] Hongyi Zhang, Moustapha Cisse, Yann Dauphin, and David Lopez-Paz. mixup: Beyond empirical risk minimization. 10 2017.
[0098]
[0030] Xu Zhang, Svebor Karaman, and Shih-Fu Chang. Detecting and simulating artifacts in gan fake images (extended ver- sion).
[0099]
[0031] Sm Zobaed, Md Fazle Rabby, Md Hossain, Ekram Hossain, Md Sazib Hasan, Asif Karim, and Khan Hasib. DeepFakes: Detecting Forged and Synthetic Media Content Using Machine Learning, pages 177-201. 09 2021.
[0100] 7
Claims
P258779WO00 - 41 - CLAIMS:
1. A method for detecting synthetic visual media, the method comprising the steps of: receiving visual media, wherein the visual media is synthetic visual media or genuine visual media; 5 identifying at least one feature within the visual media; identifying a first portion of the visual media and second portion of the visual media, the first portion of the visual media containing the at least one feature and the second portion of the visual media not including the at least one feature; providing the first portion and the second portion to a trained machine learning (ML) 10 model; and the trained ML model providing data indicating the visual media to be synthetic and / or genuine.
2. The method of claim 1, wherein the visual media is an image or a video. 15 3. The method of claim 1 or claim 2, wherein the data indicating the visual media to be synthetic and / or genuine comprises a probability value.
4. The method of any previous claim further comprising the steps of: 20 comparing the data indicating the visual media to be synthetic and / or genuine to a threshold; when the data is above the threshold then providing an output stating that the visual media was synthetically generated; and when the indication is below the threshold then providing an output stating that the 25 visual media is genuine.
5. The method of any previous claim further comprising the step of: extracting the visual media from an application comprising an identity verification request; 30 when the provided data indicates the visual media is synthetic, flagging the visual media as anomalous and / or flagging the application as anomalous.
6. The method of any previous claim further comprising the step of before being provided to the trained ML model generating embeddings for the first portion of the visual 35 media and the second portion of the visual media, wherein the step of generating 16216834.JAC.HSSP258779WO00 - 42 - embeddings further comprises the step of converting the first portion of the visual media and the second portion of the visual media into a non-spatial domain.
7. The method of any previous claim further comprising the step of before being 5 provided to the trained ML model generating embeddings for the first portion of the visual media and the second portion of the visual media, wherein the step of generating embeddings further comprises encoding features of the first portion of the visual media and the second portion of the visual media in a non-spatial domain. 10 8. The method of claims 6 or 7, wherein the non-spatial domain is the frequency domain and the step of converting the first portion of the visual media and the second portion of the visual media into the frequency domain uses Fast Fourier Convolution, the method further comprising the step of combining the embeddings before they are provided to the trained ML model, wherein the embeddings are combined by concatenation. 15 9. The method of any previous claim wherein the ML model is a combination of at least one or more ML models comprising: at least one encoder model mapping the first and the second portions to corresponding embeddings; and 20 a classifier model mapping the first portion embedding and the second portion embedding to an indication that the visual media is synthetic.
10. The method of claim 9 further comprising at least two encoder models, wherein: at least one encoder model is configured to map the first portion to the first portion 25 embedding; and at least one encoder model is configured to map the second portion to the second portion embedding, wherein the at least one encoder model is a ML model configured to generate embeddings in a non-spatial domain. 30 11. The method of claim 10, wherein the non-spatial domain is a frequency domain and at least one encoder model is a Fast Fourier Convolution Network.
12. The method of any previous claim further comprising the step of providing one or more descriptor of the visual media to a trained ML classifier of the ML model, wherein the 35 one or more descriptor is an output of a second trained ML model. 16216834.JAC.HSSP258779WO00 - 43 - 13. The method of claim 12 further comprising the step of combining the visual media descriptor with the first and the second portion embeddings, wherein the visual media descriptor and the first and the second portion embeddings are combined by 5 concatenation.
14. The method of any previous claim further comprising the steps of: before the trained ML model provides data indicating the visual media to be synthetic and / or genuine identifying one or more additional features in the visual media; 10 for each of the one or more additional features: identifying a first portion of the visual media and second portion of the visual media, the first portion of the visual media containing the one or more additional features and the second portion of the visual media not including the one or more additional features; and providing the first portion and the second portion to the trained ML model. 15 15. The method of claim 14 further comprising selecting a subset of the identified features to be separated into the first portion and the second portion.
16. The method of any previous claim, wherein the first portion is a foreground of the 20 visual media and second portion of the visual media is a background of the visual media, and wherein the first portion corresponds to one or more subject features of the visual media and second portion of the visual media corresponds to the features other than the subject features of the visual media. 25 17. The method of any previous claim, wherein the visual media is a video and the first and second portion of the visual media comprise different frames of the video.
18. The method of claim 17, wherein the first and second portion of the visual media comprise consecutive frames of the video. 30 19. The method of claim 18, wherein the first and second portion of the visual media comprise sequential frames of the video.
20. The method of any previous claim, wherein the visual media is a video and the 35 method further comprises the step of aggregating frames of the video. 16216834.JAC.HSSP258779WO00 - 44 - 21. A method for training a machine learning (ML) model to detect synthetic visual media, the method comprising the steps of: providing a plurality of genuine visual media; 5 providing a plurality of synthetic visual media; for each visual media in the plurality of synthetic visual media and genuine visual media: identifying at least one feature within the visual media; identifying a first portion of the visual media and second portion of the visual 10 media, the first portion of the visual media containing the at least one feature and the second portion of the visual media not including the at least one feature ; providing to an ML model the first and second portions of the visual media; obtaining data from the ML model indicating that the visual media is a synthetic and / or that the visual media is genuine; and 15 based on the data obtained from the ML model and data stating that the visual media is synthetic or genuine, updating the ML model.
22. The method of claim 21, wherein the data obtained from the ML model indicating the visual media to be synthetic and / or genuine comprises a probability value, 20 wherein for each visual media in the plurality of synthetic visual media and genuine visual media before being provided to the ML model transforming the first portion and the second portion of the visual media into a non-spatial domain to form embeddings, wherein the ML model is a combination of at least one or more ML models comprising: 25 at least one encoder model mapping the first and second portions to corresponding embeddings; and a classifier model mapping the first portion embedding and the second portion embedding to an indication on if the visual media is synthetic. 30 23. The method of claim 22, wherein each of the one or more ML models may be updated based on the data obtained from the ML model and data stating that the visual media is synthetic or genuine. 16216834.JAC.HSSP258779WO00 - 45 - 24. The method of any of claims 21 – 23, wherein for each visual media in the plurality of synthetic visual media and genuine visual media one or more descriptor of the visual media from a trained ML classifier is provided to the ML model. 5 25. A method for segmenting visual media, the method comprising the steps of: receiving a visual media; generating a descriptor of the image; detecting at least one feature in the visual media; and generating a first portion and second portion of the visual media, the first 10 portion of the visual media containing the at least one feature.
26. The method of claim 25, wherein the visual media is an image or video.
27. The method of claim 25 or claim 26, further comprising the step of generating an 15 embedding of the visual media in a non-spatial domain.
28. The method of claim 27, wherein the step of generating an embedding and generating a descriptor is performed by the same model. 20 29. The method of any of claims 25 – 28, wherein the step of detecting at least one feature in the visual media is based on the at least one descriptor.
30. The method of any of claims 25 – 29, wherein the step of detecting at least one feature in the visual media further comprises creating at least one bounding box indicating 25 the location of the at least one feature.
31. The method of claim 30, further comprising the step of reducing the number of bounding boxes. 30 32. The method of any of claims 25 – 31, further comprising a step of generating at least one mask corresponding to the at least one feature and the step of generating a first portion and second portion of the visual media is based on the at least one mask.
33. The method of claim 32, wherein the step of creating at least one mask is based on 35 the at least one bounding box. 16216834.JAC.HSSP258779WO00 - 46 - 34. The method of claim 33, wherein the at least one mask accurately delineates the at least one feature. 5 35. The method of any of claims 6 – 20, 22 – 24, or 27 – 34, wherein the non-spatial domain is a frequency domain.
36. The method of any of claims 6 – 20, 22 – 24, or 27 – 34, wherein the non-spatial domain is a spectral domain. 10 37. The method of any of claims 6 – 20, 22 – 24, or 27 – 34, wherein the non-spatial domain is a multi-resolution domain.
38. The method of any of claims 6 – 20, 22 – 24, or 27 – 34, wherein the non-spatial 15 domain is a geometric domain.
39. The method of any of claims 6 – 20, 22 – 24, or 27 – 34, wherein the non-spatial domain is a statistical and dimensionality reduction domain. 20 40. The method of any of claims 6 – 20, 22 – 24, or 27 – 34, wherein the non-spatial domain is an adaptive signal decomposition domain.
41. The method of any of claims 6 – 20, 22 – 24, or 27 – 34, wherein the non-spatial domain is a fractal domain. 25 42. The method of any of claims 6 – 20, 22 – 24, or 27 – 34, wherein the non-spatial domain is a nonlinear domain.
43. The method of any of claims 6 – 20, 22 – 24, or 27 – 34, wherein the non-spatial 30 domain is a basis domain.
44. A non-transitory computer-readable medium storing instructions that, when read by an apparatus, cause the apparatus to: on receipt of visual media, wherein the visual media is synthetic visual media or 35 genuine visual media; 16216834.JAC.HSSP258779WO00 - 47 - identifying at least one feature within the visual media; identifying a first portion of the visual media and second portion of the visual media, the first portion of the visual media containing the at least one feature and the second portion of the visual media not including the at least one feature; 5 providing the first portion and the second portion to a trained machine learning (ML) model; and the trained ML model providing data indicating the visual media to be synthetic and / or genuine. 16216834.JAC.HSS
Citation Information
Patent Citations
DeepFake detection method and device, computer equipment and storage medium
CN115100722A
Method and device for real-time deepfake detecting
KR102523372B1
Audiovisual deepfake detection
US20220121868A1
KR20230070660A