Systems and methods for preventing manipulations

By embedding adversarial examples and using enhancement techniques, the systems and methods described herein effectively prevent deepfake manipulation by challenging AI algorithms and maintaining media integrity, addressing the inadequacies of existing technologies in combating deceptive content creation.

WO2025171166A1PCT designated stage Publication Date: 2025-08-14IDENTIFAI INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/014841
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-07
Filing Date
2025-02-06
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Existing technologies are inadequate in preventing the unauthorized generation and proliferation of deepfake images and videos, as current methods such as machine learning detection models, steganographic techniques, and signature authentication fail to effectively counter the sophisticated manipulation of media by malicious actors.

Method used

Implementing adversarial examples and advanced AI techniques to embed subtle perturbations into media items, making them resistant to deepfake manipulation while maintaining visual and auditory integrity, and using enhancement models to restore quality, thus creating a robust protective layer against unauthorized use.

Benefits of technology

The systems and methods proactively prevent deepfake generation by challenging AI algorithms, maintaining media integrity and quality, and safeguarding against unauthorized use in machine learning training, thereby reducing the likelihood of deceptive content creation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025014841_14082025_PF_FP_ABST
    Figure US2025014841_14082025_PF_FP_ABST
Patent Text Reader

Abstract

A method comprising: (i) receiving an input media; (ii) analyzing the input media to determine one or more locations to add one or more perturbations; (iii) generating a first adversarial example using at least the input media and the one or more locations; (iv) generating a second adversarial example using at least the first adversarial example, wherein the second adversarial example retains one or more perturbations from the first adversarial example and the one or more perturbations are unnoticeable; and (v) maintaining one or more aspects of the input media in the second adversarial example.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR PREVENTING MANIPULATIONSCROSS REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Application No. 63 / 551,047, filed on February 7, 2024, the disclosure of which is herein incorporated by reference in its entirety.BACKGROUND

[0002] The advent of artificial intelligence (“Al”) models that have the ability to generate or alter image, video, and audio content has presented substantial challenges globally for both individuals and businesses. With the continuous advancement and scaling of generative Al (“genAI”) technologies, the capacity for malevolent entities to produce deceptive content, including “deepfake” images and videos, is not only becoming more cost-effective but also exponentially more accessible.SUMMARY

[0003] Provided herein are examples related to systems and methods for preventing manipulations. In some examples, these manipulations may include photographic and videographic manipulations.

[0004] A method comprising: (i) receiving an input media; (ii) analyzing the input media to determine one or more locations to add one or more perturbations; (iii) generating a first adversarial example using at least the input media and the one or more locations; (iv) generating a second adversarial example using at least the first adversarial example, wherein the second adversarial example retains one or more perturbations from the first adversarial example and the one or more perturbations are unnoticeable; and (v) maintaining one or more aspects of the input media in the second adversarial example.

[0005] A system comprising: (i) an adversarial example generation module to: (a) analyze an input media to determine one or more locations to add one or more perturbations and generate a first adversarial example; and (b) generate a second adversarial example using at least the first adversarial example, wherein the second adversarial example retains one or more perturbations from the first adversarial example and maintains one or more aspects of the input media, and the one or more perturbations are unnoticeable; and (ii) one or more data management servers, wherein the one or more data management servers house one or more input medias, first adversarial examples, and second adversarial examples.

[0006] A system comprising: (i) one or more computing devices, wherein the one or more computing devices: (a) send an input media to a network server for analysis to determine one or more locations to add one or more perturbations and generate an adversarial example; and (b) receive a generated adversarial example, wherein the adversarial example retains one or more perturbations and maintains one or more aspects of the input media, and the one or more perturbations are unnoticeable; and (ii) the network server.

[0007] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein may be employed to achieve the benefits described herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The accompanying drawings illustrate a number of example implementations and are a part of the specification. Together with the following description, these drawings demonstrate and explain various principles of the instant disclosure.

[0009] FIG. 1 is an illustration of an example application of an adversarial example to an image.

[0010] FIG. 2 is a flow diagram of an example process for preventing videographic manipulation of media items.

[0011] FIG. 3 is a flow diagram of an example process for preventing photographic manipulation of media items.

[0012] FIG. 4 is a diagram illustrating an example system for preventing manipulation of media items.

[0013] FIG. 5 is a flow diagram of an example process for preventing manipulation of media items.

[0014] FIG. 6 is a diagram illustrating an example process for generating deepfake media items.

[0015] FIG. 7 is a diagram illustrating an example process for image editing and / or tampering via generative Al models.

[0016] FIG. 8 shows an example flow chart of training and inference pipelines for machine learning in accordance with some implementations described herein.

[0017] FIG. 9 shows an implementation of a neural network that may be employed herein.

[0018] FIG. 10 shows an implementation of a computing device for use in various implementations that may be employed here.

[0019] Throughout the drawings, identical reference characters and descriptions indicate similar, but not necessarily identical, elements. While the implementations described herein are susceptible to various modifications and alternative forms, specific implementations have been shown by way of example in the drawings and will be described in detail herein. However, the implementations described herein are not intended to be limited to the particular forms disclosed. Rather, the instant disclosure covers all modifications, equivalents, and alternatives falling within the scope of the appended claims.

[0020] Features from any of the implementations described herein may be used in combination with one another in accordance with the general principles described herein. These and other implementations, features, and advantages will be more fully understood upon reading the following detailed description in conjunction with the accompanying drawings and claims.DETAILED DESCRIPTION

[0021] Before describing various implementations of the present disclosure in detail, it is to be understood that this disclosure is not limited to the parameters of the example systems, methods, apparatus, products, processes, and / or kits, which may, of course, vary. Thus, while certain implementations of the present disclosure will be described in detail, with reference to specific configurations, parameters, components, elements, etc., the descriptions are illustrative and are not to be construed as limiting the scope of the claimed implementations. In addition, the terminology used herein is for the purpose of describingthe implementations and is not necessarily intended to limit the scope of the claimed implementations.

[0022] A current challenge in digital security involves the potential for malicious entities to exploit publicly available images from businesses and individuals to create deceptive content (such as deepfakes). This manipulation may occur in two ways. The first approach involves malicious actors that may employ advanced generative Al models for image tampering. As illustrated in FIG. 7, a typical array of business media 702, including images 704, videos 706, and audio 708, may be altered. In some examples, generative Al models may enable these bad actors to modify this media based on either natural language descriptions of desired changes or through “inpainting” techniques 710, where specific sections of an image are targeted for alteration. FIG. 7 demonstrates how these sophisticated Al tools, when misused, may produce highly convincing and misleading content 712.

[0023] The second approach involves the development of bespoke models focused on image tampering. Referencing FIG. 6, this second approach entails incorporating specific business-related media 602 into the dataset of a Generative Adversarial Network (GAN) 604. These GANs, often described as “blackbox” models 606 due to their opaque decisionmaking processes, may then generate media that is both highly realistic and business specific 608.

[0024] A concern is the inability to trace the media that is both highly realistic and business specific 608 back to their original data sources, as shown in FIG. 6. This lack of transparency and traceability in the model’s functioning may exacerbate the risk of producing indistinguishable, fake business and / or business-adjacent media. For example, amalicious actor may use a generative adversarial network and / or other generative machine learning models to create a deepfake video of a corporate executive to give the illusion that the executive in question is making a certain business decision when in reality no such decision has been made. As an alternate example, a malicious actor may create a deepfaked image of a corporate executive in a potentially compromising situation (e.g., to create a scandal). The malicious actor may use such deepfaked video and / or images in combination with social media platforms to damage the business’s reputation, artificially manipulate stock prices, or other acts that may directly or indirectly harm the business in question. In many instances, these above-described methodologies empower malicious actors to fabricate content that may harm businesses, while current market solutions offer limited solutions to counteract these threats effectively.

[0025] Some traditional methods exist for combatting deepfakes and unwanted manipulation of media. Examples of deepfake detection approaches include machine learning based detection models, steganographic techniques for overlaying an image, and signature or authentication-based approaches that may recognize an image that has been manipulated by a machine learning model. These techniques, despite years of development and growth, have all fallen short of preventing the ever-growing proliferation and creation of deepfakes. The term “prevent” herein need not mean to eliminating entirely; rather it may refer to reduction, for example, in some examples reduction in magnitude significantly enough to be noticeable. In one example, it may refer to minimizing - e g., at least substantially eliminating, including completely eliminating. The terms “substantially” and “about” used throughout this Specification are used to describe and account for small fluctuations. For example, they may refer to less than or equal to ±5%, such as less thanor equal to ±2%, such as less than or equal to ±1%, such as less than or equal to ±0.5%, such as less than or equal to ±0.2%, such as less than or equal to ±0.1%, such as less than or equal to ±0.05%; in some instances it may include 0 fluctuation.

[0026] In some examples, machine learning based detection modes, for example, suffer from scalability concerns. For example, many detection models work by being trained on a dataset of Al-generated images, and in turn, learn to detect Al-generated images. However, as new generative models are produced, detection models in some instances become less accurate, and in turn, are rebuilt or retrained on the outputs of these new models. Similarly, steganographic techniques have been thwarted by both researchers and malicious adversaries. Signature authentication techniques, while viable in certain scenarios when authentication is needed, may lack the ability to directly prevent the creation and proliferation of these images and videos, failing to solve the largest problem. The systems and methods described in these documents, in contrast, overcome many of the shortcomings of conventional approaches by providing a proactive measure to prevent the unwanted and / or unauthorized generation of deepfake images.

[0027] Adversarial examples represent another possible attack vector on machine learning models. In some examples, when a machine learning model or neural network is trained, it accepts a set of inputs, transforms the inputs in some fashion, and looks to identify patterns in transformation that allow the model or neural network to perform its task. Malicious actors have realized that, by taking binary noise and overlaying a machine learning input, the modified input may cause the machine learning model to make mistakes. While this “bug” with machine learning models has primarily been used for negativepurposes since its discovery, adversarial examples may be used to prevent the tampering or training of an image or video by a machine learning model, as will be described in greater detail below. However, adversarial examples overlaid onto other inputs may modify the input in a way that may be detected by the human eye, rendering traditional techniques unsuitable for protecting media intended to be shared with other individuals.

[0028] The systems and methods disclosed herein may have the benefits of overcoming the aforementioned shortcomings by leveraging an approach that uses both adversarial examples and advanced artificial intelligence techniques for image and video enhancement to establish a robust protective layer for images and videos. In one implementation, the protective layer is a digital firewall. Deepfake images and videos often proliferate at an alarming rate across the Internet, with their detection becoming increasingly challenging due to the sophistication of the technology used to create them. The methods and systems provided herein have the benefits of proactively elevating the difficulty of generating AI- manipulated media. In some implementations, the systems and methods described herein may achieve this by, for example, embedding an adversarial example into every image and video uploaded by a user.

[0029] As described above, adversarial examples may subtly alter digital media that effectively “trick” Al algorithms into recognizing an input as something it is not, making it more challenging for these algorithms to generate convincing deepfakes. However, the introduction of these adversarial examples in some instances may degrade the quality of the original media, causing user frustration. The systems and methods described herein in some implementations overcome this challenge by using Al-driven and / or other enhancement techniques to restore and, in some implementations even improve, the quality of the media,thereby maintaining the integrity and clarity of the original images and videos. In some implementations, these enhancement techniques may help avoid identification of protection techniques applied to the image, reducing the likelihood that a malicious actor will be able to successfully use the media to generate a deepfake or similar image / video. In some implementations, these techniques may help prevent using media content for training purposes without consent and improve Al training data security. The systems and methods described herein in some implementations may also cause machine learning training systems to misidentify objects, places, people, or other features of an image or other piece of media, thus rendering protected media items unsuitable for training machine learning algorithms. In some implementations, these techniques may help prevent the production of fake business and / or business-adjacent media. For example, the systems and methods disclosed herein may prevent a malicious actor from being able to successfully train a generative machine learning model such that the malicious actor would be rendered incapable of generating a deepfake video or image of a corporate executive.

[0030] By implementing this dual strategy of protection and enhancement, the methods and systems provided herein may provide a benefit by safeguarding media against unauthorized use in Al-driven technologies. Not only may this protection act as a deterrent against the unauthorized use of their images and videos for deepfake creation, but it may also maintain the quality of the media. In essence, the concepts presented herein represent a forward-thinking approach to digital media protection, offering a robust defense against Al-generated media manipulation. The principles described herein may also help prevent or deter copyright infringement through image theft by preventing safeguarded images from being used to generate derivative media using generative machine learning models.

[0031] As described above, the systems and methods described herein adopt a proactive strategy to preventing deepfakes. While the majority of existing technologies in this field concentrate on a reactive methodology (e.g., identifying deepfakes after their creation), the systems and methods described herein may preemptively disrupt the generation of deepfakes at the source. Furthermore, in some implementations, the techniques employed by the systems and methods described herein diverge significantly from traditional steganographic methods. In one implementation, unlike the embedding of a “digital watermark” (as used in steganographic approaches), the use of adversarial examples by the disclosed systems and methods present a more complex challenge for potential manipulators. These adversarial examples are inherently more resilient than mere watermarks; they are not only harder to detect but also significantly more challenging to eliminate without destroying the original image. This fundamental difference in approach and technology underscores one of the unique approaches used by the disclosed systems and methods.

[0032] The term “deepfake,” as used herein, may generally refer to an image or video generated or tampered with by an artificial intelligence model, machine learning model, neural network, or other digital generative process. Deepfakes may be used to create images involving individuals to create the illusion that the affected individuals, for example, endorse a product that they do not, or are involved in a situation that they are not actually present for. Some deepfakes refer to manipulations of images; for example, replacing the face of an individual in a photo or video with another individual’s face. Other deepfakes may be wholly generated images. In some implementations, a deepfake may be a digital image that purportsto have been created by a particular individual despite being created by a generative machine learning model or similar.

[0033] The term “adversarial example,” as used herein, may generally refer to systematically engineered inputs designed to induce error in the operational efficacy of a neural network or machine learning model when used as an input, potentially resulting in erroneous classification of the adversarial example input. These inputs may be indistinguishable from ordinary inputs to the human eye yet compromise a neural network or machine learning model’s capacity to discern and categorize elements within the adversarial example. Machine learning models that are inadvertently trained on adversarial examples may be unable to properly identify features of other, non-adversarial inputs.

[0034] The term “media” as used herein generally refers to any medium for transferring or sharing information electronically. In some examples, a piece of media may be a still image such as a two-dimensional photograph or digital painting. Other examples of media include video, which may be with or without audio, standalone audio, three-dimensional models, three-dimensional images, animations, and the like. Media items may be represented in a variety of file formats, such as JPEG, MP3, MP4, MPEG, PDF, AVI, FLV, MOV, OGG, WAV, PNG, FLAC, SVG, AAC, BMP, GIF, TIFF, or any other suitable file format for representing a media item.

[0035] The systems and methods described herein may use adversarial examples and advanced machine learning techniques for image and video enhancement to target deepfakes. In some implementations, the systems and methods described herein may embed subtly altered digital media, such as adversarial examples as described above, into images and videos. In some implementations, these embedded adversarial examples maybe tailored to deceive machine learning algorithms, particularly those used in generating deepfakes, making it challenging for these algorithms to manipulate the media.

[0036] In some implementations, the systems and methods described herein may be configured to transform standard images or videos into fortified adversarial examples resistant to deepfake manipulation. This enhancement then serves to address the aforementioned security concerns. The procedure may vary between images and videos, but is founded on the same core principles.Example Process for Preventing Photographic Manipulation of Media Items

[0037] As shown in the example photographic aspect shown in FIG. 3, an example process 300 initiates with the upload of an image 302. This may be done through an Application Programming Interface (“API”) or a User Interface (“UI”) specifically designed for this purpose. In some cases, entities such as large corporate enterprises may also leverage a Software Development Kit (“SDK”) to implement the systems and methods presented herein on their own local networks and / or computing systems rather than transferring data to a third-party entity, thereby allowing the systems and methods described herein to protect the images without compromising privacy. Once uploaded, the image may undergo initial processing on a server 304. This stage may include resizing the image for compatibility and converting it into a format that is accessible for machine learning algorithms 306. In some examples, the resized image 308 may be converted into a tensor to prepare it for adversarial example embedding 310. After these preparatory processes, the image may be fed into an advanced adversarial example generation algorithm to embed an adversarial example into the resized and converted image 312. This adversarial example generation algorithm may introduce calculated perturbations to the image, effectively altering it in a way that, while itmay be imperceptible to human eyes, it may be confusing for machine learning models. Following this, the image 314 may undergo a secondary phase of processing using an AI- driven enhancement model. In some examples, this enhancement model (which is described in greater detail below) may be engineered to restore the image’s quality to its original state, thereby maintaining the visual integrity while maintaining the underlying perturbations and / or alterations 316. The final process in this example may involve resizing the image back to its original dimensions 318, while preserving the perturbations that render the image resistant to use in the production of deepfakes and other manipulated imagery. In examples where the original image is converted into a tensor, this final process may also include converting the image data from a tensor back into a coherent image that is visually meaningful to humans. In one implementation, the end product is an image that may not only maintain its visual fidelity but also may be inherently resistant to unauthorized tampering and misuse in machine learning model training. This process effectively shields the image from being exploited for generating misleading or deceptive content. In some implementations, the enhancement model may be completed as part of the adversarial example embedding 310, 312.Working Example

[0038] One objective of the systems and methods described herein is to exploit a model’s inherent weaknesses, causing it to misinterpret the data. For instance, as demonstrated in FIG. 1 , an image of a cat 100 may be recognized by a machine learning model as a lynx with a confidence interval of 0.3181, or 31.81% certainty, that the model has correctly identified the image. However, in this working example when this input image underwent subtle yet precise alterations, such as introducing a pattern of noise 102 that wascalculated (i.e., the middle image), the model’s interpretation shifted dramatically. In this working example, the model then identified the image as a ping pong ball 104 with a confidence interval of 0.0188, corresponding to 1.88% certainty, that the model correctly identified the image (the rightmost image in FIG. 1). In this working example, the identification of the image as being of a ping pong ball 104 with very low certainty occurred because the model was unable to identify any other classifications for the image with a higher confidence. This working example illustrates the effectiveness of adversarial examples in manipulating machine learning models by exploiting their algorithmic tendencies, leading to unexpected and often incorrect classifications, often with very low confidence intervals. In the systems and methods described herein, a standard adversarial example model may be utilized to perform these perturbations on given images and videos that users wish to protect against deepfake manipulation.Example Process for Preventing Videographic Manipulation of Media Items

[0039] With regards to videographic content, as illustrated in FIG. 2, the disclosed method 200 aligns closely with that of the method for images 300, albeit with some additional processes to accommodate the complexities of video. The initial phase of uploading a video 202 (or collection of videos) may mirror that of an image, utilizing either an API or a specialized User Interface designed for this purpose. Once the video is uploaded, it may undergo preliminary processing on a server 204. In some examples, the disclosed system may deconstruct the video into a plurality of video frames 206, effectively creating a series of images that may be stored in a directory of video frames 208. The plurality of video frames in the directory may then be converted into a tensor to prepare them for adversarial example embedding 210. This frame-by-frame breakdown allows for the application of the sameadversarial example enhancement process as detailed previously for photographs. Each frame of the plurality of video frames may be individually processed, applying a perturbation algorithm to impart the necessary adversarial characteristics 212 to create a directory of adversarial frames 214. The perturbation algorithm may include determining one or more locations for perturbations to be added to a video frame and / or adding perturbations to the video frame. This treatment may, for example, facilitate that every frame is subtly altered to enhance its security against misuse 216. The adversarial frames may also be enhanced in the same manner as the individual images as described above to minimize apparent visual differences in the final protected video. After each frame has been processed and enhanced, the video may be reconstructed by reassembling these frames and reintegrating the original audio track with the newly composed video 218. This reassembly may help facilitate retaining the original narrative and auditory elements of the enhanced video 220 while being safeguarded against tampering and exploitation in machine learning models. The result is a video that may not only maintain its visual and auditory integrity but also stand resilient against unauthorized manipulation and training misuse. In some implementations, the enhancement model may be completed as part of the adversarial example embedding 210, 212.

[0040] In some implementations, a method for protecting video against unwanted manipulation may be applied directly to the video file as a whole rather than frame by frame as described above. In such an example, the initial input phase would be much the same as described above, with a client or user submitting a video for protection via an API, user interface, or similar. Once the video is submitted for protection, the systems and methods described herein may perform preliminary processing on the video to render it suitable foruse as an input to the appropriate machine learning models without breaking the video file down into a plurality of video frames. The perturbation algorithm may analyze the video as a whole and, in some implementations, create a perturbed video file. This perturbed video file may then be used to generate a difference video that represents the differences between the original video and the perturbed video, and the difference video may be used to impart the necessary adversarial characteristics into the original video file without unduly impacting the final quality of the protected video as described above. The final process in this example may involve the reintegration of the original audio track with the newly composed video. This reassembly helps facilitate that the enhanced video retains its original narrative and auditory elements while being safeguarded against tampering and exploitation in machine learning models. The result is a video that not only maintains its visual and auditory integrity but also stands resilient against unauthorized manipulation and training misuse.

[0041] In some implementations, the audio track may be protected separately using a similar method to those described above, albeit directed towards audio files rather than video or still-image files. A method for protecting an audio file may include receiving an audio file; processing the audio file for analysis by an adversarial example module and / or perturbation algorithm; generating one or more adversarial examples based on the audio file; and maintaining one or more aspects of the audio file. In some implementations, maintaining one or more aspects of the audio file may occur as part of the generation of the one or more adversarial examples. In some implementations, maintaining one or more aspects of the audio file may occur after the generation of the one or more adversarial examples. For audio files, perturbation may include frequency shifts; time stretching;amplitude variations; noise addition, or other modifications such that the modifications are unnoticeable to human ears.

[0042] One specific example system for protecting media against unwanted manipulation includes an adversarial generation module and an enhancement module. The term “module” as referred to herein may generally refer to a computer-implemented algorithm, a local server, a cloud server, or another distinct functional unit. The adversarial generation module may use an iterative fast gradient sign method to create images that are misclassified by a pre-trained classification model. In this example, the enhancement model blends the adversarial example with an input media item to maintain visual quality of the input media item while simultaneously preserving the adversarial characteristics of the images generated by the adversarial generation module.

[0043] Adversarial image generation 412, 512 may include several processes. In one example, an input media item is processed to be compatible with the pre-trained classification model, which may then attempt to predict a category of the original input media item. The adversarial generation module may then apply an iterative fast gradient sign method to the processed input media item to iteratively perturb the image with specified parameters to maximize the classification errors when the perturbed media item is used as an input to the pre-trained classification model. During each iteration, the perturbed media item (which is an example of an adversarial example) may be saved in association with the classification model’s attempt to predict the category of the perturbed image. In this example, the iteration continues for a predetermined number of processes or until the classification model’s confidence or accuracy on predicting the category of theperturbed image falls below a certain threshold, which may be for example less than or equal to about 20% - e.g., less than or equal to about 15%, about 10%, or less.

[0044] Image enhancement and / or quality preservation 416, 516 may likewise include several processes. In some implementations, the enhancement module identifies structural similarities between the two images and computes a structural similarity index (“SSIM”). Next, the enhancement module creates a difference image that highlights areas of significant change between the original input media item and the perturbed or modified image. The enhancement module then determines a threshold using at least the top percentage (e.g., the top 10%, 20%, 30% or other suitable threshold percentage) of differences between the original and perturbed image and creates an image mask to isolate those differences rather than indiscriminately blending the entirety of the perturbed image into the final image and unnecessarily degrading image quality. In other words, the systems and methods described herein might use the top 20% of pixels that differ the most between the perturbed image and the original image for generating the final adversarial image, using an image mask to blend only those pixels in the image blending process. Finally, the enhancement module blends the original image and the perturbed image using the image mask, weighting the blending such that areas with larger differences favor the original media item and areas with lower differences favor the perturbed image. In some implementations, weighting the blending includes blending the original image and perturbed image such that a weighted amount of the original image remains after blending, that a weighted of the perturbed image remains after blending, or other ratios. As mentioned above, the systems and methods described herein might only use the top 20% of differences to generate the blended and enhanced final image. By only using the perturbations that are most likely to change how a machine learning model,classifier, or similar interprets the final image, the above-described process may maintain that adversarial properties of the final blended image while likewise maintaining that the final image’s visual quality is similar to that of the original image.

[0045] In one example, the end result of the above process is a visually enhanced image that retains adversarial characteristics of the perturbed image necessary to deceive classifier models and / or other artificial intelligence / machine learning (“AI / ML”) models. Thus, the systems and methods described herein may provide a balance between maintaining effectiveness of the adversarial properties of the final image and maintaining that the final image is visually coherent for practical applications such as publication and display on digital services. The above-described method may also be used to generate a dataset of adversarial examples to safeguard against the unauthorized use of images in creating deepfakes. The systems and methods described herein may also be extended or modified to accommodate different classifier models, different adversarial generation modules, multiple classifiers and / or adversarial generation modules, blending techniques, etc.System for Preventing Manipulation of Media Items

[0046] FIG. 4 illustrates one implementation of an example system 400 for preventing manipulation of media items and implementing the methods described herein. Such systems are referred to as manipulation prevention systems in the present disclosure. In some implementations, a manipulation prevention system may include one or more processors and one or more storage servers. The one or more processors may run one or more AI / ML models to generate one or more adversarial examples, enhance image and / or video quality of one or more adversarial examples, or perform other functions. The one or more processors may include an adversarial example generation module, an enhancementmodule, cloud server, local server, and / or other system. The one or more storage servers may house one or more store input medias, one or more store adversarial examples comprising, one or more maps of perturbations, and / or other data. The one or more storage servers may also run predetermined data management protocols or otherwise manage data. The one or more store input medias may include one or more input medias previously uploaded by a user. The one or more store adversarial examples may include one or more first adversarial examples previously generated and one or more second adversarial examples previously generated. The stored one or more first adversarial examples previously generated may be referred to as one or more store first adversarial examples. The stored one or more second adversarial examples previously generated may be referred to as one or more store second adversarial examples. The one or more storage servers may include one or more data management servers, one or more data storage servers, cloud servers, local servers, and / or other storage servers. An adversarial example generation server 402 may house one or more AI / ML models that may be used to generate one or more adversarial examples. AI / ML models housed in the adversarial example generation server 402 may include a Stable Diffusion adversarial example model; a facial classification adversarial example model; other relevant AI / ML models, or a combination thereof. A Stable Diffusion adversarial example model may include a system that leverages a Stable Diffusion generative Al model to create adversarial examples. A facial classification adversarial example model may include a model designed to generate slightly altered facial images that, while appearing identical, or substantially identical, to the original to human eyes, may trick a facial recognition system into misclassifying a person in an image and / or video, essentially creating a deceptive or adversarial version of the face. This may beachieved by adding subtle, calculated perturbations to the image and / or video that exploit vulnerabilities in one or more facial recognition algorithms. Adversarial generation server 402 may send and receive communication, images, and / or videos via network 410. An enhancement model may be housed on the adversarial example generation server 402, on a local server, or on a cloud server. The enhancement model may include a convolutional neural network, generative adversarial network, and / or other AI / ML models. The enhancement model may enhance image quality while minimizing the reduction in security that comes with enhancing image quality of an image with perturbations. One or more data management servers 406 may house a data management system, AI / ML functionalities, or other components for purposes of communicating with other components of FIG. 4. The one or more data management servers 406 may store input media, varying adversarial examples, user information, organization information, and / or other relevant data. An AI / ML engine may be stored or operated at various computing devices within manipulation prevention system 400. Any or all of servers 402 and 406 may comprise an AI / ML engine(s). AI / ML engines may comprise or perform AI / ML functionality as described herein.

[0047] In some implementations, a user may use one or more computing devices, which may be any devices comprising a processor (e.g., computers, tablets, mobile devices, etc.), to send and receive communication, images, and / or videos using network 410 (e.g., Internet, cellular, Bluetooth™, Wi-Fi, satellite, enterprise, private network, similar networks, or combinations of the foregoing).Process for Preventing Manipulation of Media Items

[0048] FIG. 5 illustrates one implementation of an example process 500 for preventing manipulation of media items. A user may first upload an input media 502 such that network 410 may send the input media to the adversarial example generation server 402. The adversarial example generation server 402 may analyze the input media to determine the one or more locations on the input media to add perturbations 504. The one or more locations to add perturbations may be the locations on the input media where the perturbations are least likely to be noticeable to the naked eye; most likely to cause errors in deepfake generation models, classification models, and / or other models; a weighted preference of both; and / or another goal. Then the adversarial example generation server 402 may generate a first adversarial example 506. In some implementations, generating a first adversarial example may include feeding the input media as an input to a facial classification adversarial example model and generating the first adversarial example based on the input media and the determined perturbation one or more locations. The facial classification adversarial example model may run iteratively until the result of the loss function meets a predetermined threshold that may be set by a user or adjusted automatically based on user input. Then the adversarial example generation server 402 may generate a second adversarial example 508 based on the first adversarial example. To the extent applicable, the terms “first,” “second,” “third,” etc. in this disclosure are merely employed to show the respective objects described by these terms as separate entities and are not meant to connote a sense of chronological order, unless stated explicitly otherwise herein.

[0049] The second adversarial example may retain one or more perturbations from the first adversarial example. The one or more perturbations may be unnoticeable. In someimplementations herein, when something is unnoticeable, it is referred to being not observable to the naked human eye. In some implementations, generating the second adversarial example may include feeding the first adversarial example as an input to a Stable Diffusion adversarial example model and generating the second adversarial example based on the first adversarial example and the determined perturbation one or more locations. The Stable Diffusion adversarial example model may run iteratively for a preset number of rounds. Next, the enhancement model may maintain key aspects of the input media in the second adversarial example 510. Key aspects of the input media may be preset by a user, may be determined by the enhancement model, or otherwise determined. In some implementations, this process may be completed during the generation of the first adversarial example and / or second adversarial example. In some implementations, this process may be completed as part of an iterative process to improve output media quality. The iterative process may be the generation of the first adversarial example and / or second adversarial example. When generating the first adversarial example and / or second adversarial example, the adversarial example generation server 402 may incorporate maintaining key aspects of the input media as part of the process. In some implementations, this method 500 may be used to secure a specific element of an input media rather than the entirety of the input media itself.

[0050] In some implementations, the second adversarial example may be automatically uploaded to a third-party storage location, social media app / website, or other third-party application / website. This may be done by the manipulation prevention system 400 using the third-party’s application programming interface (“API”). In some implementations,this may allow a media to be protected through a method described herein as the media is uploaded to a third-party application / website.

[0051] In some implementations, a security preference and a quality preference may be presented to a user such that a user may configure the security to quality ratio for the adversarial example generated by the manipulation prevention system 400. This configuration may include manipulating varying parameters which may include number of iterations, an alpha value, a target label, an epsilon value which define perturbation strength and quality levels for an adversarial example, and / or other parameters. There may be more optionality in the future, but for now those are the primary options.AI / ML Implementations

[0052] Various implementations under the present disclosure may incorporate AI / ML functionality. For example, for purposes of the present disclosure, adversarial example generation server 402 of FIG. 4 may be said to comprise an AI / ML model, either separately or together. Each component may comprise a separate instance of an identical AI / ML model. Or a “central” AI / ML model may be running at any location, such as adversarial example generation server 402, and others of the foregoing devices may function like an output / input interface to the central AI / ML model, allowing user input, data collection, user interface for a user, etc. As described above, adversarial example generation server 402 may collect data from network 410, or other devices, receive input media, generate adversarial examples, communicate with an enhancement model, or otherwise receive or use a variety of other data. This data may be used to determine the one or more locations to add perturbations to an input media, generate one or more adversarial examples, enhanceoutput media quality, share output media via one or more third-party channels, or other actions. This data may also be used to train an AI / ML model or may be analyzed by a previously trained AI / ML model.

[0053] The term artificial intelligence commonly is used to refer to an entire system that achieves intelligence-like outcomes while using multiple sub-systems, such as multiple machine learning algorithms. But both ML and Al have been used to identify a variety of functionalities or types of systems that utilize various combinations of specific ML algorithms. As used herein, AI / ML model is intended to denote a variety of AI / ML functionalities that fall under the category of Al or ML algorithms and systems that utilize such functionalities. Examples of AI / ML models may comprise of supervised learning, reinforcement learning, natural language processing such as LLMs, neural networks, computer vision, facial recognition, chatbots, virtual assistants, unsupervised learning, generative Al, other Al or ML models, or combinations of any of the foregoing.

[0054] In manipulation prevention system 400 of FIG. 4, multiple AI / ML models may be used. For example, one AI / ML engine may comprise an adversarial example generation model, an enhancement model, and / or another AI / ML model. It may be that multiple different cognitive neural networks are used as part of the adversarial example generation server 402. A different AI / ML model may be stored or implemented at various of the components shown in FIG. 4. Alternatively, there may be a smaller number of AI / ML models, and various of the components of FIG. 4 may function as user interfaces for a remote AI / ML model stored at e.g., adversarial example generation server 402. Data used to train, retrain, or implement any of AI / ML models may be stored at any one or more components shown in FIG. 4. Any other suitable variations may be employed.

[0055] The architecture of an AI / ML model (e.g., structure, number of layers, nodes per layer, activation function etc.) may be customized for each particular use case. For example, properties that vary may include e.g.: input media, quality preferences, security preferences, perturbation suggestions, and a variety of other factors. These may all be considered when designing an AI / ML model architecture.

[0056] Building an AI / ML model may include several development processes where the actual training of a ML model or algorithm is just one process in a training pipeline. An important part in AI / ML development is AI / ML model lifecycle management. One implementation of a model lifecycle management procedure 2700 is illustrated in FIG. 8. The model lifecycle management may in some implementations comprise two pipelines: a training pipeline 2705 and an inference pipeline 2750.

[0057] At 2710 in the training pipeline 2705, data ingestion 2710 occurs, which includes gathering raw (training) data from a data storage. For example, in some implementations, an adversarial generation module may gather input images and adversarial examples corresponding to the input images from one or more data managements servers. In some implementations, an enhancement module may gather input images and output images corresponding to the input images from one or more data management servers. After data ingestion 2710, there may also be a process that controls the validity of the gathered data. At 2715 data pre-processing occurs, which may include feature engineering applied to the gathered data. This may involve, e.g., data normalization or data formatting or transformation required for the input data to the AI / ML model. In some of the previously discussed implementations, this may include resizing images and / or a plurality of image frames or other processing of images and / or videos for AI / ML analysis. After the MLmodel’s architecture is fixed, it should be trained on one or more datasets. At 2720 model training is performed in which the AI / ML model is trained with the raw training data. To achieve good performance during live operation in a system (the so-called inference phase), the training datasets should be representative of actual data the ML model will encounter during live operation. For example, an adversarial generation module may be trained using a dataset comprising input images and corresponding adversarial examples, and an enhancement module may be trained using a dataset comprising input images and corresponding output images. The training process often involves numerically tuning the ML model’s trainable parameters (e.g., the weights and biases of the underlying neural network (NN)) to minimize a loss function on the training datasets.

[0058] The loss function may be, for example, based on minimizing the quality loss of an input media; maximizing security; or other metrics. The purpose of the loss function is to meaningfully quantify the reconstruction error for the particular use case at hand. At 2725 model evaluation may be performed where the performance is benchmarked to some baseline. Model training 2720 and evaluation 2725 may be iterated until an acceptable level of performance is achieved. At 2730 model registration occurs, in which the AI / ML model is registered with any corresponding data on how the AI / ML model was developed, and e.g., AI / ML model evaluation data. At 2735 model deployment occurs, wherein the trained / re-trained AI / ML model is implemented in the inference pipeline 2750.

[0059] Data ingestion 2755 in the inference pipeline 2750 refers to gathering raw(inference) data from a data source. Data pre-processing 2760 may be essentially identical, or similar, to the data pre-processing 2715 of the training pipeline 2705. At 2765, the operational model received from the training pipeline 2705 is used to process new datareceived during operation of e.g., manipulation prevention system 400 of FIG. 4 or components thereof. At 2770 data and model monitoring is performed. Here the inference data is analyzed to determine whether the inference data are from a distribution that aligns with the training data, as well as monitoring model outputs for detecting any performance, or operational, variance or drifts. The variance or drift is used at 2745 (drift detection) to update the AI / ML model registration.

[0060] The training process may be based on some variant of a gradient descent algorithm, which, at its core, typically comprises three components: a feedforward component, a back propagation component, and a parameter optimization component. These components may be described using a dense ML model (i.e., a dense NN with a bottleneck layer) as an example.

[0061] Feedforward: A batch of training data, such as a mini-batch, (e.g., several downlink-channel estimates) is pushed through the ML model, from the input to the output. The loss function is used to compute the reconstruction loss for all training samples in the batch. The reconstruction loss may be an average reconstruction loss for all training samples in the batch.

[0062] Back propagation (BP): The gradients (partial derivatives of the loss function, L, with respect to each trainable parameter in the ML model) are computed. The back propagation algorithm sequentially works backwards from the ML model output, layer-by- layer, back through the ML model to the input. The back propagation algorithm is built around the chain rule for differentiation: When computing the gradients for layer n in the ML model, it uses the gradients for layer n + 1.

[0063] Parameter optimization: The gradients computed in the back propagation process are used to update the ML model’s trainable parameters. An approach is to use the gradient descent method with a learning rate hyperparameter (a) that scales the gradients of the weights and biases. It is preferred to make small adjustments to each parameter with the aim of reducing the average loss over the (mini) batch. It is common to use special optimizers to update the ML model’s trainable parameters using gradient information. The following optimizers are widely used to reduce training time and improving overall performance: adaptive sub-gradient methods (AdaGrad), RMSProp, and adaptive moment estimation (ADAM).

[0064] The above process (feedforward, back propagation, parameter optimization) may be repeated many times until an acceptable level of performance is achieved on the training dataset. An acceptable level of performance may refer to the ML model achieving a predefined average reconstruction error over the training dataset (e.g., normalized MSE of the reconstruction error over the training dataset is less than, say, 0.1). Alternatively, it may refer to the ML model achieving a pre-defined value chosen by a user.

[0065] In some implementations, a function F(-) may be generated by a ML process, such as, for example, supervised learning, reinforcement learning, and / or unsupervised learning. It should further be understood that supervised learning may be done in various ways, such as, for example, using random forests, support vector machines, neural networks, and the like. By way of non-limiting example, any of the following types of neural networks that may be utilized, including, deep neural networks (“DNNs”), convolutional neural networks (“CNNs”), and recurrent neural networks (“RNNs”), or any other known or future neural network that satisfies the needs of the system. In an implementation using supervisedlearning the neural networks may be easily integrated into the hardware described in manipulation prevention system 400 of FIG. 4 (e.g., in the form of simple vector-matrix multiplications).

[0066] Referring now to FIG. 9, an example NN 2900 (e.g., DNN) is shown. In some implementations, and as shown, the neural network 2900 may include two hidden layers represented by dashed boxes 2901 and 2902. In one implementation, the inputs 2903 may be fed into the NN 2900. Next, the inputs 2403 may go through a set of hidden layers (e.g., 2901 and / or 2902). Once the inputs 2903 pass though the hidden layers 2901 and / or 2902, they may be output (e.g., as an output layer) as outputs 2904, 2905. Outputs 2904, 2905 may be, e.g., an adversarial example, graph of suggested locations for perturbations, or another output valuable. Possible inputs may include e.g. : an input media, a first adversarial example, a graph of suggested locations for perturbations, or other variables.

[0067] In order for the NN 2900 to output a proper analysis, it often involves proper training (e.g., with a collection of samples) to accurately extract the likelihood values. If not trained properly, overfitting (e.g., when the NN memorizes the structure of the preambles but is unable to generalize to unseen preamble characteristics) or underfitting (e.g., when the NN is unable to learn a proper function even on the data that it was trained on) may happen. Thus, implementations may exist that prevent overfitting or underfitting, involving a set of well-engineered features that must be extracted from the preamble characteristics.Additional Implementations

[0068] FIG. 10 illustrates an implementation of various computing devices within manipulation prevention system 400 of FIG. 4, or components thereof e.g., adversarialgeneration server 402, which may comprise e.g., computers, tablets, servers, databases, mobile devices, or other computing or smart devices described herein. FIG. 10 shows a schematic block diagram of a computing device 3500 (or components thereof) according to certain implementations of the present disclosure. System 3500 may be used to analyze and / or optimize: the functionalities described with respect to manipulation prevention 400 of FIG. 4 and its components, or to perform other methods, such as Al or ML-related tasks and analyses as described herein.

[0069] Computing device 3500 includes processor 3501 that is operatively coupled via a bus 3502 to an input / output interface 3505, a power source 3513, a memory 3515, a RF interface 3509, network communication interface 3511, and / or any other component, or any combination thereof. The level of integration between the components may vary from one implementation to another. Further, certain computing devices 3500 (or components thereof) may contain multiple instances of a component, such as multiple processors, memories, transceivers, transmitters, receivers, etc.

[0070] The processor 3501 is configured to process instructions and data and may be configured to implement any sequential state machine operative to execute instructions stored as machine-readable computer programs in memory 3515. Processor 3501 may be implemented as one or more hardware-implemented state machines (e.g., in discrete logic, field-programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), etc ); programmable logic together with appropriate firmware; one or more stored computer programs, general-purpose processors, such as a microprocessor or digital signal processor (DSP), together with appropriate software; or any combination of the above. For example, the processor 3501 may include multiple central processing units (CPUs).

[0071] In the example, input / output interface 3505 may be configured to provide an interface or interfaces to an input / output device(s) 3506, such as a screen, keyboard, indicator light, keypad, touchscreen, or other input or output device. Other examples of an output device include a speaker, a sound card, a video card, a display, a monitor, a printer, an actuator, an emitter, a smartcard, another output device, or any combination thereof. An input device may allow a user to capture information into system 3500. Other examples of an input device include a touch-sensitive or presence-sensitive display, a camera (e.g., a digital camera, a digital video camera, a web camera, etc.), a microphone, a sensor, a mouse, a trackball, a directional pad, a trackpad, a scroll wheel, a smartcard, and the like. The presence-sensitive display may include a capacitive or resistive touch sensor to sense input from a user. A sensor may be, for instance, an accelerometer, a gyroscope, a tilt sensor, a force sensor, a magnetometer, an optical sensor, a proximity sensor, a biometric sensor, etc., or any combination thereof. An output device may use the same type of interface port as an input device. For example, a Universal Serial Bus (USB) port may be used to provide an input device and an output device.

[0072] In some implementations, the power source 3513 is structured as a battery or battery pack. Other types of power sources, such as an external power source (e.g., an electricity outlet), photovoltaic device, or power cell, may be used. The power source 3513 may further include power circuitry for delivering power from the power source 3513 itself, and / or an external power source, to the various parts of computing device 3500 via input circuitry or an interface such as an electrical power cable.

[0073] Memory 3515 may be configured to include memory such as random access memory (RAM) 3517, read-only memory (ROM) 3519, programmable read-only memory(PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic disks, optical disks, hard disks, removable cartridges, flash drives, other storage medium 3521, and so forth. In one example, the memory 3515 includes one or more application programs 3525, an operating system 3523, web browser application, a widget, gadget engine, or other application, and corresponding data 3527. Memory 3515 may store, for use by the computing device 3500, any of a variety of various operating systems or combinations of operating systems. An article of manufacture, such as one including a simulation system or communication system may be tangibly embodied as or in memory 2515, which may be or comprise a device- readable storage medium.

[0074] Processor 3501 may be configured to communicate with an access network or other network using the RF interface 3509 or network connection interface 3511. The RF interface 3509 or network connection interface 3511 may comprise one or more communication subsystems and may include or be communicatively coupled to an antenna. In the illustrated implementation, communication functions of the RF interface 3509 or network connection interface 3511 may include cellular communication, Wi-Fi communication, LPWAN communication, data communication, voice communication, multimedia communication, short-range communications such as Bluetooth, near-field communication, location-based communication such as the use of the global positioning system (GPS) to determine a location, another like communication function, or any combination thereof.

[0075] The manipulation prevention system 400 of FIG. 4, or computing devices 3500 as described above or in regard to FIG. 4, may perform a variety of method implementationsunder the present disclosure. Several example method implementations are given above but these examples are non-limiting and are only meant to illustrate certain implementations.

[0076] Although the computing devices described herein (e.g., servers, computing devices, etc. of manipulation prevention system 400 of FIG. 4) may include the illustrated combination of hardware components, other implementations may comprise computing devices with different combinations of components. It is to be understood that these computing devices may comprise any suitable combination of hardware and / or software needed to perform the tasks, features, functions and methods disclosed herein. Determining, calculating, obtaining or similar operations described herein may be performed by processing circuitry, which may process information by, for example, converting the obtained information into other information, comparing the obtained information or converted information to information stored in the network node, and / or performing one or more operations based on the obtained information or converted information, and as a result of said processing making a determination. Moreover, while components are depicted as single boxes located within a larger box, or nested within multiple boxes, in practice, computing devices may comprise multiple different physical components that make up a single illustrated component, and functionality may be partitioned between separate components. For example, a communication interface may be configured to include any of the components described herein, and / or the functionality of the components may be partitioned between the processing circuitry and the communication interface. In another example, non-computationally intensive functions ofany of such components may be implemented in software or firmware and computationally intensive functions may be implemented in hardware.

[0077] In certain implementations, some or all of the functionality described herein may be provided by processing circuitry executing instructions stored in memory, which in certain implementations may be a computer program product in the form of a non- transitory computer-readable storage medium. In alternative implementations, some or all of the functionality may be provided by the processing circuitry without executing instructions stored on a separate or discrete device-readable storage medium, such as in a hard-wired manner. In any of those particular implementations, whether executing instructions stored on a non-transitory computer-readable storage medium or not, the processing circuitry may be configured to perform the described functionality. The benefits provided by such functionality are not limited to the processing circuitry alone or to other components of the computing device, but are enjoyed by the computing device as a whole, and / or by end users and a wireless network generally.

[0078] It will be appreciated that computer systems are increasingly taking a wide variety of forms. In this description and in the claims, the terms “controller,” “computer system,” or “computing system” are defined broadly as including any device or system — or combination thereof — that includes at least one physical and tangible processor and a physical and tangible memory capable of having thereon computer-executable instructions that may be executed by a processor. By way of example, not limitation, the term “computer system” or “computing system,” as used herein is intended to include personal computers, desktop computers, laptop computers, tablets, hand-held devices (e.g., mobile telephones, PDAs, pagers), microprocessor-based or programmable consumer electronics,minicomputers, mainframe computers, multi-processor systems, network PCs, distributed computing systems, datacenters, message processors, routers, switches, and even devices that conventionally have not been considered a computing system, such as wearables (e.g., glasses).

[0079] The computing system also has thereon multiple structures often referred to as an “executable component.” For instance, the memory of a computing system may include an executable component. The term “executable component” is the name for a structure that is well understood to one of ordinary skill in the art in the field of computing as being a structure that may be software, hardware, or a combination thereof. For instance, when implemented in software, one of ordinary skill in the art would understand that the structure of an executable component may include software objects, routines, methods, and so forth, that may be executed by one or more processors on the computing system, whether such an executable component exists in the heap of a computing system, or whether the executable component exists on computer-readable storage media. The structure of the executable component exists on a computer-readable medium in such a form that it is operable, when executed by one or more processors of the computing system, to cause the computing system to perform one or more functions, such as the functions and methods described herein. Such a structure may be computer-readable directly by a processor — as is the case if the executable component were binary. Alternatively, the structure may be structured to be interpretable and / or compiled — whether in a single stage or in multiple stages — so as to generate such binary that is directly interpretable by a processor.

[0080] The terms “component,” “service,” “engine,” “module,” “control,” “generator,” or the like may also be used in this description. As used in this description and in this case,these terms — whether expressed with or without a modifying clause — are also intended to be synonymous with the term “executable component” and thus also have a structure that is well understood by those of ordinary skill in the art of computing.

[0081] In terms of computer implementation, a computer is generally understood to comprise one or more processors or one or more controllers, and the terms computer, processor, and controller may be employed interchangeably. When provided by a computer, processor, or controller, the functions may be provided by a single dedicated computer or processor or controller, by a single shared computer or processor or controller, or by a plurality of individual computers or processors or controllers, some of which may be shared or distributed. Moreover, the term “processor” or “controller” also refers to other hardware capable of performing such functions and / or executing software, such as the example hardware recited above.

[0082] In general, the various example implementations may be implemented in hardware or special purpose chips, circuits, software, logic, or any combination thereof. For example, some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, microprocessor, or other computing device, although the disclosure is not limited thereto. While various aspects of the example implementations of this disclosure may be illustrated and described as block diagrams, flow charts, or using some other pictorial representation, it is well understood that these blocks, apparatus, systems, techniques, or methods described herein may be implemented in, as non-limiting examples, hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.

[0083] While not all computing systems require a user interface, in some implementations a computing system includes a user interface for use in communicating information from / to a user. The user interface may include output mechanisms as well as input mechanisms. The principles described herein are not limited to the precise output mechanisms or input mechanisms as such will depend on the nature of the device. However, output mechanisms might include, for instance, speakers, displays, tactile output, projections, holograms, and so forth. Examples of input mechanisms might include, for instance, microphones, touchscreens, projections, holograms, cameras, keyboards, stylus, mouse, or other pointer input, sensors of any type, and so forth.

[0084] The principles described herein may be implemented in the form of an API or software development kit (SDK) infrastructure depending on whether a client individual or entity wishes to execute the systems and methods described herein via cloud infrastructure or on their own computing system(s), as well as a more Ul-centric platform infrastructure. The disclosed systems and methods may be used by both individuals and businesses to proactively protect their media before releasing the same to websites, news agencies, social media platforms, or other modes of digital sharing. By providing both an APVSDK infrastructure that may be integrated into an existing infrastructure as well as a platform that is usable for non-technical users, the disclosed systems and methods may serve a broad user base while also enabling clients to perform image processing on-premise (e.g., on computing systems entirely under their own control and / or physically on-site) should they wish to preserve their privacy by not transmitting images to a third-party or cloud-based processing system.

[0085] The disclosed systems and methods may also be applied to a variety of different forms of content. While one objective of the disclosed system is to safeguardphotographs and videos, this system may also be used to safeguard audio samples and other file formats. In the current digital environment, photographs and videos are frequently used for disseminating disinformation. However, it is increasingly likely that audio files and various document formats, such as PDFs and Word Documents, will also become prevalent mediums for the spread of misinformation. Fortunately, the systems and methods described herein may also be used to protect these alternative forms of content. The systems and methods disclosed herein thus represent versatile tools in the fight against the evolving landscape of digital misinformation. By encompassing a broader range of media formats, the disclosed technology may offer comprehensive protection, remaining effective and relevant as the methods of spreading misinformation continue to diversify and evolve.

[0086] The preceding description has been provided to enable others skilled in the art to best utilize various aspects of the example implementations disclosed herein. This example description is not intended to be exhaustive or to be limited to any precise form disclosed. Many modifications and variations are possible without departing from the spirit and scope of the present disclosure. The implementations disclosed herein should be considered in all respects illustrative and not restrictive. Reference should be made to the appended claims and their equivalents in determining the scope of the present disclosure. Logistics and receiving personnel may later view these logs to verify that the cap lock system is being used correctly.

[0087] Unless otherwise noted, the terms “connected to” and “coupled to” (and their derivatives), as used in the specification and claims, are to be construed as permitting both direct and indirect (i.e., via other elements or components) connection. In addition, the terms “a” or “an,” as used in the specification and claims, are to be construed asmeaning “at least one of.” Finally, for ease of use, the terms “including” and “having” (and their derivatives), as used in the specification and claims, are interchangeable with and have the same meaning as the word “comprising.”Additional Examples

[0088] Example 1

[0089] A method for protecting media against unwanted manipulation, the method comprising: formatting an input media item for processing via a machine learning algorithm, the input media item being received from a user device; providing the input media item as an input to an adversarial example generation algorithm to add perturbations to the input media item and produce an adversarial media item; providing the adversarial media item as an input to an image enhancement machine learning model to produce an enhanced adversarial media item that retains the perturbations applied to the input media item while matching a visual fidelity of the enhanced adversarial media item to the input media item; and processing the enhanced adversarial media item to match one or more aspects of the adversarial media item to one or more original aspects of the input media item.

[0090] Example 2

[0091] The method of Example 1, wherein the input media item comprises a static image.

[0092] Example 3

[0093] The method of any of Example 1, wherein: the input media item comprises video; formatting the input media item comprises processing the video into a plurality of video frames; and providing the input media item as the input to the adversarial examplegeneration algorithm comprises providing the plurality of video frames to the adversarial example generation algorithm.

[0094] Example 4

[0095] The method of Example 1, wherein the adversarial example generation algorithm adds the perturbations to the input media item by overlaying a noise pattern over the input media item.

[0096] Example 5

[0097] The method of Example 1, wherein the input media item comprises at least one of: a static image; a segment of video; an audio file; or a multimedia document.

[0098] Example 6

[0099] The method of Example 1, wherein the one or more original aspects comprise one or more of: an image size; an image resolution; or a color profile.

[0100] Example 7

[0101] The method of Example 1, wherein the enhanced adversarial media item, when used as an input into a machine learning training process, introduces errors into the machine learning training process.

[0102] Example 8

[0103] A method comprising: receiving an input media; analyzing the input media to determine one or more locations to add one or more perturbations; generating a first adversarial example using at least the input media and the one or more locations; generating a second adversarial example using at least the first adversarial example, wherein the second adversarial example retains one or more perturbations from the first adversarialexample and the one or more perturbations are unnoticeable; and maintaining one or more aspects of the input media in the second adversarial example.

[0104] Example 9

[0105] The method of Example 8, wherein the input media is a static image.

[0106] Example 10

[0107] The method of Example 8, wherein the input media is a video.

[0108] Example 11

[0109] The method of Example 8, wherein: the input media comprises a video and the method further comprises processing the video into a plurality of video frames; analyzing the input media further comprises analyzing the plurality of video frames; generating a first adversarial example further comprises generating a first adversarial example using at least the plurality of video frames and the one or more locations; and maintaining one or more aspects of the input media in the second adversarial example further comprises recombining the plurality of video frames from the second adversarial example into a second video with original audio from the video.

[0110] Example 12[OHl] The method of any of Examples 8-11, wherein generating the first adversarial example further comprises adding the one or more perturbations to the input media by overlaying a noise pattern over the input media.

[0112] Example 13

[0113] The method of any of Examples 8-12, wherein the one or more aspects of the input media comprise an image size, an image resolution, a color profile, or a combination thereof.

[0114] Example 14

[0115] The method of any of Examples 8-13, wherein the input media is selected from a group consisting of: a static image, a video, an audio file, and a multimedia document.

[0116] Example 15

[0117] The method of any of Examples 8-14, wherein generating a first adversarial example further comprises generating a facial recognition adversarial example by a facial classification adversarial example model.

[0118] Example 16

[0119] The method of any of Examples 8-15, wherein generating a second adversarial example further comprises generating an enhanced adversarial example by a stable diffusion model.

[0120] Example 17

[0121] The method of any of Examples 8-16, further comprising: presenting an option for a user to configure one or more quality preferences and one or more security preferences; and adjusting the first adversarial example to comply with the user preferences.

[0122] Example 18

[0123] A system comprising: an adversarial example generation module to: (i) analyze an input media to determine one or more locations to add one or more perturbations and generate a first adversarial example; and (ii) generate a second adversarial example using at least the first adversarial example, wherein the second adversarial example retains one or more perturbations from the first adversarial example and maintains one or more aspects of the input media, and the one or more perturbations are unnoticeable; and one or more data management servers, wherein the one or more data management servers house one or more input medias, first adversarial examples, and second adversarial examples.

[0124] Example 19

[0125] The system of Example 18, wherein the input media is a static image.

[0126] Example 20

[0127] The system of Example 18, wherein the adversarial example generation module further to process a video into a plurality of video frames and the input media is the plurality of video frames.

[0128] Example 21

[0129] The system of any of Examples 18-20, wherein the adversarial example generation module is further to generate the first adversarial example by adding the one or more perturbations to the input media by overlaying a noise pattern over the input media.

[0130] Example 22

[0131] The system of any of Examples 18-21, wherein the one or more aspects of the input media comprise an image size, an image resolution, a color profile, or a combination thereof.

[0132] Example 23

[0133] The system of Example 18, wherein the input media is selected from a group consisting of: a static image, a video, an audio file, and a multimedia document.

[0134] Example 24

[0135] The system of any of Examples 18-23, wherein the adversarial example generation module implements a facial classification adversarial example model to generate a facial recognition adversarial example from the first adversarial example.

[0136] Example 25

[0137] The system of any of Examples 18-24, wherein the adversarial example generation model implements an enhancement module to generate an enhanced adversarial exampleand wherein the second adversarial example maintains one or more aspects of the input media.

[0138] Example 26

[0139] The system of any of Examples 18-25, wherein the one or more data management servers are further to automatically export the second adversarial example to a third-party.

[0140] Example 27

[0141] The system of any of Examples 18-26, further comprising one or more user computing devices, wherein the one or more user computing devices are to receive and display the second adversarial example.

[0142] Example 28

[0143] A system comprising: one or more computing devices, wherein the one or more computing devices: (i) send an input media to a network server for analysis to determine one or more locations to add one or more perturbations and generate an adversarial example; and (ii) receive a generated adversarial example, wherein the adversarial example retains one or more perturbations and maintains one or more aspects of the input media, and the one or more perturbations are unnoticeable; and the network server.

[0144] Example 29

[0145] The system of Example 28, wherein the one or more user computing devices are further to: present an option to configure one or more quality and security preferences; and communicate the preferences.

Claims

WHAT IS CLAIMED:

1. A method comprising: receiving an input media; analyzing the input media to determine one or more locations to add one or more perturbations; generating a first adversarial example using at least the input media and the one or more locations; generating a second adversarial example using at least the first adversarial example, wherein the second adversarial example retains one or more perturbations from the first adversarial example and the one or more perturbations are unnoticeable; and maintaining one or more aspects of the input media in the second adversarial example.

2. The method of claim 1, wherein the input media is a static image.

3. The method of claim 1, wherein the input media is a video.

4. The method of claim 1, wherein: the input media comprises a video and the method further comprises processing the video into a plurality of video frames; analyzing the input media further comprises analyzing the plurality of video frames; generating a first adversarial example further comprises generating a first adversarial example using at least the plurality of video frames and the one or more locations; and maintaining one or more aspects of the input media in the second adversarial example further comprises recombining the plurality of video frames from the second adversarial example into a second video with original audio from the video.

5. The method of any of claims 1 -4, wherein generating the first adversarial example further comprises adding the one or more perturbations to the input media by overlaying a noise pattern over the input media.

6. The method of any of claims 1-5, wherein the one or more aspects of the input media comprise an image size, an image resolution, a color profile, or a combination thereof.

7. The method of any of claims 1-6, wherein the input media is selected from a group consisting of: a static image, a video, an audio file, and a multimedia document.

8. The method of any of claims 1-7, wherein generating a first adversarial example further comprises generating a facial recognition adversarial example by a facial classification adversarial example model.

9. The method of any of claims 1-8, wherein generating a second adversarial example further comprises generating an enhanced adversarial example by a stable diffusion model.

10. The method of any of claims 1-9, further comprising: presenting an option for a user to configure one or more quality preferences and one or more security preferences; and adjusting the first adversarial example to comply with the user preferences.

11. A system comprising: an adversarial example generation module to: analyze an input media to determine one or more locations to add one or more perturbations and generate a first adversarial example; and generate a second adversarial example using at least the first adversarial example, wherein the second adversarial example retains one or more perturbationsfrom the first adversarial example and maintains one or more aspects of the input media, and the one or more perturbations are unnoticeable; and one or more data management servers, wherein the one or more data management servers house one or more store input medias comprising the input media, one or more store first adversarial examples comprising the first adversarial example, and one or more store second adversarial examples comprising the second adversarial examples.

12. The system of claim 11, wherein the input media is a static image.

13. The system of claim 11, wherein the adversarial example generation module further to process a video into a plurality of video frames, and the input media to be analyzed is the plurality of video frames.

14. The system of any of claims 11-13, wherein the adversarial example generation module is further to generate the first adversarial example by adding the one or more perturbations to the input media by overlaying a noise pattern over the input media.

15. The system of any of claims 11-14, wherein the one or more aspects of the input media comprise an image size, an image resolution, a color profile, or a combination thereof.

16. The system of claim 11, wherein the input media is selected from a group consisting of: a static image, a video, an audio file, and a multimedia document.

17. The system of any of claims 11-16, wherein the adversarial example generation module implements a facial classification adversarial example model to generate a facial recognition adversarial example from the first adversarial example.

18. The system of any of claims 11-17, wherein the adversarial example generation model implements an enhancement module to generate an enhanced adversarial example and wherein the second adversarial example maintains one or more aspects of the input media.

19. The system of any of claims 11-18, wherein the one or more data management servers are further to automatically export the second adversarial example to a third-party.

20. The system of any of claims 11-19, further comprising one or more user computing devices, wherein the one or more user computing devices are to receive and display the second adversarial example.

21. A system comprising: one or more computing devices, wherein the one or more computing devices: send an input media to a network server for analysis to determine one or more locations to add one or more perturbations and generate an adversarial example; and receive a generated adversarial example, wherein the adversarial example retains one or more perturbations and maintains one or more aspects of the input media, and the one or more perturbations are unnoticeable; and the network server.

22. The system of claim 21, wherein the one or more user computing devices are further to: present an option to configure one or more quality and security preferences; and communicate the preferences.

Citation Information

Patent Citations

  • Utilizing multiple stacked machine learning models to detect deepfake content

    US20210406568A1