Fake video detection
By combining a face detection module and neural network with discrete Fourier transform analysis of image texture irregularities, and integrating blockchain technology and digital fingerprint verification, the problem of identifying forged videos is solved, ensuring the authenticity of video content and preventing tampering.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-13
- Publication Date
- 2026-04-03
AI Technical Summary
Existing technologies are insufficient to effectively identify and detect fake videos, especially those generated using deep learning algorithms, which pose a risk of defamation or information tampering.
The system employs a face detection module and neural network combined with Discrete Fourier Transform (DFT) to analyze texture irregularities in images. It determines whether an image has been altered by detecting brightness regions and facial lighting irregularities in the spectrum. It also combines blockchain technology and digital fingerprints to verify the authenticity of the video.
It enables efficient identification and verification of fake videos, ensuring the authenticity of video content and preventing the spread and tampering of fake videos.
Smart Images

Figure CN114600174B_ABST
Abstract
Description
Technical Field
[0001] This application generally involves unconventional solutions that are technologically innovative, and these solutions must originate from computer technology and produce specific technological improvements. Background Technology
[0002] As understood in this paper, the combination of modern digital image processing and deep learning algorithms presents the interesting and enjoyable, yet potentially insidious, ability to alter a person's video image into that of another, or to modify a person's video by having that person's voice say things that the person never actually said. While this ability can be used in easily accessible ways, it can also be used to defame an individual by making it appear as if the person has said defamatory things. Therefore, this paper presents techniques for determining whether a video is genuine or a forgery generated by machine learning. Summary of the Invention
[0003] Therefore, a system includes at least one face detection module for receiving an image and determining whether at least one texture irregularity exists on a face in the image or at least one texture irregularity exists between a face and a background in the image, or both. The system also includes at least one first neural network for receiving the image; at least one discrete Fourier transform (DFT) for receiving the image and outputting a spectrum to at least one second neural network; and at least one detection module for accessing features output by the face detection module, the first neural network, and the second neural network to determine whether the image has been altered from the original image and providing an output representing it.
[0004] Texture irregularities can include checkerboard patterns.
[0005] The detection module can determine that an image has been altered from the original image, at least in part, by detecting at least one irregularity in the spectrum.
[0006] Irregularities in the spectrum can include at least one brightness region that is brighter than the corresponding region in the original image. Brightness regions can be located along the periphery of the image in the frequency domain. In fact, irregularities in the spectrum can include multiple brightness regions located along the periphery of the image in the frequency domain.
[0007] The face detection module can be configured to output feature vectors indicating lighting irregularities on faces in an image, which indicate that the image has been altered from the original image.
[0008] In another approach, one method involves processing an image via a face detection module to output a feature vector indicating at least one illumination irregularity on a face in the image, or at least one texture irregularity in the image, or both. The method further includes processing the image via at least one discrete Fourier transform (DFT) and at least one neural network to output a feature vector indicating at least one irregularity in the image in the frequency domain, and returning an indication, at least in part, that the image has been altered from the original image based on the feature vector.
[0009] In another aspect, an apparatus includes at least one computer storage medium having instructions executable by at least one processor to process an image by an image detection module to determine whether irregularities exist in the image in the spatial domain. The instructions are executable to convert the image to the frequency domain and process the image in the frequency domain to determine whether irregularities exist in the frequency domain. The instructions are executable to output an indication of a digitally altered image from the original image, based at least in part on the determination that irregularities exist in the image.
[0010] An indication that the image has been digitally altered from the original image can be output in response to determining either irregularity in the frequency domain or irregularity in the spatial domain. Alternatively, an indication that the image has been digitally altered from the original image can be output in response to determining only both irregularity in the frequency domain and irregularity in the spatial domain.
[0011] The details of this application regarding both its structure and operation can be best understood with reference to the accompanying drawings, in which the same reference numerals refer to the same parts, and in the drawings: Attached Figure Description
[0012] Figure 1 This is a block diagram of an example system that includes examples based on the principles of the present invention;
[0013] Figure 2 This is a diagram illustrating a real video and a fake video derived from the real video;
[0014] Figure 3 This is a flowchart of example logic for detecting fake videos using both image processing and frequency domain analysis;
[0015] Figure 4 It is used for training Figure 3 A flowchart of example logic for the neural network used;
[0016] Figure 5 The illustration shows a real video frame and its corresponding forged video frame, and shows the artifacts in the forged frame;
[0017] Figure 6 It is used for execution Figure 3 A block diagram of an example neural network architecture;
[0018] Figure 7 This is a flowchart of example logic for detecting fake videos using video sequence analysis;
[0019] Figure 8 It is used for execution Figure 7 A block diagram of an exemplary neural network architecture for the logic;
[0020] Figure 9 This is a flowchart illustrating example logic for using blockchain technology to process the generation of fake videos;
[0021] Figure 10 This is a screenshot of an example user interface (UI) used to report fake videos to Internet Service Providers (ISPs) or resellers so that the ISP / reseller can remove the video from public view;
[0022] Figure 11 This is a flowchart of example logic used for recording, uploading, or downloading videos, as well as for verifying hashes embedded in the videos;
[0023] Figure 12 It is used for playback Figure 11 A flowchart of example logic for recording or accessing videos, where hashing is used to verify authenticity;
[0024] Figure 13 This is a flowchart of an example logic using hybrid logic based on the previously mentioned principles;
[0025] Figure 14 Two sets of real images and example lighting artifacts in altered images are shown;
[0026] Figure 15 Examples of Generative Adversarial Network (GAN) artifacts or irregularities are shown in the images; and
[0027] Figure 16 Using real and altered images, another artifact or irregularity related to GANs is shown. Detailed Implementation
[0028] This disclosure generally relates to the computer ecosystem, which includes various aspects of consumer electronics (CE) device networks, such as, but not limited to, computer simulation networks, such as computer gaming networks, and standalone computer simulation systems. Systems described herein may include server and client components connected via a network, enabling data exchange between client and server components. Client components may include one or more computing devices, including, for example, Sony... Game consoles, such as those made by Microsoft, Nintendo, or other manufacturers; virtual reality (VR) headsets; augmented reality (AR) headsets; portable televisions (e.g., smart TVs, internet-enabled TVs); portable computers (e.g., laptops and tablets); and other mobile devices (including smartphones and additional examples discussed below). These client devices can operate in a variety of operating environments. For example, some client computers may run on operating systems such as Linux, Microsoft's operating system, or Unix, or an operating system made by Apple or Google. These operating environments can be used to execute one or more browsing programs, such as browsers made by Microsoft, Google, or Mosla, or other browser programs that can access websites hosted by the internet servers discussed below. Furthermore, the operating environment according to the principles of the invention can be used to execute one or more computer game programs.
[0029] The server and / or gateway may include one or more processors that execute instructions to configure the server to receive and transmit data over a network, such as the Internet. Alternatively, the client and server may connect via a local intranet or virtual private network. The server or controller may be manufactured by, for example, Sony. Instantiate game consoles such as personal computers.
[0030] Information can be exchanged between clients and servers over a network. For this purpose and for security, servers and / or clients may include firewalls, load balancers, temporary storage devices, and proxies, as well as other network infrastructure for reliability and security. One or more servers may form a device that implements a method of providing network members with a secure community, such as an online social networking site.
[0031] As used herein, an instruction refers to a computer-implemented step for processing information in a system. Instructions can be implemented in software, firmware, or hardware and include any type of programmed steps performed by system components.
[0032] The processor can be any conventional general-purpose single-chip or multi-chip processor, which can perform logic by means of various lines such as address lines, data lines and control lines, as well as registers and shift registers.
[0033] The software modules described in the flowcharts and user interfaces herein may include various subroutines, programs, etc. Without limiting this disclosure, logic stated to be executed by a particular module may be reassigned to other software modules and / or combined together in a single module and / or made available in a shareable library.
[0034] The principles of the invention described herein can be implemented as hardware, software, firmware, or a combination thereof; therefore, illustrative components, frames, modules, circuits, and steps are described in accordance with their functionality.
[0035] In addition to the above, the logic blocks, modules, and circuits described below may be implemented or executed by a general-purpose processor, digital signal processor (DSP), field-programmable gate array (FPGA), or other programmable logic device (e.g., application-specific integrated circuit (ASIC), discrete gate or transistor logic, discrete hardware components, or any combination thereof) designed to perform the functions described herein. The processor may be implemented by a combination of a controller or state machine or computing device.
[0036] The functions and methods described below, when implemented in software, can be written in a suitable language such as, but not limited to, Java, C#, or C++, and can be stored on or transmitted through a computer-readable storage medium, such as random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), optical disc read-only memory (CD-ROM) or other optical disc storage devices (e.g., Digital Universal Optical Disc (DVD)), magnetic disk storage devices, or other magnetic storage devices including removable thumb drives. Connections can be established using computer-readable media. Such connections can include, for example, hardwired cables, including fiber optic and coaxial cables, as well as digital subscriber line (DSL) and twisted-pair cables. Such connections can include wireless communication connections, including infrared and radio.
[0037] Components included in one embodiment may be used in any suitable combination in other embodiments. For example, any of the various components described herein and / or depicted in the figures may be combined, interchanged, or excluded from other embodiments.
[0038] "A system having at least one of A, B and C" (similarly, "a system having at least one of A, B or C" and "a system having at least one of A, B and C") includes the following systems: having only A; having only B; having only C; having both A and B; having both A and C; having both B and C; and / or having both A, B and C, etc.
[0039] Now for specific reference Figure 1An example system 10 is illustrated, which may include one or more of the example devices mentioned above and further described below according to the principles of the invention. A first example device among the example devices included in system 10 is a consumer electronics (CE) device, such as an audio-visual device (AVD) 12, for example, but not limited to, an internet-enabled TV with a TV tuner (equivalently, a set-top box controlling the TV). However, alternatively, the AVD 12 may be a home appliance or household item, such as a computerized internet-enabled refrigerator, washing machine, or dryer. The AVD 12 may also alternatively be a computerized internet-enabled (“smart”) mobile phone, tablet computer, laptop computer, wearable computerized device (e.g., a computerized internet-enabled watch, a computerized internet-enabled bracelet), other computerized internet-enabled devices, a computerized internet-enabled music player, a computerized internet-enabled headset, a computerized internet-enabled implantable device (e.g., an implantable skin device), etc. In any case, it should be understood that AVD 12 is configured to implement the principles of the present invention (e.g., to communicate with other CE devices to implement the principles of the present invention, to perform the logic described herein, and to perform any other functions and / or operations described herein).
[0040] Therefore, in order to achieve this principle, AVD 12 can be... Figure 1 Some or all of the components shown may be constructed. For example, AVD 12 may include one or more displays 14, which may be implemented by a high-definition or ultra-high-definition "4K" or higher resolution flat screen and may support touch for receiving user input signals via touch on the display. AVD 12 may include: one or more speakers 16 for outputting audio according to the principles of the invention; and at least one additional input device 18, such as an audio receiver / microphone, for inputting audible commands to control AVD 12, for example. Example AVD 12 may also include one or more network interfaces 20 for communicating over at least one network 22, such as the Internet, WAN, LAN, etc., under the control of one or more processors 24. Graphics processor 24A may also be included. Thus, interface 20 may be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as, but not limited to, a mesh network transceiver. It will be understood that processor 24 controls AVD 12 to implement the principles of the invention, including other elements of AVD 12 described herein, such as controlling display 14 to present images on the display and to receive input from the display. Furthermore, it should be noted that network interface 20 may be, for example, a wired or wireless modem or router or other suitable interface, such as a wireless telephone transceiver, or a Wi-Fi transceiver as mentioned above.
[0041] In addition to the foregoing, AVD 12 may also include one or more input ports 26, such as a High Definition Multimedia Interface (HDMI) port or USB port for physically connecting (e.g., using a wired connection) to another CE device, and / or a headphone port for connecting headphones to AVD 12 to present audio from AVD 12 to a user via headphones. For example, input port 26 may be connected via wired or wireless cable to audio / video content or a satellite source 26a. Thus, source 26a may be, for example, a separate or integrated set-top box or satellite receiver. Alternatively, source 26a may be a game console or disk player containing content that can be selected by the user as favorite content for channel allocation purposes described further below. When implemented as a game console, source 26a may include some or all of the components described below with respect to CE device 44.
[0042] The AVD 12 may also include one or more computer memories 28 that are not transient signals, such as disk-based storage devices or solid-state storage devices. In some cases, the one or more computer memories are embodied as a stand-alone device within the AVD's housing, or as a personal video recording device (PVR) or video disk player for playing back AV programs, either inside or outside the AVD's housing, or as removable memory media. Furthermore, in some embodiments, the AVD 12 may include a location or positioning receiver, such as, but not limited to, a mobile phone receiver, a GPS receiver, and / or an altimeter 30, configured to receive geographic location information, for example, from at least one satellite or cell tower and provide said information to the processor 24 and / or, in conjunction with the processor 24, determine the altitude of the AVD 12. However, it should be understood that, according to the principles of the invention, another suitable location receiver besides a mobile phone receiver, GPS receiver, and / or altimeter may be used to, for example, determine the location of the AVD 12 in, for example, all three dimensions.
[0043] Continuing the description of AVD 12, in some embodiments, according to the principles of the invention, AVD 12 may include one or more cameras 32, which may be, for example, thermal imaging cameras, digital cameras such as webcams, and / or cameras integrated into AVD 12 and controllable by processor 24 to collect pictures / images and / or videos. Bluetooth transceivers 34 and other near-field communication (NFC) elements 36 may also be included on AVD 12 for communicating with other devices using Bluetooth and / or NFC technologies, respectively. An example NFC element may be a radio frequency identification (RFID) element.
[0044] In addition, AVD 12 may include one or more auxiliary sensors 37 (e.g., motion sensors such as accelerometers, gyroscopes, gyroscopes, or magnetometers, infrared (IR) sensors, optical sensors, speed and / or rhythm sensors, gesture sensors (e.g., for sensing gesture commands), etc.) that provide input to processor 24. AVD 12 may include an over-the-air (OTA) TV broadcast port 38 for receiving OTA TV broadcasts that provide input to processor 24. In addition to the foregoing, it should be noted that AVD 12 may also include an infrared (IR) transmitter and / or an IR receiver and / or an IR transceiver 42, such as an IR data association (IRDA) device. A battery (not shown) may be provided to power AVD 12.
[0045] Still referencing Figure 1 In addition to AVD 12, system 10 may also include one or more other CE device types. In one example, a first CE device 44 may be used to send computer game audio and video to AVD 12 via commands sent directly to AVD 12 and / or via a server described below, while a second CE device 46 may include components similar to the first CE device 44. In the example shown, the second CE device 46 may be configured as a VR headset worn by player 47, as illustrated. In the example shown, only two CE devices 44, 46 are shown, but it should be understood that fewer or more devices may be used. For example, the principles discussed below describe how multiple players 47 communicate with each other through their respective headsets while playing a computer game provided by a game console to one or more AVD 12.
[0046] In the example shown, for the purpose of illustrating the principles of the invention, it is assumed that all three devices 12, 44, and 46 are, for example, members of a home entertainment network, or at least close to each other in a location such as a house. However, unless otherwise expressly required, the principles of the invention are not limited to the specific location shown by dashed line 48.
[0047] The first CE device 44, as an example of a non-limiting device, can be constructed from any of the aforementioned devices, such as a portable wireless laptop computer, a notebook computer, or a game controller, and therefore can have one or more of the components described below. The first CE device 44 can be a remote control (RC) for issuing AV play and pause commands to the AVD 12, for example, or it can be a more complex device, such as a tablet computer, a game controller communicating with the AVD 12 and / or the game console via a wired or wireless link, a personal computer, a wireless telephone, etc.
[0048] Therefore, the first CE device 44 may include one or more displays 50, which may have touch functionality for receiving user input signals via touch on the displays. The first CE device 44 may include: one or more speakers 52 for outputting audio according to the principles of the invention; and at least one additional input device 54, such as an audio receiver / microphone, for inputting audible commands to the first CE device 44, for example, to control the device 44. Example: The first CE device 44 may also include one or more network interfaces 56 for communicating via network 22 under the control of one or more CE device processors 58. It may also include a graphics processor 58A. Therefore, the interface 56 may be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, including a mesh network interface. It should be understood that the processor 58 controls the first CE device 44 to implement the principles of the invention, including other elements of the first CE device 44 described herein, such as controlling the displays 50 to present images on the displays and receive input from the displays. In addition, it should be noted that network interface 56 may be, for example, a wired or wireless modem or router or other suitable interface, such as a wireless telephone transceiver, or a Wi-Fi transceiver as mentioned above.
[0049] In addition to the foregoing, the first CE device 44 may also include one or more input ports 60 for physically connecting (e.g., using a wired connection) to another CE device, such as an HDMI port or a USB port and / or a headphone port for connecting headphones to the first CE device 44 to present audio from the first CE device 44 to a user via headphones. The first CE device 44 may also include one or more tangible computer-readable storage media 62, such as disk-based storage devices or solid-state storage devices. Furthermore, in some embodiments, the first CE device 44 may include a location or positioning receiver, such as, but not limited to, a mobile phone and / or a GPS receiver and / or an altimeter 64, configured to receive geographic location information from at least one satellite and / or cell tower, for example, using triangulation, and to provide said information to the CE device processor 58 and / or, in conjunction with the CE device processor 58, determine the altitude of the first CE device 44. However, it should be understood that, according to the principles of the invention, another suitable location receiver besides a mobile phone and / or a GPS receiver and / or an altimeter may be used to, for example, determine the location of the first CE device 44 in, for example, all three dimensions.
[0050] Continuing the description of the first CE device 44, in some embodiments, according to the principles of the invention, the first CE device 44 may include one or more cameras 66, which may be, for example, thermal imaging cameras, digital cameras such as webcams, and / or cameras integrated into the first CE device 44 and controllable by the CE device processor 58 to collect pictures / images and / or videos. The first CE device 44 may also include a Bluetooth transceiver 68 and other near-field communication (NFC) elements 70 for communicating with other devices using Bluetooth and / or NFC technologies, respectively. An example NFC element may be a radio frequency identification (RFID) element.
[0051] In addition, the first CE device 44 may include one or more auxiliary sensors 72 (e.g., motion sensors, such as accelerometers, gyroscopes, gyroscopes, or magnetometers, infrared (IR) sensors, optical sensors, speed and / or rhythm sensors, gesture sensors (e.g., for sensing gesture commands), etc.) that provide input to the CE device processor 58. The first CE device 44 may also include other sensors that provide input to the CE device processor 58, such as one or more climate sensors 74 (e.g., barometers, humidity sensors, wind sensors, light sensors, temperature sensors, etc.) and / or one or more biometric sensors 76. In addition to the foregoing, it should be noted that in some embodiments, the first CE device 44 may also include an infrared (IR) transmitter and / or an IR receiver and / or an IR transceiver 78, such as an IR data association (IRDA) device. A battery (not shown) may be provided to power the first CE device 44. The CE device 44 may communicate with the AVD 12 via any of the communication modes and associated components described above.
[0052] The second CE device 46 may include some or all of the components shown for CE device 44. Either or both CE devices may be powered by one or more batteries.
[0053] Referring now to the aforementioned at least one server 80, the server includes at least one server processor 82, at least one tangible computer-readable storage medium 84, such as a disk-based storage device or a solid-state storage device, and at least one network interface 86, which, under the control of the server processor 82, allows communication on network 22 with... Figure 1 It can communicate with other devices, and in fact, facilitate communication between the server and client devices according to the principles of the invention. It should be noted that network interface 86 can be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other suitable interface, such as a wireless telephone transceiver.
[0054] Therefore, in some implementations, server 80 may be an internet server or an entire server "cluster," and in example implementations such as online game applications, it may include and perform "cloud" functionality, enabling devices of system 10 to access a "cloud" environment via server 80. Alternatively, server 80 may be provided by... Figure 1 Other devices shown are implemented in the same room or nearby using one or more game consoles or other computers.
[0055] The methods described herein can be implemented as software instructions executed by a processor, a suitably configured application-specific integrated circuit (ASIC) or field-programmable gate array (FPGA) module, or any other convenient means as will be understood by those skilled in the art. Where employed, the software instructions can be embodied in a non-transitory device such as a CD-ROM or flash drive. Alternatively, the software code instructions can be embodied as a transient arrangement of, for example, radio or optical signals, or via download over the Internet.
[0056] Now for reference Figures 2 to 6 This illustrates the first technique used to determine if an image has been "forged," meaning it has been digitally altered from the original image. Figure 2 In this context, the original image 200 that can be displayed on the display 202 shows a person with a face 204 speaking an audible phrase 206. The image 200 can be, for example, an image from an I-frame of a video stream, and some or all frames of the video stream can be processed as disclosed herein.
[0057] A person operating a computer 208 using an input device 210, such as but not limited to a keyboard, can alter images and / or audio to produce an altered image 212 of that person, who may be depicted as uttering an altered audible phrase 214. The principle of the invention is to detect that the altered image 212 has actually been altered from the original image 200.
[0058] Figure 3 This illustrates the logic that can be executed in the first technique, while Figure 6 Provide can reflect Figure 3 The following is an example architecture of the logic. Starting in box 300, an image is received. In box 302, the image can be directly analyzed by processing it via a first neural network (NN), such as a convolutional neural network (CNN). In box 304, the first NN outputs a feature vector representing the image.
[0059] Furthermore, in box 306, an image can be input to the face recognition module to analyze artifacts (also referred to as irregularities in this paper) in the face and / or background of the image, as well as illumination irregularities in the image. Feature vectors can be output to box 304 using the face recognition module of one or more neural networks.
[0060] For example, irregularities of a face in an image (spatial domain) can include small areas with a checkerboard appearance, indicating blurred resolution due to digital alterations.
[0061] Furthermore, the image can be transformed to the frequency domain using, for example, the Discrete Fourier Transform (DFT) of the output spectrum in box 308, and the spectrum can be analyzed using another neural network, such as a CNN, in box 310 to detect irregularities in the image in the frequency domain. A feature vector representing the spectrum is provided to box 304.
[0062] For example, irregularities in the frequency domain can include one or more bright spots along the periphery of the graphical representation of an image in the frequency domain.
[0063] Moving to decision diamond 312, the detection modules of one or more neural networks can analyze the feature vectors from box 304 to determine whether one or more irregularities exist in the spatial and / or frequency domains. If no irregularities exist, the process can end at state 314; however, in some implementations, if any irregularities exist in any domain, an indication that the image is a forgery can be returned at box 316. In other implementations, an indication that the image is a forgery can only be returned at box 316 if irregularities exist in both the spatial and frequency domains.
[0064] Brief reference Figure 4 The diagram illustrates the process used to train the neural network (NN) discussed in this paper. Starting in box 400, the original, unaltered image of the ground reality is input into the NN. Furthermore, in box 402, an altered or forged image of the ground reality is input into the NN. Designers can use "deepfake" techniques to generate forged images from the original ground reality image. The NN can be programmed to begin analysis using any or example irregularities discussed above for both the frequency and spatial domains. In box 404, the NN is trained on the ground reality input. Reinforcement learning can then be applied to refine the training of the NN at box 404.
[0065] Figure 5 Example spatial domain irregularities and frequency domain irregularities are shown. The original image 500 is shown in the original spatial domain 502 and the original frequency domain 504. A modified image 506 of the original image 500 has a modified spatial domain image 508 and a modified frequency domain depicted at 510.
[0066] As shown in the figure, region 512 in the altered spatial domain image 508 has a magnified checkerboard pattern depicted at 514. Illumination irregularities may also exist between the original image and the altered image.
[0067] One or more frequency domain irregularities 516 can also be detected in the image representation in the frequency domain 510. As shown, frequency domain irregularities 516 can include bright spots along the edges or perimeters depicted in the frequency domain graph. In the example shown, there are two bright spots on each side, indicating irregularities caused by image changes in the frequency domain.
[0068] Figure 6 This can be used to illustrate Figure 3 An example architecture for the logic is provided. An image 600, to be tested against changes, is input to a face detection module 602, which analyzes the image in the spatial domain to detect illumination irregularities in the image at a neural network (NN) 604 and to perform a face resolution / irregularity check at 606. The face detection module 602 may employ image recognition principles and may be embodied by one or more NNs.
[0069] Furthermore, image 600 can be directly input into NN 608 for direct analysis using additional rules, which can be those of a CNN. It should be noted that NN 608 extracts the feature vector of the image. Additionally, NN 604 performs image processing and is particularly advantageous when sufficient training data is lacking. However, NNs 604 and 608 can be implemented using a single NN.
[0070] Furthermore, image 600 is processed by Discrete Fourier Transform (DFT) 610, the DFT output of which represents the spectrum 612 of image 600 in the frequency domain. Spectrum 612 is sent to CNN 614 for spectrum analysis.
[0071] A face recognition module 602 (including illumination irregularity check 604 and face resolution / artifact check 606) and CNNs 608 and 614 generate a set of feature vectors 616 representing image 600 in both the spatial and frequency domains. A detection module 618, which may be implemented by one or more neural networks (e.g., recurrent neural networks, such as long short-term modules (LSTM)), analyzes the feature vectors according to the principles proposed herein to determine whether image 600 contains digital alterations from the original image. If so, an indication that image 600 may be a forgery is generated at 620.
[0072] Figure 7 A second technique for detecting altered video is shown, and Figure 8 Provided for embodiment Figure 7 An example architecture for the logic is shown below. In box 700, a video sequence, such as a video clip or other video frame sequence, is input into the NN. In box 702, the NN analyzes the sequence, and in box 704, the NN outputs a feature vector representing the video sequence.
[0073] When analyzing video sequences, neural networks (NNs) can be trained to learn natural human facial movement patterns, such as during speech. As understood in this paper, when a video sequence is altered, the alteration procedure may not accurately simulate natural movement patterns, such as those of the lips, and therefore the NN may detect slightly unnatural movement patterns in forged video sequences.
[0074] Furthermore, in box 706, the audio associated with the video sequence is input to the frequency transform. In box 708, the spectrum output from the frequency transform 706 is provided to the NN for spectrum analysis to output a feature vector representing the audio to box 704.
[0075] When analyzing accompanying audio, neural networks (NNs) can be trained to learn natural human speech characteristics such as rhythm, pitch, tone patterns, and stress. As understood in this paper, when the audio, for example, is altered, the alteration process may fail to accurately simulate natural human speech patterns. Therefore, NNs can detect slightly unnatural speech patterns, such as unnatural rhythm, pitch, or tone, in fabricated audio sequences. This can be further explored along... Figure 4 The route shown completes the training, with ground-based audio and fake ground-based audio derived from the original ground-based audio used as the training set.
[0076] The feature vector 704 can be fed to a neural network, such as an RNN 710, to analyze the feature vector and detect in decision diamond 712 whether the input video sequence and / or accompanying audio has been altered from the original. If no anomalies / irregularities are found, the process can end in state 714, but if an irregularity is detected, an indication that the video sequence may have been altered is output in box 716.
[0077] In some implementations, an indication of a forgery is output at box 716 if any irregularity is detected in the audio or video. In other implementations, an indication of a forgery is output at box 716 only if irregularities are detected in both the audio and video.
[0078] Figure 8 This can be used to illustrate Figure 7 The logical architecture is as follows: Video sequence 800 is input to a neural network (NN) 802, such as a CNN, to extract feature vectors 804. Furthermore, audio, such as speech 806, is input to a frequency transform 808, such as a short-time Fourier transform (STFT), to generate a representation of the audio in the frequency domain. This audio representation is analyzed by a neural network (NN) 810, such as a CNN, to extract feature vectors. A neural network (NN) 212, such as an RNN (e.g., LSTM), analyzes the feature vectors according to the principles described herein to detect any irregularities in the video sequence 800 and audio 806 in box 814. State 816 indicates the output of an indication that the input may be a forgery.
[0079] Turn now Figure 9 This paper illustrates a third technology for addressing the generation of counterfeit videos using blockchain technology and / or digital fingerprinting. Typically, the hash / signature of a video can be integrated into an imaging device, such as a web browser, into a smartphone or other recording device, or encoded into hardware. A digital fingerprint can be generated from data bits throughout the video or a sub-track, such that if the video content changes, the fingerprint will also change. The digital fingerprint can be generated along with metadata, such as the location and timestamp where the video was originally created. Each time a redistribution of the video is attempted, the distributor must request permission from the original video on the blockchain and link a new block for the new (copied) video, making it easy to trace back to the original video and any node on the black chain. Before uploading the video again, the fingerprint of the new video can be matched against the original fingerprint to determine if the video being added has been tampered with.
[0080] For example, video websites could incorporate video fingerprinting detectors, so that each time a video is uploaded / downloaded, it is logged and timestamped. If a video is classified as fake based on a fingerprint that doesn't match the original video fingerprint, this logging could continue throughout the entire chain. This mimics antivirus software, but in this case, all users are protected simultaneously.
[0081] Starting at box 900, the original (“real”) video, along with the video’s hash, is added to the video blockchain. This hash can serve as a digital fingerprint and is typically based on pixel values, encoded information, or other image-related values within the video. A request to copy the video can be received at box 902, and the request can be granted at box 904.
[0082] Moving to box 906, a request to add a new video, a copy of the original video, back to the blockchain can be received. To make the request valid, a hash (fingerprint) can be attached to the new video. Proceeding to decision diamond 908, the hash of the video attempting to be added to the blockchain is compared with the hash of the original video from which the video was copied, and if the hashes match, the new video can be added to the blockchain at box 910.
[0083] On the other hand, if a hash mismatch is determined in decision diamond 908, the logic can move to box 912 to either reject adding the new video to the blockchain or add the new video to the blockchain along with an indication that the new video has been altered from the original video and is therefore likely a forgery. If necessary, box 914 can make the altered video inaccessible from the blockchain or otherwise unplayable.
[0084] In addition, when a modified video is detected, the logic can move to box 916 to report to the Internet Service Provider (ISP) or other reseller of the new modified video that the video has actually been digitally altered from the original video, and therefore should be checked to determine whether the new (modified) video should be removed from public view. Figure 10 As shown.
[0085] As shown in the figure, the user interface (UI) 1000 can be presented in the execution Figure 9 The logic of the device is displayed on and / or from the execution device 1002. Figure 9 The UI 1000 may include a notification 1004 indicating that a potentially counterfeit video has been detected. The UI 1000 may also include a selector 1006 to enable a user to report the presence of a counterfeit, along with identification information, to a dealer or other organization.
[0086] Figure 11 and Figure 12 The additional fingerprint logic is shown. From Figure 11 Starting at block 1100, based on the creation of the new original video and / or during video upload or download, hashing is performed in the frame for at least some implementations of all video frames and in some implementations of all video frames. Proceeding to block 1002, the hash is then embedded in the frame from which it is derived.
[0087] In the example, the hash of a video frame can be steganographically embedded in the video frame in a way that is imperceptible to the naked eye and can be uniformly distributed across the video frame. For example, the pixels for each segment of the steganographic hash can be located at a known location, either because it is always a fixed location, or because said location is contained in the frame's video metadata (allowing it to be different for each frame). Knowing this location allows pixels representing the hash to be excluded from the video data being hashed. In other words, the original hash is created only by pixels that have not been steganographically modified. Video compression algorithms can also use this location to ensure that the pixels representing the hash are not compressed or altered in a way that would affect the hash.
[0088] Figure 12 The video playback software then reverses this process. Starting at box 1200, a hash of the steganographic embedding is extracted from the video frame. Moving to box 1202, the remaining pixels of the video frame are hashed. Proceeding to decision diamond 1204, the new hash is compared with the hash extracted from the frame. If they match, the frame has not been altered from the original source video, and therefore, if it is necessary to add the video to the blockchain (assuming all or at least a threshold number of frame hashes match), the logic moves to box 1206 to indicate this. If the hashes do not match, the logic moves to box 1208 to indicate that the video being viewed has been altered from the original video, with a red border or highlight around the altered frame, for example. The altered portion of the frame can even be outlined.
[0089] This same verification process can be performed on a backend server that detects forgeries and proactively prevents them from being published or adds warnings to the video.
[0090] If any malicious actor alters the source video in any meaningful way, the frames will be hashed differently and / or the embedded steganographic hash will be compromised. The alteration to the video can be detected as long as there are benevolent actors at both ends of this activity.
[0091] Figure 13 The diagram illustrates hybrid techniques that combine the principles described above. Box 1300 indicates that image processing / video sequence analysis combined with frequency domain analysis can be used to identify artifacts / irregularities in video. Box 1302 further indicates that speech processing can be combined with any of the above techniques to identify artifacts / irregularities in video. Box 1304 indicates that the identification of artifacts / irregularities in video can be combined with blockchain technology to track the original (real) video and its altered (forged) copies.
[0092] Figures 14 to 16 Provide other examples of artifacts or irregularities that may appear in altered images (labeled as "fake" images in the figure). Figure 14 The first real image 1400 has been modified to produce a corresponding modified image 1402, where the lighting in region 1404 appears brighter than the corresponding region in the first real image 1400. Similarly, the second real image 1406 has been modified to produce a modified image 1408, where the lighting in region 1410 on the face appears brighter than in the real image 1406. The resolution of modified images 1402 and 1408 is also lower than the resolution of the corresponding real images 1400 and 1406, meaning that the NN can learn to distinguish modified images based on one or both of illumination irregularities and resolution reduction.
[0093] Figure 15 The altered image 1500 is shown, where image irregularities or artifacts exist in a small region 1502 due to upsampling performed by a generative adversarial network (GAN) to produce the altered image 1500. As shown in the decomposition diagram 1504 of region 1502, GAN irregularities can include uniform solid color regions in the image, where non-uniform solid color subjects appear in the original image (in the example shown, grass with different shades).
[0094] Figure 16Showing a real image 1600 and a modified image 1602 derived from real image 1600 by overlaying another person's face onto the face of the object in real image 1600. As shown in 1604, this overlay results in the face being misaligned with the head or other parts of the body, in this case, the nose being misaligned with the angle at which the head is depicted.
[0095] It will be understood that although the principles of the invention have been described with reference to some exemplary embodiments, these embodiments are not intended to be limiting, and various alternative arrangements may be used to achieve the subject matter claimed herein.
Claims
1. A system comprising: At least one face detection module is configured to receive an image and determine whether there is at least one texture irregularity on a face in the image or at least one texture irregularity between the face and the background in the image, or both. At least one first neural network is used to receive the image and output features representing the image in the spatial domain; At least one Discrete Fourier Transform (DFT) is used to receive the image and output the spectrum to at least one second neural network, the second neural network being configured to output features representing the spectrum in the frequency domain; At least one detection module is configured to access features output by the face detection module, the first neural network, and the second neural network to determine whether the image has been altered from the original image and to provide an output representing it, wherein the detection module is configured to output an indication that the image has been digitally altered from the original image in response to determining that both frequency domain irregularities and spatial domain irregularities exist in the image.
2. The system of claim 1, wherein the texture irregularity includes a checkerboard pattern.
3. The system of claim 1, wherein the detection module determines that the image has been altered from the original image by detecting at least one irregularity in the spectrum.
4. The system of claim 3, wherein the irregularity in the spectrum includes at least one brightness region that is brighter than the corresponding region in the original image.
5. The system of claim 4, wherein the brightness region is located along the periphery of the image in the frequency domain.
6. The system of claim 3, wherein the irregularity in the spectrum comprises a plurality of brightness regions.
7. The system of claim 6, wherein the plurality of brightness regions are located along the periphery of the image in the frequency domain.
8. The system of claim 1, wherein the face detection module is configured to output a feature vector indicating illumination irregularities on a face in the image, the illumination irregularities indicating that the image has been altered from the original image.
9. A method comprising: An image is processed by at least one face detection module to determine whether there is at least one texture irregularity on a face in the image or at least one texture irregularity between the face and the background in the image, or both. The image is processed by a first neural network to output features representing the image in the spatial domain; The image is processed by at least one Discrete Fourier Transform (DFT) to output a spectrum, and the spectrum is further processed by a second neural network to output features representing the spectrum in the frequency domain; and The image is determined by accessing features output by the face detection module, the first neural network, and the second neural network through at least one detection module to determine whether the image has been altered from the original image and to provide an output representing it, wherein an indication that the image has been digitally altered from the original image is output in response to determining that both frequency domain irregularities and spatial domain irregularities exist in the image.
10. The method of claim 9, wherein the texture irregularity includes texture irregularity on a face in the image or texture irregularity between the face and the background in the image, or both.
11. The method of claim 9, wherein the texture irregularity includes a checkerboard pattern.
12. The method of claim 9, wherein the irregularity in the image in the frequency domain includes at least one brightness region that is brighter than the corresponding region in the original image.
13. The method of claim 12, wherein the brightness region is located along the periphery of the image in the frequency domain.
14. The method of claim 9, wherein the irregularity in the frequency domain comprises a plurality of brightness regions.
15. The method of claim 14, wherein the plurality of brightness regions are located along the periphery of the image in the frequency domain.
16. The method of claim 9, further comprising outputting a feature vector indicating illumination irregularities on a face in the image, the illumination irregularities indicating that the image has been altered from the original image.
Citation Information
Patent Citations
Face authentication system for cellphone, and program
JP2008059542A
Method of motion vector and feature vector based fake face detection and apparatus for the same
KR1020170006355A
Tampering detection and location identification of digital audio recordings
US20170200457A1