Deepfake detection using synchronous observation of machine learning residuals
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- ORACLE INT CORP
- Filing Date
- 2023-09-25
- Publication Date
- 2026-04-28
AI Technical Summary
High-fidelity deepfakes pose a security risk as they can manipulate a person's face and voice to make them say or do something they did not, which can disrupt global markets and undermine international stability, and existing detection methods are not effective in real-time and at fine resolution.
A deepfake detection system that converts audiovisual content into time-series signals, generates residual signals based on machine learning estimates, and performs sequential analysis in a two-dimensional array to detect anomalies, enabling real-time and accurate identification of deepfake content.
Enables real-time, accurate detection of deepfake content by analyzing audiovisual signals at fine resolution, identifying anomalies through multivariate spatiotemporal characterization, and generating alerts for deepfake modifications.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Background technology]
[0001] background A "deepfake" is a video in which a person's face and / or voice has been manipulated using artificial intelligence (AI) software in a way that makes the altered video appear authentic. High-fidelity deepfakes are of growing concern. Deepfakes can represent a person as saying or doing something that they did not say or do. It is possible that malicious actors could inject deepfake content into live video / audio streams. This poses an international security risk, for example, where a national or international leader speaking in a live communication or broadcast could appear and sound as if they are saying something that will disrupt global markets or undermine international stability. Summary of the Invention
[0002] overview In one embodiment, a non-transitory computer-readable medium is presented. The non-transitory computer-readable medium includes computer-executable instructions stored thereon, which, when executed by at least a processor of the computer, cause the computer to perform the operations or steps of a method. The instructions cause the computer to convert an audiovisual signal containing speech by a human speaker into a set of time-series signals including a video subset of time-series signals for video and an audio subset of time-series signals for audio. The instructions cause the computer to generate a set of residual time-series signals from the set of time-series signals and a set of estimates of the time-series signals produced by a machine learning model, the machine learning model generating the estimates to be consistent with authentic speech by the human speaker. The instructions cause the computer to place residual values from a synchronized observation of one of the set of residual time-series signals into a two-dimensional array divided into a video partition and an audio partition. The residual values generated for the video subset are placed in the video partition, and the residual values generated for the audio subset are placed in the audio partition. The instructions cause the computer to perform sequential analysis of the residual values across two dimensions of the two-dimensional array to detect anomalies in the residual values. The instructions also cause the computer, in response to detecting the anomaly, to generate an alert that deepfake content that misrepresents a human speaker or speech has been detected in the audiovisual signal.
[0003] In one embodiment, a computing system is presented. The computing system includes at least one processor, at least one memory operatively connected to the processor, and one or more non-transitory computer-readable media. The non-transitory computer-readable media includes instructions stored thereon that, when executed by at least the processor, cause the computing system to perform the operations or steps of a method. The instructions cause the computing system to convert audiovisual content of a person speaking into a set of time-series signals. The instructions cause the computing system to generate a residual time-series signal indicating the degree to which the time-series signals differ from a machine-learning estimate of the person's authentic speech. The instructions cause the computing system to organize residual values from one synchronized observation of the residual time-series signals into an array of residual values for a time point. The instructions cause the computing system to perform sequential analysis of the array of residual values to detect anomalies in the residual values for the time point. The instructions also cause the computing system to generate an alert in response to detecting the anomaly that at least one of a fake voice word or an altered movement has been detected in the audiovisual content.
[0004] In one embodiment, a computer-implemented method is presented. The method includes converting audiovisual content of a person speaking into a set of time-series signals. The method includes generating a residual time-series signal indicating the degree to which the time-series signals differ from a machine-learning estimate of the person's authentic speech. The method includes organizing residual values from one synchronous observation of the residual time-series signals into an array of residual values for a time point. The method includes performing a sequential analysis of the array of residual values to detect anomalies in the residual values for the time point. The method also includes generating an alert that deepfake content has been detected in the audiovisual content in response to detecting the anomaly.
[0005] The accompanying drawings, which are incorporated into and constitute a part of this specification, illustrate various systems, methods, and other embodiments of the present disclosure. It should be understood that the illustrated element boundaries (e.g., boxes, groups of boxes, or other shapes) in the figures represent one embodiment of the boundaries. In some embodiments, one element may be implemented as multiple elements, or the multiple elements may be implemented as one element. In some embodiments, an element shown as an internal component of another element may be implemented as an external component, and vice versa. Additionally, elements may not be drawn to scale. [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 illustrates one embodiment of a deepfake detection system associated with autonomous deepfake detection. [Figure 2] FIG. 1 illustrates one embodiment of a deepfake detection method associated with autonomous deepfake detection. [Figure 3] FIG. 1 illustrates an example two-dimensional static array for one observation associated with autonomous deepfake detection. [Figure 4A] FIG. 1 illustrates a three-dimensional plot of an exemplary video / audio surface of video and audio signal values in one observation of a synchronized, uniformly sampled database of time-series signals. [Figure 4B] FIG. 1 illustrates an exemplary three-dimensional plot of residual surfaces for audiovisual content consistent with authentic speech by a human speaker. [Figure 4C] FIG. 1 illustrates an exemplary three-dimensional plot of a residual surface for audiovisual content containing anomalies indicative of deepfake modification. [Figure 5] FIG. 1 illustrates an additional example method for deepfake detection associated with autonomous deepfake detection. [Figure 6]FIG. 1 illustrates one embodiment of a computing system configured with the disclosed example systems and / or methods. DETAILED DESCRIPTION OF THE INVENTION
[0007] Detailed Description Described herein are systems and methods that provide autonomous deepfake detection based on multivariate spatiotemporal characterization and analysis of video and integrated audio. In one embodiment, the deepfake detection system autonomously detects deepfake modifications to audiovisual content. In one embodiment, the deepfake detection system detects deepfake content at a time point in the audiovisual content based on an analysis of residual values for that time point. In one embodiment, the deepfake detection system detects deepfake content using a two-dimensional, moment-by-moment analysis of video and audio of a human speaker. For example, the deepfake detection system analyzes a two-dimensional matrix of residuals between ML estimates and actual values of the audiovisual content for a time point using two-dimensional sequential analysis to detect anomalies in the residuals. The ML estimates are consistent with authentic speech by a human speaker. Anomalies in the residuals indicate the presence of deepfake content.
[0008] In one embodiment, audiovisual content including speech by a human speaker is converted into two sets of time-series signals: a video set representing video content within the audiovisual content, and an audio set representing audio content within the audiovisual content. An ML model generates estimates for both sets of time-series signals that are consistent with authentic speech from the human speaker. Residual time-series signals are generated from the time-series signals and the estimates. Residuals in the residual time-series signals indicate the degree to which values in the time-series signals deviate from estimates that are consistent with authentic speech.
[0009] A single synchronous observation of the residual time signals—in other words, a slice one observation thick across all residual signals at one time point—selects residual values from each signal. For example, an observation in one time series signal is synchronized with an observation in another time series signal when the observations are simultaneous and appear at corresponding time stamps in the respective time series signals. The selected residual values are placed into a two-dimensional array representing frames of audiovisual content. In the two-dimensional array, values from the residual time series signal corresponding to the video time series signal are placed in the video partition of the two-dimensional array, and values from the residual time series signal corresponding to the audio time series signal are placed in the audio partition of the two-dimensional array. Sequential analysis for anomaly detection is performed on individual rows and columns of the two-dimensional array, thus analyzing the residuals in two dimensions across synchronous observations or frames. If any of the sequential analyses detects an anomaly, an alert is generated indicating the presence of deepfake content. This process can be repeated as a loop for a series of synchronous observations, analyzing the audiovisual content frame by frame in two dimensions to detect deepfake modifications to a human speaker's speech or video.
[0010] As used herein, the term "time series signal" refers to a data structure in which a series of data points (e.g., observations or sampled values) are indexed in time order. In one embodiment, the data points of a time series signal may be indexed by timestamp and / or observation number. In one embodiment, the data points of a time series signal are repeated at uniform or regular intervals. In one embodiment, the data points of a time series are repeated at irregular, non-uniform, or non-regular intervals and may then be preprocessed with analytical resampling to achieve uniform or regular spacing, for example, as discussed in more detail below with respect to resampler and synchronizer 160 and under the heading "Resampling a Time Series at a Uniform Rate."
[0011] As used herein, the term "time series database" refers to a data structure that contains one or more time series signals that share a common index (such as a series of timestamps, locations, or observation numbers).
[0012] As used herein, the term "residual" refers to the difference between a value (such as a sampled or resampled value) and an ML prediction or ML estimate of what that value will be predicted by an ML model. Thus, a residual time series signal refers to the residual values of a time series between the actual values of the time series and the ML estimate of that value of the time series.
[0013] As used herein, the term "audiovisual content" refers to video with integrated audio. As used herein, the term "audiovideo signal" (or "audiovisual signal") refers to a stream of information used to convey audiovisual content.
[0014] In one embodiment, the deepfake detection system and method as shown and described herein enable real-time analysis of live streaming audio-video for deepfake content at fine resolution. The reason for this is that, in one embodiment, the deepfake detection system and method as shown and described herein are naturally parallelizable to multi-threaded, multi-core CPUs and GPUs, while other deepfake detection techniques, such as neural networks and support vector machines, cannot be parallelized due to the stochastic optimization of weights. In one embodiment, the ability to analyze audio-video at fine resolution in real time enables more accurate and more sensitive identification of deepfake content.
[0015] No act or function described or claimed herein is performed by the human mind, and the interpretation that any act or function can be performed by the human mind is inconsistent with and contrary to this disclosure.
[0016] -An example deepfake detection system- 1 illustrates one embodiment of a deepfake detection system 100 associated with autonomous deepfake detection based on multivariate spatiotemporal characterization and analysis of video and integrated audio. Deepfake detection system 100 includes a signal converter 105, a residual signal generator 110, a sequence generator 115, a two-dimensional sequential analyzer 120, and an alert generator 125. In one embodiment, each of these components 105, 110, 115, 120, and 125 (and their respective subcomponents) of deepfake detection system 100 may be implemented as a software module.
[0017] In one embodiment, audiovisual content 130 may be received from an audiovisual content source. In one embodiment, signal converter 105 is configured to convert audiovisual content 130, including speech by a human speaker, into a set of time-series signals 135. In this manner, the audiovisual content of a person speaking may be represented by a set of time-series signals. The set of time-series signals 135 includes a video subset of time-series signals 140 for video and an audio subset of time-series signals 145 for audio. In one embodiment, signal converter 105 includes video sampler 150. Video sampler 150 is configured to convert audiovisual content 130 into a video subset of time-series signals 140 by sampling the time-series signals from pixels of frames of video in audiovisual content 130. Signal converter 105 includes audio sampler 155. The audio sampler 155 is configured to convert the audiovisual content 130 into an audio subset of the time-series signals 145 by sampling the time-series signals from a frequency range of sounds within the audio signal. In one embodiment, the signal converter 105 includes a resampler and synchronizer 160. The resampler and synchronizer 160 is configured to resample one or more of the time-series signals 135 so that the set of time-series signals 135 is sampled at a uniform rate. The resampler and synchronizer 160 is also configured to phase-shift the time-series signals so that observations of the signals are synchronized. The resampled, synchronized set of time-series signals 135 is provided to the residual signal generator 110.
[0018] In one embodiment, the residual signal generator 110 is configured to generate a set of residual time series signals 165 between the set of time series signals 135 and a set of estimates of the time series signals produced by the machine learning model 170. In this manner, a residual time series signal may be generated that indicates the degree to which the time series signals differ from the machine learning estimates of authentic speech by a human speaker. The machine learning model 170 is trained or configured to generate estimates that are consistent with authentic speech by a human speaker. In one embodiment, the machine learning model is trained on a reference dataset 172 of time series signals that represent authentic speech by a human speaker. The reference dataset 172 may be a time series signal of a human speaker making a previous speech known to be authentic, or an initial segment of the time series signals 135 that is designated to represent authentic speech. The set of residual time series signals 165 includes a subset of the video signal 140 and a video subset 175 of residual time series signals calculated from the model-generated estimates for the subset of video signals. The set of residual time series signals 165 also includes an audio subset 180 of residual time series signals calculated from the subset of audio signals 145 and model-generated estimates for the subset of audio signals. The set of residual time series signals 165 is provided to a sequence generator 115.
[0019] In one embodiment, array generator 115 is configured to place residual values from one synchronized observation of the set of residual time-series signals into array 185. In this manner, residual values from one synchronized observation (at a time point) of the residual time-series signals may be placed into an array of residual values for that time point. Array 185 is partitioned into video partition 186 and audio partition 187. Residual values in synchronized observations generated for video subset 175 are placed into video partition 186. Residual values in synchronized observations generated for audio subset 187 are placed into audio partition 187. In one embodiment, array 185 is two-dimensional or rectangular. In one embodiment, two-dimensional array 185 is a data structure of values arranged in rows along a first (e.g., horizontal) dimension and columns along a second (e.g., vertical) dimension. In one embodiment, two-dimensional array 185 is partitioned into video partition 186 and audio partition 187 along the larger dimension. In one embodiment, two-dimensional array 185 has a minor dimension of a size that encompasses the minor dimension of a pixel grid for a frame of video in the audiovisual content. In one embodiment, two-dimensional array 185 has a major dimension that encompasses both the number of columns in the audio partition and the major dimension of the pixel grid multiplied by the number of color channels per pixel. In one embodiment, array generator 115 is configured to place residual values generated for video subset 175 into the video partition in cells that correspond to the locations of pixels in the pixel grid. Array 185 is provided to sequential analyzer 120.
[0020] In one embodiment, sequential analyzer 120 is configured to perform a sequential analysis of the residual values of the array to detect anomalies in the residual values at a point in time. In one embodiment, sequential analyzer 120 is a two-dimensional (2D) sequential analyzer configured to perform a sequential analysis of the residual values across two dimensions of rectangular array 185 to detect anomalies in the residual values in the array. In one embodiment, 2D sequential analyzer 120 is configured to perform a sequential analysis for the larger dimension along one or more rows in the larger dimension of rectangular array 185. Each row in the larger dimension of rectangular array 185 includes cells in both video partition 186 and audio partition 187 of rectangular array 185. In one embodiment, sequential analyzer 120 is configured to perform a sequential probability ratio test on the residual values in the array to detect one or more anomalies in the residual values. In one embodiment, 2D sequential analyzer 120 is configured to perform a sequential probability ratio test across the rows and columns of rectangular array 185. In one embodiment, the 2D sequential analyzer 120 is configured to use parallel processors to simultaneously perform sequential probability ratio tests across the rows and columns of the rectangular array 185. In one embodiment, the 2D sequential analyzer 120 is configured to detect anomalies 190 when any one of the sequential probability ratio tests across the rows or columns identifies an anomalous residual. A report of the detected anomalies 190 is provided to the alert generator 125.
[0021] In one embodiment, alert generator 125 is configured to generate an alert 195 in response to detecting anomaly 190 that deepfake content that misrepresents a human speaker or speech has been detected in audiovisual content 130. In this manner, an alert may be generated that deepfake content has been detected in audiovisual content.
[0022] Further details regarding the deepfake detection system 100 are presented herein. In one embodiment, operation of the deepfake detection system is described with reference to the exemplary deepfake detection method shown in FIG. 2. In one embodiment, the configuration and use of the rectangular array 185 is described with reference to the diagram of a rectangular array for one observation shown in FIG. 3. In one embodiment, operation of the 2D sequential analyzer 120 on the rectangular array is described with reference to FIG. 3 and FIGS. 4A-4C. In one embodiment, an additional exemplary method for deepfake detection is shown and described with reference to FIG. 5.
[0023] 2 illustrates one embodiment of a deepfake detection method 200 associated with autonomous deepfake detection based on multivariate spatiotemporal characterization and analysis of video and integrated audio. Multivariate spatiotemporal characterization of video and integrated audio, as described herein, refers to the individual description of many distinct portions of audiovisual content as variables within a spatial structure (such as an array) across a sequence of distinct time points. Multivariate spatiotemporal analysis of video and integrated audio, as described herein, refers to the examination of distinct portions of audiovisual content across dimensions of the array structure in synchronized observation.
[0024] In overview, in one embodiment, deepfake detection method 200 converts audio-video of a speaking human into a set of time-series signals. In this conversion, the audio and video are converted into separate audio and video subsets of the set of time-series signals. A machine learning model generates estimates for the set of time-series signals of what the time-series signals should be if they were consistent with authentic human speech. From the time-series signals and the estimates, a residual time-series signal is generated that represents the degree of deviation of the signal from values consistent with authentic speech. The residual signal is then analyzed one observation (or frame) at a time to detect anomalies, and one synchronized observation of the residual values from all residual signals is organized into a rectangular array and analyzed sequentially across two dimensions of the rectangular array to detect anomalies. If an anomaly is detected, an alert is generated indicating the presence of deepfake content in the audiovisual content.
[0025] In one embodiment, deepfake detection method 200 begins at start block 205 in response to a computer processor determining one or more of: (i) an input stream or broadcast of audiovisual content including human speech has been detected; (ii) an instruction to perform deepfake detection method 200 on audiovisual content including human speech has been received; (iii) a user or administrator of deepfake detection system 100 has initiated deepfake detection method 200; (iv) the current time is a scheduled time for deepfake detection method 200 to run; or (v) deepfake detection method 200 should begin in response to the occurrence of some other condition. In one embodiment, the computer is configured with computer-executable instructions to perform the functions of deepfake detection system 100. After starting at start block 205, deepfake detection method 200 continues to processing block 210.
[0026] In processing block 210, deepfake detection method 200 converts audiovisual content of a speaking person into a set of time-series signals. For example, deepfake detection method 200 converts an audiovisual signal containing speech by a human speaker into a set of time-series signals including a video subset of the time-series signals for video and an audio subset of the time-series signals for audio. After conversion, the audiovisual content is represented by a time series of values sampled from the audiovisual content, such as a synchronized, uniformly sampled database of time-series signals TSS as discussed below. In one embodiment, the functionality of processing block 210 is performed by signal converter 105.
[0027] In one embodiment, the deepfake detection method 200 converts audiovisual content into a time-series signal by sampling values from the audiovisual content at intervals. The sampling observes or detects values of a specific portion of the audiovisual content at a specific time. In one embodiment, the specific portion of the audiovisual content includes pixels of a video frame and a range of audio frequencies for a sound waveform. The sampled values are placed into the time-series signal for the specific portion of the audiovisual content within a position for the specific time. The sampling of values is repeated at intervals. A new value for the specific portion is placed into a subsequent position in the time-series signal for the specific portion of the audiovisual content.
[0028] Thus, in one embodiment, a series of values may be sampled at intervals from the pixels of a video frame and from the audio frequency range in a sound waveform to create a time-series signal for each pixel and each audio frequency range. In one embodiment, the sampling of the video frame may be performed by a video sampler 150 to generate a video time-series signal 140. In one embodiment, the sampling of the audio frequency range may be performed by an audio sampler 155 to generate an audio time-series signal 145.
[0029] A red / green / blue (RGB) intensity value may be sampled from each pixel of a video frame in the audiovisual content. The intensity value is a value that indicates the level of brightness of a color channel. The intensity value may range from the minimum brightness to the maximum brightness of a color channel, for example, an integer ranging from 0 (minimum brightness or no output) to 255 (maximum brightness or full output).
[0030] An amplitude value may be sampled from a range of audio frequencies for a sound waveform within the audiovisual content. The amplitude value is a value that indicates the loudness of a sound within the audio frequency range. Because an audio frequency range (or bin) may include multiple frequencies within the sound waveform, the amplitude value for the audio frequency range may be a representative amplitude value selected for the frequency range. For example, the representative amplitude value for the audio frequency range may be the maximum amplitude value of the audio frequency range, the average (arithmetic mean or median) amplitude value of the audio frequency range, or the minimum amplitude value of the audio frequency range. In one embodiment, the representative value sampled from the audio frequency range is the arithmetic mean amplitude value of the audio frequency range at the time the sample is taken.
[0031] The sampling intervals may be different for video pixels and audio frequency ranges. For example, video pixels may be sampled at intervals such as the frame rate of the video signal. For example, the audio frequency range may be sampled at a rate up to approximately twice the frequency of the highest frequency sound contained in the audio signal. To ensure that the video and audio time-series signals are sampled at a uniform rate, the audio and video time-series signals may be resampled as discussed below under the heading "Resampling the Time Series at a Uniform Rate." In one example, the resampling increases or decreases the effective sampling rate by including interpolated values in the time-series signals. To ensure that the resampled audio and video time-series signals have synchronous observations, the resampled audio and video time-series signals may be synchronized as discussed below under the heading "Synchronizing the Time-Series Signals." Thus, in one embodiment, a set of time-series signals may be stored as a time-series database in which the set of time-series signals share a common index, either due to sampling at uniform and synchronized intervals for the indexes or due to resampling and / or synchronization to reach uniform and synchronized intervals for the indexes. In one embodiment, the resampling and synchronization of the time series signal may be performed by a resampler and synchronizer 160 .
[0032] In one embodiment, the audiovisual content is sound and video recording of a human speaker or person giving a speech. For example, a speaking person or human speaker is a being who is speaking or uttering words. The visual content shows the speaker's movements while the speaker is giving the speech. The speaker's movements may include head and body movements, including mouth, eye, or other facial movements, as well as gestures with the head, limbs, hands, or fingers. The human speaker and speaker movements are represented by one or more pixels of the video signal. The audio content includes the sound of the speaker's utterance of words in the speech. The speech of the human speaker is represented by one or more frequency ranges of the audio signal.
[0033] Audiovisual content is conveyed by an audio-video signal, which includes a video signal and a simultaneous audio signal integrated with the video signal. The audio-video signal may be transmitted by broadcast for simultaneous reception by multiple devices, or by one or more individual, non-simultaneous streams to one or more devices. The audio-video signal may be encoded for transmission and decoded before sampling.
[0034] A video signal is data that describes the visual portion of audiovisual content. The video signal describes intensity values for pixels of frames of visual content over time. As mentioned above, these intensity values of pixels (such as intensity values for each of the red, green, and blue channels) may be sampled from the video signal. The video signal may be parsed to identify intensity values for various pixels, and the intensity values of the pixels may be sampled and placed into a time-series signal corresponding to that pixel. The time-series signal sampled from the video signal may be referred to herein as a video time-series signal. After sampling, the human speaker and the speaker's movements are represented in the video time-series signal.
[0035] An audio signal is data that describes the audio portion of audiovisual content. The audio signal describes the sound waveform of the audio content over time. As mentioned above, amplitude values of various frequency ranges of the sound waveform may be sampled from the audio signal. The sound waveform may be decomposed into amplitudes of frequency ranges, and the amplitude values for the frequency ranges are sampled into time-series signals corresponding to the frequency ranges. The time-series signals sampled from the audio signal may be referred to herein as audio time-series signals. After sampling, speech by a human speaker is represented by the audio time-series signals.
[0036] The set of time series signals includes a time series data structure containing pixel and audio frequency range values sampled from distinct portions of the audiovisual content, with a time series signal in the set for each distinct portion of the audiovisual content sampled.
[0037] A video time-series signal is a subset of a set of time-series signals sampled from video portions (e.g., pixels) of audiovisual content, such as a set of video time-series signals VTSS as discussed below. For example, a video time-series signal includes a series of intensity values for one or more color channels of a pixel. In another example, a video time-series signal includes a series of intensity values from two or more adjacent pixels. The intensity values recorded in a video time-series signal are indexed in the order in which the samples were taken from the pixels. The video time-series signals may be collectively referred to herein as a video subset of time-series signals.
[0038] In another example, the distinct portions of the video content are not individual pixels, but blocks of multiple adjacent pixels within a video frame. The blocks of multiple adjacent pixels are sampled, and a representative intensity value (per color channel) for the blocks is selected for inclusion in the video time-series signal. For example, the representative intensity value for a block (for a given color channel) may be the maximum intensity value of the block, the average (arithmetic mean or median) intensity value of the block, or the minimum intensity value of the block. In this embodiment, the video time-series signal includes a series of representative intensity values from two or more blocks of adjacent pixels.
[0039] An audio time-series signal is a subset of a set of time-series signals sampled from an audio portion (e.g., an audio frequency range) of audiovisual content, such as the set of audio time-series signals ATSS discussed below. For example, an audio time-series signal includes a series of amplitude values for one audio frequency range (or bin) of the audio spectrum. The amplitude values recorded in the audio time-series signal are indexed in the order in which the samples were taken from the frequency range. The audio time-series signals may be collectively referred to herein as an audio subset of the time-series signals.
[0040] Thus, in one embodiment, deepfake detection method 200 converts audiovisual content into time-series signals by repeatedly sampling values from video pixels and audio frequency ranges and writing the sampled values into time-series signals corresponding to the pixels and frequency ranges. The time-series signals are sampled (or resampled, as discussed below) to have a uniform sampling rate or interval between observations across the set of time-series signals and are synchronized (or synchronized, as discussed below) to cause simultaneous observations to occur at corresponding timestamps. Processing block 210 is then complete, and deepfake detection method 200 continues at processing block 215. Additional details regarding the conversion of audio-video content into time-series signals are provided herein below, for example, under the heading "Conversion to Time-Series Signals."
[0041] Upon completion of processing block 210, deepfake detection method 200 has created a set of time series signals representing audiovisual content, such as a synchronized, uniformly sampled database of time series signals TSS as discussed below. The set of time series signals is ready for processing with a machine learning model to generate a set of residual time series signals in processing block 215.
[0042] In processing block 215, deepfake detection method 200 generates residual time series signals that indicate the degree to which the time series signals differ from machine learning estimates of authentic speech by the person. For example, deepfake detection method 200 generates a set of residual time series signals from a set of time series signals and a set of estimates of the time series signals produced by a machine learning model, where the machine learning model produces estimates that are consistent with authentic speech by a human speaker. In one embodiment, the functionality of processing block 215 is performed by residual signal generator 110.
[0043] As discussed above, a residual value is the difference between a value and a machine learning estimate or prediction of what that value is predicted to be. A residual time series signal is a time series of residual values. A residual time series signal for a variable may be generated by calculating the difference between the values in the time series signal for that variable and the machine learning estimate for that value.
[0044] The machine learning estimation of a value is an estimation of the authentic speech of a person who is the subject of the audiovisual content. When used herein in connection with a person speaking or a human speaker giving a speech, the term "authentic" refers to audiovisual content (and a time-series signal derived therefrom) in which the words spoken by the speaker are not faked or altered, and the movements of the person while speaking are not faked or altered. In other words, the machine learning estimation of a value is an estimation or prediction of what value the time-series signal should be, or what value it should be, provided that the words spoken by the person and the movements of that person (especially their mouth movements) are not faked or altered, but rather represent authentic and unmodified audiovisual content.
[0045] The machine learning model is trained to generate estimates consistent with authentic speech by a speaking person (human speaker) (as discussed in more detail below). The machine learning model is a multivariate pattern recognition model that accepts values for multiple variables as inputs and generates estimates for each variable based on the input values of other correlated variables. The time-series signals, including both the video and audio subsets, are assigned as inputs for the variables of the machine learning model. For example, values of a first time-series signal in the set of time-series signals are provided as a sequence of values for a first variable of the multivariate machine learning model, values of a second time-series signal in the set of time-series signals are provided as a sequence of values for a second variable of the multivariate machine learning model, and so on. Because the set of time-series signals includes both the audio and video subsets of the time-series signals, the input to the machine learning model therefore includes an audio time-series signal representing audible words spoken by a person in the audiovisual content and a video time-series signal representing the person's appearance and movements while speaking the audible words in the audiovisual content.
[0046] A machine learning model generates estimates of values one observation at a time. For example, the machine learning model accepts one value for each variable from each input signal and generates an estimate for each variable based on the input values for the other variables. The input values for the variables are correlated with each other by sharing an index position (e.g., an observation number or a timestamp) within their respective time series signals. A set of residual time series signals corresponding to a set of time series signals can be generated from the values of the set of time series signals and machine learning estimates of those values. Because a machine learning estimate of a value approximates what that value should be, provided that the value is derived from authentic audiovisual content that accurately represents a human speaker, the residual between that value and the estimate of the value indicates the amount by which the value deviates from the ML-estimated authentic speech of a human speaker. Thus, the residual time series signals generated for the variables indicate, for each observation, the degree to which the value in the time series signal differs from the estimate representing the person's authentic speech. A set of residual time series signals generated from a set of time series signals sampled from audiovisual content indicates how much the set of time series signals deviate from what the time series signals should be if the audiovisual content is authentic.
[0047] Residual values for each variable are generated from the input values for that variable and the machine learning estimates for that variable. Thus, residual values for a variable are determined from the sampled (or resampled) values in the time series signal for that variable and the machine learning estimates for those values. The machine learning model iterates estimating and generating residuals for the number of observations or the length of the time series signal, thus generating a series of residuals for each variable. The residuals for a variable may be stored in a time series data structure, creating a residual time series signal for that variable. The residuals for a variable are stored in the residual time series signal for that variable in the same order as the values from which the residuals are generated appear in the time series signal.
[0048] Thus, in one embodiment, deepfake detection method 200 generates, for each time series value in the set of time series signals, an estimate of what the value should be if the audiovisual content is authentic, calculates the residual between that value and the estimate, and stores the residual values in a set of residual time series signals, thereby generating residual time series signals of residual values that indicate the degree to which the time series signals differ from a machine learning estimate of the person's authentic speech. Processing block 215 is then complete, and deepfake detection method 200 continues at processing block 220. Additional details regarding the generation of residual time series signals are provided below in the specification, for example, under the heading "ML Generation of Residuals."
[0049] Upon completion of processing block 215, a set of residual time series signals has been generated that indicate how much the audiovisual content differs from what would be expected if the audiovisual content were predicted. Each residual time series signal indicates the degree of difference between the actual authentic value and the predicted authentic value for a particular portion of the audiovisual content (e.g., pixel or frequency range) as it varies over time. The residual values may be analyzed to determine whether the corresponding portion of the audiovisual content includes deepfake modifications that are inconsistent with the authentic content.
[0050] In processing block 220, deepfake detection method 200 places residual values from one synchronous observation of the residual time series signal into an array of residual values for a point in time. For example, deepfake detection method 200 places residual values from one synchronous observation of a set of residual time series signals into a two-dimensional array that is divided into a video partition and an audio partition. Residual values generated for the video subset are placed in the video partition, and residual values generated for the audio subset are placed in the audio partition. In one embodiment, the functionality of processing block 220 is performed by array generator 115.
[0051] An array is a data structure in which data values can be stored in cells that are addressable by an index. Each cell in the array corresponds to or is assigned to a particular portion of the audiovisual content (a pixel or pixel color channel, or a frequency range). Thus, each cell in the array can be populated with a residual value from a residual time-series signal for a particular portion of the audiovisual content. The cells in the array are divided into two separate partitions of contiguous cells: an audio partition and a video partition. Cells corresponding to pixels (or pixel color channels) are included in the video partition (but not in the audio partition). Cells corresponding to frequency ranges (bins) are included in the audio partition (but not in the video partition).
[0052] The array is populated with residual values for a single specific time point. These are the residual values of one synchronized observation of a set of residual time series signals. Observations are "synchronized" if the values of the observations share a common timestamp. The residual values of that time point or synchronized observation share a common index position within their respective residual time series signals. The residual value for a specific time point, observation, or index position is selected from each of the residual time series signals in the set and placed into the array in a cell corresponding to the particular portion of the audiovisual content represented by the residual time series signal. Because the array contains values for one specific time point, observation, or index position across the set of residual time series signals, the array may be referred to herein as a "static" array. Once populated, the array indicates the degree to which each particular portion of the audiovisual content deviates from values consistent with authentic speech at that particular time point.
[0053] As used herein, "putting" a value into an array refers to storing a value in a cell of the array, for example, by writing the value to a memory location for the cell. An array is populated by putting values into the cells of the array. In one embodiment, the array is divided into two partitions: a video partition and an audio partition. A partition of the array is a distinct region or range of cells that does not overlap with another partition. A partition may be defined by a selected range of index values for the cells of the array. Cells of the array that have index values within the selected range are within the partition. In one embodiment, if the array is a two-dimensional array, the two-dimensional array is divided into a video partition and an audio partition along the larger of the two dimensions. In this configuration, the video partition spans one distinct range of the larger dimension, and the audio partition spans another distinct range of the larger dimension.
[0054] A video partition is a region or range of cells in an array reserved for values representing pixels (or color channels of pixels) of video content. When populating an array (i.e., when putting or storing values into an array), values representing pixels (or color channels of pixels) are placed in cells of the video partition and not in cells of the audio partition. Thus, residual values generated for a video subset of a time-series signal are placed in the video partition. An audio partition is a region or range of cells in an array reserved for values representing an audio frequency range of audio content. When populating an array, values representing frequency ranges (i.e., "bins") are placed in cells of the audio partition and not in cells of the video partition. Thus, residual values generated for an audio subset of a time-series signal are placed in the audio partition. As used herein, values (or time-series signals of values) "represent" a particular portion (pixel or color channel of a pixel, or frequency range) of audiovisual content if they are values sampled from a particular portion, ML estimated values for a particular portion, or residual values generated for a particular portion.
[0055] In one embodiment, an array may be a two-dimensional array, a data structure for storing data values in a matrix or grid of cells. The index value for a cell in the two-dimensional array is a tuple of two values: a position along the first dimension of the array and a position along the second dimension of the array. A two-dimensional array may also be referred to herein as a "rectangular" array.
[0056] As discussed below, the spatial correspondence between pixel locations in a video frame and cell locations in the video partition enhances the interpretability of the alert, since when deviations from authentic speech are detected, the corresponding pixel locations in the frame are easily identifiable from the cell locations. In one embodiment, the two-dimensional array has dimensions that accommodate the spatial layout of a pixel grid for a frame of video content. The video partition of the two-dimensional array has dimensions that allow values representing pixels (or color channels of pixels) to be placed in cells that correspond to the pixel's location in the frame.
[0057] For example, the smaller dimension of the array may have a size that encompasses or includes the smaller dimension of the pixel grid. Here, the smaller dimension of the array may have as its size in cells the number of pixels along the smaller dimension of the pixel grid for a frame of video content. Also, for example, the larger dimension of the array may have a size that encompasses or includes both the audio partition and the larger dimension of the pixel grid multiplied by the number of color channels per pixel. For example, the larger dimension of the array may have as its size in cells the number of pixels along the larger dimension of the pixel grid for the frame multiplied by the number of color channels per pixel in addition to the cell for the audio partition.
[0058] More generally, a frame of video content may have a pixel grid of M pixels in a first dimension by N pixels in a second dimension. A pixel has K color channels, e.g., if a pixel has red, green, and blue channels, then K=3. In one embodiment, the K values representing the color channels of a pixel are placed into a sequence of K cells in a video partition of the array. The sequence of K cells in the video partition corresponds to the location of the pixel in the pixel grid of the frame.
[0059] In one exemplary configuration of a two-dimensional array, the dimensions of the video partition are (at least) K×M times N cells. K values representing the color channels of a pixel are placed in a sequence of K cells in the first dimension, e.g., horizontally in rows along the x-axis of the two-dimensional array. A pixel P in a video frame for a cell Cx,y at coordinates (x,y) in the video partition of the two-dimensional array is m,n The corresponding location of is (m,n). In this example configuration, pixel coordinate m is the integer quotient of array coordinate x divided by the number of color channels K (m=x / K, use the division principle or modular division to get the integer quotient and discard the remainder), and pixel coordinate n is equal to array coordinate y (n=y). Also, in this example, pixel P m,n Cells Cx for K color channels in the video partition for 1,y ,...,C xK,y The corresponding locations in the sequence or rows of are (x1,y),...,(x K ,y), and x1=K×M-(K-1),...,x K = K×M−(KK). An exemplary use of this configuration, with K=3 color channels, is a 3M×N video partition 305, as shown and described below with reference to FIG.
[0060] In one exemplary configuration of a two-dimensional array, the dimensions of the video partition are M times K×N cells. The K values representing the color channels of a pixel are placed in a sequence of K cells in the second dimension, e.g., vertically in columns along the y-axis of the two-dimensional array. As noted above, a pixel P in a video frame m,n and cell C in the video partition x,y is the corresponding location. In this example configuration, the corresponding pixel P m,n and Cell C x,y , pixel coordinate m is equal to array coordinate x (m=x), and pixel coordinate n is the integer quotient of array coordinate y divided by the number of color channels K (n=y / K, using division principle or modular division). Also, in this exemplary configuration, pixel P m,nfor the K color channels in the video partition, x,y1 ,...C x,yK The corresponding locations in the sequence or column of are (x,y1),...(x,y K ), and y1=K×N-(K-1),...,y K =K×N-(KK).
[0061] An audio partition in a two-dimensional array may be positioned adjacent to a video partition along one edge of the video partition. Thus, the audio partition has one dimension equal to the dimension of the video partition along the edge where it is adjacent to the video partition. In one embodiment, the audio partition may be a one-dimensional row or column positioned along the edge of the array. In one embodiment, the audio partition is two-dimensional, with another dimension extending away from the edge where it is adjacent to the video partition. In one embodiment, the dimensions of the audio partition are equal, and the audio partition is a square region of cells having a dimension equal to the dimension of the video partition along the edge where the audio and video partitions are adjacent. In one embodiment, the audio content is subdivided into as many audio frequency ranges as there are cells in the audio partition. Thus, the number of audio frequency ranges is selected based on the number of cells in the audio partition.
[0062] In one embodiment, the audio partition is positioned along the highest edge (or "right edge") in the x dimension of the video partition and has one dimension that is the length of the y dimension of the video partition, as is the case for example with the N×N audio partition 310 shown and described below with reference to FIG. 3. Other alternative arrangements of audio partitions within a two-dimensional array are contemplated by the present invention. In another embodiment, the audio partition is positioned along the smallest edge (or "left edge") in the x dimension of the video partition and has one dimension that is the length of the y dimension of the video partition. In another embodiment, the audio partition is positioned along the smallest edge (or "top edge") in the y dimension of the video partition and has one dimension that is the length of the x dimension of the video partition. In another embodiment, the audio partition is positioned along the highest edge (or "bottom edge") in the y dimension of the video partition and has one dimension that is the length of the x dimension of the video partition. When the audio partition is positioned on the left edge, the cell coordinates (as discussed above) within the two-dimensional array are adjusted by adding the length of the audio partition in the x dimension to the x coordinate of the cell. If the audio partition is placed on the top edge, the cell coordinate in the two-dimensional array is adjusted by adding the length of the audio partition in the y dimension to the y coordinate of the cell.
[0063] Values representing audio frequency ranges are placed in cells of an audio partition. In one embodiment, the values are placed in the audio partition in ascending or descending order of range. For example, the value representing the lowest audio frequency range is placed in the first cell of the audio partition, the value representing the next lowest range is placed in the second cell of the audio partition, and so on until the value representing the highest audio frequency range of the audio content. In one embodiment, if an audio partition spans more than one row or column in a two-dimensional array, the sequence of values representing ranges may wrap across the rows and columns. For example, values may wrap from left to right (low x-dimension to high x-dimension), top to bottom (low y-dimension to high y-dimension), top to bottom (low y-dimension to high y-dimension), left to right (low x-dimension to high x-dimension), right to left (high x-dimension to low x-dimension), top to bottom (low y-dimension to high y-dimension), top to bottom (low y-dimension to high y-dimension), right to left (high x-dimension to low x-dimension), left to right (low x-dimension to high x-dimension), bottom to top (high y-dimension to high y-dimension), The audio may wrap from bottom to top (high y to low y), from left to right (low x to high x), from right to left (high x to low x), from bottom to top (high y to low y), from bottom to top (high y to low y), from right to left (high x to low x), from spiral inwards or outwards, from back and forth, or otherwise be placed into audio partitions in frequency range order.
[0064] Residual values from one synchronous observation of a set of residual time series signals are placed into an array. The synchronous observation of the residual values is the residual value at a particular time point. The synchronous observation of one of the set of residual time series signals is the set of residual values that appear at a particular index value within the residual time series signals within the set of residual time series signals. Just as time series signals share a common index across signals in a set, the set of residual time series signals share a common index across residual signals in the set. A particular index value indicates the residual value within the set of residual time series signals that occurs at that time point.
[0065] One residual value at an index value is read or extracted from each of the residual time-series signals. The extracted residual values are placed into an array. The cell locations into which the extracted residuals are placed are cell locations corresponding to the portions of the audiovisual content represented by the respective residual time-series signals from which values are extracted. Thus, a synchronous observation of the residual values is a one-value-thick "slice" across the set of residual time-series signals occurring at the time indicated by the index. A synchronous observation of the residual values is a collection of residual values occurring at one instant in time. The array may be referred to herein as a "static" array because it contains all values for one specific time point, rather than values for multiple time points.
[0066] Thus, in one embodiment, deepfake detection method 200 populates an array of residual values for a time point with a residual value from one synchronized observation of the residual time-series signals by reading a value at a particular index value for a particular time point from each of the residual time-series signals and writing that value to the array in a cell corresponding to the particular portion of the audiovisual signal represented by the residual time-series signal. Processing block 220 is then complete, and deepfake detection method 200 continues at processing block 225. Additional details regarding the structure of the array and the placement of residual values in the array are provided herein below, for example, under the heading "Residual Array Structure for Deepfake Detection."
[0067] Upon completion of processing block 220, an array of residuals has been created that indicates the magnitude of the difference between the recorded values and the ML estimate of the true values at one moment for all individual portions of the audiovisual content. This array of residuals can be analyzed sequentially to detect excessive deviations from the true values in the audiovisual content. In one embodiment where the array is two-dimensional, the rows and columns of the array can be analyzed by simultaneous parallel analysis.
[0068] In processing block 225, deepfake detection method 200 performs a sequential analysis of the residual values of the array to detect anomalies in the residual values for that time. For example, deepfake detection method 200 performs a sequential analysis of the residual values across two dimensions of a two-dimensional array to detect anomalies in the residual values. In other words, a series of residual values in the array is checked for values that deviate significantly from other residual values in the series. Residual values that deviate significantly from others indicate portions of the audiovisual signal where the difference between the actual value and the estimated authentic value is abnormally large. Portions of the audiovisual signal where such anomalies occur are likely modified to misrepresent the movements of the person speaking and / or the words spoken by the person. In one embodiment, the functionality of processing block 225 is performed by sequential analyzer 120.
[0069] If the array is a one-dimensional array, the sequential analysis is performed on the series of residual values placed in the array. If the array is a two-dimensional array, the method 200 operates to perform sequential analysis of the residual values across two dimensions of the two-dimensional array. The sequential analysis is performed on the series of residual values in rows and / or columns. Consider an example of a two-dimensional array having dimensions of K×M+N×N cells. This exemplary two-dimensional array has a video partition that is K×M×N cells and an audio partition adjacent to the right edge of the video partition that is N×N cells (similar to the configuration of static array 300 as shown and described below with reference to FIG. 3). In this exemplary two-dimensional array, the K×M+N cell (horizontal) dimension of the array is the larger dimension of the array, and the N cell (vertical) dimension of the array is the smaller dimension of the array.
[0070] In one embodiment, the sequential analysis of the residual values is performed across two dimensions of the two-dimensional array. Performing sequential analysis across two or more rows of the two-dimensional array, two or more columns of the two-dimensional array, or both two or more rows and two or more columns of the two-dimensional array are examples of performing sequential analysis across two dimensions of the array. In one embodiment, the sequential analysis is performed on each of the N rows in the two-dimensional array. In one embodiment, the sequential analysis is performed on each of the K×M+N columns in the two-dimensional array. In one embodiment, the sequential analysis is performed on each of the rows and columns of the two-dimensional array. The sequential analyses on individual rows and individual columns of the array may be performed in parallel with each other.
[0071] In one embodiment, the sequential analysis of the residual values of the sequence may be performed using a sequential probability ratio test (SPRT), as discussed in further detail below. Sequential analyses such as the SPRT may be used to detect anomalies within a sequence of values—i.e., values that deviate from other values in a manner that satisfies a threshold calculation. Such anomalous deviations within a sequence may be defined by a pre-established threshold calculation. In the SPRT, for example, a value within a sequence that satisfies a user-selected threshold of the cumulative sum of log-likelihood ratios over the sequence of values is an anomaly, as discussed further below. A user or administrator may configure a threshold to set a sensitivity level for the detection of anomalies. In one embodiment, the sensitivity level is set relatively high to detect as anomalies those residual values that are almost certainly due to manipulation of audio-video content, with a relatively lower false alert. In one embodiment, the sensitivity level is set lower to detect as anomalies those residual values that are almost certainly due to manipulation of audio-video content, with a relatively higher false alert.
[0072] The sequential analysis of the residual values of the array is used to detect anomalies in the residual values for a particular time point. Because the residual values in the array are all from one particular time point (or observation) across each residual time series signal, the sequential analysis examines sequences of residual values from multiple separate time series signals for one particular time point for anomalies, rather than examining a sequence of residual values across multiple time points within one time series signal for anomalies. Thus, residual values representing the same instant in time for multiple portions of audiovisual content are analyzed to detect whether the residual for any portion is anomalous relative to the other residuals at that time point. It is possible that the sequential analysis of the residual values may detect multiple anomalies in the sequence of residual values.
[0073] A residual value representing a portion of audiovisual content at a time that is anomalous relative to the residual values for other portions at that time indicates that portion has been altered away from what is authentic at that time. A large residual value for a particular portion of audiovisual content at a particular time is indicative of the presence of counterfeit content in that portion and time. Sequential analysis of residuals representing multiple portions of audiovisual content for the same time point thus detects or identifies portions of the content that have been altered or manipulated.
[0074] An identifier of the manipulated portion of the content may be recorded. Upon detection of an anomalous residual value, an identifier of the portion of the audiovisual content having the anomalous residual may be stored or recorded for subsequent processing. Such an identifier may include the array index location of the cell in which the anomalous residual was detected and / or the pixel location or audio frequency range of the portion corresponding to the cell containing the anomalous residual. The corresponding portion of the audiovisual content may then be derived from the index location of the cell containing the anomalous residual as discussed above.
[0075] In one embodiment, deepfake detection method 200 performs a sequential analysis of the residual values of the array to detect anomalies in the residual values for a time point by accessing or obtaining a sequence of residual values for a time point from the array, analyzing the sequence of residual values with SPRT, and detecting that a residual value in the series is anomalous because it deviates from other residual values to a degree that indicates that the portion of the content represented by the residual value may have been manipulated. In one embodiment, the sequential analysis of residual values is performed in two dimensions by analyzing the sequence of residuals from one or more rows of the array and one or more columns of the array with SPRT, and detecting anomalous residual values in any one or more of the columns or rows. Processing block 225 is then complete, and deepfake detection method 200 continues at processing block 230. Additional details regarding the sequential analysis of an array of residual values for anomaly detection are provided herein below, for example, under the heading "Static Array Residual Analysis."
[0076] Upon completion of processing block 225, anomalous residual values have been detected in the array of values from the residual time series signal at one time point or observation, subject to the presence of any such anomalous residual values. The presence of an anomalous residual value in a cell of the array indicates that the portion of the audiovisual content represented in that cell has satisfied a threshold, indicating that the portion of the content has been modified.
[0077] At processing block 230, in response to detecting the anomaly, deepfake detection method 200 generates an alert that deepfake content has been detected in the audiovisual content. For example, in response to detecting the anomaly, deepfake detection method 200 generates an alert that deepfake content that misrepresents a human speaker or speech has been detected in the audiovisual content. In one embodiment, the functionality of processing block 230 is performed by alert generator 125.
[0078] In one embodiment, deepfake content refers to the modification of audiovisual content or instances when false words are inserted and / or mouth or other facial movements are altered to correspond with speaking the false words. The false words may be inserted into speech content conveyed by an audiovisual signal by replacing words as spoken by the speaker with the sound of other words. The inserted words are false and may not actually be spoken by the recorded human speaker, but may be spoken in a very similar voice. In one embodiment, the detected deepfake content includes altered mouth movements of the human speaker to accompany the false words. The human speaker's appearance, i.e., face or features, is maintained in the detected deepfake content to misrepresent to viewers of the audiovisual content that the inserted words were spoken by the speaker. This deepfake content misrepresents a human speaker or the human speaker's speech by altering or modifying the audiovisual content away from an accurate representation of what the human speaker actually said and presenting the altered content as if it were unaltered. Thus, in one embodiment, the method detects false words that are seemingly spoken by a human speaker who is the subject of the audiovisual content. The false words that are added to the speech can then be pointed out to a viewer or audience of the audiovisual content.
[0079] To determine whether to generate an alert for a time point, deepfake detection method 200 determines whether an anomaly is detected in the sequence of residuals for that time point. Insertion of false words, altered mouth movements, or other modifications that misrepresent speech will show up as anomalies in the sequential analysis (described above in processing block 225) of observations or time points. Deepfake manipulation of audiovisual content can be inferred from the presence of anomalous residuals because the portion of the content represented by the anomaly is so inconsistent with authentic speech when compared to the other portions that it could not have been modified in the same way that the other portions were not modified.
[0080] In one embodiment, generating an alert is performed in response to detecting an anomaly in the residual by generating the alert, at least in part, directly or indirectly, because the anomaly is detected. Thus, when an observation includes an anomaly, the presence of deepfake content is inferred from the anomaly and an alert is automatically generated. When an observation does not include an anomaly, the absence of deepfake content is inferred from the absence of the anomaly and no alert is generated.
[0081] In one embodiment, the generation of an alert may be based on the detection of a single anomaly in the residual array for a time point. In one embodiment, the generation of an alert may be based on the detection of a warning threshold number or percentage of anomalies in the residual array for a time point. For example, the alert may be based on a warning threshold number or percentage of cells in the residual array that hold anomalous residuals, such as triggering an alert if an anomaly is detected in 1% or more of the cells in the residual array for that time point. Thus, the generation of an alert may be based on detecting (at processing block 225) an amount of anomalous residuals greater than a warning threshold number or percentage of residuals. Thus, in one embodiment, if the number of anomalous residuals detected for a time point is greater than the warning threshold, the presence of deepfake content is inferred and an alert is automatically generated. If the number of anomalous residuals detected does not exceed the warning threshold, no alert is generated. The generation of an alert in response to the detection of a single anomalous residual in the array for a time point may be considered a warning threshold of 1.
[0082] In one embodiment, the alert is an electronic message. In one embodiment, deepfake detection method 200 creates an electronic message indicating or communicating the presence of deepfake content in the audiovisual content. The alert may include a description of the portion of the audiovisual content where the deepfake content occurs (e.g., pixel location and audio frequency range) and the time at which the deepfake content occurs. An identifier for the manipulated portion of the content (recorded as described above in processing block 225) is retrieved from storage and written to the alert to indicate the pixel location and audio frequency range of the deepfake content. In one embodiment, the alert includes one or more pixel locations or coordinates where the deepfake content is present as seen in the observation of the visual content. In one embodiment, the alert includes one or more audio frequency ranges where the deepfake content is present as seen over a time period of the audio content. In one embodiment, the alert includes the time (e.g., timestamp, time range, observation, or frame) at which the detected deepfake content is present in the audiovisual content.
[0083] An alert may be generated and then transmitted for subsequent presentation on a display or other action. An alert may be configured to be presented by a display in a graphic user interface. An alert may be configured as a request (such as a REST request) that is used to trigger the initiation of some other function. An alert may be presented by extracting the content of the alert through a REST API.
[0084] In one embodiment, the alert may be transmitted to a GUI audiovisual player for simultaneous display with the audiovisual content. The audiovisual player may be a multimedia player embedded in a web browser or a separate software application. In one embodiment, the audiovisual player is configured to display the alert message along with the audiovisual content. In one embodiment, the GUI is configured to display the alert message (or a human-readable interpretation of the alert message) along with the audiovisual content. In one embodiment, the GUI is configured to parse pixel locations from the alert message and highlight them in the visual content by changing their color, outlining them with a colored border, or otherwise indicating where the pixels are deepfake content. In one embodiment, the GUI is configured to add a visual alert message to the visual content indicating that deepfake content is present in the audiovisual content. In this manner, the deepfake content may be visually indicated to a viewer of the audiovisual content.
[0085] In one embodiment, the generation and transmission of the alert may be performed in real time (or near real time), such that the alert is presented at the time the deepfake content occurs within a live transmission (e.g., a broadcast or stream) of the audiovisual content. In one embodiment, as used herein, "real time" refers to substantially immediate operation, with the generation and availability of the alert incurring only minimal delay, as is acceptable within the context of a live audiovisual transmission. In this manner, an alert that deepfake content is present within the audiovisual content may be presented simultaneously with the display of the deepfake within the audiovisual content. The alert may be used to draw the attention of a viewer of the audiovisual content to the fact that the deepfake content is being used to misrepresent what a speaking person is saying.
[0086] Thus, in one embodiment, deepfake detection method 200 generates an alert that deepfake content has been detected in the audiovisual content by determining that a warning threshold amount of anomalous residuals has been detected in the sequence, creating an alert indicating that deepfake content that misrepresents speech or a speaker is present in the audiovisual content, and transmitting the alert for presentation with the audiovisual content, and processing block 230 is then complete.
[0087] Upon completion of processing block 230, an electronic alert message has been generated indicating that the deepfake content included a synchronized observation (for a particular point in time) of the audiovisual content. In one embodiment, those portions of the observation where the audiovisual content was modified are identified as deepfake content in the alert. In one embodiment, the alert describes the audio frequency range of the audiovisual content where the fake speech sounds are inserted and / or the pixel locations where the speaker's mouth, face, or other movements of the fake speech were inserted or the mouth or other facial movements were altered to match speaking the fake speech. The alert can be used to warn viewers of the audiovisual content about deepfake misrepresentations of speech content.
[0088] In one embodiment, processing blocks 220-230 may repeat in a loop for a sequence of observations of the residual time series signal. If there is a subsequent observation (representing a next time point or frame) in the residual adjusted sequence signal, processing returns to processing block 220 and repeats for the subsequent observation. If there are no more observations in the residual time series signal, deepfake detection method 200 continues to END block 235 and completes.
[0089] Upon completion of deepfake detection method 200, a series of deepfake alerts are provided as companion signals to an audiovisual signal carrying the audiovisual content. The deepfake alerts correspond to the audiovisual content based on time (e.g., by timestamp), and as a result, the deepfake alerts describe the time at which they are applicable to the audiovisual content. In one embodiment, the deepfake alerts may be provided in real time, simultaneously with the audiovisual content to which the alerts correspond. This allows viewer attention to be drawn to the presence of a deepfake misrepresentation in a live audiovisual transmission at the time the misrepresentation occurs, even if the deepfake misrepresentation would otherwise be imperceptible to a human viewer.
[0090] As discussed above, audiovisual content may be converted into time-series signals, including audio time-series signals representing audio frequency ranges or bins and video time-series signals representing intensity values for color channels of pixels. In one embodiment of processing block 210, deepfake detection method 200 converts the audiovisual signal into a video subset of the time-series signals by sampling the time-series signals from pixels of frames of video in the audiovisual signal. Deepfake detection method 200 also converts the audiovisual signal into an audio subset of the time-series signals by sampling the time-series signals from a frequency range of the audio signal. In another embodiment of processing block 210, deepfake detection method 200 samples pixels of video in the audiovisual content to create a video subset of the time-series signals representing a video of a person. Deepfake detection method 200 also samples a frequency range of audio in the audiovisual content to create an audio subset of the time-series signals representing the sound of speech. Residual values generated from the video subset are placed in the video partition of the array, and residual values generated from the audio subset are placed in the audio partition of the array.
[0091] As discussed above, the sequential analysis of anomalies in the array may be performed on a sequence of residual values, including residuals generated from video intensity values and corresponding estimates, and residuals generated from audio amplitude values and corresponding estimates. Thus, in one embodiment of processing block 225, deepfake detection method 200 performs the sequential analysis along one or more rows in the larger dimension of the two-dimensional array, where each row in the larger dimension of the two-dimensional array includes cells in both the video and audio partitions of the rectangular array.
[0092] As discussed above with reference to processing block 220, the array may be two-dimensional and include a video partition having a dimension for accommodating a frame of video content, as well as an audio partition adjacent to the video partition along one edge. In one embodiment, after processing block 215, deepfake detection method 200 generates a two-dimensional array. The two-dimensional array is generated to include a video partition. The video partition has a first dimension of the video partition that is the first size of the first dimension of a pixel grid for a frame of video in the audiovisual signal. The video partition has a second dimension of the video partition that is the second size of the second dimension of the pixel grid multiplied by the number of color channels per pixel. In one embodiment of processing block 220, residual values generated for the video subset are placed in the video partition in cells that correspond to the location of the pixel in the pixel grid.
[0093] As discussed herein, the rows and columns of the two-dimensional array allow for parallelism in performing the sequential analysis of the residual values. In one embodiment of processing block 225, deepfake detection method 200 has parallel processors perform sequential probability ratio tests simultaneously across the rows of the two-dimensional array and the columns of the two-dimensional array. An anomaly is detected when any one of the sequential probability ratio tests across the rows or columns identifies an anomalous residual. In one embodiment, the sequential probability ratio tests are performed simultaneously across the rows of the two-dimensional array and the columns of the two-dimensional array by parallel processors. The parallel processors may be, for example, multiple CPUs or GPUs.
[0094] As noted above in the detailed discussion of converting audiovisual content into time-series signals in processing block 210, the sampling intervals across the set of signals need not necessarily be uniform or synchronized. Thus, in one embodiment of processing block 210, deepfake detection method 200 resamples one or more of the time-series signals so that the set of time-series signals is sampled at a uniform rate (as discussed below under the heading "Resampling the Time Series at a Uniform Rate"). Deepfake detection method 200 also phase-shifts the time-series signals so that observations of the signals are synchronized (as discussed below under the heading "Synchronizing the Time Series Signals").
[0095] In one embodiment, prior to generating the residual time series signals as discussed in processing block 215, a machine learning model is trained to generate estimates consistent with authentic speech by a speaking person in the audiovisual content. In one embodiment, the audiovisual content itself may be used as a reference for authentic speech by designating as the reference a portion of speech, such as an introduction or greeting, that is highly likely to be authentic and not modified by the deepfake content. In one embodiment, while streaming the audiovisual signals, the deepfake detection method 200 designates a reference segment of the set of time series signals to represent authentic speech by a human speaker. For example, the reference dataset 172 may be a segment of the time series signal 135 derived from the audiovisual content 130. Prior to generating the set of residual time series signals, the deepfake detection method 200 trains a machine learning model to generate estimates of the time series signals consistent with authentic speech based on the reference segment of the set of time series signals.
[0096] In one embodiment, other audiovisual content of other speech by the same human speaker may be used as a reference set of time series signals for authentic speech. In one embodiment, the other audiovisual content is verified as authentic audio and video of speech by the human speaker. In one embodiment, deepfake detection method 200 obtains a reference set of time series signals that represent authentic speech by the human speaker in the reference audiovisual content. For example, reference dataset 172 may be time series signals derived from a reference audio-video recording of a speaker speaking in a different scene than that shown in audiovisual content 130. Prior to generating the residual time series signals, deepfake detection method 200 trains a machine learning model to generate machine learning estimates of authentic speech based on the reference set of time series signals. Additional details regarding training a machine learning model are discussed below under the heading "ML Generation of Residuals."
[0097] In one embodiment of processing block 230, deepfake detection method 200 includes in the alert one of the pixel location where the anomaly occurred or the frequency range where the anomaly occurred. Deepfake detection method 200 also includes in the alert a timestamp when the anomaly occurred.
[0098] In another embodiment of processing block 230, deepfake detection method 200 includes in the alert the location of the pixel where the anomaly occurred and the time the anomaly occurred, and using that information, deepfake detection method 200 highlights the pixel where the anomaly occurred in the audiovisual content to visually indicate the deepfake content.
[0099] In another embodiment of processing block 230, deepfake detection method 200 includes in the alert an identifier for the frequency range in which the anomaly occurred and the frame of the audiovisual content in which the anomaly occurred, and using that information, deepfake detection method 200 adds a warning to the audiovisual content at the frame in which the anomaly occurred to visually indicate the deepfake content.
[0100] In one embodiment, the deepfake detection systems and methods described herein exhibit artificial intelligence or machine learning that can autonomously analyze streaming audiovisual content of a speaking person and detect the possible insertion of fake audible words and / or altered mouth movements accompanying the words. In one embodiment, such analysis and detection can be performed in real time while the audiovisual content is being streamed live.
[0101] In one embodiment, the deepfake detection system and method described herein employ a novel integrated frequency-domain-time-domain framework. This framework systematically decomposes video signals into high-resolution spatiotemporal (i.e., both spatially within the video frame and temporally within the video signal) time-series signals and audio signals into fine-grained frequency bins to create the time-series signals. The framework then synchronizes and fuses distinct clusters of the time-series signals together with a method to ensure synchronized, uniform sampling intervals. The framework then consumes the time-series signals through nonlinear, nonparametric pattern recognition to detect the insertion of deepfake segments.
[0102] -Conversion to time series signals- In one embodiment, the deepfake detection systems and methods described herein begin with separate techniques for converting streaming megapixel video of a speaking person and a simultaneously streaming audio signal into two large clusters of time-series signals, producing a video cluster of time-series signals for the video signal and an audio cluster for the audio signal.
[0103] For video signals, a space-time transformation is performed on a fine pixel granularity video signal by turning each pixel into a 3-tuple time series. The video signal is represented by an array (e.g., a megapixel array) of individual pixels that make up the video signal. The pixel array includes a pixel for each location within a frame of the video signal. Each individual pixel in the pixel array is transformed into three time series signals VTSS by tracking a red / green / blue (RGB) 3-tuple metric over time. This creates three video time series signals VTSS per pixel, and the number of pixels in the array N Pixels Video N, which is three times VTSS (N VTSS =3×N Pixels ) results in a number of time series signals representing
[0104] For an audio signal, the transformation from the frequency domain to the time domain is performed by transforming the continuous frequency waveform of the audio signal into a set of time series for the discrete frequency ranges of the audio signal. The frequency spectrum of an audio signal is a number N Bins The frequency may be subdivided into unbroken ranges, or frequency "bins", of N. For example, Binsmay be the square of the number of pixels in the smaller dimension of the video frame (as discussed below). Or, for example, the number of bins may be another positive integer multiple of the number of pixels in the smaller dimension of the video frame. Or, in another example, the number of bins may be a positive integer multiple of the number of pixels in the larger dimension of the video frame. The amplitude values for each of the bins may be sampled at intervals to generate a set of audio time-series signals ATSS representing the audio signal. This is where the size of the number of bins, audio N, is taken as the number of bins. ATSS (N ATSS =N Bins ) to create a number of time series signals representing
[0105] -Resampling a time series at a uniform rate- The two databases of video and audio time-series signals resulting from the above transformation may have completely different sampling rates and are therefore asynchronous. ML pattern recognition cannot analyze time-series signals with different sampling intervals. Therefore, the two databases of video and audio time-series signals are converted into synchronized, uniformly sampled databases of time-series signals through an analytical resampling process. First, the two databases of video and audio time-series signals are resampled to generate uniformly sampled video and audio time-series signals. Then, the uniformly sampled video and audio time-series signals are phase-shifted to synchronize the signals.
[0106] To resample the two databases of video and audio time-series signals, a common sampling interval is selected for the synchronized, uniformly sampled databases. (As used herein, "common sampling interval" refers to a sampling interval commonly used for both the video and audio time-series signals.) The video and audio time-series signals are resampled at the common sampling interval using an interpolation algorithm. The interpolation algorithm calculates values at the common sampling interval that fall between existing sample points of the video and audio time-series signals. The interpolation algorithm may include linear interpolation or higher-order interpolation. The existing sample rate of the signals may be upsampled or downsampled, as appropriate, by the interpolation algorithm. In one embodiment, rather than downsampling the higher sampling rate time series because downsampling involves some information loss, the existing sampling rate of the time-series signal having a slower sampling rate is upsampled to match the higher sampling rate time series. In one embodiment, the common sampling interval is an integer multiple of the frame interval of the video signal, reducing the need for interpolation calculations for the video time-series signal. The resampling results in uniformly sampled video and audio time-series signals from the two databases of video and audio time-series signals. In this manner, one or more of the time series signals in the two databases are resampled so that the set of time series signals in the two databases is sampled at a uniform rate across both the time series signals for the pixels of the video signal and the time series signals for the frequency range of the audio signal.
[0107] -Synchronizing time series signals- To synchronize uniformly sampled video and audio time-series signals, the signals are phase shifted to maximize the correlation between the signals. The uniformly sampled video and audio time-series signals are synchronized with each other using synchronization techniques such as correlogram, cross power spectral density, or genetic algorithm techniques.
[0108] In the correlogram technique, one of the uniformly sampled video and audio time-series signals is selected as a reference signal. All of the other time-series signals are then aligned to this reference signal by calculating pairwise cross-correlation coefficients and adjusting the lags of the individual signals to optimize the cross-correlation coefficients relative to the reference signal.
[0109] In the cross-power spectral density technique, each pair of uniformly sampled video and audio time-series signals is analyzed using a fast Fourier transform to infer the phase angle between the two signals. An estimate of the lag time is then calculated from the phase angle, and the signals are adjusted to make the lag time zero.
[0110] In the genetic algorithm technique, uniformly sampled video and audio time-series signals are randomly adjusted over several iterations. In each iteration, the time-series signals are adjusted in a positive or negative direction, and the overall synchronization score for the time-series signals is evaluated. Adjustments that improve the overall synchronization score are retained for subsequent iterations, while adjustments that do not improve the overall synchronization score are discarded. The iterations are repeated until a threshold indicating a satisfactory overall synchronization score is met. The adjustments are gradually reduced in size from iteration to iteration to prevent oscillations that exceed a satisfactory overall synchronization score.
[0111] Once the synchronization technique is complete, the uniformly sampled video and audio time series signals are synchronized and ready for analysis using ML pattern recognition. The resulting synchronized, uniformly sampled database of time series signals TSS can be used to analyze several time series signals N TSS In one embodiment, the synchronized, uniformly sampled database of time series signals TSS includes signals resampled from signals representing pixels VTSS and audio ATSS, and the number N of time series signals TSS is the number of time series signals representing some audio, N ATSS , and the number of time series signals representing pixels N VTSS Contains (N TSS =N ATSS +N VTSS ).
[0112] -ML generation of residuals- The deepfake detection systems and methods described herein analyze a synchronized, uniformly sampled database of time-series signals TSS using an ML model. In one embodiment, the ML model is an ML pattern recognition model. In one embodiment, the ML model is implemented as one or more nonlinear nonparametric (NLNP) regression algorithms used for multivariate anomaly detection, e.g., similarity-based modeling (SBM), such as multivariate state estimation technology (MSET) (including Oracle's proprietary multivariate state estimation technology (MSET2)). Thus, in one embodiment, the ML model is an NLNP model, such as an MSET model.
[0113] In one embodiment, the deepfake detection system and method inputs synchronously sampled regions of the transformed dynamic time-series signals into an ML model (e.g., in real time), analyzes the input time-series signals with NLNP pattern recognition, and creates a database of synchronized digitized residuals. The ML model estimates or predicts what each of the database of derived and transformed time-series A / V signals should or is expected to be based on training segments of video of a person delivering broadcasted speech.
[0114] Training of the ML model can be performed as an initial process prior to analysis of the audio / visual broadcast based on a library of past videos of the video subject. Training can also be performed in real time to detect deepfake insertions into real-time audio / visual broadcasts for subjects for which no library of past videos exists.
[0115] An ML model is a multivariate model that accepts input values for multiple variables and generates an estimate for the value of one variable based on the input values for the other variables. In one embodiment, the ML model includes a variable for each time-series signal in the set of time-series signals discussed above with respect to processing block 210. In other words, in one embodiment, the ML model includes a variable for each color channel of each pixel in a frame of video content and a variable for each audio frequency range (bin) in the audio content. For example, if a frame of video content is 426 x 240 three-channel pixels, the number of variables is 57,600 (as discussed below with reference to Table 1). Thus, the ML model is configured to predict the value of variable 1 based on one or more of variables 2 through 57,600, and similarly for all variables. Training on the variables can be completed in parallel on multiple processing devices.
[0116] The ML model is trained to generate estimates of what the values of variables should be based on training with a reference set of time series signals from authentic audiovisual content of a human speaker speaking. To train the ML model, a reference set of time series signals for each variable is provided to the ML model. During training, a series of sets of reference values for the variables, each set including one reference value from each of the reference time series signals in the set, is also provided to the ML model. The configuration of the correlation patterns between the variables of the ML model is automatically adjusted based on the reference values, causing the ML model to generate accurate estimates for each variable based on the inputs to the other variables. Because the reference set of time series signals is taken from audiovisual content of a human speaker truly speaking without deepfake modifications, the ML model is therefore configured to generate estimates that are consistent with authentic speech by the human speaker. The ML model therefore has learned correlation patterns between variables that indicate when the speech by the human speaker is authentic and free of deepfake insertion of words or alterations to the speaker's mouth, face, or other movements.
[0117] In one embodiment, the ML model may be trained from a previous video of a person. In this case, the reference set of time-series signals is from a reference audiovisual recording of a speaking human speaker that has been verified to be authentic speech and movements by the human speaker, free of both inserted words and adjustments to the speaker's mouth, face, or other movements. The reference set of time-series signals is transformed from the reference audiovisual recording, for example, in the same manner as described above with respect to processing block 210. Such a reference set of time-series signals may be provided as inputs to the variables of the ML model during training so that the ML model is automatically configured to generate estimates that are consistent with the authentic speech and movements exhibited by the speaker in the reference audiovisual recording.
[0118] Alternatively, in one embodiment, the ML model can be trained from the beginning or early portion of a live transmission of audiovisual content. In this case, training is gradually built according to a bootstrap technique. Early portions of speech often contain preambles that are unlikely to be targets of fraudulent deepfake modulation and provide examples of speech by human speakers. Therefore, such speech by a speaker near the beginning of the audiovisual content can be assumed to be authentic speech and movements by a human speaker, without any inserted words or adjustments to the speaker's movements. A time range of audiovisual content starting at or shortly after the person begins speaking can be designated as the training portion of the audiovisual content. For example, 2500 video frames of a human speaker speaking may cover only 104 seconds of speech (at a video frame rate of 24 frames per second), but may be sufficient to train an ML model. When the training portion of the audiovisual content arrives, the training portion of the audiovisual content is converted into reference values for the variables to generate a set of streaming reference time-series signals for the variables (e.g., as described in processing block 210). The ML model is trained with the streaming reference time series signal until the end of a training range of audiovisual content. At the completion of the training range, the ML model is configured to generate estimates that are consistent with authentic speech and movements exhibited by the speaker at the beginning of the speech. The ML model then transitions from training to monitoring or real-time inference of deepfake presence. The ML model proceeds to generate estimates of authentic content for a monitoring range of audiovisual content following the training range.
[0119] After training and during real-time inference, each signal in the comprehensive database of transformed, fused, and resampled time series signals is compared to the other N TSS The estimates are estimated using the learned correlation patterns with all N signals. The estimates produced by the ML model for the variables are used to generate residual values for the variables. TSSOnce generated for all N signals, the estimates are subtracted from the actual transformed signals, yielding point-by-point differences, referred to herein as "residuals." The residual values indicate how significantly the signal values input to the ML model differ from values consistent with authentic speech and movements by a human speaker. For a synchronized, uniformly sampled database of time-series signals TSS, the successive element-by-element pairwise differences between predicted (estimated) and real-time (observed) values are TSS Create a database of "residual signals".
[0120] - Residual array structure for deepfake detection - The residuals at any one time point or time observation may be stored in a rectangular grid. TSS The residual values for one observation of the signal may be placed into a two-dimensional (2D) rectangular array (or matrix or grid) for processing. At this step in the analysis, the 2D array of residuals is "static," i.e., it represents a specific point in time or a single observation. For example, a static 2D array of residuals may correspond to a frame of a video signal. Residual processing of the present invention is performed for each fixed point in time as each frame of video is captured. Thus, for this residual processing step, there is a static 2D array of residual values. The analysis for the static 2D array representing the frame may be conceptually represented as an inner loop in one embodiment of the deepfake detection method. The temporal phase of the deepfake detection method may be conceptually represented in an outer loop that is triggered by each new incoming video frame (accompanied by audio content that is then synchronized to the frame processing rate of the video stream as described above). The temporality or aspect of the deepfake detection method that relates to changes over time occurs in the outer loop recursive processing. However, at this point in the residual processing, N TSS The 2D array of signal values is static in time and represents the residual for one concurrent observation of all signals in a synchronized, uniformly sampled database of time series signals TSS.
[0121] 3, an example 2D static array 300 for one observation associated with autonomous deepfake detection is shown. The static array 300, in one embodiment, is an inner-loop "static array" residual analysis that analyzes N TSS A specific XY orientation of the residuals is used for each signal. The static array represents one synchronous observation for all the various video and audio signals. The layout of the fused and transformed visual and audio signals in the static array is such that the transformed video time series retains the original spatial layout of pixels. In one embodiment, the layout of the static array is a matrix array with two partitions: a 3M×N video partition 305 and an N×N audio partition 310. Thus, the static array has dimensions of (3M+N)×N.
[0122] The video partition 305 of the static array 300 preserves the spatial layout of the video frames of the video signal. Recall that a frame of video in the video signal has a particular rectangular grid of M×N pixels. When a spatiotemporal transformation of these M×N pixels is performed (creating 3×M×N time series), each pixel in the video signal in the original M×N rectangular layout is represented by three time series signals that quantitatively represent red-green-blue metrics. Within the transformed spatiotemporal video signal, the exact spatial grid is preserved as the original M×N video frame layout using the 3M×N cells of the video partition 305.
[0123] Each pixel is represented by a sequence of three cells—a red cell for the pixel, a green cell for the pixel, and a blue cell for the pixel—at a grid location corresponding to the pixel location. For example, pixel (1,1) in a video signal is represented by cell C for the red value of pixel (1,1). 1,1 315, cell C for the green value of pixel (1,1) 2,1 320, and cell C for the blue value of pixel (1,1) 3,1325 in the video partition 305, pixel (M,N) of the video signal is represented by C for the red value of pixel (M,N). 3M-2,N 330, cell C for the green value of pixel (M,N) 3M-1,N 335, and cell C for the blue value of pixel (M,N). 3M,N 340. Other pixels of the video signal are also represented by three RGB signal values in video partition 305 at locations similarly corresponding to the locations of the original pixels. In general, the pixel of the video frame that corresponds to a cell in video partition 305 is the pixel in the video frame that has the same N coordinate value in the video frame as the cell in video partition 305, and that has an M coordinate value in the video frame that is the integer quotient of the M coordinate value for the cell in video partition 305 divided by three.
[0124] Retaining the N×M video frame layout enables root cause explainability of deepfake alerts. When a deepfake alert occurs, the pixel location or region within the video frame that caused the alert can be indicated. If, instead of retaining the N×M video frame layout, the deepfake detection system were to simply place all video residuals within a one-dimensional (1D) array or any convenient, unstructured blob of residual values, the deepfake alert would still be triggered, but explainability of the deepfake alert would be difficult or cumbersome. Thus, when an X-Y 2D static array 300 (or matrix or grid) of residuals is formed, the video partition 305 of such a rectangular grid of residuals preserves the rectangular structure of the original pixel grid.
[0125] Static array 300 is an array diagram of synchronous observations of a synchronized, uniformly sampled time-series database 345. Time-series database 345 is comprised of individual component time-series signals, such as video time-series signal (1,1) 350 and audio time-series signal (N,3M+N) 355. As discussed above, static array 300 represents synchronous observations 360 or "slices" across the time-series signals for all the various video and audio signals. Thus, static array 300 is comprised of values at commonly time-stamped synchronous observations in each of the component video and audio time-series signals in time-series database 345.
[0126] The audio partition 310 of the static array 300 stores audio values in a 2D rectangular (in this case, square) N×N matrix structure. While the audio values could be processed in a linear array, creating an N×N matrix greatly facilitates analysis of the residual, as explained below. The transformed audio signal (frequency domain to time domain transform discussed above) is stored in a square N×N array, where N is the number of pixels on the smaller side of the original N×M rectangular grid of pixels. In the frequency domain to time domain transform discussed above, the frequency axis is divided into several bins for analysis.
[0127] Each bin is represented by one cell. The bins may be ordered within the cells of the audio partition 310 from left to right along the X-axis of the static array 300, wrapping from top to bottom along the Y-axis of the static array 300, or following another order or arrangement as is convenient. For example, bin 1 of the audio spectrum is represented by cell C in the audio partition 310. 1,3M+1 By 365, Bin 2 is Cell C 1,3M+2 370, which is cell C N,3M+N Continues up to bin NxN at 375.
[0128] Although there is no spatial relationship to the cell locations for the transformed audio time series, there are computational advantages by creating a "square" NxN grid of residuals for the transformed audio time series and partitioning this square grid adjacent to the previously defined NxM rectangular array (with the shared partition boundary being shorter than the shorter "N" dimension). For the deepfake detection systems and methods described herein, a significant overall computational boost (i.e., reduction in processor time) can be achieved by increasing the number of bins, N Bins N×N or N 2 where N is the smaller of the NxM pixel layouts for the video. The computational boost is also possible for the number of bins, N, which is another positive integer multiple of N. Bins A number of bins smaller than NxN may exhibit lower sensitivity to deepfakes than NxN.
[0129] At this point in the process, a rectangular array of residuals has been created, with one side of the rectangle having N values and the longer side of the rectangle having (3M+N) values, for example, as shown and described for static array 300. Thus, static array 300 is a rectangular array with multiple cells along each dimension. The "left side" of this residual array is a video partition that holds the transformed video spatiotemporal values with an N x 3M layout (preserving the original N x M pixel spatial layout). The "right side" of this residual array is an audio partition that holds the frequency-time transformed audio signals in a square N x N array. As discussed above, a spatial layout for audio signals is not necessarily required for deepfake detection, but constructing the residual array in this manner facilitates analysis of the residuals to detect deepfakes in the next step.
[0130] Therefore, the dimensions of the static array 300 and the number of bins N Bins is determined by the frame size of the video signal. Exemplary dimensions of static array 300 and the number of audio frequency bins N Bins are given in Table 1.
[0131] [Table 1]
[0132] The array dimensions and number of bins can therefore become quite large. -Static array residual analysis- Because the values in static array 300 are the residuals between the estimated signal values and the actual signal values at a given observation, the values in static array 300 represent a three-dimensional residual surface for that observation. The residual surface can be examined with a two-dimensional sequential probability ratio test (2D SPRT) (or other sequential analysis in two dimensions) to determine whether the residual surface indicates the presence of deepfake content. As discussed above, an observation corresponds to a frame of a video signal. Thus, the residual surface represents how far the actual video and audio time-series signal values are from the predicted video and audio time-series signal values for one frame. Thus, the residuals are analyzed a single frame (or static array) at a time with a 2D SPRT.
[0133] In general, a sequential probability ratio test (SPRT) is a form of sequential analysis for analyzing samples of an unfixed size until it finds a result that satisfies a predetermined threshold for a significant result. The sequential probability ratio test (SPRT) detects abnormal deviations from normal residuals by calculating a cumulative sum of log-likelihood ratios for each successive residual between the actual signal value and the estimated signal value and comparing this cumulative sum against a threshold to determine that an anomaly has been detected. If the cumulative sum of log-likelihood ratios exceeds the threshold, an alarm is issued.
[0134] To detect deepfake content in static array 300, the deepfake detection system performs SPRT across the residual values in both the x and y dimensions of static array 300. In the x dimension, SPRT is applied to the residual values in each row of the static array. For example, if row 380 in the x dimension, y=1, SPRT is applied to the residual values in the y-dimension cell C 1,1 ,C 1,2 ,C1,3 ,...C 1,3M+N This process is applied to the sequence of residual values in static array 300. This process is repeated for rows 1 through N (y=1, 2, ..., N) in static array 300. The process may be repeated (serial or parallel) as a separate SPRT test per row. Alternatively, the process may be performed as a single SPRT test for values that wrap from row to row by incrementing the value of y and resetting the value of x when the final value of x for that row (i.e., 3M+N) is reached.
[0135] As with the y-dimension, SPRT is applied to the residual values in each column of the static array. For example, if column 385x=1 in the y-dimension, SPRT is applied to the residual values in x-dimension cell C. 1,1 ,C 2,1 ,...C N,1 The process is applied to the sequence of residual values in x. This process is repeated for columns 1 through 3M+N (x=1, 2, ..., 3M+N). The process may be repeated (serially or in parallel) as a separate SPRT test per column. Alternatively, the process may be performed as a single SPRT test for values that wrap from column to column by incrementing the value of x and resetting the value of y when the final value of y for that row (i.e., N) is reached. Thus, all values in the residual plane are evaluated by two separate SPRTs: an x-dimension SPRT and a y-dimension SPRT. If the SPRT in either dimension raises an alarm, deepfake activity has been detected.
[0136] 2D SPRT is not a function of time. As discussed above, the array of residuals is static in time and represents the results of the analysis for only one observation or frame. Instead of SPRT on the time dimension, N simultaneous SPRT tests are performed across the horizontal array (x-dimension) of values within the static array of residuals (as discussed above). Additionally, 3M+N SPRT tests are performed simultaneously across the y-dimension. Because SPRT tests are performed on residual values for one observation or frame, the SPRT tests can be performed simultaneously, i.e., concurrently in time, in parallel.
[0137] Computational performance using SPRT is good, SPRT is computationally lightweight, and SPRT computation is highly parallel. Thus, N and 3M+N can scale to very large sizes (e.g., for megapixel high-resolution video quality) on parallel CPU or GPU processors.
[0138] FIG. 4A shows a 3D plot 400 of an exemplary video / audio plane 405 of video and audio signal values in one observation (or frame) of a synchronized, uniformly sampled database or set of time-series signals (such as the database TSS discussed above). The video / audio plane 405 shows the spatiotemporally transformed video and frequency-time transformed audio time series in an array frozen at one time step for the purpose of illustrating 2D SPRT analysis. To form the video / audio plane 405, the signal values for the time-series signals are arranged in a 2D rectangular array, as discussed above with reference to FIG. 3. Thus, the video / audio plane 405 is 3M+N cells or positions wide in the x-dimension 410 and N cells or positions deep in the y-dimension 415. The signal value at each cell or position is plotted against a value axis 420. The video / audio plane 405 has a video partition 425 that extends 3M cells or positions wide along the x-dimension 410. The video / audio plane 405 includes an audio partition 430 that extends N cells or positions wide along the x dimension 410 .
[0139] FIG. 4B is a 3D plot 435 of an exemplary residual surface 440 of audiovisual content consistent with authentic speech by a human speaker. The residual surface 440 is an illustration of what the residual array might look like for authentic, real-time streaming audio / visual of a person on whom an ML pattern recognition model is being trained. While the person speaking in the video / audio is still a real person speaking their authentic voice (i.e., not an impersonation or a synthesized voice), the residual surface of each observation, as shown by the residual surface 440, is flat, with some small random noise on it. The residual surface 440 is composed of residual values between the actual and ML-predicted values for each video and audio signal in one observation (or frame) of a synchronized, uniformly sampled database of time-series signals. In the residual surface 440, the residuals are arranged in a 2D rectangular array, as discussed above with reference to FIG. 3. Residual surface 440 is 3M+N cells or positions wide in x dimension 410 and N cells or positions deep in y dimension 415. The residual value at each cell or position is plotted against residual axis 445. Residual surface 440 has a video partition 450 that extends 3M cells or positions wide along x dimension 410, and an audio partition 455 that extends N cells or positions wide along x dimension 410.
[0140] 4C is a 3D plot 460 of an example residual surface 465 of audiovisual content containing anomalies indicative of deepfake modification. Similar to residual surface 440, the residuals of residual surface 465 are arranged in a 2D rectangular array that is 3M+N cells or positions wide in the x-dimension 410 and N cells or positions deep in the y-dimension 415, with the residual value at each cell plotted against residual axis 445. Residual surface 465 has a video partition 470 that extends 3M cells or positions wide along the x-dimension 410, and an audio partition 455 that extends an additional N cells or positions wide along the x-dimension 410.
[0141] Residual plane 465 is an example of what a residual array might look like in real-time streaming audio / visuals of a person containing deepfake content. When video and / or audio insertions are present to create a deepfake, one or more "humps" or spikes appear in the residual plane. In a hump or spike, the residual values are significantly larger than a background plane with small random noise. For example, a first hump 480 occurs in video partition 470, indicating the presence of deepfake video content at the pixels represented by the cells included in hump 480. And, for example, a second hump 485 occurs in audio partition 475, indicating the presence of deepfake audio content in the frequency bin represented by the cells included in hump 485.
[0142] A 2D SPRT interrogation of the residual values within the rows and columns of residual surface 465 generates an alert for residual values at either hump 480 or hump 485. While thresholds above and below the residual surface can be established to alert whenever a hump, bump, spike, or other pattern in the residual exceeds the high / low threshold, the use of dual-dimensional concurrent SPRT has a significant advantage, as it has been proven to achieve the lowest false alarm and false alarm probabilities mathematically possible. Low false alarm and false alarm probabilities enhance the functional ability to make real vs. deepfake determinations.
[0143] -An example deepfake detection process- 5 illustrates an additional exemplary method for deepfake detection 500. In one embodiment, the method for deepfake detection 500 begins at start block 505 in response to a computer processor determining that method 500 should begin, for example, in response to the occurrence of a condition discussed above for initiating method 200. In one embodiment, the computer is configured with computer-executable instructions to perform the functions of deepfake detection system 100.
[0144] At processing block 510, method 500 converts pixels of the video portion of the audio-video signal into a first set of time-series signals, e.g., as shown and described above with respect to processing block 210 and under the heading "Conversion to Time-Series Signals." At processing block 515, method 500 converts frequency ranges of the audio portion of the audio-visual signal into a second set of time-series signals, also as shown and described above with respect to processing block 210 and under the heading "Conversion to Time-Series Signals."
[0145] At process block 520, method 500 adjusts one or more time-series signals from the first and second sets of time-series signals so that the first and second sets of time-series signals are sampled at a uniform rate and synchronously. This adjustment may include, for example, resampling the time-series signals as shown and described above under the heading "Resampling a Time Series at a Uniform Rate." This adjustment may also include synchronizing the resampled signals as shown and described above under the heading "Synchronizing the Time-Series Signals."
[0146] In processing block 525, method 500 generates residual time series signals from the time series signals belonging to the video and audio sets and machine learning estimation of the time series signals, e.g., as shown and described above with respect to processing block 215 and under the headings "ML Generation of Residuals" and "ML Generation of Residuals."
[0147] At processing block 530, method 500 places one synchronized observation of residual values from the residual time series signal into a rectangular array, e.g., as shown and described above with respect to processing block 220 and under the heading "Residual Array Structure for Deepfake Detection." At processing block 535, method 500 performs a two-dimensional sequential probability ratio test on the values in the rectangular array, e.g., as shown and described above with respect to processing block 225 and under the heading "Static Array Residual Analysis." At processing block 540, in response to an indication of the presence of an anomaly from the two-dimensional sequential probability ratio test, method 500 generates an alert that deepfake content has been detected in the audio-video signal, e.g., as shown and described above with respect to processing block 230. Processing blocks 530-540 may repeat in a loop over the sequence of synchronized observations of the residual time series signal while subsequent observations of the residual time series signal remain. When there are no more observations of residual values available from the residual time-series signal, the method 500 continues to end block 545 and completes.
[0148] -Selected Benefits- The deepfake detection systems and methods described herein offer several advantages. Prototype analysis of the present deepfake detection systems and methods shows clear advantages over neural networks (NNs) that attempt to detect deepfake audio / video insertions.
[0149] In one embodiment, the deepfake detection system and method detects the insertion of deepfake segments with much higher fidelity and much lower false positive and false negative rates than can be achieved by neural network-based approaches to deepfake detection. Also, in one embodiment, the deepfake detection system and method far surpasses human experts viewing an audio / video presentation. For example, the deepfake detection system and method detects inserted deepfake content, or the video frame in which it appears, at first observation, and the detection is not subjective.
[0150] Applying neural network-based solutions to megapixel or other high-resolution fine-grained video / audio analysis is not computationally feasible. Neural networks require large computer clusters to scale even to small pixel grids. This is especially true at the kHz sampling intervals, frame rates required to keep up with audio speech recognition. Neural network-based tools for multivariate anomaly detection are often limited or capped at the number of input time series signals, e.g., no more than 300. Neural networks in general, and very short-term memory (LSTM) neural networks in particular, "clog up" and cannot process more than a few hundred time series signals, making them unsuitable for analyzing anything other than very low-resolution video. In contrast, the deepfake detection systems and methods described herein scale to millions of time series signals.
[0151] Thus, there is a scalability gap of orders of magnitude between the deepfake detection systems and methods described herein and other ML approaches. Neural networks (NNs) and support vector machines (SVMs) have very limited scalability because these solutions use stochastic optimization of weights and therefore cannot be finely parallelized on multi-threaded, multi-core CPUs or GPUs. NN and SVM ML analysis of megapixel-correlated RGP pixel content requires inter-process communication within ML and leaves no room for finely multivariate parallelization. In contrast, the deepfake detection systems and methods described herein use deterministic mathematical algorithms (MSET or other SBMs) that naturally parallelize on modern multi-threaded, multi-core CPUs and GPUs. Simply put, while NN and SVM approaches to identifying deepfake content are not scalable to the number of pixels in even the lowest resolution video, in one embodiment, the deepfake detection systems and methods described herein are scalable to very high resolutions, e.g., up to 8K video and above.
[0152] In one embodiment, the novel and inventive transformations described herein for spatially and temporally transforming high-resolution pixel granularity and frequency-to-time domain transforming high-fidelity acoustic audio patterns into distinct clusters of time-series signals, followed by an analytical resampling process to fuse the clusters into a synchronized, uniformly sampled time-series database, enable frame-by-frame 2D SPRT analysis to detect deepfake content by its location within video frames and audio frequency bins. Thus, unlike any other ML techniques, in one embodiment, the deepfake detection system and method described herein can simultaneously perform a frequency-to-time domain transform of the acoustic signal, a spatiotemporal transform of the high-density video content into pixel-level RGB time series, and generate a master array of transformed time series that can then be analyzed by the novel 2D SPRT technique in a final, continuous real-vs.-fake determination process with ultra-low false positives and false negatives.
[0153] Due to the easily parallelizable nature of the deepfake detection systems and methods described herein, deepfake content is detected in streaming audio / video in real time and is scalable to high video resolutions.
[0154] -Cloud or Enterprise Implementation- In one embodiment, the system (such as deepfake detection system 100) is a computing / data processing system including a computing application or a collection of distributed computing applications for access and use by other client computing devices communicating with the system over a network. In one embodiment, deepfake detection system 100 is a component of a time-series data service configured to collect, provide, and perform operations on time-series data. The applications and computing system may be configured to operate with or be implemented as a cloud-based network computing system, an Infrastructure-As-A-Service (IAAS), Platform-As-A-Service (PAAS), or Software-As-A-Service (SAAS) architecture, or other type of networked computing solution. In one embodiment, the system provides at least one or more of the functions disclosed herein and a graphical user interface for accessing and operating the functions. In one embodiment, deepfake detection system 100 is a centralized server-side application that provides at least the functionality disclosed herein and is accessed by many users using computing devices / terminals that communicate over a computer network with a computer (which functions as one or more servers) of deepfake detection system 100. In one embodiment, deepfake detection system 100 may be implemented by a server or other computing device configured with hardware and software to implement the functions and features described herein.
[0155] In one embodiment, the components of deepfake detection system 100 may be implemented as a set of one or more software modules executed by one or more computing devices specially configured for such execution. In one embodiment, the components of deepfake detection system 100 are implemented on one or more hardware computing devices or hosts interconnected by a data network. For example, the components of deepfake detection system 100 may be executed by one or more computational hardware forms, such as central processing units (CPUs), or networked computing devices in general-purpose forms, dense input / output (I / O) forms, graphics processing units (GPUs), and high-performance computing (HPC) forms.
[0156] In one embodiment, components of deepfake detection system 100 communicate with each other through electronic messages or signals. These electronic messages or signals may be configured as calls to functions or procedures, such as application programming interface (API) calls, that access features or data of the components. In one embodiment, these electronic messages or signals are transmitted between hosts in a format compatible with Transmission Control Protocol / Internet Protocol (TCP / IP) or other computer networking protocols. Components of deepfake detection system 100 may (i) generate or create electronic messages or signals to issue commands or requests to another component, (ii) transmit the messages or signals to other components of deepfake detection system 100, (iii) parse the content of received electronic messages or signals to identify commands or requests that the component can implement, and (iv) automatically implement or execute the commands or requests in response to identifying the commands or requests. Electronic messages or signals may include queries against a database. The queries may be created and executed in a query language compatible with the database and may be executed in a runtime environment compatible with the query language.
[0157] In one embodiment, a remote computing system may access information or applications provided by deepfake detection system 100, for example, through a web interface server. In one embodiment, the remote computing system may send requests to and receive responses from deepfake detection system 100. In one example, access to information or applications may be enabled through the use of a web browser on a personal computer or mobile device. In one example, communications exchanged with deepfake detection system 100 may take the form of, for example, remote Representational State Transfer (REST) requests using JavaScript Object Notation (JSON) as the data exchange format, or Simple Object Access Protocol (SOAP) requests to and from an XML server. The REST or SOAP requests may include API calls to components of deepfake detection system 100.
[0158] -Computing Device Embodiment- 6 illustrates an example computing system 600 configured and / or programmed as a special-purpose computing device with one or more of the example systems and methods described herein and / or equivalents. The example computing device may be a computer 605 including one or more processors 610, memory 615, and input / output ports 620 operably connected by a bus 625. In one example, the computer 605 may include deepfake detection logic 630 configured to facilitate autonomous deepfake detection based on multivariate spatiotemporal characterization and analysis of video and integrated audio, similar to the logic, systems, and methods shown and described with reference to FIGS. 1-5. In different examples, the logic 630 may be implemented in hardware, a non-transitory computer-readable medium having instructions 637 stored thereon, firmware, and / or a combination thereof. While logic 630 is illustrated as a hardware component attached to bus 625, it should be understood that in other embodiments, logic 630 may be implemented in processor 610, stored in memory 615, or stored on disk 635. In one embodiment, multiple processors 610 and / or multiple logics 630 may operate in parallel to perform tasks simultaneously (such as parallel execution of SPRT on sequences of values in individual rows and columns of a two-dimensional array, as described above with respect to processing block 225).
[0159] In one embodiment, the logic 630 or computer is a means (e.g., constructs such as hardware, non-transitory computer-readable media, firmware, etc.) for performing the described acts. In some embodiments, the computing device may be a server operating within a cloud computing system, a server configured within a Software as a Service (SaaS) architecture, a smartphone, a laptop, a tablet computing device, etc.
[0160] The means may be implemented, for example, as an ASIC that is programmed to autonomously detect deepfake modifications to audiovisual content. The means may also be implemented as stored computer-executable instructions that are temporarily stored in memory 615 and presented to computer 605 as data 640 that are then executed by processor 610.
[0161] Logic 630 may also provide means (e.g., hardware, non-transitory computer-readable media storing executable instructions, firmware) for autonomously detecting deepfake modifications to audiovisual content.
[0162] To generally describe an exemplary configuration of computer 605, processor 610 may be a wide variety of different processors, including dual microprocessors and other multi-processor architectures. Memory 615 may include volatile memory and / or non-volatile memory. Non-volatile memory may include, for example, ROM, PROM, etc. Volatile memory may include, for example, RAM, SRAM, DRAM, etc.
[0163] Storage disk 635 may be operatively connected to computer 905, for example, through an input / output (I / O) interface (e.g., card, device) 645 and input / output port 620 controlled by at least an input / output (I / O) controller 647. Disk 635 may be, for example, a magnetic disk drive, solid-state disk drive, floppy disk drive, tape drive, Zip drive, flash memory card, memory stick, etc. Furthermore, disk 635 may be a CD-ROM drive, CD-R drive, CD-RW drive, DVD ROM, etc. Memory 615 may store, for example, processes 650 and / or data 640. Disk 635 and / or memory 615 may store an operating system that controls computer 605 and allocates its resources.
[0164] In one embodiment, storage / disk 635 is configured for structured storage and retrieval of one or more collections of information or data in a non-transitory computer-readable medium, for example, as one or more data structures. In one embodiment, storage / disk 635 includes one or more databases configured to store and serve information used by deepfake detection system 100. In one embodiment, storage / disk 635 includes one or more time-series databases configured to store and serve time-series data. In one embodiment, the time-series databases are not solely SQL (NOSQL) databases.
[0165] Computer 605 interacts with, controls, and / or can be controlled by input / output (I / O) devices via input / output controller 647, I / O interface 645, and input / output ports 620. The input / output devices may include one or more displays 670, printers 672 (such as inkjet, laser, or 3D printers), and audio output devices 674 (such as speakers or headphones), character input devices 680 (such as a keyboard), pointing and selection devices 682 (such as a mouse, trackball, touchpad, touchscreen, joystick, pointing stick, stylus mouse), audio input devices 684 (such as a microphone), video input devices 686 (such as video and still cameras), video card (not shown), disk 635, network devices 655, sensors (not shown), etc. The input / output ports 620 may include, for example, serial ports, parallel ports, and USB ports.
[0166] In one embodiment, computer 605 may be connected to an audiovisual content source 690. Audiovisual content source 690 may be a streaming service for transmitting a stream of audiovisual signals to the computing device. Audiovisual content source 690 may also be a broadcast receiver for collecting transmitted broadcasts of audiovisual signals.
[0167] Computer 605 may operate in a networked environment and thus may be connected to network device 655 via I / O interface 645 and / or I / O port 620. Through network device 655, computer 605 may interact with network 660. Through network 660, computer 605 may be logically connected to remote computer 665, as well as to live, real-time broadcast transmissions and / or streams of audiovisual content from audiovisual content source 690. Networks with which computer 605 may interact include, but are not limited to, LANs, WANs, and other networks.
[0168] -Definitions and Other Embodiments- In another embodiment, the described methods and / or their equivalents may be implemented by computer-executable instructions. Thus, in one embodiment, a non-transitory computer-readable / storage medium is configured with stored computer-executable instructions of an algorithm / executable application that, when executed by the machine, causes the machine (and / or associated components) to perform the method. Exemplary machines include, but are not limited to, processors, computers, servers operating in a cloud computing system, servers configured in a Software as a Service (SaaS) architecture, smartphones, etc. In one embodiment, a computing device is implemented with one or more executable algorithms configured to perform any of the disclosed methods.
[0169] In one or more embodiments, one or more of the components described herein are configured as program modules stored on a non-transitory computer-readable medium, the program modules being comprised of stored instructions that, when executed by at least a processor, cause a computing device to perform the corresponding functions described herein.
[0170] In one or more embodiments, the disclosed methods or their equivalents are implemented by either computer hardware configured to perform the methods or by computer instructions embodied in modules stored on a non-transitory computer-readable medium, the instructions configured as an executable algorithm that, when executed by at least a processor of a computing device, is configured to perform the methods.
[0171] In one embodiment, each step of a computer-implemented method described herein may be performed by a processor of one or more computing devices configured with logic to (i) access memory and (ii) cause the system to perform the method steps. For example, the processor accesses, reads from, or writes to memory to perform the computer-implemented method steps described herein. These steps may include (i) obtaining any necessary information, (ii) calculating, determining, generating, classifying, or otherwise creating any data, and (iii) storing any calculated, determined, generated, classified, or otherwise created data for subsequent use. References to storing or storing refer to storage as a data structure in the memory or storage / disk of the computing device.
[0172] In one embodiment, each subsequent step of the method begins automatically in response to parsing a received signal or obtained stored data that indicates that the previous step has been performed at least to the extent necessary for the subsequent step to begin. Typically, the received signal or obtained stored data indicates completion of the previous step.
[0173] For ease of explanation, the illustrated methods in the figures are shown and described as a series of algorithmic blocks, but it should be understood that the method is not limited by the order of the blocks. Some blocks may occur in a different order than shown and described and / or concurrently with other blocks than shown and described. Furthermore, less than all of the illustrated blocks may be used to implement the example method. Blocks may be combined or separated into multiple acts / components. Furthermore, additional and / or alternative methods may use additional acts not illustrated in the blocks.
[0174] The following contains definitions of selected terms used herein. The definitions include various examples and / or forms of components that fall within the scope of the term and that may be used for implementation. The examples are not intended to be limiting. Both singular and plural forms of a term may be within the scope of the definition.
[0175] References to "one embodiment," "one embodiment," "one example," "one example," etc. indicate that the embodiment or example so described may include a particular feature, structure, characteristic, property, element, or limitation, and that not all embodiments or examples necessarily include that particular feature, structure, characteristic, property, element, or limitation. Furthermore, repeated use of the phrase "in one embodiment" does not necessarily refer to the same embodiment, although it may.
[0176] A "data structure," as used herein, is an organization of data within a computing system stored in memory, a storage device, or other computerized system. A data structure may be, for example, any one of a data field, a data file, a data array, a data record, a database, a data table, a graph, a tree, a linked list, etc. A data structure may be formed from and may contain many other data structures (e.g., a database contains many data records). According to other embodiments, other examples of data structures are possible as well.
[0177] As used herein, "computer-readable medium" or "computer storage medium" refers to a non-transitory medium that stores instructions and / or data that, when executed, are configured to perform one or more of the disclosed functions. Data may function as instructions in some embodiments. Computer-readable media may take many forms, including, but not limited to, non-volatile media and volatile media. Non-volatile media may include, for example, optical disks, magnetic disks, and the like. Volatile media may include, for example, semiconductor memory, dynamic memory, and the like. Common forms of computer-readable media may include, but are not limited to, floppy disks, flexible disks, hard disks, magnetic tape, other magnetic media, application-specific integrated circuits (ASICs), programmable logic devices, compact disks (CDs), other optical media, random access memory (RAM), read-only memory (ROM), memory chips or cards, memory sticks, solid-state storage devices (SSDs), flash drives, and other media with which a computer, processor, or other electronic device can function. Each type of media, when selected for implementation in one embodiment, may include stored instructions of an algorithm configured to perform one or more of the disclosed and / or claimed functions.
[0178] "Logic," as used herein, refers to components implemented using computer or electrical hardware, non-transitory media having executable application or program module instructions stored thereon, and / or combinations thereof, to perform any of the functions and acts as disclosed herein and / or to cause a function or act from another logic, method, and / or system to be performed as disclosed herein. Equivalent logic may include firmware, a microprocessor programmed with an algorithm, discrete logic (e.g., an ASIC), at least one circuit, analog circuit, digital circuit, programmed logic device, memory device containing algorithmic instructions, etc., any of which may be configured to perform one or more of the disclosed functions. In one embodiment, logic may include one or more gates, combinations of gates, or other circuit components configured to perform one or more of the disclosed functions. Where multiple logics are described, it may be possible to combine the multiple logics into one logic. Similarly, where a single logic is described, it may be possible to distribute the single logic among multiple logics. In one embodiment, one or more of these logics are corresponding structures associated with performing the disclosed and / or claimed functions. The selection of what type of logic to implement may be based on desired system requirements or specifications. For example, if faster speed is a consideration, hardware may be selected to implement the function. If lower cost is a consideration, stored instructions / executable applications may be selected to implement the function.
[0179] An "operable connection," or a connection in which entities are "operably connected," is one in which signals, physical communications, and / or logical communications may be sent and / or received. An operable connection may include a physical interface, an electrical interface, and / or a data interface. An operable connection may include different combinations of interfaces and / or connections sufficient to enable operable control. For example, two entities may be operably connected to communicate signals with each other directly or through one or more intermediate entities (e.g., a processor, an operating system, logic, a non-transitory computer-readable medium). Logical and / or physical communication channels may be used to create an operable connection.
[0180] "User", as used herein, includes, but is not limited to, one or more people, computers, or other devices, or combinations thereof.
[0181] While the disclosed embodiments have been illustrated and described in considerable detail, it is not intended to limit, or in any way restrict, the scope of the appended claims to such detail. It is, of course, not possible to describe every conceivable combination of components or methodologies for purposes of describing various aspects of the subject matter. Accordingly, the present disclosure is not limited to the specific details or illustrative examples shown and described. It is therefore intended that the present disclosure embrace all such changes, modifications, and variations that fall within the scope of the appended claims.
[0182] To the extent the terms "include" or "comprising" are used in the detailed description or claims, they are intended to be inclusive in a manner similar to how the term "comprising" is interpreted when used as a transitional word in the claims.
[0183] To the extent the term "or" is used in the detailed description or claims (e.g., A or B), it is intended to mean "A or B or both." When applicants intend to indicate "A or B only, but not both," "A or B only, but not both" is used. Thus, the use of "or" herein is inclusive, not exclusive.
Claims
1. A method by which a computer performs an action. Converting an audiovisual signal, including speech by a human speaker, into a set of time-series signals that include a video subset of a time-series signal for video and an audio subset of a time-series signal for audio, The process involves generating a set of residual time series signals from the set of time series signals and the set of estimations of the time series signals created by the machine learning model, Includes, The machine learning model generates the estimation so as to match the authentic speech of the human speaker. The aforementioned method, The method further includes placing the residual values from one synchronous observation of the set of residual time-series signals into a two-dimensional array divided into video and audio partitions. The residual values generated for the video subset are placed in the video partition, and the residual values generated for the audio subset are placed in the audio partition. The aforementioned method, To detect anomalies within the residual values, a sequential analysis of the residual values is performed across the two dimensions of the two-dimensional array, A method performed by a computer, further comprising, in response to the detection of the anomaly, generating an alert that deepfake content falsely conveying the human speaker or the speech has been detected in the audiovisual signal.
2. The audiovisual signal is converted into a video subset of the time-series signal by sampling a time-series signal from the pixels of the video frames within the audiovisual signal. The method performed by a computer according to claim 1, further comprising converting the audiovisual signal into an audio subset of the time-series signal by sampling the time-series signal from the frequency range of the audio signal.
3. A computer-operated method according to claim 1 or 2, further comprising using multiple processors to perform the sequential analysis in parallel along (i) one or more rows in the larger dimension of the two-dimensional array and (ii) one or more columns in the smaller dimension of the two-dimensional array, wherein each row in the larger dimension of the two-dimensional array includes cells in both the video partition and the audio partition of the rectangular array, and the anomaly is detected when any one of the sequential probability ratio tests across rows or columns identifies an anomaly residual.
4. Converting the audiovisual content of a person giving a speech into a set of time-series signals, The process involves generating a residual time series signal of the residuals that shows the degree to which the aforementioned time series signal differs from the machine learning estimation of the actual speech by the person, The residual value from one synchronous observation of the aforementioned residual time series signal is placed into an array of residual values for a given point in time. In order to detect anomalies within the residual values at a certain point in time, a sequential analysis of the residual values of the sequence is performed. A method performed by a computer, which in response to the detection of the anomaly, includes generating an alert that at least one of a false audio word or an altered movement has been detected in the audiovisual content.
5. Converting the audiovisual content of a person giving a speech into a set of time-series signals, The process involves generating a residual time series signal of the residuals that shows the degree to which the aforementioned time series signal differs from the machine learning estimation of the actual speech by the person, The residual value from one synchronous observation of the aforementioned residual time series signal is placed into an array of residual values for a given point in time. In order to detect anomalies within the residual values at a certain point in time, a sequential analysis of the residual values of the sequence is performed. A method performed by a computer, which includes, in response to the detection of the anomaly, generating an alert that deepfake content has been detected in the audiovisual content.
6. To create a video subset of the time-series signal representing the video of the person, the pixels of the video in the audiovisual content are sampled, The method further includes sampling the frequency range of the audio within the audiovisual content in order to create an audio subset of the time-series signal representing the audio of the speech, The method performed by a computer according to claim 4 or 5, wherein residual values generated from the video subset are placed in the video partition of the array, and residual values generated from the audio subset are placed in the audio partition of the array.
7. The method performed by the computer according to claim 6, wherein the array is a two-dimensional array, and the method performed by the computer further includes placing residual values generated for the video subset into the video partition in a cell corresponding to a location in a pixel grid represented by the residual values.
8. The aforementioned array is a two-dimensional array, and performing the sequential analysis of the residual values in order to detect anomalies within the residual values at a given time point is, Perform a sequential probability ratio test across the rows and columns of the two-dimensional array, A computer-based method according to claim 4 or 5, further comprising detecting an anomaly when any one of the sequential probability ratio tests across rows or columns identifies an anomaly residual.
9. To ensure that the set of time-series signals is sampled at a uniform rate, one or more of the time-series signals are resampled. A method performed by a computer according to any one of claims 1, 4, and 5, further comprising phase-shifting the time-series signal so that the observation of the time-series signal is synchronized.
10. During the streaming of the audiovisual signal, a reference segment of the time-series signal set representing the authentic speech by the human speaker is specified. The computer method according to claim 1, further comprising training the machine learning model to generate estimates of the time series signals that match the authentic speech, based on the reference segments of the set of time series signals, before generating the set of residual time series signals.
11. The alert includes the location of the pixel where the anomaly occurred, or one of the frequency ranges in which the anomaly occurred. A method performed by a computer according to any one of claims 1, 4, and 5, further comprising including the timestamp of the occurrence of the anomaly in the alert.
12. The location of the pixel where the anomaly occurred and the time when the anomaly occurred are included in the alert, A computer method according to any one of claims 1, 4, and 5, further comprising highlighting the pixel in which the anomaly occurred within the audiovisual content in order to visually represent the deepfake content.
13. The alert includes the frequency range in which the anomaly occurred and an identifier for the frame of the audiovisual content in which the anomaly occurred. A computer method according to any one of claims 1, 4, and 5, further comprising adding a warning to the audiovisual content in the frame in which the anomaly occurred in order to visually indicate the deepfake content.
14. A program for causing one or more computers to execute the method according to any one of Claims 1, 4, and 5.
15. A computing system comprising one or more computers, each having at least one processor and at least one memory storing a program for causing the at least one processor to perform the method according to any one of claims 1, 4, and 5.