RGB Video Authentication Using Infrared and Depth Sensing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current technologies lack effective solutions to authenticate red green blue (RGB) video content in video conferencing settings, particularly against deep fake attacks, which can lead to unauthorized information disclosure and verbal abuse, due to the remote nature of video conferences making it difficult to verify the authenticity of video feeds in real-time.
Innovation Solution
A system that utilizes a combination of RGB video content, infrared (IR) video content, and depth sensor data to authenticate video feeds by comparing frames and determining a match within a threshold level of confidence, potentially using a side channel for increased security, and issuing challenges via an IR illuminator to verify authenticity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If RGB video content is authenticated using only RGB data, then the authentication process is simple and fast, but it is vulnerable to deep fake attacks and cannot reliably verify video authenticity
Solution Approach 1:
The patent combines RGB video data with infrared (IR) video data and depth sensor data into a unified authentication system. By merging these three different data types, the system creates a multi-modal authentication approach that significantly improves reliability against deep fake attacks compared to using RGB data alone, while distributing the complexity across multiple sensor inputs rather than requiring a single complex authentication mechanism
Solution Approach 2:
The patent adds new dimensions to video authentication by incorporating infrared spectral data and depth spatial data alongside the traditional RGB visual data. This multi-dimensional approach transforms the authentication process from a two-dimensional RGB analysis to a three-dimensional verification system that includes spectral dimension (IR) and spatial dimension (depth), making it computationally intensive for deep fake AI to generate matching frames across all dimensions simultaneously
2Measurement precision
If multiple sensor data types (IR and depth) are used for authentication, then deep fake detection accuracy improves, but computational requirements and processing time increase
Solution Approach 1:
The system performs preliminary actions by continuously capturing and pre-processing IR video data and depth sensor data in parallel with RGB video data. This allows the authentication system to have pre-prepared multi-modal data ready for comparison, reducing the computational burden during the actual authentication moment and enabling faster processing despite the complexity of analyzing multiple data types
Solution Approach 2:
The patent implements a threshold-based authentication system where the comparison of multi-sensor data accepts a match within a threshold level of confidence rather than requiring perfect alignment. This partial action approach balances detection accuracy with processing efficiency, allowing the system to achieve high precision in deep fake detection while avoiding the need for exhaustive computational analysis of every pixel across all sensor data types
3Reliability
If real-time authentication of video feeds is implemented, then unauthorized access is prevented, but system latency and processing overhead increase
Solution Approach 1:
The patent implements continuous authentication by constantly comparing incoming RGB video frames with corresponding IR video frames and depth sensor data throughout the video conference session. This continuous verification process maintains real-time authentication without requiring periodic interruptions, ensuring that video feed authenticity is verified continuously while minimizing latency through ongoing parallel processing of multi-sensor data streams
Data Source
AI summary
In one aspect, a device may include at least one processor and storage accessible to the at least one processor. The storage may include instructions executable by the at least one processor to access a first frame of RGB video content corresponding to a first time, access a first frame of IR video content corresponding to the first time, and access data from a depth sensor corresponding to the first time. The instructions may also be executable to determine whether at least a portion of the first frame of the RGB video content correlates to at least a portion of the first frame of the IR video content and/or the data from the depth sensor. Responsive to a determination that it does, the instructions may be executable to authenticate the RGB video content and indicate the RGB video content as being authenticated via a graphical user interface.


