Document Identification via Audio-Visual Impact Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current eKYC processes face challenges in verifying ownership of ID documents using scanned or photographed copies, as computer vision algorithms struggle to identify manipulated documents and require extensive training data and resources, leading to potential identity theft.
Innovation Solution
A document identification method and system that extracts image frames and audio signals from a video clip of an object impacting a surface, using a trained model to generate scores and identify if the object is a physical document, providing a more robust verification of ownership.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If computer vision algorithms are used to validate ownership factor by detecting and identifying documents, then the authentication process can proceed automatically, but the system cannot reliably distinguish between genuine physical documents and manipulated copies
Solution Approach 1:
The patent transitions from analyzing only visual characteristics (2D image data) to incorporating audio characteristics (sound wave data from document handling). This adds a new dimension of verification that is difficult to replicate with manipulated images, thereby maintaining automation while improving reliability in distinguishing genuine documents from copies
Solution Approach 2:
The patent introduces audio signal analysis as an intermediary verification layer between the user submitting the document and the authentication system. The audio characteristics of genuine document handling serve as a mediator that confirms the authenticity before the system proceeds with automatic validation, enhancing reliability without compromising automation
2Measurement precision
If computer vision algorithms are trained to identify manipulated documents, then identification accuracy improves, but large training data sets, resources and time are required
Solution Approach 1:
The patent replaces the computationally intensive mechanical process of training large computer vision models with audio-based verification. Instead of training algorithms to recognize visual patterns of manipulation, the system uses inherent audio characteristics of genuine document handling, which require minimal training data and computational resources while achieving high identification accuracy
Solution Approach 2:
The patent changes the verification parameter from visual characteristics (requiring extensive training to detect manipulations) to audio characteristics (which have inherent properties that distinguish genuine from fake documents). This parameter change reduces the need for large training datasets and extensive computational resources while maintaining or improving identification accuracy
3Ease of operation
If scanned or photographed copies of ID documents are accepted for eKYC, then the authentication process becomes convenient for users, but malicious actors can use these copies to commit identity theft
Solution Approach 1:
The patent performs preliminary audio verification of the document's authenticity before completing the authentication process. Users still conveniently submit scanned or photographed copies, but the system first analyzes audio characteristics to confirm the document is genuine before proceeding with the authentication, thereby preventing identity theft while maintaining ease of operation
Solution Approach 2:
The patent implements a feedback mechanism where audio analysis results are used to validate whether the submitted document copy is authentic. The system provides feedback by verifying the audio characteristics match those of genuine document handling, allowing convenient user submission while preventing misuse by malicious actors through this additional verification layer
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A document identification method and a document identification system are provided. The method includes extracting, using an image frame extraction device, a sequence of image frames from a video clip, the video clip capturing impact of an object against a surface, extracting, using an audio signal extraction device, a stream of audio signals from the video clip, and generating, using a processing device, a first score based on the sequence of image frames and a second score based on the stream of audio signals, using a trained document identification model. The document identification model is trained with a plurality of historical video clips, each of the plurality of historical video clips capturing impact of a document against a surface. The method also includes generating, using the processing device, an identification score based on the first score and the second score, and identifying, using the processing device, if the object in the video clip is a document based on a comparison between the identification score and an identification threshold.