Multi-Modal Deception Detection via Audio-Video Feature Aggregation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing deception detection systems in digital and distributed transactions are unreliable due to high false positive rates and failure to accurately analyze multiple modalities such as video, audio, and micro-expressions, making it difficult to identify deception in real-time, especially in high-value industries like finance and identity verification.
Innovation Solution
A distributed computing architecture that enables real-time deception detection in Audio-Video responses by extracting Natural Language Processing features, speech features, and micro-expressions using a multi-modal algorithm, aggregating these features, and employing a machine learning model to classify responses as deceptive or truthful.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If existing physiological signal-based deception detection approaches are used, then deception detection can be performed, but false positive rates increase and reliability decreases
Solution Approach 1:
The patent combines multiple independent detection modalities (facial micro-expressions, voice tone analysis, body language detection, physiological signals) into a unified deception detection system. By merging these different sensing approaches, the system achieves more reliable and accurate deception detection while reducing false positives compared to single-modality systems.
Solution Approach 2:
The system employs a multi-functional approach by integrating various detection capabilities (visual, auditory, physiological) into a single platform that can analyze multiple aspects of human behavior simultaneously. This universal system adapts to different deception scenarios and provides comprehensive analysis across multiple dimensions.
2Reliability
If manual verification processes are used in digital transactions, then security can be maintained, but productivity and transaction speed decrease
Solution Approach 1:
The deception detection system operates autonomously without requiring manual verification intervention. The automated analysis of facial expressions, voice patterns, and physiological signals enables the system to independently assess transaction authenticity, maintaining security while dramatically improving processing speed and eliminating human involvement in routine verification.
Solution Approach 2:
The patent replaces manual verification processes with automated computational analysis systems. Machine learning algorithms and computer vision techniques substitute human reviewers, enabling rapid automated deception detection that maintains security standards while increasing transaction throughput and eliminating bottlenecks associated with manual processing.
3Device complexity
If single-modality deception detection systems are used, then system complexity is reduced, but measurement precision and detection accuracy decrease
Solution Approach 1:
The system segments deception detection into distinct analytical modules: facial micro-expression recognition, voice tone analysis, body language detection, and physiological signal processing. Each module independently analyzes specific behavioral aspects and feeds results to a centralized decision-making algorithm, managing complexity through modular organization while achieving high detection precision.
Solution Approach 2:
The patent transitions from single-dimensional analysis to multi-dimensional assessment by incorporating diverse sensing modalities (visual, auditory, physiological). This dimensional expansion allows the system to capture deception indicators across multiple independent dimensions, significantly improving detection precision while organizing complexity through structured multi-modal integration.
Data Source
AI summary
A method for (of) detecting deception in an Audio-Video response of a user, using a server, in a distributed computing architecture, characterized in that the method including: enabling an Audio-Video connection with a user device upon receiving a request from a user; obtaining, from the user device, an Audio-Video response of the user corresponding to a first set of questions that are provided to the user by the server; extracting audio signals and video signals from the Audio-Video response; detecting an activity of the user by determining a plurality of Natural Language Processing (NLP) features from the extracted audio signals by (i) performing a speech to text translation and (ii) extracting the plurality of NLP features from the translated text, and determining a plurality of speech features from the extracted audio signals by (i) splitting the extracted audio signals into a plurality of short interval audio signals and (ii) extracting the plurality of speech features from the plurality of short interval audio signals; aggregating (i) the plurality of NLP features to obtain a plurality of temporal NLP features and (ii) the plurality of speech features to obtain a plurality of temporal speech features; aggregating the plurality of temporal NLP features and the plurality of temporal speech features to obtain first temporal aggregated features; detecting a plurality of micro-expressions of the user by splitting extracted video signals into a plurality of short fixed-duration video signals, detecting a plurality of Region Of Interest (ROI) in the plurality of short fixed-duration video signals, and comparing the plurality of detected ROI with video signals annotated with micro-expression labels that are stored in a database to detect the plurality of micro-expressions of the user in the plurality of short fixed-duration video signals; tracking and determining a gesture of the user from the extracted video signals; aggregating the plurality of micro-expressions and the gesture of the user to obtain second temporal aggregated features; aggregating the first temporal aggregated features and the second temporal aggregated features to obtain final temporal aggregated features; and detecting, using a machine learning model, a deception in the Audio-Video response based on the final temporal aggregated features.


