Spatiotemporal Transformer for Long-Range Multi-Person Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Smart home technologies face challenges in accurately identifying multiple users at a distance and preventing spoofing attacks, particularly with face recognition systems that are not foolproof and require proper alignment, while behavioral biometrics are challenging to implement due to the need for extensive labeled training data.
Innovation Solution
A spatiotemporal transformer machine learning model is used to generate facial and pose features over time, combining face and gait identification with lip reading and anti-spoofing techniques to identify multiple users robustly at a distance, using synthetic datasets for training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If face recognition systems are used for user identification, then identification speed is improved, but reliability deteriorates due to spoofing attacks and alignment requirements
Solution Approach 1:
The patent combines face recognition, gait recognition, and lip reading into a unified multi-modal identification system. The fusion of these different biometric modalities allows the system to maintain high identification speed while improving reliability by cross-validating multiple independent features, making spoofing attacks significantly more difficult to execute successfully.
Solution Approach 2:
The system creates a composite biometric identification approach by integrating multiple recognition techniques (face, gait, lip reading) into a single robust identification mechanism. This composite approach leverages the strengths of each individual method while compensating for their respective weaknesses, particularly regarding spoofing vulnerability and alignment requirements.
2Reliability
If behavioral biometrics are implemented for improved security, then reliability is improved, but device complexity increases due to extensive labeled training data requirements
Solution Approach 1:
The spatiotemporal transformer model serves multiple functions simultaneously: it performs gait recognition, lip reading, and face analysis within a single unified framework. This multi-functionality reduces the need for separate complex training pipelines for each biometric modality, simplifying the overall system while maintaining high security reliability through diverse behavioral biometric analysis.
Solution Approach 2:
The system uses synthetic datasets to pre-train the spatiotemporal transformer model with diverse gait and lip reading patterns before deployment. This preliminary training action prepares the model to handle real-world variations effectively, reducing the need for extensive on-site labeled training data collection and simplifying the deployment complexity.
3Area of stationary object
If multi-person identification at long range is implemented, then area of coverage is improved, but measurement precision deteriorates due to distance and size constraints
Solution Approach 1:
The system transitions from relying solely on 2D facial features to incorporating 3D spatiotemporal information through gait analysis and lip reading sequences. By adding the temporal dimension and utilizing body-wide motion patterns, the system maintains measurement precision at long ranges where facial details become too small to resolve accurately.
Solution Approach 2:
The identification system segments the biometric analysis into multiple independent components: face recognition, gait recognition, and lip reading. This segmentation allows each component to operate optimally at different scales and distances, with gait providing robust long-range identification and face/lip reading providing closer-range verification, thereby maintaining precision across varying distances.
Data Source
AI summary
A method includes obtaining image frames capturing one or more people in at least one scene and identifying features of the image frames. The method also includes providing the identified features to a trained spatiotemporal transformer machine learning model configured to generate a set of features for each of the one or more people. The set of features for each person includes facial features of the person and pose features of the person over time. The method further includes performing face identification using the facial features to generate one or more first embeddings representing at least one face of at least one person and performing gait identification using the pose features to generate one or more second embeddings representing at least one gait of at least one person. In addition, the method includes identifying at least one of the one or more people based on the first and second embeddings.


