User Image Retrieval via Pose Data and 3D Scene Reconstruction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Users face difficulties in locating images of themselves taken by others at specific geographic locations, especially when they are partially depicted or in the background, due to the lack of efficient methods for image retrieval and identification.
Innovation Solution
A system and method that utilize time and location indicators, pose data, and 3D reconstruction to identify and select images of a user from a set of candidate images, employing facial recognition to determine the presence and visibility of the user in the images, allowing for the retrieval of images taken at specific locations and times.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If users manually search through publicly shared photos to locate images of themselves, then they can find images depicting themselves, but the process becomes time-consuming and difficult especially when they are partially depicted or in the background
Solution Approach 1:
The patent replaces manual mechanical searching with automated computer vision and machine learning systems. The system automatically analyzes photo metadata (location, time), performs 3D reconstruction of scenes, estimates user pose and position, and identifies images containing the user without manual intervention, thereby eliminating time loss while maintaining high retrieval accuracy
Solution Approach 2:
The patent creates a digital twin or 3D reconstruction copy of the physical scene where the photo was taken. By reconstructing the 3D environment from multiple photos and comparing it with the target photo's metadata and pose data, the system can accurately determine whether the user appears in the image without manually examining each photo, thus improving retrieval efficiency
2Reliability
If the system analyzes all publicly shared photos to ensure complete coverage, then all images of the user are found, but the computational complexity and data processing requirements increase significantly
Solution Approach 1:
The patent segments the large-scale photo analysis task into manageable components: filtering photos by geographic location and time metadata, performing 3D reconstruction only for relevant locations, estimating pose data selectively, and conducting facial recognition only on candidate images. This segmentation maintains retrieval completeness while reducing overall system complexity by processing only necessary subsets of data
Solution Approach 2:
The patent performs preliminary filtering of photos based on metadata (location, time) before conducting complex 3D reconstruction and pose estimation. By pre-screening photos that match the user's known visit locations and time periods, the system reduces the dataset size significantly before applying computationally intensive algorithms, thereby maintaining reliability while managing complexity
3Measurement precision
If the system uses detailed pose data and 3D reconstruction to accurately determine user visibility, then image identification accuracy improves, but the computational resources and processing time required increase
Solution Approach 1:
The patent applies partial pose estimation and 3D reconstruction only to candidate images that pass initial metadata filtering, rather than processing all photos. By performing detailed analysis only on a subset of relevant images based on location and time matches, the system achieves high detection accuracy for visible users while reducing computational energy consumption compared to exhaustive analysis of all publicly shared photos
Data Source
AI summary
The aspects described herein include receiving a request for available images depicting a user. One or more time and location indicators indicating one or more locations visited by the user are determined. Based on at least in part the one or more time and location indicators, a set of candidate images may be identified. The set of candidate images depict one or more locations at a time corresponding to at least one of the time indicators. Pose data related to the user may be obtained based on the location indicators. The pose data indicates a position and orientation of the user during a visit at a given location depicted in the set of candidate images. One or more images from the set of candidate images may be selected based on the pose data and the 3D reconstruction. The selected images include at least a partial view of the user.


