Cross-Modal Sensor Fusion for Mobile Device Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for tracking mobile devices and their users in indoor settings face challenges due to the small, shiny, and dark nature of smartphones, making them difficult to image clearly, and require a direct line of sight or active markers, which are rare and limited in application.

Innovation Solution

A cross-modal sensor fusion technique that matches motion features captured by sensors on a mobile device with motion features in images, allowing for the tracking of the device and its user without requiring a model of the appearance or direct line of sight, using inertial sensors and depth cameras to compare device and image accelerations in a 3D coordinate frame.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If direct line of sight and active markers are used for tracking, then tracking reliability is improved, but device complexity and ease of operation deteriorate due to requirements for markers and visibility

Engineering Contradiction:
Improvetracking reliabilityVSAvoidtracking system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and removes the requirement for active markers and direct line of sight from the tracking system. By using inertial sensors onboard the mobile device combined with image processing of natural features, the system achieves reliable tracking without needing additional markers or visual contact with the device, thereby reducing system complexity while maintaining reliability

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary approach by using inertial measurement units (IMUs) as a bridge between the mobile device and the tracking system. The IMUs provide motion data that can be correlated with video footage, enabling tracking through indirect means rather than requiring direct visual observation of the device itself

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If inertial sensors and image processing are fused, then tracking capability is improved for obscured devices, but computational complexity increases

Engineering Contradiction:
Improvetracking capabilityVSAvoidcomputational complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the tracking problem into two independent but complementary parts: inertial sensor data processing and image processing. By dividing the computational task and processing motion features and image features separately before fusing them, the system reduces overall computational complexity while improving adaptability to handle devices that are obscured or not directly visible

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal tracking framework that can handle both visible and obscured devices using the same sensor fusion approach. The system processes motion features from inertial sensors and correlates them with image features from video, enabling the same system to track devices whether they are visible or hidden in pockets, thereby improving versatility

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3077992B1Process and system for determining the location of an object by fusing motion features and iamges of the object
Publication Date: 2019.11.06 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3077992B1 patent drawingFigure 1
  • EP3077992B1 patent drawingFigure 2
  • EP3077992B1 patent drawingFigure 3

AI summary

The cross-modal sensor fusion technique described herein tracks mobile devices and the users carrying them. The technique matches motion features from sensors on a mobile device to image motion features obtained from images of the device. For example, the acceleration of a mobile device, as measured by an onboard internal measurement unit, is compared to similar acceleration observed in the color and depth images of a depth camera. The technique does not require a model of the appearance of either the user or the device, nor in many cases a direct line of sight to the device. The technique can operate in real time and can be applied to a wide variety of ubiquitous computing scenarios.