Deep Neural Network Style Transfer for Mixed Reality Passthrough

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current mixed-reality systems require multiple cameras to generate passthrough visualizations, leading to increased weight, cost, and battery usage, while failing to optimize the viewing experience with enhanced data.

Innovation Solution

A deep neural network (DNN) is used to transition the style of images from one camera type to another, allowing a single camera to capture and process multiple types of data, such as thermal and low-light images, by learning and modifying image styles to match different camera perspectives.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple cameras are used to generate passthrough visualizations, then the viewing experience is enhanced with diverse data, but the weight, cost, and battery usage increase

Engineering Contradiction:
Improveviewing experience qualityVSAvoiddevice weight
Core Design Contradiction:
ReliabilityVSWeight of moving object

Solution Approach 1:

The patent uses a deep neural network to learn the imaging characteristics of different camera types and generate synthetic images that copy the appearance and data characteristics of thermal, low-light, or other specialized camera outputs. A single visible light camera captures images, and the DNN processes these images to create multiple style variations, effectively replacing multiple physical cameras with one camera plus computational processing.

Inventive Principle:
Principle #26Copying

2Reliability

If multiple cameras are used to generate passthrough visualizations, then the viewing experience is enhanced with diverse data, but the device cost increases

Engineering Contradiction:
Improveviewing experience qualityVSAvoiddevice cost
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The deep neural network serves as a universal processor that can generate multiple types of image data (thermal, low-light, enhanced visible light) from a single camera input. This multi-functional approach allows the system to provide diverse data streams without requiring multiple specialized camera hardware components, thereby reducing device cost while maintaining enhanced viewing experience quality.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If multiple cameras are used to generate passthrough visualizations, then the viewing experience is enhanced with diverse data, but the battery consumption increases

Engineering Contradiction:
Improveviewing experience qualityVSAvoidbattery consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system captures images with a single visible light camera and uses a deep neural network to computationally generate synthetic thermal and low-light images by copying and transforming the visual information. This approach consumes less battery power than operating multiple power-hungry specialized cameras simultaneously, as the DNN processing is more energy-efficient than multiple camera sensors and processing pipelines.

Inventive Principle:
Principle #26Copying

4Weight of moving object

If a single camera is used, then the device weight, cost, and battery usage are reduced, but the ability to capture multiple types of data is limited

Engineering Contradiction:
Improvedevice weightVSAvoiddata capture capability
Core Design Contradiction:
Weight of moving objectVSAdaptability or versatility

Solution Approach 1:

The deep neural network dynamically changes the parameters and characteristics of captured images through style transfer and transformation. By adjusting computational parameters rather than physical camera parameters, the system can generate thermal-like, low-light-like, and enhanced visible light images from a single camera, thereby achieving multi-type data capture capability with reduced device weight.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12002226B2Using machine learning to selectively overlay image content
Publication Date: 2024.06.04 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12002226B2 patent drawing
  • US12002226B2 patent drawing
  • US12002226B2 patent drawing

AI summary

Modifications are performed to cause a style of an image to match a different style. A first image is accessed, where the first image has the first style. A second image is also accessed, where the second image has a second style. Subsequent to a deep neural network (DNN) learning these styles, a copy of the first image is fed as input to the DNN. The DNN modifies the first image copy by transitioning the first image copy from being of the first style to subsequently being of the second style. As a consequence, a modified style of the transitioned first image copy bilaterally matches the second style.