Gaze-Guided Camera Capture for Context-Aware Training Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing deep learning methods for image processing in immersive extended-reality technologies rely heavily on high-quality training data, which is often not optimized for specific use cases, leading to inefficient and resource-intensive manual auditing and potential biases, especially when transitioning to use-case specific training data.

Innovation Solution

A method and display apparatus that utilize gaze-tracking and pose-tracking to adjust camera settings for capturing high-quality real-world images, generating context-aware training data by ensuring images meet predefined quality criteria, thereby eliminating the need for manual labor and resource-intensive processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual auditing of training data is performed, then data quality can be ensured, but significant human effort and time are required

Engineering Contradiction:
Improvedata qualityVSAvoidhuman effort
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system uses automated quality assessment mechanisms where the neural network itself and auxiliary evaluation modules assess the quality of training data without human intervention. The system self-evaluates whether captured images meet quality criteria through automated detection of blur, resolution, and other quality metrics.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Manual auditing processes are replaced with automated computational systems including neural networks and quality assessment algorithms. The mechanical human effort of reviewing and validating training data is substituted with electronic processing and automated evaluation systems.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of manufacture

If generic training datasets are used, then data collection is simplified, but the neural network produces factually incorrect results in specific use cases

Engineering Contradiction:
Improvedata collectionVSAvoidaccuracy
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The system transitions from uniform generic data collection to context-specific data capture. Quality criteria and capture parameters are customized according to specific use cases (e.g., medical imaging, navigation, entertainment). The neural network is trained with data that has local quality characteristics optimized for particular applications rather than general-purpose data.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The training data collection process becomes dynamic and adaptive to different use cases. The system adjusts capture parameters, quality criteria, and data selection based on the specific application requirements, enabling the same system to generate specialized training data for multiple different purposes.

Inventive Principle:
Principle #15Dynamics

3Reliability

If reinforcement learning with human feedback (RLHF) is used, then neural network behavior can be improved, but significant human effort and resources are required

Engineering Contradiction:
Improvemodel behaviorVSAvoidefficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system replaces human feedback loops with automated evaluation mechanisms. Quality assessment is performed automatically by computational systems that evaluate whether training data meets predefined criteria, eliminating the need for continuous human review and feedback in the training process.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

Automated feedback mechanisms are implemented where the system continuously monitors and evaluates the quality of captured images and training data. Quality metrics are automatically calculated and used to adjust capture parameters and select appropriate training samples without human intervention.

Inventive Principle:
Principle #23Feedback

4Measurement precision

If high-quality training data is manually curated, then training accuracy improves, but the process becomes resource-intensive and non-scalable

Engineering Contradiction:
Improvetraining accuracyVSAvoidscalability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

Manual data curation processes are replaced with automated systems that capture, evaluate, and select training data at scale. The mechanical process of human reviewers examining and selecting images is substituted with electronic capture devices and automated quality assessment algorithms that can process vast quantities of data simultaneously.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

Quality criteria are established beforehand and applied automatically during the data capture and selection process. The system pre-defines what constitutes high-quality training data for each use case and automatically filters and selects samples that meet these criteria, eliminating the need for post-capture manual review.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250218129A1Method and display apparatus incorporating generation of context-aware training data
Publication Date: 2025.07.03 VARJO TECH OY
  • US20250218129A1 patent drawing
  • US20250218129A1 patent drawing
  • US20250218129A1 patent drawing

AI summary

Disclosed is a method including determining a gaze point and a gaze depth; controlling camera(s) for capturing a real-world image, by adjusting camera settings according to the gaze point and the gaze depth; determining a pose of the camera(s) at a time of capturing the real-world image; identifying region(s) of the real-world environment represented in the real-world image; determining whether a representation of region(s) satisfies quality criteria; when the representation fails to satisfy the quality criteria, capturing a reference real-world image such that the representation fulfills the quality criteria; generating training data comprising reference data and input data wherein reference data comprises reference real-world image, and input data with real-world image and/or previously-captured real-world image; sending training data to a processor to train a first neural network.