Multimodal Inputs in Adaptive Computer-Generated Reality Recording

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing augmented reality technologies lack efficient methods for integrating multimodal inputs, such as facial expressions, gestures, and speech, to enhance user interaction and annotation in computer-generated reality environments.

Innovation Solution

A computer-generated reality system that utilizes facial expressions, gestures, and speech as input modalities to record, track regions of interest, and provide annotations in real-time, with the ability to operate in standalone, wireless tethered, or connected modes, leveraging network and base device resources for processing and power management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple sensors and processing modes are integrated into the electronic device, then the functionality and versatility of the device is improved, but the power consumption and processing load increase

Engineering Contradiction:
ImprovefunctionalityVSAvoidpower consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The system dynamically adjusts its operating mode (standalone, wireless tethered, or connected) based on real-time conditions, allowing the device to optimize power consumption by leveraging network and base device resources when available, while operating independently when needed

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If multiple sensors and processing modes are integrated into the electronic device, then the functionality and versatility of the device is improved, but the processing capability requirements increase

Engineering Contradiction:
ImprovefunctionalityVSAvoidprocessing capability
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system uses network and base device resources as intermediaries to offload processing tasks, allowing the electronic device to access enhanced processing capabilities without permanently increasing its own hardware complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The electronic device is designed to perform multiple functions across different operating modes (standalone, wireless tethered, connected), allowing a single device to serve multiple purposes by leveraging both local and remote resources

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Speed

If real-time processing and rendering are performed locally, then the response time is improved, but the power consumption increases

Engineering Contradiction:
Improveresponse timeVSAvoidpower consumption
Core Design Contradiction:
SpeedVSUse of energy by moving object

Solution Approach 1:

The system dynamically determines the appropriate processing location (local vs. remote) based on real-time conditions, balancing response time requirements with power consumption by selecting the optimal operating mode

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP4028862B1Multimodal inputs for computer-generated reality
Publication Date: 2025.09.03 APPLE INC
  • EP4028862B1 patent drawingFigure 1
  • EP4028862B1 patent drawingFigure 2
  • EP4028862B1 patent drawingFigure 3A~3C

AI summary

Implementations of the subject technology provide determining an operating mode of an electronic device based at least in part on whether the electronic device is communicatively coupled to an associated base device. Based on the determined operating mode, the subject technology identifies a set of input modalities for initiating a recording of content within a field of view of the electronic device. The subject technology monitors sensor information generated by at least one sensor included in, or communicatively coupled to, the electronic device. Further, the subject technology initiates the recording of content within the field of view of the electronic device when the monitored sensor information indicates that at least one of the identified set of input modalities has been triggered.