Voice Annotation for Asynchronous Mixed Reality

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Mixed-reality systems lack the capability for efficient asynchronous communication between users who are not simultaneously present in a shared mixed-reality environment.

Innovation Solution

Implementing a voice annotation system that allows users to generate and associate voice annotations with virtual objects or environments, enabling other users to access and playback these annotations later, even if they are not simultaneously present, through a computing system with voice annotation components and metadata management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real-time communication is implemented in mixed-reality systems, then users can interact simultaneously in shared environments, but communication between users not present at the same time cannot be supported

Engineering Contradiction:
Improvecommunication capabilityVSAvoidasynchronous communication support
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary actions by recording and storing voice annotations, environmental data, and user actions during a user's presence in the mixed-reality environment. These pre-captured elements are then made available for later playback to other users, enabling asynchronous communication without requiring simultaneous presence.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates copies of the mixed-reality environment including voice annotations, spatial audio, visual elements, and contextual metadata. These copies are stored and transmitted to other users who can playback the reconstructed environment, allowing them to experience and communicate about the original environment at a different time.

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If voice annotations are recorded and stored for asynchronous access, then communication beyond real-time interactions is enabled, but system complexity increases

Engineering Contradiction:
Improveasynchronous communication capabilityVSAvoidvoice annotation system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system implements multi-functionality by using a unified voice annotation mechanism that serves multiple purposes: capturing user commentary, recording environmental sounds, storing spatial relationships, and creating communicative content. This single system handles recording, storage, retrieval, and playback functions, reducing overall system complexity compared to separate specialized systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The server acts as an intermediary that manages the complexity of voice annotation storage, retrieval, and synchronization. By offloading these complex management tasks to the server, the client devices can maintain simpler local implementations while still accessing the full asynchronous communication capability through the server-mediated interface.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of information

If voice annotations are associated with virtual objects and environments, then contextual communication is improved, but data management complexity increases

Engineering Contradiction:
Improvecontextual information preservationVSAvoidmetadata management complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system implements nesting by organizing metadata hierarchically: voice annotations are nested within specific virtual objects, which are nested within environments, which are in turn nested within user sessions. This nested structure allows contextual information to be preserved at multiple levels while enabling efficient retrieval by accessing only the relevant nested level for a given query.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentEP3903177B1Asynchronous communications in mixed-reality
Publication Date: 2023.09.13 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3903177B1 patent drawingFigure 1
  • EP3903177B1 patent drawingFigure 2A~2B
  • EP3903177B1 patent drawingFigure 3

AI summary

A system and method include presentation of a plurality of virtual objects to a first user, reception, from the first user, of a command to associate a voice annotation with one of the plurality of virtual objects, reception of audio signals of a first voice annotation from the first user, and storage the received audio signals in association with metadata indicating the first user and the one of the plurality of virtual objects.