Voice Annotation for Asynchronous Mixed Reality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Mixed-reality systems lack the capability for efficient asynchronous communication between users who are not simultaneously present in a shared mixed-reality environment.
Innovation Solution
Implementing a voice annotation system that allows users to generate and associate voice annotations with virtual objects or environments, enabling other users to access and playback these annotations later, even if they are not simultaneously present, through a computing system with voice annotation components and metadata management.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If real-time communication is implemented in mixed-reality systems, then users can interact simultaneously in shared environments, but communication between users not present at the same time cannot be supported
Solution Approach 1:
The system performs preliminary actions by recording and storing voice annotations, environmental data, and user actions during a user's presence in the mixed-reality environment. These pre-captured elements are then made available for later playback to other users, enabling asynchronous communication without requiring simultaneous presence.
Solution Approach 2:
The system creates copies of the mixed-reality environment including voice annotations, spatial audio, visual elements, and contextual metadata. These copies are stored and transmitted to other users who can playback the reconstructed environment, allowing them to experience and communicate about the original environment at a different time.
2Adaptability or versatility
If voice annotations are recorded and stored for asynchronous access, then communication beyond real-time interactions is enabled, but system complexity increases
Solution Approach 1:
The system implements multi-functionality by using a unified voice annotation mechanism that serves multiple purposes: capturing user commentary, recording environmental sounds, storing spatial relationships, and creating communicative content. This single system handles recording, storage, retrieval, and playback functions, reducing overall system complexity compared to separate specialized systems.
Solution Approach 2:
The server acts as an intermediary that manages the complexity of voice annotation storage, retrieval, and synchronization. By offloading these complex management tasks to the server, the client devices can maintain simpler local implementations while still accessing the full asynchronous communication capability through the server-mediated interface.
3Loss of information
If voice annotations are associated with virtual objects and environments, then contextual communication is improved, but data management complexity increases
Solution Approach 1:
The system implements nesting by organizing metadata hierarchically: voice annotations are nested within specific virtual objects, which are nested within environments, which are in turn nested within user sessions. This nested structure allows contextual information to be preserved at multiple levels while enabling efficient retrieval by accessing only the relevant nested level for a given query.
Data Source
Figure 1
Figure 2A~2B
Figure 3
AI summary
A system and method include presentation of a plurality of virtual objects to a first user, reception, from the first user, of a command to associate a voice annotation with one of the plurality of virtual objects, reception of audio signals of a first voice annotation from the first user, and storage the received audio signals in association with metadata indicating the first user and the one of the plurality of virtual objects.