AR Transcript Visualization for Selective Document Inclusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In collaborative environments, such as classrooms or meetings, existing technologies fail to efficiently capture and selectively include contextually relevant spoken content into documents, leading to asymmetry in information distribution among participants.

Innovation Solution

A system and method using a head-mounted Augmented Reality (AR) device to capture, transcribe, and visualize spoken content, allowing users to selectively copy relevant parts of the transcript into documents, with features like participant identification, direction, and timestamping, and permission management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If all spoken content from multiple participants is captured and transcribed, then complete information is recorded, but information asymmetry increases and document complexity increases

Engineering Contradiction:
Improvecompleteness of discussion recordVSAvoiddocument complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the complete transcript by participant identity, allowing the system to capture all spoken content while organizing it into distinct, manageable sections attributed to specific speakers. This segmentation reduces the perceived complexity by structuring the information hierarchically.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts only the contextually relevant portions of the transcript for inclusion in the final document, separating essential information from unnecessary content. This extraction process maintains information completeness for reference purposes while reducing document complexity by including only what is needed.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If manual transcription and selection of relevant content is performed, then information accuracy is maintained, but time consumption increases

Engineering Contradiction:
Improveinformation accuracyVSAvoidtime for transcription and selection
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs automatic transcription of all participant speech into a structured format before the user needs to select relevant content. This preliminary action captures and organizes the information accurately, reducing the time required for manual transcription while maintaining precision through subsequent user review and selection.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If selective copying of transcript portions is enabled, then document efficiency improves, but system complexity increases

Engineering Contradiction:
Improvedocument creation efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary layer between the complete transcript and the final document - a selection interface that allows users to efficiently choose relevant portions. This intermediary maintains system simplicity by providing a straightforward selection mechanism while enabling efficient document creation through targeted content inclusion.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12190886B2Selective inclusion of speech content in documents
Publication Date: 2025.01.07 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US12190886B2 patent drawing
  • US12190886B2 patent drawing
  • US12190886B2 patent drawing

AI summary

In an approach for enabling a user to visualize a transcript of a discussion via a head-mounted AR device and to selectively copy one or more parts of the transcript that is contextually relevant for inclusion in a new document file or a previously created document file, a processor captures audio of a spoken content of a first participant of a discussion via an AR device worn by a user. A processor analyzes the audio of the spoken content of the first participant. A processor converts the audio of the spoken content of the first participant to text to create a transcript. A processor creates a visualization of the transcript. A processor presents the visualization of the transcript to the user via the AR device. A processor enables the user to copy one or more parts of the transcript into a document file via a selection support.