Voice Tagging Video Recording with AR Display

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video recording technologies require manual tagging after recording, which is inefficient and does not allow for real-time annotation during the recording process.

Innovation Solution

Implementing a system where users can add tags to a video in real-time by voice commands, using a trigger word to activate the tag function, allowing for duration or point-in-time tagging, and displaying suggested tags through an augmented reality display on a head-mounted device, which can be selected based on context and location.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual tagging is performed after video recording, then tagging accuracy can be ensured through careful review, but time consumption increases significantly and real-time annotation is not achieved

Engineering Contradiction:
Improvetagging accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary tagging actions during the video recording process itself, rather than after completion. Users can speak tags during recording, and the system automatically associates them with the corresponding video timestamps, achieving both real-time annotation and maintaining accuracy through immediate contextual association.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The manual mechanical process of reviewing and tagging videos after recording is replaced with an automated speech recognition system. The system captures voice commands during recording, automatically transcribes them, and associates tags with video timestamps without requiring manual intervention, thus eliminating time consumption while maintaining tagging accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Productivity

If voice commands are used for real-time tagging, then tagging efficiency improves and real-time annotation is achieved, but the system complexity increases due to speech recognition and processing requirements

Engineering Contradiction:
Improvetagging efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The head-mounted display device performs multiple functions: video recording, voice capture, speech recognition, and tag association all within a single integrated system. This multi-functionality reduces the need for separate dedicated devices for each task, managing system complexity while achieving high tagging efficiency through unified real-time processing.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system automatically processes voice commands during recording without requiring external intervention. The speech recognition system autonomously transcribes tags, determines timestamps, and associates them with video content, making the tagging process self-service and efficient while keeping the complexity contained within the automated workflow.

Inventive Principle:
Principle #25Self-service

3Ease of operation

If tags are displayed through augmented reality display during recording, then user awareness and workflow guidance improve, but the device complexity and power consumption increase

Engineering Contradiction:
Improveuser awarenessVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The augmented reality display acts as an intermediary between the user and the video recording process. It provides real-time visual feedback about active tags and recording status without requiring the user to manually check or manage tags, enhancing ease of operation. The complexity of tag management is handled by the system while the AR display provides simple, intuitive visual guidance.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11074292B2Voice tagging of video while recording
Publication Date: 2021.07.27 REALWEAR INC
  • US11074292B2 patent drawing
  • US11074292B2 patent drawing
  • US11074292B2 patent drawing

AI summary

Aspects of the technology described herein allow a user to add tags to a video as the video is being recorded. The tags can be added by capturing the user's voice as the video is recorded. Aspects can be performed by a head-mounted display. The head-mounted display can include an augmented reality display. In one aspect, a list of tags are displayed. The tags can be selected from a curated list of tags associated with a project the user is filming. For example, a project could comprise a building inspection during construction. In another aspect, the most commonly used tags in a given context can be shown. The most commonly used tags associated with a particular context can be determined through machine learning process. At a high level, a machine learning process can sort through historical tag data and associated contexts to determine a pattern correlating a context to a tag.