Video Summary Device for Virtual Event Transcription

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Transcriptions of virtual meetings are lengthy and time-consuming to process, consuming network, storage, and computing resources, and lack visual insight, making it inefficient to review the content.

Innovation Solution

A video summary device generates a video summary by converting transcriptions into phonemic transcriptions, text embeddings, audio embeddings, and image embeddings, combining them to create a visual representation of the virtual event, preserving resources and providing visual insight.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If full transcriptions of virtual meetings are processed and stored, then complete information is preserved, but network, storage, and computing resources are consumed excessively

Engineering Contradiction:
Improveinformation completenessVSAvoidresource consumption
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The system extracts only the essential information from full meeting transcriptions by generating summaries that capture key points, decisions, and action items. This extraction process removes redundant content while preserving the core information needed for understanding meeting outcomes.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The transcription processing is segmented into multiple stages: initial transcription generation, key information identification, summary creation, and visual representation. This segmentation allows the system to process only relevant portions of the full transcription at each stage, reducing overall resource consumption.

Inventive Principle:
Principle #1Segmentation

2Loss of information

If full transcriptions are provided to participants, then complete meeting content is available, but review time becomes substantial and inefficient

Engineering Contradiction:
Improvecontent completenessVSAvoidreview time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The system extracts and highlights only the most important segments of meeting content, presenting them in a condensed visual format that allows participants to quickly grasp key points without reviewing entire transcriptions. This extraction maintains information completeness while dramatically reducing review time.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system transforms linear text transcriptions into visual representations with multiple dimensions including time stamps, speaker identification, key topic markers, and hierarchical organization. This dimensional transformation enables participants to navigate and understand meeting content more efficiently.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Reliability

If detailed transcriptions are stored and processed, then comprehensive records are maintained, but storage and computing resources are consumed

Engineering Contradiction:
Improverecord accuracyVSAvoidstorage resources
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The system extracts essential meeting information and stores it in a condensed summary format rather than preserving complete transcriptions. This extraction maintains the reliability of key records while significantly reducing the quantity of data that requires storage resources.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system changes the parameter of data representation from full-text format to structured summary format with metadata tags, time stamps, and hierarchical organization. This parameter change reduces storage requirements while maintaining the accuracy and accessibility of essential information.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12200322B2Systems and methods for generating a video summary of a virtual event
Publication Date: 2025.01.14 VERIZON PATENT & LICENSING INC
  • US12200322B2 patent drawing
  • US12200322B2 patent drawing
  • US12200322B2 patent drawing

AI summary

A video summary device may generate a textual summary of a transcription of a virtual event. The video summary device may generate a phonemic transcription of the textual summary and generate a text embedding based on the phonemic transcription. The video summary device may generate an audio embedding based on a target voice. The video summary device may generate an audio output of the phonemic transcription uttered by the target voice. The audio output may be generated based on the text embedding and the audio embedding. The video summary device may generate an image embedding based on video data of a target user. The image embedding may include information regarding images of facial movements of the target user. The video summary device may generate a video output of different facial movements of the target user uttering the phonemic transcription, based on the text embedding and the image embedding.