Video Transcription Using Visual Context Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing transcription engines perform poorly when transcribing audio files with specific context, such as those from enterprise environments, due to the use of uncommon words and expressions, and manual input of context information is cumbersome and time-consuming for large volumes of video files.

Innovation Solution

A method that extracts text content from visual content in video files using OCR techniques to generate context information for improving the automatic transcription of audio content, with optional iterations and post-processing to enhance accuracy and efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing transcription engines are used for enterprise video files with specific context, then transcription speed is maintained, but transcription accuracy deteriorates due to uncommon words and expressions

Engineering Contradiction:
Improvetranscription accuracyVSAvoidadaptability to specific context
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent extracts text from visual content (slides, captions, on-screen text) before transcription to generate context information that is provided to the transcription engine in advance. This preliminary extraction of contextual text from the video's visual elements enables the transcription engine to better understand domain-specific terminology and improve accuracy without requiring manual intervention.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If manual context information is input to improve transcription accuracy, then transcription quality improves, but processing time increases significantly

Engineering Contradiction:
Improvetranscription accuracyVSAvoidtime for context input
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system automatically extracts text from the visual content of the video file itself (slides, captions, on-screen text) to generate context information, eliminating the need for manual context input. The video file provides its own contextual information through its visual elements, enabling automated context generation that improves transcription accuracy without requiring human intervention or additional time for manual preparation.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical process of context input with an automated optical recognition system. OCR technology automatically extracts text from visual content frames, substituting the manual typing or copying of context information with an automated image-to-text conversion process that significantly reduces time while maintaining or improving accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If manual context input is performed for each video file, then transcription accuracy improves, but productivity decreases due to repetitive manual work

Engineering Contradiction:
Improvetranscription accuracyVSAvoidprocessing throughput
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

Each video file automatically generates its own context information by extracting text from its visual content, eliminating the need for repetitive manual context input for each file. This self-service approach maintains high transcription accuracy while enabling batch processing of multiple video files, significantly improving productivity and processing throughput.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces repetitive manual context input operations with automated OCR-based text extraction that processes visual content frames. This substitution transforms a manual, sequential process into an automated, parallel process that can handle multiple video files simultaneously, maintaining accuracy while dramatically increasing processing throughput and productivity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11990131B2Method for processing a video file comprising audio content and visual content comprising text content
Publication Date: 2024.05.21 BULL SA
  • US11990131B2 patent drawing
  • US11990131B2 patent drawing
  • US11990131B2 patent drawing

AI summary

This invention relates to a computer implemented method (10) for processing a video file, said video file comprising audio content and visual content, the visual content comprising text content, wherein the method comprises:(S11) extracting the text content in the visual content;(S12) generating a context information for the audio content based on the text content extracted from said visual content; and(S13) converting the audio content into text by using the context information generated based on the text content extracted from the visual content of the video file.