Video Transcription Using Visual Context Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing transcription engines perform poorly when transcribing audio files with specific context, such as those from enterprise environments, due to the use of uncommon words and expressions, and manual input of context information is cumbersome and time-consuming for large volumes of video files.
Innovation Solution
A method that extracts text content from visual content in video files using OCR techniques to generate context information for improving the automatic transcription of audio content, with optional iterations and post-processing to enhance accuracy and efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing transcription engines are used for enterprise video files with specific context, then transcription speed is maintained, but transcription accuracy deteriorates due to uncommon words and expressions
Solution Approach 1:
The patent extracts text from visual content (slides, captions, on-screen text) before transcription to generate context information that is provided to the transcription engine in advance. This preliminary extraction of contextual text from the video's visual elements enables the transcription engine to better understand domain-specific terminology and improve accuracy without requiring manual intervention.
2Measurement precision
If manual context information is input to improve transcription accuracy, then transcription quality improves, but processing time increases significantly
Solution Approach 1:
The system automatically extracts text from the visual content of the video file itself (slides, captions, on-screen text) to generate context information, eliminating the need for manual context input. The video file provides its own contextual information through its visual elements, enabling automated context generation that improves transcription accuracy without requiring human intervention or additional time for manual preparation.
Solution Approach 2:
The patent replaces the manual mechanical process of context input with an automated optical recognition system. OCR technology automatically extracts text from visual content frames, substituting the manual typing or copying of context information with an automated image-to-text conversion process that significantly reduces time while maintaining or improving accuracy.
3Measurement precision
If manual context input is performed for each video file, then transcription accuracy improves, but productivity decreases due to repetitive manual work
Solution Approach 1:
Each video file automatically generates its own context information by extracting text from its visual content, eliminating the need for repetitive manual context input for each file. This self-service approach maintains high transcription accuracy while enabling batch processing of multiple video files, significantly improving productivity and processing throughput.
Solution Approach 2:
The patent replaces repetitive manual context input operations with automated OCR-based text extraction that processes visual content frames. This substitution transforms a manual, sequential process into an automated, parallel process that can handle multiple video files simultaneously, maintaining accuracy while dramatically increasing processing throughput and productivity.
Data Source
AI summary
This invention relates to a computer implemented method (10) for processing a video file, said video file comprising audio content and visual content, the visual content comprising text content, wherein the method comprises:(S11) extracting the text content in the visual content;(S12) generating a context information for the audio content based on the text content extracted from said visual content; and(S13) converting the audio content into text by using the context information generated based on the text content extracted from the visual content of the video file.


