Video Frame OCR Classification for Searchable Communication Sessions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing digital communication platforms lack the ability to automatically extract textual content from video recordings of communication sessions, making it difficult and time-consuming to analyze and search for specific information such as presentation titles and slide content.

Innovation Solution

A system that extracts frames from video content, classifies them based on image analysis, identifies frames containing text, detects titles within these frames using optical character recognition (OCR), and transmits the extracted textual content to client devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If video content is manually reviewed to extract textual content, then accuracy of text extraction is improved, but time consumption and productivity deteriorate

Engineering Contradiction:
Improveaccuracy of text extractionVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent replaces manual mechanical review with an automated computer-based system that uses optical character recognition (OCR) technology to extract text from video frames. The system automatically captures frames, converts them to images, recognizes text through OCR, and stores the extracted content in a searchable database, eliminating the need for manual viewing and transcription while maintaining high accuracy through sophisticated image processing and recognition algorithms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If all video frames are processed for text extraction, then completeness of text extraction is improved, but computational resources and processing time worsen

Engineering Contradiction:
Improvecompleteness of text extractionVSAvoidcomputational resources
Core Design Contradiction:
Loss of informationVSUse of energy by moving object

Solution Approach 1:

The patent segments the video processing task by extracting and analyzing only key frames rather than every single frame. The system identifies frames that contain potential text content based on motion detection and frame difference analysis, processing only those segments that are likely to contain textual information. This selective approach maintains completeness of text extraction while significantly reducing computational resource requirements compared to processing all frames.

Inventive Principle:
Principle #1Segmentation

3Reliability

If video content is stored in original format, then quality of video is preserved, but searchability and accessibility of textual information deteriorate

Engineering Contradiction:
Improvequality of videoVSAvoidsearchability of text
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent creates a textual copy of the video content through OCR extraction without altering the original video file. The system generates text transcripts and metadata that are stored separately in a database, allowing users to search and access textual information independently while the original high-quality video remains unchanged. This copying approach enables full-text search, indexing, and quick retrieval of specific content without compromising video quality or requiring video playback.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12430932B2Video frame type classification for a communication session
Publication Date: 2025.09.30 ZOOM COMMUNICATIONS INC
  • US12430932B2 patent drawing
  • US12430932B2 patent drawing
  • US12430932B2 patent drawing

AI summary

Methods and systems provide for providing video frame type classification in a communication session. In one embodiment, the system receives video content of a communication session with a number of participants; extracts frames from the video content; classifies the frames of the video content based on image analysis; and transmits, to one or more client devices, the classification of the frames of the video content.