Speech-Based Video Segmentation for Automated Semantic Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video production methods are expensive, time-consuming, and require technical expertise, and there is a need for systems and computerized methods to automatically segment video data for easier organization and use.

Innovation Solution

A method involving a series of machine learning models to receive and process video segments, producing text data, categorized text data, and semantic vectors, which are used to store and retrieve video segments based on metadata and search queries.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional video production methods are used, then video quality and control are maintained, but production cost and time consumption increase significantly

Engineering Contradiction:
Improvevideo production efficiencyVSAvoidsystem complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system automatically segments video data, extracts speech content, generates text transcripts, and organizes video segments without human intervention. The automated video generation system selects and assembles segments based on text inputs, enabling self-service video production that eliminates the need for manual editing and production workflows

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces traditional mechanical video production processes (manual editing, scripting, and assembly) with machine learning models and automated systems. The first ML model transcribes speech to text, the second ML model generates video segments from text, and the third ML model assembles final videos, substituting human-operated mechanical processes with automated intelligent systems

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If manual video organization is used, then video content can be arranged, but time and effort required for organization increase

Engineering Contradiction:
Improvevideo organization easeVSAvoidorganization time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system automatically generates text transcripts from video speech, extracts key information, and organizes video segments into searchable databases with metadata. This self-organizing capability eliminates manual video cataloging and makes retrieval instant through text-based queries rather than manual browsing

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent introduces text data as an intermediary between video content and user queries. The text transcript serves as a searchable index that mediates between the user's text input and the video database, enabling efficient retrieval without manual organization or direct video search

Inventive Principle:
Principle #24Intermediary (Mediator)

3Productivity

If automated video segmentation is implemented, then productivity increases, but technical complexity of the system increases

Engineering Contradiction:
Improvevideo segmentation speedVSAvoidmachine learning system complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent divides the complex video processing task into three distinct ML model segments: Model 1 for speech-to-text transcription, Model 2 for text-based video segment generation, and Model 3 for video assembly. This segmentation of the AI system itself makes the overall complexity manageable by breaking it into specialized, independent components

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The machine learning models perform multiple functions: they transcribe speech, generate text descriptions, create video segments, and assemble final videos. This multi-functionality reduces the need for separate specialized systems, managing complexity by having versatile models handle diverse tasks

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12347462B1Methods and systems for segmenting video content based on speech data and for retreiving video segments to generate videos
Publication Date: 2025.07.01 VIDEOFORCEAI INC
  • US12347462B1 patent drawing
  • US12347462B1 patent drawing
  • US12347462B1 patent drawing

AI summary

A method includes receiving a series of video segments and providing the series of video segments as input to a first machine learning model to produce text data. The text data is provided as input to a second machine learning model to produce categorized text data that includes a classification indication. The classification indication is added to metadata of the video segment, and the categorized text data is provided as input to a third machine learning model to produce a semantic vector. The method also includes causing the video segment and the metadata that includes the classification indication to be stored at a location of a database based on the semantic vector, the database being configured to be searched based on a search query associated with the semantic vector.