Video Cutting via Speech-to-Text Semantic Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional video cutting methods require professional knowledge and are time-consuming, as they rely on image analysis which is ineffective for extracting course content from videos featuring a person lecturing with minimal image changes.

Innovation Solution

A method that extracts text information from videos using speech-to-text technology and semantic segmentation models to identify candidate paragraph segmentation positions, timestamps, and cut videos into relevant clips based on preset content, automating the video cutting process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If traditional image analysis methods are used for video cutting, then the cutting process can be automated to some extent, but the method is ineffective for extracting course content from videos with minimal image changes

Engineering Contradiction:
Improvevideo cutting automationVSAvoidcourse content extraction accuracy
Core Design Contradiction:
Extent of automationVSReliability

Solution Approach 1:

The patent replaces image analysis (visual/mechanical system) with speech-to-text conversion and semantic analysis (acoustic and linguistic systems). By converting video audio to text and analyzing text semantics, the system effectively extracts course content from lectures with minimal visual changes, resolving the contradiction between automation and reliability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If professional editing staff manually cut videos based on visual content, then the cutting quality can be ensured, but the process is time-consuming and laborious

Engineering Contradiction:
Improvevideo cutting qualityVSAvoidvideo cutting efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent implements self-service by enabling the video cutting system to automatically identify course content segments through speech-to-text conversion and semantic analysis without requiring professional editors. The system autonomously determines cutting points based on text semantics, maintaining quality while dramatically improving productivity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent changes the analysis parameter from visual features to text semantic features. By extracting and analyzing text content from video speech, the system identifies meaningful segments based on semantic boundaries rather than visual changes, achieving both high quality and efficiency in video cutting.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If speech-to-text technology is used to extract text from video, then text-based content analysis becomes possible, but additional processing steps are required

Engineering Contradiction:
Improvecontent analysis capabilityVSAvoidprocessing pipeline complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by performing speech-to-text conversion at the beginning of the processing pipeline. The extracted text is then reused throughout subsequent semantic analysis and segment identification steps, avoiding repeated transcription and simplifying the overall processing despite the initial conversion step.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11954912B2Method for cutting video based on text of the video and computing device applying method
Publication Date: 2024.04.09 PING AN TECH (SHENZHEN) CO LTD
  • US11954912B2 patent drawing
  • US11954912B2 patent drawing
  • US11954912B2 patent drawing

AI summary

A method for cutting or extracting video clips from a video, including the audio content relevant to points of particular interest, and combining the same for instruction or training on particular points; a computing device applying the method extracts text information from the spoken audio content of a video to be cut and obtains multiple paragraph segmentation positions as candidates for inclusion in a desired and finished presentation by analyzing the information from text representing the spoken audio content, the analysis being carried out by a semantic segmentation model. Candidate items of text are obtained by isolating pieces of text according to the paragraph segmentation positions. Time stamps of the candidate text segments are acquired, and candidate video clips are obtained by cutting the video according to the acquired time stamps.