Video Editing Automation via Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video editing methods are inefficient in identifying and removing ineffective segments from oral broadcast videos, such as pauses and stutters, as they rely on audio waveform analysis, which is time-consuming.

Innovation Solution

A method and apparatus for video editing that uses speech recognition to determine ineffective text and its timeline position in a video, allowing for precise identification and deletion of these segments on an editing interface, improving editing efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If audio waveform analysis is used to identify ineffective segments, then the editing can be performed, but the editing efficiency is low and time consumption is high

Engineering Contradiction:
Improvevideo editing efficiencyVSAvoidtime consumption for identifying ineffective segments
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces the mechanical audio waveform analysis method with an automated speech recognition system. The speech recognition module automatically identifies ineffective segments (pauses, stutters, verbiage) by processing speech text and timing information, eliminating the need for manual waveform inspection and significantly reducing time consumption while improving editing efficiency

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service editing by automatically generating speech text, identifying ineffective segments, and marking them on the timeline without requiring manual intervention. The automated identification process uses speech recognition results and timing data to autonomously locate and highlight ineffective segments for deletion

Inventive Principle:
Principle #25Self-service

2Ease of operation

If manual editing methods are used to remove ineffective segments, then the editing can be performed, but the editing process is complex and time-consuming

Engineering Contradiction:
Improveease of video editing operationVSAvoidtime required for manual editing operations
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent replaces manual editing operations with automated speech recognition-based identification. The system automatically processes audio to generate speech text, determines timing information, identifies ineffective segments, and marks them on the timeline, eliminating complex manual editing steps and reducing the time required for video editing operations

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent introduces speech text and timing information as intermediary elements that bridge the gap between raw audio and visual editing. These intermediaries enable automated identification of ineffective segments by providing structured data about speech content and temporal positioning, simplifying the editing process

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12051445B1Method and apparatus of video editing, and electronic device and storage medium
Publication Date: 2024.07.30 BEIJING ZITIAO NETWORK TECH CO LTD
  • US12051445B1 patent drawing
  • US12051445B1 patent drawing
  • US12051445B1 patent drawing

AI summary

Embodiments of the present disclosure provide a method and apparatus of video editing, an electronic device and a storage medium. The method comprises: determining an ineffective text in a speech text of a target video and a timeline position of the ineffective text by performing speech recognition on an audio in the target video; presenting an editing track segment of the target video on an editing interface of the target video, and identifying a timeline interval of the ineffective text on the editing track segment based on the timeline position of the ineffective text; in response to an adjustment operation on the ineffective text, adjusting the timeline interval of the ineffective text on the editing track segment of the target video; in response to a video segment deleting operation on the ineffective text, deleting a video segment of the editing track segment within the timeline interval of the ineffective text from the target video.