Video Editing Automation via Speech Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video editing methods are inefficient in identifying and removing ineffective segments from oral broadcast videos, such as pauses and stutters, as they rely on audio waveform analysis, which is time-consuming.
Innovation Solution
A method and apparatus for video editing that uses speech recognition to determine ineffective text and its timeline position in a video, allowing for precise identification and deletion of these segments on an editing interface, improving editing efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If audio waveform analysis is used to identify ineffective segments, then the editing can be performed, but the editing efficiency is low and time consumption is high
Solution Approach 1:
The patent replaces the mechanical audio waveform analysis method with an automated speech recognition system. The speech recognition module automatically identifies ineffective segments (pauses, stutters, verbiage) by processing speech text and timing information, eliminating the need for manual waveform inspection and significantly reducing time consumption while improving editing efficiency
Solution Approach 2:
The system enables self-service editing by automatically generating speech text, identifying ineffective segments, and marking them on the timeline without requiring manual intervention. The automated identification process uses speech recognition results and timing data to autonomously locate and highlight ineffective segments for deletion
2Ease of operation
If manual editing methods are used to remove ineffective segments, then the editing can be performed, but the editing process is complex and time-consuming
Solution Approach 1:
The patent replaces manual editing operations with automated speech recognition-based identification. The system automatically processes audio to generate speech text, determines timing information, identifies ineffective segments, and marks them on the timeline, eliminating complex manual editing steps and reducing the time required for video editing operations
Solution Approach 2:
The patent introduces speech text and timing information as intermediary elements that bridge the gap between raw audio and visual editing. These intermediaries enable automated identification of ineffective segments by providing structured data about speech content and temporal positioning, simplifying the editing process
Data Source
AI summary
Embodiments of the present disclosure provide a method and apparatus of video editing, an electronic device and a storage medium. The method comprises: determining an ineffective text in a speech text of a target video and a timeline position of the ineffective text by performing speech recognition on an audio in the target video; presenting an editing track segment of the target video on an editing interface of the target video, and identifying a timeline interval of the ineffective text on the editing track segment based on the timeline position of the ineffective text; in response to an adjustment operation on the ineffective text, adjusting the timeline interval of the ineffective text on the editing track segment of the target video; in response to a video segment deleting operation on the ineffective text, deleting a video segment of the editing track segment within the timeline interval of the ineffective text from the target video.


