Audio-Visual Video Editing for Accurate Singing Segment Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing video editing methods struggle to accurately extract and generate short videos focused on singing segments from longer videos, particularly in live broadcasts, without relying on fixed song durations and distinguishing between singing, speaking, and background music.
Innovation Solution
A video editing method that involves segmenting a video into labeled segments for singing, speaking, background music, and other segments, and then generating a new video based on consecutive singing segments, with optional boundary adjustments and extensions using audio and visual features to enhance accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional video editing methods are used to extract singing segments, then the editing process can be completed, but the accuracy of identifying singing segments is low and cannot distinguish between singing, speaking, and background music
Solution Approach 1:
The video is divided into multiple segments, and each segment is independently labeled as singing, speaking, background music, or other content. This segmentation approach enables precise identification of singing segments without requiring complex global analysis of the entire video.
Solution Approach 2:
The patent introduces an audio-visual analysis module as an intermediary that processes audio and video features separately and combines them to determine segment labels. This intermediary approach improves identification accuracy while keeping the overall system structure manageable.
2Adaptability or versatility
If fixed song durations are used for video editing, then the editing process is simplified, but the flexibility to handle varying song lengths and live broadcast content is reduced
Solution Approach 1:
The video editing system dynamically adjusts segment boundaries based on actual audio-visual content rather than using fixed durations. The labeling process adapts to varying song lengths and content types, enabling the system to handle both standard songs and live broadcast content with flexibility.
3Manufacturing precision
If detailed audio and visual analysis is performed to accurately identify singing segments, then the extraction precision is improved, but the processing time increases
Solution Approach 1:
The video is divided into multiple segments, and each segment is independently labeled as singing, speaking, background music, or other content. This segmentation approach enables precise identification of singing segments without requiring complex global analysis of the entire video.
Solution Approach 2:
The system performs detailed audio-visual analysis on segmented portions of the video rather than the entire video at once. This partial action approach maintains high precision in singing segment extraction while reducing overall processing time through parallel processing of segments.
Data Source
AI summary
Provided are a video editing method, an electronic device and a medium. The video editing method includes: acquiring a first video; cutting the first video to obtain a plurality of segments; determining a plurality of labels respectively corresponding to the plurality of segments, each label among the plurality of labels selected from one of a first label, a second label, a third label or a fourth label, where the first label indicates singing, where the second label indicates speaking, where the third label indicates background music, and where the fourth label indicates a segment that does not correspond to the first label, the second label or the third label; determining a singing segment set based on the plurality of labels, the singing segment set including consecutive segments among the plurality of segments that correspond to the first label; and generating a second video based on the singing segment set.


