Audio-Visual Video Editing for Accurate Singing Segment Extraction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video editing methods struggle to accurately extract and generate short videos focused on singing segments from longer videos, particularly in live broadcasts, without relying on fixed song durations and distinguishing between singing, speaking, and background music.

Innovation Solution

A video editing method that involves segmenting a video into labeled segments for singing, speaking, background music, and other segments, and then generating a new video based on consecutive singing segments, with optional boundary adjustments and extensions using audio and visual features to enhance accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional video editing methods are used to extract singing segments, then the editing process can be completed, but the accuracy of identifying singing segments is low and cannot distinguish between singing, speaking, and background music

Engineering Contradiction:
Improveaccuracy of identifying singing segmentsVSAvoidcomplexity of video editing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The video is divided into multiple segments, and each segment is independently labeled as singing, speaking, background music, or other content. This segmentation approach enables precise identification of singing segments without requiring complex global analysis of the entire video.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an audio-visual analysis module as an intermediary that processes audio and video features separately and combines them to determine segment labels. This intermediary approach improves identification accuracy while keeping the overall system structure manageable.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If fixed song durations are used for video editing, then the editing process is simplified, but the flexibility to handle varying song lengths and live broadcast content is reduced

Engineering Contradiction:
Improveflexibility to handle varying song lengthsVSAvoidsimplicity of video editing process
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The video editing system dynamically adjusts segment boundaries based on actual audio-visual content rather than using fixed durations. The labeling process adapts to varying song lengths and content types, enabling the system to handle both standard songs and live broadcast content with flexibility.

Inventive Principle:
Principle #15Dynamics

3Manufacturing precision

If detailed audio and visual analysis is performed to accurately identify singing segments, then the extraction precision is improved, but the processing time increases

Engineering Contradiction:
Improveprecision of singing segment extractionVSAvoidvideo processing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The video is divided into multiple segments, and each segment is independently labeled as singing, speaking, background music, or other content. This segmentation approach enables precise identification of singing segments without requiring complex global analysis of the entire video.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs detailed audio-visual analysis on segmented portions of the video rather than the entire video at once. This partial action approach maintains high precision in singing segment extraction while reducing overall processing time through parallel processing of segments.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12387762B2Video editing method and apparatus, electronic device and medium
Publication Date: 2025.08.12 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US12387762B2 patent drawing
  • US12387762B2 patent drawing
  • US12387762B2 patent drawing

AI summary

Provided are a video editing method, an electronic device and a medium. The video editing method includes: acquiring a first video; cutting the first video to obtain a plurality of segments; determining a plurality of labels respectively corresponding to the plurality of segments, each label among the plurality of labels selected from one of a first label, a second label, a third label or a fourth label, where the first label indicates singing, where the second label indicates speaking, where the third label indicates background music, and where the fourth label indicates a segment that does not correspond to the first label, the second label or the third label; determining a singing segment set based on the plurality of labels, the singing segment set including consecutive segments among the plurality of segments that correspond to the first label; and generating a second video based on the singing segment set.