Video Segment Positioning via Image-Text Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video clipping techniques fail to accurately and quickly position the video segment containing a specific item during the video clip process, particularly in electronic commerce live broadcasts with multiple scenes and redundant information.

Innovation Solution

A method utilizing an image-text matching model to determine a target video frame and corresponding audio segment, combining image and audio data to identify the target video segment, thereby improving positioning accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If current video clipping technique is used, then the video clip process can be performed, but the video segment containing a particular item cannot be accurately and quickly positioned

Engineering Contradiction:
Improvepositioning accuracy of video segmentVSAvoidtime for video segment positioning
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the video processing task into multiple independent components: audio feature extraction, image feature extraction, and fused feature matching. By dividing the complex video segment positioning task into separate audio and visual streams that are processed independently and then combined, the system achieves both speed (through parallel processing) and accuracy (through multi-modal verification) in locating video segments containing particular items

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a temporal dimension by extracting audio features from video audio tracks and comparing them against a database of audio fingerprints. This adds a time-based matching dimension to the traditional spatial/image-only matching approach, enabling faster and more accurate identification of video segments through temporal audio cues that complement visual information

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If current video clipping technique is used, then the video clip process can be performed, but the positioning of video segment containing specific item is inaccurate

Engineering Contradiction:
Improvepositioning accuracy of video segmentVSAvoidcomplexity of video processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system separates audio and image processing into distinct modules with specialized algorithms for each modality. Audio features are extracted through spectrum analysis and compared against audio fingerprints, while image features are extracted through frame analysis and object detection. This segmentation allows each module to be optimized independently, improving overall accuracy without requiring a complete redesign of the entire system

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a feature fusion module that acts as an intermediary between audio and image processing streams. This mediator combines the results from both modalities, weighting and integrating their contributions to produce a final positioning decision. The intermediary layer manages the complexity by providing a standardized interface that coordinates the separate audio and image processing subsystems

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If electronic commerce live broadcast with multiple scenes and redundant information is processed, then comprehensive product coverage is achieved, but video segment positioning becomes slower and less accurate

Engineering Contradiction:
Improvepositioning accuracy in complex videoVSAvoidspeed of video clip process
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent extracts key distinguishing features from the complex live broadcast content by comparing audio and visual features against a database of known product signatures. Instead of analyzing every frame and audio sample in detail, the system extracts salient features that uniquely identify product segments, filtering out redundant information from multiple scenes and hosts. This extraction approach maintains high speed by focusing only on discriminative features while achieving accurate positioning of target video segments

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20240078807A1Method, apparatus,electronic device and storage medium for video processing
Publication Date: 2024.03.07 DOUYIN VISION CO LTD
  • US20240078807A1 patent drawing
  • US20240078807A1 patent drawing
  • US20240078807A1 patent drawing

AI summary

Embodiments of the present disclosure provide method, apparatus, electronic device and storage medium for video processing. The method for video processing comprises: obtaining a plurality of video frames in a video to be processed and audio data corresponding to the video to be processed; determining a target video frame comprising a target object from the plurality of video frames through an image-text matching model; determining a target audio segment matching the target object from the audio data; determining a target video segment comprising the target video frame from the video to be processed in the case that a video corresponding to the target audio segment comprises the target video frame.