Video Processing via Semantic Label Combinations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face difficulties in quickly locating and finding videos of interest due to the vast number of videos available online, necessitating advanced video interpretation and processing technologies to extract key information and generate video highlights.

Innovation Solution

A method and apparatus for video processing that involves semantic recognition across multiple dimensions, generating candidate label combinations, and creating target video clips based on user selections, utilizing video label data and playback time periods to filter and prioritize content for efficient content representation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If semantic recognition is performed across multiple dimensions to improve video content representation precision, then the comprehensiveness of video analysis is improved, but the processing time and computational complexity increase

Engineering Contradiction:
Improvevideo content representation precisionVSAvoidvideo processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The video processing system segments the video into multiple clips based on different semantic dimensions (e.g., action types, objects, scenes). Each dimension is analyzed independently to generate specific label combinations, which are then integrated to form comprehensive video highlights. This segmentation allows parallel processing of different dimensions, improving efficiency while maintaining precision.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary semantic recognition on the entire video to identify potential key moments and generate candidate label combinations before final highlight generation. This preliminary analysis pre-processes the video content, organizing semantic information by dimensions and time periods, which reduces the computational burden during the final highlight compilation stage.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If multiple candidate label combinations are generated to improve content selection accuracy, then the precision of video highlights is improved, but the system complexity and processing overhead increase

Engineering Contradiction:
Improvecontent selection accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system generates multiple candidate label combinations beyond what is strictly necessary, including some redundant or overlapping combinations. This excessive action ensures that the most accurate content selections are captured, even if it means processing additional candidates. The system then filters and ranks these candidates to select the best matches, prioritizing accuracy over minimal processing.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system introduces an intermediary ranking and selection mechanism that evaluates multiple candidate label combinations against the video content. This intermediary layer filters the excessive candidates, ranking them by relevance and selecting the top matches. This mediator manages the complexity by systematically organizing and evaluating candidates rather than requiring direct comparison of all possibilities.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If video clips are extracted based on multiple semantic dimensions and user selections, then the relevance of video highlights to user interests is improved, but the processing time and computational resources increase

Engineering Contradiction:
Improveuser interest matchingVSAvoidvideo processing efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system dynamically adjusts the semantic dimensions and label combinations based on user selections and preferences. When users select specific dimensions or keywords, the system reconfigures the processing to focus on those areas, generating candidate label combinations that are tailored to user interests. This dynamic adaptation maintains high relevance while optimizing processing efficiency by avoiding unnecessary analysis of unrelated dimensions.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11436831B2Method and apparatus for video processing
Publication Date: 2022.09.06 ALIBABA GROUP HOLDING LTD
  • US11436831B2 patent drawing
  • US11436831B2 patent drawing
  • US11436831B2 patent drawing

AI summary

Embodiments of the disclosure provides methods and apparatuses for video processing. In one embodiment, the video processing method comprises: obtaining at least one video from a video repository as a video to be processed; performing semantic recognition on the video in one or more semantic recognition dimensions to obtain one or more video label data items corresponding to the video in the one or more semantic recognition dimensions; generating at least one candidate label combination based on at least one of the one or more video label data items; determining, based on a target label combination selected by a user from the at least one candidate label combination, one or more video clips in the video corresponding to at least one video label in the target label combination; and generating at least one target video clip corresponding to the target label combination based on at least one of the one or more video clips.