Multi-Granularity Feature Extraction for Video Content Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current content retrieval models face challenges in precision due to significant semantic deviations between retrieval content features and result content features, leading to decreased retrieval accuracy, especially when there are large differences between these features.

Innovation Solution

A method and apparatus for improving model training in content retrieval models by performing feature extraction and quantification across multiple granularities, calculating semantic loss, and aligning features to reduce modal differences, thereby enhancing retrieval precision.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If feature extraction is performed using conventional content retrieval models, then the model can retrieve content based on content features, but significant semantic deviations occur between retrieval content features and result content features, leading to decreased retrieval precision

Engineering Contradiction:
Improveretrieval precisionVSAvoidsemantic deviation
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent segments the feature extraction process into multiple granularities (coarse-grained and fine-grained features). By dividing the feature representation into different levels of detail, the model can capture both overall semantic meaning and specific details, reducing semantic deviation between retrieval content and result content features while maintaining retrieval precision.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new dimension of feature representation by extracting features at multiple granularities simultaneously. This multi-granularity approach adds dimensional complexity to the feature space, allowing the model to better align retrieval content features with result content features and reduce semantic information loss.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If single-granularity feature extraction is used, then the model structure remains simple, but retrieval accuracy decreases when large differences exist between retrieval content feature and result content feature

Engineering Contradiction:
Improveretrieval accuracyVSAvoidfeature extraction complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments feature extraction into multiple granularity levels (coarse and fine), allowing the model to handle large differences between retrieval and result content features by matching at appropriate granularity levels. This segmentation improves retrieval accuracy without requiring excessive complexity in the feature extraction process.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different extraction strategies for different granularity levels, with coarse-grained features capturing overall semantics and fine-grained features capturing specific details. This local quality approach allows the model to adapt to varying feature differences in retrieval scenarios, improving accuracy while maintaining manageable complexity through structured processing.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20240256601A1Model training method and apparatus, computer device, and storage medium
Publication Date: 2024.08.01 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US20240256601A1 patent drawing
  • US20240256601A1 patent drawing
  • US20240256601A1 patent drawing

AI summary

Embodiments of this application disclose a video content retrieval method performed by a computer device. The method includes: obtaining a query text; performing feature extraction processing on the query text through a video content retrieval model, to obtain a plurality of text content features at different feature granularities; calculating, based on the text content feature of each feature granularity, a similarity corresponding to the query and a candidate video content retrieval result at the corresponding feature granularity; and determining, based on the similarities at different feature granularities, a video content retrieval result corresponding to the query text. The solution may improve model training for the content retrieval model and improve content retrieval precision of the content retrieval model.