Multi-Granularity Feature Extraction for Video Content Retrieval
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current content retrieval models face challenges in precision due to significant semantic deviations between retrieval content features and result content features, leading to decreased retrieval accuracy, especially when there are large differences between these features.
Innovation Solution
A method and apparatus for improving model training in content retrieval models by performing feature extraction and quantification across multiple granularities, calculating semantic loss, and aligning features to reduce modal differences, thereby enhancing retrieval precision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If feature extraction is performed using conventional content retrieval models, then the model can retrieve content based on content features, but significant semantic deviations occur between retrieval content features and result content features, leading to decreased retrieval precision
Solution Approach 1:
The patent segments the feature extraction process into multiple granularities (coarse-grained and fine-grained features). By dividing the feature representation into different levels of detail, the model can capture both overall semantic meaning and specific details, reducing semantic deviation between retrieval content and result content features while maintaining retrieval precision.
Solution Approach 2:
The patent introduces a new dimension of feature representation by extracting features at multiple granularities simultaneously. This multi-granularity approach adds dimensional complexity to the feature space, allowing the model to better align retrieval content features with result content features and reduce semantic information loss.
2Reliability
If single-granularity feature extraction is used, then the model structure remains simple, but retrieval accuracy decreases when large differences exist between retrieval content feature and result content feature
Solution Approach 1:
The patent segments feature extraction into multiple granularity levels (coarse and fine), allowing the model to handle large differences between retrieval and result content features by matching at appropriate granularity levels. This segmentation improves retrieval accuracy without requiring excessive complexity in the feature extraction process.
Solution Approach 2:
The patent applies different extraction strategies for different granularity levels, with coarse-grained features capturing overall semantics and fine-grained features capturing specific details. This local quality approach allows the model to adapt to varying feature differences in retrieval scenarios, improving accuracy while maintaining manageable complexity through structured processing.
Data Source
AI summary
Embodiments of this application disclose a video content retrieval method performed by a computer device. The method includes: obtaining a query text; performing feature extraction processing on the query text through a video content retrieval model, to obtain a plurality of text content features at different feature granularities; calculating, based on the text content feature of each feature granularity, a similarity corresponding to the query and a candidate video content retrieval result at the corresponding feature granularity; and determining, based on the similarities at different feature granularities, a video content retrieval result corresponding to the query text. The solution may improve model training for the content retrieval model and improve content retrieval precision of the content retrieval model.


