Video Feature Extraction Model Training Using Temporal Sample Images

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing feature extraction models for video processing fail to accurately capture changes in video data over time, leading to poor anti-noise performance and low identification accuracy, especially when videos deform.

Innovation Solution

The method involves selecting at least two images from a sample video that depict the same object at different times, allowing the model to learn temporal changes and enhancing both global and partial information, thereby improving anti-noise performance in the time dimension.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional feature extraction models are used without considering temporal changes, then the model training is simple and fast, but the anti-noise performance in time dimension is poor and identification accuracy is low

Engineering Contradiction:
Improveanti-noise performanceVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent transitions from spatial-only feature extraction to spatiotemporal feature extraction by adding the time dimension. Sample images are selected from multiple time points to capture temporal changes of objects, transforming the feature extraction from 2D spatial features to 3D spatiotemporal features, thereby improving anti-noise performance without excessive complexity increase

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent performs preliminary action by pre-selecting sample images that capture temporal changes before model training. The sample image selection process identifies images at different time points showing object changes in advance, preparing quality training data that enables the model to learn temporal patterns, improving reliability before actual training occurs

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If sample images are selected without considering temporal dimension, then the training process is simple, but the accuracy of extracted video features is affected

Engineering Contradiction:
Improvefeature extraction accuracyVSAvoidtraining time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial action by selecting only key sample images at specific time points rather than using all video frames. This selective sampling captures essential temporal changes while avoiding redundant processing, improving feature extraction accuracy without proportionally increasing training time

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent extracts temporal information by selectively taking out sample images that represent temporal changes from the video sequence. The sample image selection module extracts key frames showing object changes at different times, separating temporal features from spatial features, thereby improving measurement precision while controlling training time

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If at least two images including the same object indicating temporal changes are selected as sample images, then anti-noise performance in time dimension is improved, but the complexity of sample image selection increases

Engineering Contradiction:
Improvetemporal anti-noise performanceVSAvoidsample image selection difficulty
Core Design Contradiction:
ReliabilityVSDifficulty of detecting and measuring

Solution Approach 1:

The sample image selection module performs self-service by automatically selecting sample images based on temporal change detection without manual intervention. The system autonomously identifies images showing object changes at different time points and selects them as training samples, improving temporal anti-noise performance while managing selection complexity through automation

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent implements feedback by using temporal change detection results to guide sample image selection. The system detects temporal changes in video sequences, uses this feedback information to identify relevant sample images, and continuously refines the selection process, thereby improving reliability while making the complex selection process more systematic and manageable

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11538246B2Method and apparatus for training feature extraction model, computer device, and computer-readable storage medium
Publication Date: 2022.12.27 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US11538246B2 patent drawing
  • US11538246B2 patent drawing
  • US11538246B2 patent drawing

AI summary

Aspects of the disclosure provide a method and an apparatus for training a feature extraction model, a computer device, and a computer-readable storage medium that belong to the field of video processing technologies. The method can include detecting a plurality of images in one or more sample videos and obtaining at least two images including the same object. The method can further include determining the at least two images including the same object as sample images, and training according to the determined sample images to obtain the feature extraction model, where the feature extraction model is used for extracting a video feature of the video.