Adaptive Shift for Text-Video Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-video retrieval systems perform shift operations in a parameter-free manner, neglecting the semantic context of different videos and leading to sub-optimal shift performance.

Innovation Solution

Implementing adaptive shift machine learning models with a candidate selector and a step selector to dynamically adjust shift granularities based on spatial and temporal dimensions under different visual contexts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If parameter-free shift operations are used, then device complexity is reduced, but video representation quality deteriorates

Engineering Contradiction:
Improveshift operation complexityVSAvoidvideo representation quality
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent introduces learnable shift parameters (shift ratio and shift step) that are dynamically adjusted based on video content semantics. The candidate selector model determines the shift ratio by analyzing frame embeddings, and the step selector model determines the shift step, replacing the fixed parameter-free approach with adaptive parameter control to improve video representation quality

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediate adaptive shift module between frame encoding and video representation. This module includes a candidate selector and step selector that process frame embeddings and generate optimized shift parameters, acting as a mediator that enhances the quality of video representation without directly modifying the core encoding architecture

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If adaptive shift models are implemented, then video representation quality is improved, but device complexity increases

Engineering Contradiction:
Improvevideo representation qualityVSAvoidmodel architecture complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the adaptive shift functionality into two independent models: a candidate selector model that determines shift ratio and candidate positions, and a step selector model that determines shift step. This segmentation allows each model to specialize in specific aspects of shift optimization, improving video representation quality while enabling modular implementation that manages complexity

Inventive Principle:
Principle #1Segmentation

3Reliability

If semantic context is considered, then shift performance is improved, but computation time increases

Engineering Contradiction:
Improveshift performanceVSAvoidcomputation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by computing frame embeddings and analyzing semantic context before executing the shift operation. The candidate selector and step selector models process the embeddings in advance to determine optimal shift parameters, ensuring high shift performance while preparing all necessary computations before the actual video representation generation

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12216709B1Computer-implemented methods for machine learning model based spatial-temporal adaptive shift for end-to-end text-video retrieval
Publication Date: 2025.02.04 AMAZON TECH INC
  • US12216709B1 patent drawing
  • US12216709B1 patent drawing
  • US12216709B1 patent drawing

AI summary

Techniques for performing a machine learning model based spatial-temporal adaptive shift for end-to-end text-video retrieval are described. According to some examples, a computer-implemented method includes receiving a video comprising a plurality of frames at a content delivery service; generating, by the content delivery service, a set of embeddings for each of a plurality of sections of each frame of the plurality of frames; determining, by a candidate selector machine learning model of the content delivery service, a proper subset of the plurality of sections of each frame of the plurality of frames for a time shift based on the set of embeddings; time shifting, by the content delivery service, the proper subset of the plurality of sections of each frame of the plurality of frames to generate time shifted frames; generating, by the content delivery service, an updated set of embeddings based on the time shifted frames; receiving a search request comprising input text from a user device; determining the video is a match for the search request based on the input text and the updated set of embeddings for the time shifted frames; and sending the video to the user device based on the match.