Topic-Based Video Navigation Using Semantic Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video compression techniques are computationally intensive and struggle to keep up with demands for higher video quality and resolution on memory-constrained devices, while also lacking efficient navigation methods that allow users to move through videos based on parameters other than time.

Innovation Solution

A computer-implemented method and system for navigating through video content using a video navigation system that receives an encoded video file and a transcript, automatically identifies topics using a natural language model, and generates a user interface with topic-based navigation options, allowing users to select topics and playback sessions based on associated video segments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If conventional video compression techniques are used, then video size is reduced, but computational complexity increases and video quality deteriorates on memory-constrained devices

Engineering Contradiction:
Improvevideo sizeVSAvoidcomputational complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The video is segmented into discrete codes representing individual frames or groups of frames. Each code captures essential visual information in a compressed form, allowing selective reconstruction of video segments without processing the entire video stream, thus reducing computational complexity while maintaining video size reduction benefits

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of storing and processing complete video frames, the system creates compressed representations (codes) that copy only the essential visual information. These codes can be stored efficiently and reconstructed when needed, reducing both storage requirements and computational processing while preserving video quality on memory-constrained devices

Inventive Principle:
Principle #26Copying

2Quantity of substance

If conventional video compression techniques are used, then video size is reduced, but video quality and resolution deteriorate on memory-constrained devices

Engineering Contradiction:
Improvevideo sizeVSAvoidvideo quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The compression technique applies different levels of detail to different parts of the video content. Important regions and features are preserved with higher fidelity while less critical areas are compressed more aggressively, maintaining overall video quality while achieving size reduction suitable for memory-constrained devices

Inventive Principle:
Principle #3Local quality

3Ease of operation

If users navigate through video using time-based methods, then navigation is simple, but navigation efficiency deteriorates when searching for specific content

Engineering Contradiction:
Improvenavigation simplicityVSAvoidnavigation time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The system introduces an intermediary indexing layer between the user and the video content. Codes are generated with metadata and descriptors that act as intermediaries, allowing users to search and navigate to specific video segments based on content characteristics rather than time positions, reducing navigation time while maintaining ease of use through intuitive search interfaces

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12288570B1Conversational AI-encoded language for video navigation
Publication Date: 2025.04.29 NVIDIA CORP
  • US12288570B1 patent drawing
  • US12288570B1 patent drawing
  • US12288570B1 patent drawing

AI summary

Systems and methods of compressing video content as encoded data and selectively reconstructing portions of the content are disclosed. The proposed systems provide a computer-implemented process configured to classify a person's behavior(s) during a video and encode the behaviors as a representation of the video. When playback of the video is requested, a video navigation assistant will allow the end-user to select specific segments of the video based on topics discussed in the video and the codes that were generated to represent the video. The user is then able to move through segments of the video in a sequence that aligns with their viewing preferences.