Text-Driven Video Processing for Intent-Aware Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video processing technologies struggle to accurately interpret user intent and generate metadata that reflects user requests for video editing, leading to inefficiencies in video processing.

Innovation Solution

A video processing device that includes an interpretation unit to generate query information from text input, a search unit to find relevant video elements, and a generation unit to create digest videos based on user-defined metadata, allowing for precise video editing and generation of digest videos that reflect user intent.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional video processing technology is used, then basic video editing is possible, but the system cannot accurately interpret user intent and generate metadata that reflects user requests

Engineering Contradiction:
Improveaccuracy of user intent interpretationVSAvoidreliability of metadata generation
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent introduces an intermediary mechanism between user input and video processing. The interpretation unit acts as a mediator that receives user requests, generates query information, and passes it to the search unit. This intermediary layer enables accurate intent interpretation by transforming vague user requests into structured search queries, thereby resolving the contradiction between basic editing capability and accurate intent understanding.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback through the metadata generation process. The search unit retrieves video elements based on query information, and the metadata generation unit creates metadata that reflects the actual search results and user intent. This feedback loop ensures that the metadata accurately represents both the user's original request and the processed video content, improving both measurement precision and reliability.

Inventive Principle:
Principle #23Feedback

2Manufacturing precision

If comprehensive video element search is performed, then accurate metadata can be generated, but the processing time and complexity increase

Engineering Contradiction:
Improveprecision of metadata generationVSAvoidvideo processing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent segments the video processing task into distinct functional units: interpretation unit for query generation, search unit for video element retrieval, and metadata generation unit for metadata creation. Each unit handles a specific portion of the processing workload, allowing parallel execution and reducing overall processing time while maintaining precision through specialized processing at each stage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The interpretation unit performs preliminary action by generating query information from user requests before the actual video search begins. This preliminary processing of user intent into structured queries optimizes the subsequent search process, enabling faster and more accurate retrieval of video elements without increasing overall processing time.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If user-specific query generation is implemented, then video processing accuracy improves, but device complexity increases

Engineering Contradiction:
Improveadaptability to user requestsVSAvoidcomplexity of processing system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The interpretation unit is designed with multi-functionality to handle various types of user requests and generate appropriate query information. This universal component can process different user intents through a unified mechanism, improving adaptability without proportionally increasing complexity. The same interpretation framework adapts to different query types by adjusting parameters rather than requiring separate processing paths.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250335505A1Video processing device, video processing method, and computer program product
Publication Date: 2025.10.30 KK TOSHIBA
  • US20250335505A1 patent drawing
  • US20250335505A1 patent drawing
  • US20250335505A1 patent drawing

AI summary

According to an embodiment, a video processing device receives text data representing a scene included in input first video data. The video processing device includes one or more hardware processors configured to function as an interpretation unit and a search unit. The interpretation unit interprets the text data and generates query information used to search for a video element of the first video data. The search unit searches for the video element of the first video data using the query information, generates metadata of the first video data based on a search result, and stores the metadata in a storage unit.