Video Segmentation for Task-Specific Step Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users face challenges in finding step-by-step video instructions for completing tasks, as they are often presented with multiple sources of information that are not specifically tailored to their needs.

Innovation Solution

A method and apparatus for identifying a video for completing a task and determining a plurality of video segments based on attributes of the task, involving how-to queries, confidence measures, and association of video segments with task attributes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If multiple video sources are presented to users, then information completeness is improved, but user confusion and time to find relevant information increases

Engineering Contradiction:
Improveinformation completenessVSAvoidtime to find relevant information
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent segments the full video into multiple discrete time segments based on task steps, allowing users to access specific portions of the video corresponding to particular tasks without watching the entire video. This is achieved by dividing the video timeline into segments associated with different task steps, enabling selective information retrieval.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts relevant video segments from the full video based on task attributes and presents only the necessary portions to users. The system identifies and extracts specific time segments that correspond to task steps matching the user's query, removing unnecessary content while preserving essential information.

Inventive Principle:
Principle #2Taking out (Extraction)

2Ease of operation

If videos are segmented into multiple parts, then user navigation and relevance are improved, but system complexity increases

Engineering Contradiction:
Improveuser navigationVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent performs preliminary segmentation of the video into task-step associated segments during video processing, before the user needs to access the video. The system pre-divides the video into segments and pre-associates them with task steps, so that when a user queries for video information, the segmented structure is already in place and can be quickly retrieved without complex real-time processing.

Inventive Principle:
Principle #10Preliminary action

3Loss of information

If video content is customized to task attributes, then information relevance is improved, but processing requirements and time increase

Engineering Contradiction:
Improveinformation relevanceVSAvoidprocessing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent performs preliminary analysis of video content and association with task attributes during video processing, before user queries are received. The system pre-processes videos to identify task steps and create segmented structures with metadata, so that when users make queries, the relevant information is already organized and can be quickly retrieved without extensive real-time processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12271420B1Video segments for a video related to a task
Publication Date: 2025.04.08 GOOGLE LLC
  • US12271420B1 patent drawing
  • US12271420B1 patent drawing
  • US12271420B1 patent drawing

AI summary

Methods and apparatus related to identifying a video for completing a task and determining a plurality of video segments of the identified video based on one or more attributes of the task. A task and a plurality of how-to videos related to the task may be identified. A how-to video may be selected and a plurality of video segments of the selected how-to video may be determined. One or more video segments may be associated with one or more task attributes that relate to performing the task. The selected video may be provided to a user and segmented, indexed, and/or annotated based on the associated video segments. In some implementations a given object utilized in performing the task may be identified and one or more video segments corresponding to the given object may be identified and/or provided to the user.