Weakly Supervised Action Segmentation for Unseen Video Sequences

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing supervised and unsupervised learning methods for action segmentation in instructional videos are costly, time-consuming, and require significant manual effort for labeling, while unsupervised learning often produces erroneous results, and existing weakly-supervised methods struggle to predict unseen action sequences and detect anomalies in real-time.

Innovation Solution

A system utilizing a hybrid segmentation model based on an unconstrained Viterbi algorithm that performs feature extraction, generates predicted action scores, and segments videos into sequences of actions, allowing for real-time detection of unseen anomalies by exploring a universe of possible actions when discrepancies arise between anticipated and current representations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If supervised learning is used for action segmentation, then labeling precision is improved, but cost and time consumption increase significantly

Engineering Contradiction:
Improvelabeling precisionVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables self-service by allowing the model to automatically generate action labels and segmentations without requiring human annotators. The weakly-supervised approach uses automatic pseudo-labeling where the system trains on automatically generated labels rather than requiring manual labeling, thereby eliminating time consumption while maintaining reasonable labeling precision through the hybrid segmentation model that combines multiple prediction sources.

Inventive Principle:
Principle #25Self-service

2Loss of time

If unsupervised learning is used for action segmentation, then cost and time are reduced, but result accuracy deteriorates

Engineering Contradiction:
Improvetime consumptionVSAvoidresult accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The system introduces an intermediary approach by using a hybrid segmentation model that mediates between unsupervised and supervised learning methods. The model combines unsupervised feature extraction with weakly-supervised learning using automatic pseudo-labels as intermediaries, allowing the system to benefit from both approaches: the efficiency of unsupervised learning and the accuracy of supervised learning, thereby achieving reasonable result accuracy without the time consumption of fully supervised methods.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Loss of time

If existing weakly-supervised methods are used, then cost and time are reduced, but ability to predict unseen actions deteriorates

Engineering Contradiction:
Improvetime consumptionVSAvoidability to predict unseen actions
Core Design Contradiction:
Loss of timeVSAdaptability or versatility

Solution Approach 1:

The system applies dynamics by implementing a dynamic action transition model that can adapt its behavior based on the input. The model uses an anticipation network that dynamically adjusts predictions based on the current video frame and historical context, allowing it to handle unseen actions by exploring the universe of possible actions when discrepancies arise between anticipated and current representations. This dynamic adaptation enables the system to maintain versatility while reducing time consumption through efficient weakly-supervised learning.

Inventive Principle:
Principle #15Dynamics

4Productivity

If constrained Viterbi algorithm is used, then computation efficiency is improved, but adaptability to unseen sequences is reduced

Engineering Contradiction:
Improvecomputation efficiencyVSAvoidadaptability to unseen sequences
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The system applies segmentation by dividing the action prediction task into multiple segments: feature extraction, action score generation, action transition prediction, and sequence assembly. The hybrid segmentation model processes video frames sequentially, maintaining computation efficiency through structured processing while enabling adaptability to unseen sequences by allowing the action transition model to explore possible actions at each segment rather than being constrained to predetermined sequences.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12548335B2Weakly supervised action segmentation
Publication Date: 2026.02.10 HONDA MOTOR CO LTD
  • US12548335B2 patent drawing
  • US12548335B2 patent drawing
  • US12548335B2 patent drawing

AI summary

According to one aspect, weakly-supervised action segmentation may include performing feature extraction to extract one or more features associated with a current frame of a video including a series of one or more actions, feeding one or more of the features to a recognition network to generate a predicted action score for the current frame of the video, feeding one or more of the features and the predicted action score to an action transition model to generate a potential subsequent action, feeding the potential subsequent action and the predicted action score to a hybrid segmentation model to generate a predicted sequence of actions from a first frame of the video to the current frame of the video, and segmenting or labeling one or more frames of the video based on the predicted sequence of actions from the first frame of the video to the current frame of the video.