Sign Language Video Gloss Segmentation for Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current deep learning-based sign language recognition techniques fail to effectively recognize sign language sentences as they treat the sentence as a continuous sequence, leading to unsatisfactory recognition performance, despite requiring extensive training data.

Innovation Solution

A method for segmenting sign language videos by gloss using an AI model trained on segmented video data, where the model estimates segmentation probability distributions to identify and confirm gloss boundaries, generating a video sequence segmented by gloss, and utilizing both real and virtual training data to enhance recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If sign language recognition uses End-to-End training method to directly generate sign language from video, then the system can process continuous video input, but the recognition performance is unsatisfactory and requires extensive training data

Engineering Contradiction:
Improverecognition performanceVSAvoidtraining data volume
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments the continuous sign language video into multiple video clips based on gloss boundaries. Each video clip corresponds to a specific gloss unit, enabling the model to process discrete linguistic units rather than continuous video streams. This segmentation approach improves recognition performance by aligning video processing with the linguistic structure of sign language, while reducing the need for extensive training data through more efficient learning of segmented units

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces gloss as an intermediary layer between video input and sign language recognition. The model first recognizes gloss units from video clips, then combines these gloss sequences to achieve sign language recognition. This intermediary approach decouples the complex video-to-sign-language mapping into two simpler stages: video-to-gloss and gloss-to-sign-language, improving overall recognition performance while reducing training data requirements

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If sign language sentence is treated as continuous sequence, then the video can be processed in lump, but the recognition accuracy is insufficient

Engineering Contradiction:
Improverecognition accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent divides the continuous sign language video into discrete video clips, each corresponding to a gloss unit. This segmentation enables precise recognition at the gloss level, which then contributes to improved overall sentence recognition accuracy. The segmented approach processes smaller, manageable units rather than attempting to recognize the entire continuous sentence at once

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary segmentation of the video into gloss-based clips before final sign language recognition. By pre-segmenting the video according to linguistic boundaries, the system prepares structured input that facilitates more accurate recognition in subsequent processing stages, improving overall accuracy without proportionally increasing complexity

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11798255B2Sign language video segmentation method by gloss for sign language sentence recognition, and training method therefor
Publication Date: 2023.10.24 KOREA ELECTRONICS TECH INST
  • US11798255B2 patent drawing
  • US11798255B2 patent drawing
  • US11798255B2 patent drawing

AI summary

There are provided a method for segmenting a sign language video by gloss to recognize a sign language sentence, and a method for training. According to an embodiment, a sign language video segmentation method receives an input of a sign language sentence video, and segments the inputted sign language sentence video by gloss. Accordingly, there is suggested a method for segmenting a sign language sentence video by gloss, analyzing various gloss sequences from the linguistic perspective, understanding meanings robustly in spite of various changes in sentences, and translating sign language into appropriate Korean sentences.