Sign Language Video Gloss Segmentation for Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current deep learning-based sign language recognition techniques fail to effectively recognize sign language sentences as they treat the sentence as a continuous sequence, leading to unsatisfactory recognition performance, despite requiring extensive training data.
Innovation Solution
A method for segmenting sign language videos by gloss using an AI model trained on segmented video data, where the model estimates segmentation probability distributions to identify and confirm gloss boundaries, generating a video sequence segmented by gloss, and utilizing both real and virtual training data to enhance recognition accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If sign language recognition uses End-to-End training method to directly generate sign language from video, then the system can process continuous video input, but the recognition performance is unsatisfactory and requires extensive training data
Solution Approach 1:
The patent segments the continuous sign language video into multiple video clips based on gloss boundaries. Each video clip corresponds to a specific gloss unit, enabling the model to process discrete linguistic units rather than continuous video streams. This segmentation approach improves recognition performance by aligning video processing with the linguistic structure of sign language, while reducing the need for extensive training data through more efficient learning of segmented units
Solution Approach 2:
The patent introduces gloss as an intermediary layer between video input and sign language recognition. The model first recognizes gloss units from video clips, then combines these gloss sequences to achieve sign language recognition. This intermediary approach decouples the complex video-to-sign-language mapping into two simpler stages: video-to-gloss and gloss-to-sign-language, improving overall recognition performance while reducing training data requirements
2Measurement precision
If sign language sentence is treated as continuous sequence, then the video can be processed in lump, but the recognition accuracy is insufficient
Solution Approach 1:
The patent divides the continuous sign language video into discrete video clips, each corresponding to a gloss unit. This segmentation enables precise recognition at the gloss level, which then contributes to improved overall sentence recognition accuracy. The segmented approach processes smaller, manageable units rather than attempting to recognize the entire continuous sentence at once
Solution Approach 2:
The patent performs preliminary segmentation of the video into gloss-based clips before final sign language recognition. By pre-segmenting the video according to linguistic boundaries, the system prepares structured input that facilitates more accurate recognition in subsequent processing stages, improving overall accuracy without proportionally increasing complexity
Data Source
AI summary
There are provided a method for segmenting a sign language video by gloss to recognize a sign language sentence, and a method for training. According to an embodiment, a sign language video segmentation method receives an input of a sign language sentence video, and segments the inputted sign language sentence video by gloss. Accordingly, there is suggested a method for segmenting a sign language sentence video by gloss, analyzing various gloss sequences from the linguistic perspective, understanding meanings robustly in spite of various changes in sentences, and translating sign language into appropriate Korean sentences.


