Sign Language Recognition Model Training via Data Extraction and Anonymization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in collecting and utilizing data for sign language recognition while addressing privacy concerns, particularly in environments like video relay services where data is private and sensitive.
Innovation Solution
A method involving obtaining video data from video communication sessions, procuring sign language objects from the video data, associating these objects with communication data, and constructing a sign language recognition model using these elements. The method ensures privacy by anonymizing and encrypting the data, and deleting original content to prevent reconstruction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If video data from communication sessions is collected for sign language recognition, then the quality and quantity of training data is improved, but privacy concerns and data security risks worsen
Solution Approach 1:
The patent extracts only the essential sign language features (hand gestures, facial expressions, body movements) from the video data while removing or anonymizing personally identifiable information. This extraction process separates the useful training content from the privacy-sensitive elements, allowing data collection without compromising user privacy.
Solution Approach 2:
The patent creates synthetic sign language data through computer-generated avatars and virtual environments. These synthetic copies replicate real sign language patterns and linguistic structures without using actual user video footage, thereby eliminating privacy concerns while providing abundant training data for model development.
2Measurement precision
If comprehensive video data is processed to extract sign language features, then the accuracy of recognition is improved, but computational complexity and processing time worsen
Solution Approach 1:
The patent divides the video processing task into multiple independent stages: initial frame sampling, key feature detection (hands, face, body), feature extraction, and model training. By segmenting the processing pipeline and applying different levels of analysis to different parts of the video data, the system achieves high recognition accuracy without overwhelming computational requirements.
Solution Approach 2:
The patent implements progressive processing where only essential features are extracted at each stage. Rather than analyzing every pixel and motion vector in the video, the system focuses on partial but sufficient features (hand positions, facial expressions) that provide adequate recognition accuracy with reduced computational burden.
Data Source
AI summary
A method may include obtaining, at a first device, video data of a video communication session between a first user of the first device and a second user of a second device. The video data may include sign language content. The method may also include procuring multiple sign language objects. Each of the sign language objects may corresponding to a video segment of the video data that includes multiple video frames, and including features of the video segment that conveys one or more word. The method may further include obtaining, at the first device, communication data representing the sign language content in the video data and associating, each of the sign language objects with a different portion of the communication data. The method may further include constructing a sign language recognition model using the sign language objects, the communication data, and the association therebetween.


