AR Sign Language Animation for Multi-Speaker Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sign language translation technologies face difficulties in distinguishing speaking objects in multi-person discussions, leading to poor user experience and impaired communication between individuals with and without hearing impairments.
Innovation Solution
A method and apparatus that utilize real-time voice and video information processing to identify speaking objects by recognizing face images and sound attributes, superimposing augmented reality sign language animations on a gesture area corresponding to the speaking object, enabling clear identification of speaking content in a sign language video.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If existing voice-sign language translation method is used, then translation function is provided, but speaking objects cannot be distinguished in multi-person discussions
Solution Approach 1:
The system segments the translation output by associating each sign language animation with a specific speaking object identified through face recognition and sound attribute analysis. This segmentation allows the hearing-impaired user to distinguish which speaker is corresponding to which sign language translation, resolving the information loss in multi-person discussions.
Solution Approach 2:
The system introduces an intermediary mechanism that captures video information of speaking objects, performs face recognition, extracts sound attributes, and matches them to voice information. This intermediary processing chain enables reliable identification and association of speaking objects with their corresponding speech content.
2Ease of operation
If binoculus pays attention to translated text information at all times, then translation is received, but ability to discern view of each speaking object is lost
Solution Approach 1:
The system transitions from traditional text-based translation to augmented reality sign language animations that are superimposed on the video feed of speaking objects. This dimensional change from 2D text to 3D spatial AR animation allows users to perceive both the speaker's visual presence and the translation simultaneously, eliminating the need to switch attention between text and speakers.
Solution Approach 2:
The system merges the sign language animation with the video information of the speaking object in an augmented reality display. This combination allows the hearing-impaired user to see both the speaker and the translation together in one view, maintaining the ability to discern speaking objects while receiving accurate translation information.
3Measurement precision
If face recognition and sound attribute matching is performed, then speaking object identification is achieved, but processing complexity increases
Solution Approach 1:
The system performs preliminary actions by pre-processing video information to extract face images and pre-processing audio information to extract sound attributes. These preliminary processing steps organize the data in advance, making the subsequent matching process more efficient and manageable, thus reducing overall system complexity.
Solution Approach 2:
The system employs universal processing modules that handle both face recognition and sound attribute extraction using standardized algorithms. These multi-functional modules can process different types of input data through unified processing pipelines, reducing the need for separate specialized processing systems and thereby lowering overall device complexity.
Data Source
AI summary
Sign language information processing method and apparatus, an electronic device and a readable storage medium provided by the present disclosure, achieve real-time collection of language data in a current communication of a user by obtaining voice information and video information collected by a user terminal in real time; and then match a speaking person with his or her speaking content by determining, in the video information, a speaking object corresponding to the voice information; and finally, make it possible for the user to clarify the corresponding speaking object when the user sees AR sign language animation in a sign language video by superimposing and displaying an augmented reality AR sign language animation corresponding to the voice information on a gesture area corresponding to the speaking object to obtain a sign language video. Therefore, it is possible to provide a higher user experience.


