AR Sign Language Animation for Multi-Speaker Identification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing sign language translation technologies face difficulties in distinguishing speaking objects in multi-person discussions, leading to poor user experience and impaired communication between individuals with and without hearing impairments.

Innovation Solution

A method and apparatus that utilize real-time voice and video information processing to identify speaking objects by recognizing face images and sound attributes, superimposing augmented reality sign language animations on a gesture area corresponding to the speaking object, enabling clear identification of speaking content in a sign language video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If existing voice-sign language translation method is used, then translation function is provided, but speaking objects cannot be distinguished in multi-person discussions

Engineering Contradiction:
Improvespeaking object identificationVSAvoidcommunication reliability
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The system segments the translation output by associating each sign language animation with a specific speaking object identified through face recognition and sound attribute analysis. This segmentation allows the hearing-impaired user to distinguish which speaker is corresponding to which sign language translation, resolving the information loss in multi-person discussions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system introduces an intermediary mechanism that captures video information of speaking objects, performs face recognition, extracts sound attributes, and matches them to voice information. This intermediary processing chain enables reliable identification and association of speaking objects with their corresponding speech content.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If binoculus pays attention to translated text information at all times, then translation is received, but ability to discern view of each speaking object is lost

Engineering Contradiction:
Improvecommunication easeVSAvoidspeaking object view
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The system transitions from traditional text-based translation to augmented reality sign language animations that are superimposed on the video feed of speaking objects. This dimensional change from 2D text to 3D spatial AR animation allows users to perceive both the speaker's visual presence and the translation simultaneously, eliminating the need to switch attention between text and speakers.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system merges the sign language animation with the video information of the speaking object in an augmented reality display. This combination allows the hearing-impaired user to see both the speaker and the translation together in one view, maintaining the ability to discern speaking objects while receiving accurate translation information.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If face recognition and sound attribute matching is performed, then speaking object identification is achieved, but processing complexity increases

Engineering Contradiction:
Improvespeaking object identification precisionVSAvoidprocessing system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-processing video information to extract face images and pre-processing audio information to extract sound attributes. These preliminary processing steps organize the data in advance, making the subsequent matching process more efficient and manageable, thus reducing overall system complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system employs universal processing modules that handle both face recognition and sound attribute extraction using standardized algorithms. These multi-functional modules can process different types of input data through unified processing pipelines, reducing the need for separate specialized processing systems and thereby lowering overall device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11580983B2Sign language information processing method and apparatus, electronic device and readable storage medium
Publication Date: 2023.02.14 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US11580983B2 patent drawing
  • US11580983B2 patent drawing
  • US11580983B2 patent drawing

AI summary

Sign language information processing method and apparatus, an electronic device and a readable storage medium provided by the present disclosure, achieve real-time collection of language data in a current communication of a user by obtaining voice information and video information collected by a user terminal in real time; and then match a speaking person with his or her speaking content by determining, in the video information, a speaking object corresponding to the voice information; and finally, make it possible for the user to clarify the corresponding speaking object when the user sees AR sign language animation in a sign language video by superimposing and displaying an augmented reality AR sign language animation corresponding to the voice information on a gesture area corresponding to the speaking object to obtain a sign language video. Therefore, it is possible to provide a higher user experience.