Sign language synthesis service method based on multiple modes

CN120823294APending Publication Date: 2025-10-21ZHEJIANG UNIVERSITY OF MEDIA AND COMMUNICATIONS
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510842014.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-21

AI Technical Summary

Technical Problem

Existing sign language translation systems lack emotional expression, and the synchronization of dynamic expressions and emotional intensity is insufficiently adapted, resulting in the inability of hearing-impaired and hearing people to accurately understand emotions; when generating sign language movements, the skeletal trajectory and emotional intensity are not dynamically matched, resulting in a high movement breakage rate; edge-cloud collaboration is inefficient, and services fail under network fluctuations.

Method used

Adopting an end-cloud collaborative rendering strategy, a lightweight model is deployed on the edge to generate basic actions, and MoMask and a 3D rendering engine are deployed on the cloud. Through multimodal data collection, emotion feature extraction, cross-modal feature fusion, and sign language action generation, an emotion-adapted 3D skeleton sequence is achieved.

Benefits of technology

It improves the accuracy and adaptability of emotional expression in sign language synthesis, reduces the asynchrony of cross-modal features, realizes real-time and efficient sign language synthesis services on smart terminals, and enhances the communication experience between the hearing-impaired and the normal population.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120823294A_ABST
    Figure CN120823294A_ABST
Patent Text Reader

Abstract

The invention discloses a sign language synthesis service method based on multi-modality, and the method comprises the following steps: S10, carrying out end-cloud collaborative rendering setting: deploying an edge end as a lightweight model to generate a basic action, and deploying a cloud end as operating a MoMask and 3D rendering engine; s20, performing multi-modal data acquisition: acquiring a voice signal through a microphone, acquiring a 48 * 48 pixel face grayscale image through a camera, and acquiring action posture data through an IMU sensor; s30, performing emotion feature extraction on the collected multi-modal data: performing voice emotion recognition to output six types of emotion probability distributions, and performing facial expression recognition to output seven types of expression probability distributions; s40, performing cross-modal feature fusion: dynamically fusing the voice and facial features based on a confidence weighting strategy, and generating a three-dimensional emotion intensity vector; and S50, sign language action generation is carried out, and a 3D skeleton sequence matched with emotion is generated through RVQ layering quantification and MoMask Transformer.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Cited By

  • Multi-modal emotion recognition method and system for service-oriented robot

    CN120995416A

  • Companion robot emotional intelligence analysis method based on multi-modal data analysis

    CN122527842A