Generating synthetic sign language videos through text input (text-to-sign) and more

A machine learning-based system translates sign language into text and vice versa, using a photorealistic avatar to generate synthetic sign language videos, addressing the challenge of communication barriers for deaf individuals and improving accessibility.

DE102023005368A1Inactive Publication Date: 2025-07-03LAW SIU KUEN KAY +1

Patent Information

Application Number
DE102023005368
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-31
Publication Date
2025-07-03
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies fail to facilitate effective, bidirectional communication between sign language users and non-users, limiting accessibility for deaf individuals.

Method used

A method and system utilizing machine learning to translate sign language into text and vice versa, employing a photorealistic avatar to generate synthetic sign language videos, integrating modules for text-to-sign language conversion, sign glossary to skeletal positions, and skeletal positions to video generation, with cloud-based deployment and web/mobile applications.

Benefits of technology

Enables accurate, bidirectional communication between sign language users and non-users, enhancing accessibility across various platforms and environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The invention proposes, inter alia, a method for automated translation using a sign language, the method comprising at least three of the following steps: - Assignment between a text and a sign glossary, - Assignment between a sign glossary and a skeleton position, - Generation of a gesture execution from a skeleton position, - Generation of a skeleton position from a gesture execution. In addition, an overall device, a module, computer program products, and various applications are proposed.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The following invention relates to a method for automated translation by means of a sign language, an overall device therefor, a first module therefor, a mobile device, a stationary device, a first and a second computer program product therefor and various applications for the method.

[0002] Sign language is the native language of the majority of deaf people. It is a visual language used to convey messages through hand gestures, hand and finger movements, and facial expressions. However, only a minority of the population is able to understand any of the many sign languages. Researchers and developers have focused on using machine learning (ML) to enable better participation for those who rely on sign language.

[0003] The object of the present invention is to create a way to enable at least one-directional, but preferably also bidirectional, communication between sign language users and non-users.

[0004] This object is achieved with a method having the features of claim 1, with an overall device having the features of claim 4, with a module having the features of claim 5, with a mobile device having the features of claim 6, with a stationary device having the features of claim 7, with a first computer program product having the features of claim 8, with a second computer program product having the features of claim 9 and with an application having the features of claim 10. Further advantageous embodiments and features emerge from the subclaims, the following description and the following figures. However, the present invention is not limited to these, and in particular not to the present set of claims. Rather, these serve to describe the invention in a first attempt.Therefore, one or more features from the claims may be supplemented or replaced by one or more features from the subclaims, the following description, and the following figures. Likewise, further developments may result from a combination of individual features that emerge from the subclaims, the following description, and the following figures.

[0005] A method for automated translation using a sign language is proposed, the method comprising at least three of the following steps: - Assignment between a text and a sign glossary, - Assignment between a sign glossary and a skeleton position, - Generation of a gesture execution from a skeleton position, - Generation of a skeleton position from a gesture execution.

[0006] It is further proposed that automated translation be used to create an avatar that represents sign language.

[0007] One further training course involves using a camera to record the gestures of a sign speaker and translate them into text.

[0008] Furthermore, an overall device is proposed. The overall device comprises at least a first module, a second module, and a third module, wherein at least one of the modules is located separately from the other modules at a different location and can be executed there. The first module, the second module, and the third module can exchange one or more pieces of information with each other. The overall device implements a method according to the invention as described above and below.

[0009] A further embodiment of the invention provides a first module for a device for implementing a method according to the invention with a connection to the second and third modules, wherein the first module is integrated in an app.

[0010] Also proposed is a mobile device having a screen and / or a camera, implemented with a method according to the invention for using an automated translation of sign language into text and / or of text into sign language.

[0011] A stationary device is also proposed, comprising a screen and / or a camera, implemented with a method according to the invention for using an automated translation of sign language into text and / or of text into sign language.

[0012] Furthermore, a first computer program product is provided, executable on a mobile or stationary device, by means of which a first module according to the invention can be executed.

[0013] A second computer program product which is executable on a mobile or stationary device, by means of which a method according to the invention can be carried out, as described above and below.

[0014] Furthermore, various applications of a method according to the invention and / or a device according to the invention in a field selected from a group comprising at least education and e-learning, customer support, healthcare, entertainment and media, government and public services, retail and e-commerce, accessibility services, training and personnel development, culture and tourism, legal services, events, in particular music events, metaverse.

[0015] The following sections describe, by way of example, the architecture and methods of both machine learning (ML) and web development (e.g., frontend, backend, and database) to generate sign language videos from text (text-to-sign language).

[0016] In particular, the invention proposes to develop and optimize three key tasks in order to then apply them together: 1. Translation of written or spoken language into sign language (signgloss); 2. Translation of the sign glossary into skeleton poses; 3. Generation of synthetic sign language videos presented by a lifelike AI sign language interpreter (avatar).

[0017] The proposed method, which preferably includes at least all of these processes to create synthetic sign language videos with high accuracy in reproducing hand and finger movements, enables application in both mobile and stationary devices. This is illustrated below with an example of how a user might use their mobile phone, which has a corresponding app installed that enables the proposed method.

[0018] Example steps from a user's perspective for video generation 1. The user enters one or more sentences in written / spoken language and clicks the "Generate" button in the web or mobile application interface. 2. The first segment of the process (step 1) translates the entered text into the corresponding sign language (sign language glossary). 3. The second segment of the process (step 2) generates skeleton positions that visualize the hand and finger movements corresponding to the sign glossary. 4. Based on the skeleton positions, the third segment of the process (step 3) creates a sequence of image frames in which a lifelike AI avatar replicates the hand and finger postures and movements. 5. A final output video consisting of a sequence of image frames can be both viewed and downloaded. Proposed technology:

[0019] A method or workflow is proposed for generating accessible AI videos presented by a lifelike sign language interpreter (photorealistic avatar) from text—text-to-sign. Using the proposed method, a user can enter text, for example, in German. The model will automatically: 1) translate the written or spoken language into the corresponding sign language (sign glossary), 2) generate the skeletal positions and movements for each sign glossary, and 3) create an AI video in which a photorealistic avatar replicates the hand and finger postures and movements corresponding to the sign glossary. Method and architecture of ML development

[0020] An example involves the use of an ML method that has three steps, preferably consisting of three steps: • Text to (sign language) glossary (text-to-glossary) • (Sign) Glossary of Skeleton Position (Glossary of Pose) • Skeleton position to sign language video (pose-to-video)

[0021] These will be discussed in more detail below, taking into account the respective exemplary processes as shown in the respective figures. However, the invention is not limited to these. Rather, other processes may also be provided. They show: Fig. 1 an exemplary implementation of text-to-gloss, Fig. 2 an exemplary NMT architecture for text-to-gloss, Fig. 3 an exemplary architecture of a “HamNoSys-to-Skeleton Pose” method, Fig. 4 an exemplary architecture of a possible “pose-to-video” method, and Fig. 5 an exemplary representation of an overall device. Step 1: Text-to-Glossary, also shown in Figures 1 and 2

[0022] The goal of step 1 is to translate written / spoken language into sign language. The text-to-glossary model uses a language method such as natural language processing (NLP) transformer models or neural machine translation (NMT) models. These NLP or NMT models convert each word into a corresponding sign glossary. When using an NLP transformer model, an attention encoder-decoder mechanism can be applied to transform the text into a glossary.

[0023] Words that aren't needed in sign language are removed. For example, the German sentence "Hello, I'm in school now" is converted to "Hello, I'm at school now" in German Sign Language (DGS). Step 2: Glossary-to-pose, also shown in Figure 3

[0024] The goal of step 2 is to convert the sign glossary into skeletal positions. Sign notations, which include a detailed list of corresponding facial and body movements, can be used to link each glossary to the corresponding skeletal positions. A phonetic transcription system such as the Hamburg Sign Language Notation System (HamNoSys) (Hanke, 2004) can be used to ensure that each glossary has a ham notation. A ham notation consists of a sequence of symbols to enable high accuracy in pose generation. Based on the ham notation of the glossary, a skeletal mapping can be created and visualized. Step 3: Pose-to-video, also shown in Figure 4

[0025] The goal of Step 3 is to create synthetic sign language videos presented by a photorealistic avatar based on the skeletal poses generated in Step 2. Since the shape and size of each skeleton can vary from person to person, another machine learning-based method can be applied to normalize the skeletal poses. The normalized skeletal poses are then used to generate image frames and videos in which a photorealistic avatar performs DGS.

[0026] A ML method such as deep learning methods such as Generative Adversarial Networks (GANs), U-Net image segmentation, or the (Stable) Diffusion Model can be used for image generation. The skeleton poses are input into the selected deep learning model. A photorealistic avatar replicates the hand and finger positions and movements corresponding to the skeleton positions. The final result is a video sequence consisting of images generated by the selected deep learning model. Method and architecture of web development

[0027] The architecture of the client-side application preferably consists of three modules: • The user interface of the web application / platform, developed with JavaScript and relevant frameworks and libraries such as ReactJS (frontend technology) • Backend technology • Database Module 1: Web and / or mobile application

[0028] In Module 1, the web application provides users with an attractive yet intuitive user interface. With JavaScript as its core programming language, ReactJS can be the primary framework for building the user interface. ReactJS excels at creating dynamic and responsive single-page applications that ensure a smooth user experience. To improve the performance and functionality of the web application, we integrate additional frameworks and libraries. Of particular note are Redux for state management, which enables efficient management of application state across components, and Axios for handling asynchronous requests to ensure optimal data retrieval and interaction. The use of these modern frameworks ensures the scalability and maintainability of a potential web application, for example, a SaaS-based one.

[0029] Other frameworks such as FabricJS and KonvaJS can also be used for web development. Users can create a video background or upload existing files (e.g., MP4 and PNG) to customize the output video. These frameworks are suitable for canvas-based web applications.

[0030] Additionally, the web application can use a microservices architecture, with containerization facilitated by Docker. This enables modular development, scalability, and efficient deployment across different environments. Continuous integration (CI) and continuous deployment (CD) practices, for example, are used to optimize the development process and maintain code quality. CI / CD tools such as Jenkins or GitLab CI can be integrated to automate testing, build processes, and deployment, thus ensuring a robust and reliable web application.

[0031] For example, text-to-sign technology can also be accessed via a mobile app to reach a wider audience and improve accessibility. In addition to the developed web application, the method can be extended to the creation of native mobile applications for iOS and Android platforms to ensure a seamless user experience across devices.

[0032] For native mobile app development, Swift is used for iOS applications and Kotlin for Android applications. Native development provides optimal performance and access to platform-specific features, enabling a highly responsive and intuitive user interface. This allows users to seamlessly interact with the proposed text-to-sign language technology on their mobile devices and leverage the unique features of each platform.

[0033] Alternatively, technologies such as Flutter or React Native can be considered to streamline the development of cross-platform mobile apps and reduce development time. Google's Flutter provides a comprehensive framework for building natively compiled applications for mobile, web, and desktop from a single code repository. React Native, based on ReactJS, enables the development of cross-platform mobile applications with a native look and feel. This ensures consistent functionality and user experience across iOS and Android devices and minimizes the need for platform-specific code.

[0034] By integrating mobile app development into Module 1, the proposed solution expands the reach of text-to-sign technology and provides users with the flexibility to seamlessly interact with our platform across both web and mobile interfaces. This multi-platform approach enhances the accessibility of state-of-the-art sign language synthesis technology and appeals to a diverse user base with varying device usage preferences. Module 2: Backend technology for the web platform and deployment of the ML model

[0035] In Module 2, the backend infrastructure is designed to seamlessly integrate with the cloud-based deployment of ML models and ensure efficient generation of sign language videos. For a modern, SaaS-based web application, Django is one of the robust backend frameworks to consider, providing a structured and scalable approach to application development.

[0036] To enable deployment of the ML model in the cloud, for example, a custom API is implemented. This API serves as a bridge between the web backend and the cloud-based ML model, enabling asynchronous and secure communication. Technologies such as the Django REST framework can be used to streamline the creation of RESTful APIs and ensure standardized communication between the web application and the cloud-based ML model.

[0037] For cloud-based deployment of the ML model, consider a serverless architecture leveraging platforms such as AWS Lambda or Google Cloud Functions. This enables on-demand execution of the ML model, optimizing resource utilization and minimizing operating costs. The ML model can be packaged as a microservice, which represents a scalable and self-contained unit accessible by the web backend via the custom API.

[0038] Additionally, a load balancing mechanism can be implemented to efficiently distribute incoming requests across multiple instances of the ML model, ensuring responsiveness and reliability. To improve security, access controls and authentication mechanisms are integrated into the API to protect against unauthorized access to the ML model.

[0039] By adopting a cloud-based approach with a custom API, the proposed backend technology facilitates seamless communication between the web application and the ML model, ensuring dynamic and responsive generation of sign language videos within the proposed state-of-the-art text-to-sign technology. Module 3: Database management for user data

[0040] For Module 3, selecting a suitable database is crucial for efficient user data management. Due to the relational nature of user data, a database-based solution based on SQL, such as PostgreSQL, is a suitable option. PostgreSQL offers ACID compliance, scalability, and extensibility, making it suitable for managing complex data relationships and ensuring data integrity. Support for the JSONB data type also enables efficient storage and querying of JSON-formatted data, which can be valuable for handling different user preferences and application settings.

[0041] To improve data security, encryption protocols can be implemented within the database and access controls can be configured to restrict unauthorized access. Regular database backups and maintenance procedures can be implemented to ensure the reliability and availability of user data and strengthen the robustness of our SaaS-based web application.

[0042] A further embodiment provides for the additional support of lip reading. Lip reading enables support for sign language, particularly to convey greater accuracy. For this purpose, lip movement can be recorded by camera and analyzed accordingly, analogous to the system proposed here. Accordingly, the created avatar can not only perform the gestures that can be specified via the text. Rather, by generating lip and facial muscle movements, the sign reader can be provided with additional assistance via the screen, for example, by the avatar including a corresponding representation and playback.

[0043] Fig.Figure 5 shows an exemplary representation with the first module 1, the second module 2, and the third module 3, which together form an overall device 4. A screen 5 for displaying the data and a recording device 6 are also provided. The recording device can be a camera, a voice input device, a text input option, or something else. For example, a camera or voice input, for example via a microphone, can be possible. The spoken text can be recorded and displayed as text on the screen, where it can be checked if necessary and corrected, for example, via a text input option. Example use cases

[0044] Our AI sign language videos, presented by a photorealistic avatar, can benefit numerous companies and industries by improving communication and accessibility for deaf and hard of hearing people and sign language users: 1) Education and e-learning a. Sign language tutorials: Online education platforms can use AI-driven avatars to create interactive sign language courses for deaf and hearing people. b. Inclusive learning: Educational institutions can make their online courses more inclusive for deaf students by providing sign language support. 2) Customer support a. Virtual customer service representatives: Companies can use sign language avatars to provide customer support and assistance in sign language, thus improving accessibility for deaf customers. b. Multilingual support: Avatars can bridge language barriers and provide support in multiple sign languages to facilitate global customer interaction. 3) Healthcare a. Telemedicine: Healthcare providers can use avatars to facilitate remote consultations for deaf patients and ensure effective communication with medical staff. b. Health education: Avatars can provide health information and instructions in sign language for deaf patients. 4) Entertainment and Media a. Accessible content: Streaming platforms can offer sign language avatars for movies, TV shows, and online content to make entertainment more accessible to deaf viewers. b. Interactive Storytelling: Interactive video games and storytelling experiences can incorporate sign language avatars for a more immersive gaming experience. 5) Government and public services a. Accessible government information: Government agencies can use avatars to provide essential information and services in sign language, ensuring accessibility for all citizens. b. Emergency alerts: Sign language avatars can be used to convey emergency alerts and safety instructions during crises. 6) Retail and e-commerce a. Inclusive shopping: E-commerce websites can offer sign language support for deaf customers during online shopping, ensuring a seamless shopping experience. b. Virtual shopping assistants: Avatars can assist customers with product information and recommendations in sign language. 7) Accessibility services a. Sign language interpreters: Organizations such as companies can offer sign language interpreting services via avatars for events, conferences, and meetings. b. Accessible websites: Websites and apps can integrate avatars to provide sign language translations of content and services. 8) Training and personnel development a. Workplace training: Companies can use avatars to offer sign language training and development programs for deaf employees. b. Onboarding: Avatars can support the onboarding process by communicating company policies and procedures in sign language. 9) Culture and Tourism a. Cultural institutions: Museums, historical sites and tourist attractions can offer sign language tours via avatars and thus appeal to deaf tourists. b. Language learning: Avatars can be used to teach tourists the basics of sign language and to describe their travel experiences, 10) Legal services a. Legal advice: i. Law firms can use our technology to facilitate access to legal advice for deaf clients. ii. Avatars can serve as an interface to answer legal questions in sign language and to interpret documents. iii. This ensures that deaf persons have equal access to legal services. 11) Music events a. Poetry interpreter (in real time): i. Our technology enables real-time interpretation of song lyrics in sign language during live music events. ii. A virtual interpreter can be displayed on screens to enable deaf participants to access the musical content. iii. This creates an inclusive environment at music festivals and concerts. 12) Metaverse a. VR or AR application: i. Companies can integrate our technology into virtual or augmented reality applications to enable barrier-free interaction in different contexts. ii. Examples can include customer support in VR, e-learning modules, HR training, marketing presentations, or even integration into virtual movies. iii. This opens up new possibilities for immersive and inclusive experiences in the metaverse.

Claims

[1] A method for automated translation using a sign language, the method comprising at least three of the following steps: - Assignment between a text and a sign glossary, - Assignment between a sign glossary and a skeleton position, - Generation of a gesture execution from a skeleton position, - Generation of a skeleton position from a gesture execution. [2] Method according to claim 1, characterized by that automated translation creates an avatar that represents sign language. [3] Method according to claim 1 or 2, characterized by that the gestures of a sign speaker are recorded using a camera and translated into text. [4] Overall device (4) comprising at least a first module (1), a second module (2) and a third module (3), wherein at least one of the modules (1, 2, 3) is separate from the other modules at a different location and can be executed there, wherein the first module (1), the second module (2) and the third module (3) can exchange one or more pieces of information with each other, wherein the overall device (4) implements a method according to one of claims 1 to 3. [5] First module (1) for a device for implementing a method according to one of claims 1 to 3 with a connection to the second and the third module, wherein the first module is integrated in an app. [6] Mobile device comprising a screen (5) and / or a recording device (6), in particular a camera, implemented with a method according to one of claims 1 to 3 for using an automated translation of sign language into text and / or of text into sign language. [7] Stationary device comprising a screen (5) and / or a recording device (6), in particular a camera, implemented with a method according to one of claims 1 to 3 for using an automated translation of sign language into text and / or of text into sign language. [8] First computer program product executable on a mobile or stationary device by means of which a first module according to claim 4 is executable. [9] Second computer program product executable on a mobile or stationary device by means of which a method according to one of claims 1 to 3 can be executed. [10] Application of a method and / or a device according to one of the claims in a field selected from a group comprising at least education and e-learning, customer support, healthcare, entertainment and media, government and public services, retail and e-commerce, accessibility services, training and staff development, culture and tourism, legal services, events, in particular music events, metaverse.

Citation Information

Patent Citations

  • Automatic translation between sign language and spoken language

    US20230085161A1

Cited By

  • Sign language translation method and system based on pre-training diffusion large language model

    CN121527813A