Efficient sign language video synthesis system

Through an efficient sign language video synthesis system, combining multimodal data and generative adversarial network optimization sign language actions, the movement stiffness and semantic understanding limitations of the existing system are solved, and high-quality and personalized sign language video generation and interactive improvement are achieved.

CN120279145APending Publication Date: 2025-07-08HARBIN INST OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411231671.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-04
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

现有手语翻译系统动作自然性不足、语义理解局限性和用户交互性低的问题,导致生成的手语视频僵硬、缺乏准确性和灵活性。

Method used

The efficient sign language video synthesis system is adopted, including data upload processing module, natural language processing module, action code generation module, action optimization module, virtual human animation system and user interface and display module. It uses Transformer model, LSTM timing model, generation adversarial network and physics engine to generate high-precision sign language videos and provide interactive feedback.

Benefits of technology

It improves the naturalness and accuracy of sign language videos, enhances user interaction, makes the generated sign language videos closer to real communication, provides a personalized user experience and feedback mechanism, and improves the flexibility of the system and user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279145A_ABST
    Figure CN120279145A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video synthesis, and discloses an efficient sign language video synthesis system which comprises a data uploading and processing module, a natural language processing module, an action code generation module, an action optimization module, a virtual human animation system, a user interface and a display module. The data uploading and processing module is used for receiving and preprocessing text, voice input, keyboard input and handwriting input provided by a user; the natural language processing module is used for performing word segmentation, part-of-speech tagging and semantic understanding on an input text, translating the input text into a sign language symbol sequence, and enhancing understanding in combination with multi-modal data; and the action code generation module is used for predicting and generating a sign language action sequence by using a time sequence model and a generative adversarial network, and generating expressions and body languages based on text emotion. By combining text and video data, richness and accuracy of sign language translation are remarkably improved, and the generated sign language video can be closer to a real communication scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video synthesis, and particularly to an efficient sign language video synthesis system. Background Art

[0002] In the fields of sign language translation and video synthesis, the existing technologies mainly rely on basic speech recognition and animation technologies to convert spoken language into sign language actions. These systems usually use predefined sign language action libraries and simple rule engines to simulate sign language expressions, including the sequence and combination of actions. However, due to the complexity of sign language and the richness of expressions, these systems often have difficulty accurately and naturally expressing complete semantic content and emotions.

[0003] The deficiencies of the existing technologies include: Lack of naturalness of actions: Existing sign language synthesis systems often lack sufficient dynamic fluency and rich expressions, making the generated sign language videos appear rigid and unnatural. This is mainly because existing systems usually use simple animation models and cannot accurately simulate complex human actions and subtle expression changes.

[0004] Limitations in semantic understanding: Many existing systems cannot fully understand complex semantics and context information and can only perform literal translations. This results in inaccurate sign language translations and cannot fully convey the intention and emotion of the original text, especially when involving metaphors and cultural background elements.

[0005] Low user interactivity: Traditional sign language video synthesis systems often do not provide or only provide limited user interaction functions, such as feedback and personalized settings. Users cannot adjust the video content or expression according to their own needs, which limits the flexibility and user experience of the system.

[0006] Therefore, those skilled in the art provide an efficient sign language video synthesis system to solve the problems raised in the above background art. Summary of the Invention

[0007] In view of the deficiencies of the existing technologies, the present invention provides an efficient sign language video synthesis system, which solves the problems of lack of naturalness of actions, limitations in semantic understanding, and low user interactivity existing in the existing technologies.

[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions: An efficient sign language video synthesis system includes a data upload and processing module, a natural language processing module, an action code generation module, an action optimization module, a virtual human animation system, and a user interface and display module; The data upload and processing module is used to receive and preprocess text, voice input, keyboard input, and handwriting input provided by users; The natural language processing module is used to tokenize, perform part-of-speech tagging, understand semantics on the input text, translate it into a sign language symbol sequence, and enhance understanding by combining multi-modal data; The action code generation module is used to predict and generate a sign language action sequence using a time series model and a generative adversarial network, and generate expressions and body languages based on text sentiment; The action optimization module is used to optimize sign language actions through a physics engine and an adaptive smoothing algorithm to ensure natural and accurate actions; The virtual human animation system is used to create a high-precision 3D virtual human model, drive it to perform sign language shows, and generate sign language videos; The user interface and display module is used to display, download, and share sign language videos, synchronously display sign language subtitles, and allow users to give feedback and make adjustments.

[0009] Preferably, the data upload and processing module includes: A data upload interface, which is used to receive text, voice input, keyboard input, and handwriting input provided by users; A data preprocessing unit, which is used to remove noise from the input data and standardize the data format.

[0010] Preferably, the natural language processing module includes: A language model loading unit, which is used to load a pre-trained Transformer model specifically for sign language grammar and syntactic structures; A text tokenization unit, which is used to tokenize the input text and break it into individual words or phrases; A part-of-speech tagging unit, which is used to tag the part of speech of each word, analyze the syntactic structure, and identify the grammatical features of the text; A semantic understanding unit, which is used to analyze the meaning of the input text through a semantic understanding model to determine the context and core information; A sign language translation unit, which is used to map the analysis result into a sign language symbol sequence to ensure that the symbol sequence conforms to the grammar and expression of sign language; A multi-modal data fusion unit, which is used to combine text and video data to enhance the understanding of text context information and improve the accuracy of translation.

[0011] Preferably, the action code generation module includes: An action prediction unit, which is used to predict a sign language action sequence using an LSTM time series model and generate an action sequence; A generative adversarial network unit, which is used to generate sign language actions, discriminates the authenticity of the actions through a discriminator, and continuously optimizes the generated sign language actions; An emotion-driven animation generation unit, which is used to generate corresponding facial expressions and body languages according to the text sentiment analyzed by the natural language processing module.

[0012] Preferably, the virtual human animation system includes: A virtual human modeling unit for creating a high-precision 3D virtual human model using Blender, including detailed modeling of hands and face; An action driving unit for driving the virtual human to perform sign language based on the optimized sign language action sequence, using skeletal animation technology.

[0013] A high-detail rendering unit for generating sign language videos through physically based rendering (PBR).

[0014] Preferably, the user interface and display module include: A video output interface unit for displaying the generated sign language video and providing download and sharing options.

[0015] A sign language subtitle synchronization unit for synchronously displaying sign language subtitles in the sign language video to help users understand the sign language content.

[0016] An interactive feedback unit for allowing users to provide feedback and adjustments to the generated video and optimizing the system based on the feedback.

[0017] Preferably, the calculation steps of the action prediction unit are as follows: Data preparation: Input the sign language symbol sequence ; Convert the symbol sequence into an embedding vector through the embedding layer LSTM layer: Initial state: Set the initial hidden state and cell state; Time step calculation: Output layer calculation: Map the hidden state to the sign language action space through the fully connected layer.

[0018] Preferably, the generative adversarial network unit includes a generator, a discriminator, loss calculation, backpropagation, and optimization. Among them, the discriminator loss expression in loss calculation is: ; The generator loss expression is: .

[0019] Preferably, the data upload and processing module is electrically connected to the natural language processing module, the natural language processing module is network-connected to the action code generation module, the action code generation module is network-connected to the action optimization module, the action code generation module is electrically connected to the virtual human animation system, the virtual human animation system is electrically connected to the user interface and display module, and the user interface and display module is electrically connected to the data upload and processing module.

[0020] An efficient sign language video synthesis method for an efficient sign language video synthesis system, comprising the following steps: S1. Data upload and processing: The user inputs text, voice, keyboard input, or handwritten input through the system interface. The data upload interface receives these inputs, performs denoising processing and standardization processing on the input data, and converts data in different formats into a unified text format. S2. Natural language processing: The natural language processing module first loads a Transformer model trained specifically for sign language grammar and syntactic structures, then tokenizes the input text data, decomposing it into individual words or phrases. Next, the system tags the part of speech of each word, analyzes the syntactic structure, identifies the grammatical features of the text, analyzes the meaning of the input text through a semantic understanding model, determines the context and core information of the text. The system maps the analysis results to a sign language symbol sequence, combines the text and video data to enhance the understanding of the text context information, and finally transmits the generated sign language symbol sequence to the action code generation module. S3. Action code generation: Use an LSTM time series model to input the sign language symbol sequence, predict and generate a preliminary sign language action sequence, including the sign language action representation at each time step. The preliminary sign language action sequence is then input into the generator of the generative adversarial network to generate an improved sign language action sequence. The discriminator simultaneously receives the real sign language action sequence and the generated sign language action sequence, and outputs a true / false discrimination result. The system calculates the losses of the discriminator and the generator, optimizes the parameters of the generator and the discriminator through the backpropagation algorithm, and generates corresponding facial expressions and body languages according to the sentiment analysis result of the text. The optimized sign language action sequence is transmitted to the action optimization module. S4. Action optimization: The action optimization module uses a physics engine to smooth the generated sign language actions to ensure natural action transitions, and dynamically adjusts the smoothing parameters according to the characteristics of the sign language actions. S5. Virtual human animation: Based on the optimized sign language action sequence, the system uses skeletal animation technology to drive the virtual human to perform sign language, and uses physically based rendering technology to generate sign language videos. S6. User interface and display: Display the generated sign language video and provide download and share options. Synchronize the sign language subtitles in the sign language video to help users understand the sign language content. The system allows users to provide feedback and adjust the generated video, and optimize the system based on the feedback.

[0021] The present invention provides an efficient sign language video synthesis system. It has the following beneficial effects: 1. The present invention combines text and video data to enhance the understanding of context and improve the accuracy of translation. This fusion enables the system to rely not only on text information, but also on non-verbal information in the video (such as expressions and gestures), so as to more comprehensively capture and reflect the speaker's intentions and emotions. This significantly improves the richness and accuracy of sign language translation, making the generated sign language video closer to the real communication scene.

[0022] 2. The present invention uses GAN technology to optimize the generated sign language movements, making the movements more natural and realistic. The introduction of this technology not only improves the quality of sign language movements, but also solves the stiffness and unnaturalness that may occur in traditional sign language animations. Through continuous training and optimization, the generator can learn how to generate movements that are more in line with actual sign language use, while the discriminator ensures the authenticity of these movements, improving the watchability and practicality of sign language videos, and providing a better visual communication tool for the deaf and mute population.

[0023] 3. The present invention provides a highly personalized user interaction experience through the interactive feedback unit of the user interface and display module. Users can not only watch sign language videos, but also provide feedback and make adjustments according to personal needs, such as modifying subtitle settings or action display speed. This personalized interaction design not only improves user satisfaction and participation, but also enables the system to self-learn and optimize according to user feedback, thereby continuously improving service quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 It is a schematic diagram of the system framework of the present invention; Figure 2 This is a schematic diagram of the data upload processing module framework of the present invention; Figure 3 This is a schematic diagram of the framework of the natural language processing module of the present invention; Figure 4 This is a schematic diagram of the action code generation module framework of the present invention; Figure 5 It is a schematic diagram of the framework of the virtual human animation system of the present invention; Figure 6 A schematic diagram of the user interface and display module framework of the present invention; Figure 7 It is a schematic diagram of the method flow of the present invention. DETAILED DESCRIPTION

[0025] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0026] Embodiment 1: Please see attached Figure 1 ,The embodiment of the present invention provides an efficient sign language video synthesis system, including a data uploading processing module, a natural language processing module, an action code generating module, an action optimization module, a virtual human animation system, a user interface and display module; The data upload processing module is electrically connected to the natural language processing module, the natural language processing module is network-connected to the action code generation module, the action code generation module is network-connected to the action optimization module, the action code generation module is electrically connected to the virtual human animation system, the virtual human animation system is electrically connected to the user interface and display module, and the user interface and display module is electrically connected to the data upload processing module.

[0027] The motion optimization module uses the PhysX engine and adaptive smoothing algorithm to simulate the physical characteristics of human motion, such as gravity, friction, and the interaction between limbs. This ensures that the motion of the virtual human model not only conforms to the principles of human biomechanics, but also appears more natural and coherent visually. This is an existing technology and will not be described in detail.

[0028] Specifically, the data upload processing module is directly connected to the natural language processing module through an electrical connection. This configuration reduces the delay of data transmission and accelerates the conversion speed from raw input to text processing. The natural language processing module is further connected to the action code generation module through a network connection, allowing flexible deployment in a cloud environment to utilize more powerful computing resources, thereby improving processing efficiency and scalability.

[0029] The motion code generation module uses LSTM and GAN technology to generate preliminary sign language motion sequences and transmits them to the motion optimization module through a network connection. This setting allows the motion data to be efficiently optimized in real time in the cloud. The optimized motion sequence is transmitted to the virtual human animation system through an electrical connection. This connection ensures the rapid transmission and high-quality output of the optimized data, thereby significantly improving the authenticity and viewing quality of the generated animation. Finally, the user interface and display module displays these finely rendered sign language videos and provides a real-time feedback mechanism, which quickly transmits the user's input and feedback back to the data upload processing module through an electrical connection, enhancing the interactivity of the system and user satisfaction.

[0030] Please refer to the appendix Figure 2 The data upload and processing module is used to receive and preprocess the text, voice input, keyboard input, and handwritten input provided by the user. The data upload and processing module includes: a data upload interface for receiving the text, voice input, keyboard input, and handwritten input provided by the user; and a data preprocessing unit for removing noise from the input data and standardizing the data format.

[0031] Specifically, the data upload interface is the primary component of the data upload and processing module and is specifically designed to receive various forms of input provided by the user. This includes text input, such as the text entered via the keyboard; voice input, which allows the user to directly provide voice data through the microphone; and handwritten input, where the user can handwrite text via the touch screen or a dedicated device. The design of this interface employs advanced reception technologies to support a wide range of input devices and formats. Through this multifunctional interface, the system can cover a wider range of user needs, especially considering the different device usage habits and physical capabilities of users, thereby increasing the accessibility and convenience of use of the system.

[0032] The data preprocessing unit is a crucial part of the data upload and processing module, and its main responsibility is to perform necessary preprocessing operations on the received data. This includes removing noise from the input data, such as background noise, irrelevant data, and interference caused by incorrect operations; and standardizing the data format to ensure the consistency and processability of the data in subsequent processing. The standardization process involves converting various formats of input into a format that the system can effectively process, for example, converting all text into a unified encoding format and converting voice data into a unified audio format. The efficient operation of this unit significantly improves the quality of data processing and the response speed of the system. By accurately removing invalid and noisy data, the system can more accurately parse the user's intent and reduce the possibility of misinterpretation. At the same time, the standardization of the data simplifies the processing process of subsequent modules, ensuring the efficiency of the processing flow and the high-quality output of data processing.

[0033] Please refer to the appendix Figure 3 The natural language processing module is used to tokenize, perform part-of-speech tagging, semantic understanding on the input text, and translate it into a sequence of sign language symbols, and combines multi-modal data to enhance understanding. It includes: a language model loading unit for loading a pre-trained Transformer model specifically for sign language grammar and syntactic structures; a text tokenization unit for tokenizing the input text into individual words or phrases; a part-of-speech tagging unit for tagging the part of speech of each word, analyzing the syntactic structure, and identifying the grammatical features of the text; a semantic understanding unit for analyzing the meaning of the input text through a semantic understanding model to determine the context and core information; and a sign language translation unit for mapping the analysis results into a sequence of sign language symbols to ensure that the symbol sequence conforms to the grammar and expressions of sign language Expression method; a multimodal data fusion unit for combining text and video data to enhance the understanding of text context information and improve the accuracy of translation.

[0034] Specifically, the language model loading unit: This unit is responsible for loading a pre-trained Transformer model, which is specifically optimized for the grammar and syntactic structure of sign language. This model is trained using a large corpus of sign language to adapt to the unique language structure of sign language, thus providing more accurate translation results. By loading a specialized language model, this unit enables the system to understand and process complex language structures, better capture the nuances of sign language expressions, and thus significantly improve the accuracy and naturalness of translation.

[0035] Text tokenization unit: This unit is used to accurately tokenize the input text, breaking the continuous text stream into individual words or phrases. This process is crucial for understanding the semantics and structure of the text. Accurate tokenization is the first step in understanding the intent and context of the text, providing the basis for subsequent part-of-speech tagging and semantic parsing, thus ensuring the accuracy of the entire translation process.

[0036] Part-of-speech tagging unit: After text tokenization, this unit is responsible for tagging the part of speech of each word and analyzing the syntactic structure of the entire sentence. This includes identifying nouns, verbs, adjectives, etc., and their functions in the sentence. Through accurate part-of-speech tagging, the system can better understand the roles and relationships of words in a sentence, which is essential for constructing a semantically complete sign language expression.

[0037] Semantic understanding unit: This unit uses a semantic understanding model to analyze the meaning of the input text, determine the context and core information. It uses context analysis algorithms to reveal the implicit meaning of the text, including metaphors, implications, and emotional colors. In-depth semantic understanding allows the system to not only translate the literal meaning of the words but also convey its emotions and context, providing sign language expressions that match the original text for deaf users and enhancing the naturalness and effectiveness of communication.

[0038] Sign language translation unit: Based on the output of the previous modules, the sign language translation unit maps the text analysis results into a sequence of sign language symbols. It ensures that these symbol sequences conform to the grammar and expression methods of sign language to accurately convey the intent and emotion of the original text. The efficient execution of this unit ensures a seamless conversion from text to sign language, making the final sign language video both natural and easy to understand.

[0039] Multimodal Data Fusion Unit: To improve the accuracy and context relevance of translation, this unit combines text and video data for analysis. By fusing these multimodal data, the system can better understand the meaning of the text in a specific context. The comprehensive utilization of multimodal data greatly enhances the system's ability to process complex contexts and non-verbal information, improving the overall quality of translation and user satisfaction.

[0040] Please refer to the appendix Figure 4 , the action code generation module is used to predict and generate sign language action sequences using a temporal model and a generative adversarial network, and generate expressions and body languages based on text sentiment; it includes: an action prediction unit, which is used to predict sign language action sequences using an LSTM temporal model and generate action sequences; a generative adversarial network unit, which is used to generate sign language actions, discriminates the authenticity of the actions through a discriminator, and continuously optimizes the generated sign language actions; an emotion-driven animation generation unit, which is used to generate corresponding facial expressions and body language actions according to the text sentiment analyzed by the natural language processing module. The action optimization module is used to optimize sign language actions through a physics engine and an adaptive smoothing algorithm to ensure natural and accurate actions; The calculation steps of the action prediction unit are as follows: Data preparation: Input sign language symbol sequence ; Convert the symbol sequence into an embedding vector through an embedding layer LSTM layer: Initial state: Set the initial hidden state and cell state; Time step calculation: Output layer calculation: Map the hidden state to the sign language action space through a fully connected layer.

[0041] The generative adversarial network unit includes a generator, a discriminator, loss calculation, backpropagation, and optimization. Among them, the discriminator loss expression in loss calculation is: ; The generator loss expression is: .

[0042] Specifically, the action prediction unit: The action prediction unit uses a long short-term memory network (LSTM) timing model to predict the sign language action sequence. By inputting a sign language symbol sequence, the unit first converts the symbol sequence into an embedding vector through an embedding layer. These vectors contain rich semantic information of the symbol. Subsequently, the LSTM network processes the embedding vector at each time step, gradually updates its hidden state and cell state, and finally maps these states to the sign language action space through a fully connected layer. This process makes the generated sign language action not only based on static symbol translation, but also takes into account the temporal dynamic characteristics of sign language, so that the generated action sequence is smoother and more natural, greatly improving the coherence and viewing experience of the animation.

[0043] Generative Adversarial Network Unit: The Generative Adversarial Network Unit consists of a generator and a discriminator. The task of the generator is to further optimize and refine the sign language movements based on the output of the LSTM to generate a more natural movement sequence. The discriminator evaluates the authenticity of these movements and determines whether they are indistinguishable from real sign language movements. Through loss calculation during the training process, the system continuously adjusts the parameters of the generator and the discriminator to optimize the generated sign language movements. The loss function of the discriminator mainly evaluates its ability to distinguish between real sign language movements and generated movements, while the loss of the generator reflects the probability that its generated movements are recognized as real movements by the discriminator. This technology ensures that the generated sign language movements are not only accurate but also difficult to distinguish from real movements, which increases the naturalness and realism of the sign language video.

[0044] Emotion-driven animation generation unit: Based on the text sentiment analysis results provided by the natural language processing module, the emotion-driven animation generation unit is responsible for generating corresponding facial expressions and body language. This unit ensures that the sign language video not only conveys the literal meaning of the text, but also expresses the speaker's emotional attitude. Through this emotional mapping, the generated sign language video can better express the tone and emotion during communication, making the communication of the deaf-mute people more vivid and expressive, and enhancing the integrity of expression and the transmission of emotions.

[0045] Please refer to the attached Figure 5 , the virtual human animation system is used to create a high-precision 3D virtual human model, drive it to perform sign language performance, and generate sign language videos; including: a virtual human modeling unit, used to create a high-precision 3D virtual human model using Blender, including detailed modeling of hands and faces; an action driving unit, used to drive the virtual human to perform sign language performance based on an optimized sign language action sequence, using skeletal animation technology, and a high-detail rendering unit, used to generate sign language videos through physically based rendering PBR; Specifically, the virtual human animation system greatly improves the quality and expressiveness of sign language videos by integrating advanced 3D modeling, motion driving, and rendering technologies. The fine model details and realistic motion expressiveness enable the virtual human to communicate in sign language in a very natural way, greatly enhancing the educational and communication value of sign language videos. In addition, the efficient rendering ability of the system ensures smooth playback even in complex scenarios, meeting the needs of a wide range of users.

[0046] Please refer to the appendix Figure 6 , the user interface and display module is used to display, download, and share sign language videos, synchronously display sign language subtitles, and allow users to provide feedback and make adjustments, including: a video output interface unit for displaying the generated sign language video and providing download and sharing options. A sign language subtitle synchronization unit for synchronously displaying sign language subtitles in the sign language video to help users understand the sign language content. An interactive feedback unit for allowing users to provide feedback and make adjustments to the generated video and optimizing the system based on the feedback; Specifically, through its multifunctional design, the user interface and display module greatly enhances the usability, interactivity, and accessibility of the sign language video synthesis system. By providing an intuitive video playback function, synchronized subtitles, and an interactive feedback mechanism, this module not only makes sign language videos more user-friendly and useful for a wide range of users but also encourages users to participate in the content creation and optimization process, providing a richer and more personalized viewing experience for all users.

[0047] Embodiment 2: Please refer to the appendix Figure 7 , the embodiment of the present invention provides an efficient sign language video synthesis method, including the following steps: S1. Data upload and processing: The user inputs text, voice, keyboard input, or handwritten input through the system interface. The data upload interface receives these inputs, performs denoising processing and normalization processing on the input data, and converts data in different formats into a unified text format; S2. Natural language processing: The natural language processing module first loads a Transformer model trained specifically for sign language grammar and syntactic structures, then tokenizes the input text data, decomposing it into individual words or phrases. Next, the system tags the part of speech of each word, analyzes the syntactic structure, identifies the grammatical features of the text, analyzes the meaning of the input text through a semantic understanding model, determines the context and core information of the text. The system maps the analysis results to a sign language symbol sequence, combines the text and video data to enhance the understanding of the text context information, and finally transmits the generated sign language symbol sequence to the motion code generation module; S3. Action Code Generation: Use the LSTM time series model to input the sign language symbol sequence, predict and generate a preliminary sign language action sequence, including the sign language action representation at each time step. The preliminary sign language action sequence is then input into the generator of the generative adversarial network to generate an improved sign language action sequence. The discriminator simultaneously receives the real sign language action sequence and the generated sign language action sequence and outputs the true / false discrimination result. The system calculates the losses of the discriminator and the generator, and optimizes the parameters of the generator and the discriminator through the backpropagation algorithm. According to the sentiment analysis result of the text, the system generates corresponding facial expressions and body languages. The optimized sign language action sequence is transmitted to the action optimization module; S4. Action Optimization: The action optimization module uses a physics engine to smooth the generated sign language actions to ensure natural action transitions and dynamically adjusts the smoothing parameters according to the characteristics of the sign language actions; S5. Virtual Human Animation: Based on the optimized sign language action sequence, the system uses skeletal animation technology to drive the virtual human to perform sign language and uses physically based rendering technology to generate sign language videos; S6. User Interface and Display: Display the generated sign language videos and provide download and sharing options. Synchronously display sign language subtitles in the sign language videos to help users understand the sign language content. The system allows users to provide feedback and make adjustments to the generated videos and optimizes the system according to the feedback.

[0048] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An efficient sign language video synthesis system, characterized in that, It includes a data upload and processing module, a natural language processing module, an action code generation module, an action optimization module, a virtual human animation system, a user interface and display module; The data upload and processing module is used to receive and preprocess the text, voice input, keyboard input, and handwritten input provided by the user; The natural language processing module is used to tokenize the input text, perform part-of-speech tagging, semantic understanding, and translate it into a sign language symbol sequence, combining multimodal data to enhance understanding; The action code generation module is used to predict and generate a sign language action sequence using a temporal model and a generative adversarial network, and generate expressions and body languages based on text sentiment; The action optimization module is used to optimize the sign language actions through a physics engine and an adaptive smoothing algorithm to ensure natural and accurate actions; The virtual human animation system is used to create a high-precision 3D virtual human model, drive it to perform sign language, and generate a sign language video; The user interface and display module is used to display, download, and share the sign language video, synchronously display sign language subtitles, and allow users to give feedback and make adjustments.

2. The high-efficiency sign language video synthesis system according to claim 1, wherein The data upload and processing module includes: A data upload interface, which is used to receive the text, voice input, keyboard input, and handwritten input provided by the user; A data preprocessing unit, which is used to remove noise from the input data and standardize the data format.

3. An efficient sign language video synthesis system according to claim 1, characterized in that, The natural language processing module includes: A language model loading unit, which is used to load a pre-trained Transformer model specifically for sign language grammar and syntactic structures; A text tokenization unit, which is used to tokenize the input text and break it down into individual words or phrases; A part-of-speech tagging unit, which is used to tag the part of speech of each word, analyze the syntactic structure, and identify the grammatical features of the text; A semantic understanding unit, which is used to analyze the meaning of the input text through a semantic understanding model, determine the context and core information; A sign language translation unit, which is used to map the analysis results to a sign language symbol sequence to ensure that the symbol sequence conforms to the grammar and expression of sign language; A multimodal data fusion unit, which is used to combine text and video data to enhance the understanding of text context information and improve the accuracy of translation.

4. An efficient sign language video synthesis system according to claim 1, characterized in that, The action code generation module includes: An action prediction unit, which is used to predict a sign language action sequence using an LSTM temporal model and generate an action sequence; A generative adversarial network unit, which is used to generate sign language actions, discriminates the authenticity of the actions through a discriminator, and continuously optimizes the generated sign language actions; An emotion-driven animation generation unit, which is used to generate corresponding facial expressions and body languages according to the text sentiment analyzed by the natural language processing module.

5. An efficient sign language video synthesis system according to claim 1, characterized in that, The virtual human animation system includes: A virtual human modeling unit, which is used to create a high-precision 3D virtual human model using Blender, including detailed modeling of the hands and face; An action driving unit, which is used to drive the virtual human to perform sign language based on the optimized sign language action sequence, using skeletal animation technology; A high-detail rendering unit, which is used to generate a sign language video through physically based rendering (PBR).

6. The efficient sign language video synthesis system according to claim 1, characterized in that, The user interface and display module includes: A video output interface unit, which is used to display the generated sign language video and provide download and sharing options; A sign language subtitle synchronization unit, which is used to synchronously display sign language subtitles in a sign language video to help users understand the sign language content; An interactive feedback unit, which is used to allow users to give feedback on and adjust the generated video, and optimize the system according to the feedback.

7. An efficient sign language video synthesis system according to claim 4, wherein, The calculation steps of the action prediction unit are as follows: Data preparation: Input sign language symbol sequence ; Convert the symbol sequence into an embedding vector through the embedding layer LSTM layer: Initial state: Set the initial hidden state and cell state; Time step calculation: Output layer calculation: Map the hidden state to the sign language action space through a fully connected layer.

8. An efficient sign language video synthesis system according to claim 4, characterized in that, The generative adversarial network unit includes a generator, a discriminator, loss calculation, backpropagation, and optimization. Among them, the discriminator loss expression in the loss calculation is: ; The generator loss expression is: 。 9. An efficient sign language video synthesis system according to claim 1, characterized in that, The data upload and processing module is electrically connected to the natural language processing module. The natural language processing module is network-connected to the action code generation module. The action code generation module is network-connected to the action optimization module. The action code generation module is electrically connected to the virtual human animation system. The virtual human animation system is electrically connected to the user interface and display module. The user interface and display module is electrically connected to the data upload and processing module.

10. An efficient sign language video synthesis method, for an efficient sign language video synthesis system according to any one of claims 1-9, characterized in that, It includes the following steps: S1. Data upload and processing: Users input text, voice, keyboard input, or handwriting input through the system interface. The data upload interface receives these inputs, performs denoising processing and standardization processing on the input data, and converts data in different formats into a unified text format; S2. Natural language processing: The natural language processing module first loads a Transformer model trained specifically for sign language grammar and syntactic structures, and then tokenizes the input text data, decomposing it into individual words or phrases. Then, the system annotates the part of speech of each word, analyzes the syntactic structure, identifies the grammatical features of the text, analyzes the meaning of the input text through a semantic understanding model, determines the context and core information of the text. The system maps the analysis results to a sign language symbol sequence, and combines the text and video data to enhance the understanding of the text context information. The finally generated sign language symbol sequence is sent to the action code generation module; S3. Action code generation: Use the LSTM time series model to input the sign language symbol sequence, predict and generate a preliminary sign language action sequence, including the sign language action representation at each time step. The preliminary sign language action sequence is then input into the generator of the generative adversarial network to generate an improved sign language action sequence. The discriminator simultaneously receives the real sign language action sequence and the generated sign language action sequence and outputs a true / false discrimination result. The system calculates the losses of the discriminator and the generator, optimizes the parameters of the generator and the discriminator through the backpropagation algorithm, and generates corresponding facial expressions and body languages according to the sentiment analysis result of the text. The optimized sign language action sequence is sent to the action optimization module; S4. Action optimization: The action optimization module uses a physics engine to smooth the generated sign language actions to ensure natural action transitions, and dynamically adjusts the smoothing parameters according to the characteristics of the sign language actions; S5. Virtual human animation: Based on the optimized sign language action sequence, the system uses skeletal animation technology to drive the virtual human to perform sign language, and uses physically based rendering technology to generate sign language videos; S6. User interface and display: Display the generated sign language videos, and provide options for downloading and sharing. Synchronously display sign language subtitles in the sign language videos to help users understand the sign language content. The system allows users to provide feedback and make adjustments to the generated videos, and optimizes the system based on the feedback.

Citation Information

Cited By

  • Multi-modal driving system based on streaming big language model output

    CN121562665A