Multi-modal interaction method and device for synchronously displaying voice and sign language
By combining the time synchronization and emotional fusion of sign language semantic features and phonological pronunciation features in digital human technology, the problem of inconsistent sign language actions and speech output is solved, emotional expression is realized, and user experience is improved.
Patent Information
- Application Number
- CN202510410044.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-18
AI Technical Summary
When existing digital human technology is displayed simultaneously in voice and sign language, sign language actions are inconsistent with voice output and it is difficult to express emotional information, resulting in poor user experience.
The distance measurement is determined based on the semantic difference loss of sign language semantic eigenvectors and phonological pronunciation eigenvectors, combined with the DTW algorithm for time synchronization, and fused the emotional eigenvectors to generate multimodal feature sequences, and controlled digital people to display sign language movements, facial expressions and lip shapes.
It achieves a high degree of consistency between digital human sign language actions and voice output, while expressing emotional information, improving user experience.
Smart Images

Figure CN120339476A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of speech and image data processing, and particularly to a multimodal interaction method and device for synchronously displaying speech and sign language. Background Art
[0002] With the rapid development of Virtual Reality (VR), Augmented Reality (AR), and Artificial Intelligence (AI) technologies, digital humans, as intelligent avatars in the virtual world, have broad application prospects in fields such as education and training, medical rehabilitation, and entertainment interaction. Digital human technology aims to simulate human facial expressions, mouth movements, and limb movements through computer-generated virtual images to achieve natural human-computer interaction.
[0003] Digital human technology has been widely used in the synchronous display of speech and sign language. However, when currently using digital human technology for the synchronous display of speech and sign language, there are still problems such as inconsistent sign language movements and speech output, and it is difficult to express emotional information, resulting in a poor user experience. Summary of the Invention
[0004] In view of this, it is necessary to provide a multimodal interaction method and device for synchronously displaying speech and sign language to solve the problems that when currently using digital human technology for the synchronous display of speech and sign language, there are still inconsistent sign language movements and speech output, and it is difficult to express emotional information.
[0005] To solve the above problems, in a first aspect, the present invention provides a multimodal interaction method for synchronously displaying speech and sign language, including: Determining a distance metric based on the semantic difference loss between the sign language semantic feature vector and the speech prosody feature vector, and performing time synchronization on the sign language semantic feature vector and the speech prosody feature vector based on the distance metric and the DTW algorithm; Fusing the emotional feature vector with the sign language semantic feature vector and the speech prosody feature vector after time synchronization to generate a multimodal feature sequence; Generating sign language movements, facial expressions, and lip shapes based on the multimodal feature sequence, and controlling the digital human to perform the display; The sign language semantic feature vector is extracted from sign language action data, the speech prosody feature vector and the emotional feature vector are extracted from speech signal data, the sign language action data includes the hand and limb action data of the signer, and the speech signal data includes the corresponding speech content read by the signer during the sign language performance.
[0006] In a possible implementation manner, the determining a distance metric based on the semantic difference loss between the sign language semantic feature vector and the speech prosody feature vector includes: Determine the distance metric between the sign language semantic feature vector and the speech prosody feature vector based on the semantic difference loss and Euclidean distance between them.
[0007] In a possible implementation, the determining the distance metric between the sign language semantic feature vector and the speech prosody feature vector based on the semantic difference loss and Euclidean distance between them includes: Determine the distance metric between the sign language semantic feature vector and the speech prosody feature vector based on the following formula:
[0008] where, represents the -th element in the sign language semantic feature vector, represents the -th element in the speech prosody feature vector, represents the distance metric between the -th element in the sign language semantic feature vector and the -th element in the speech prosody feature vector, , are weight coefficients, represents the Euclidean distance calculation, represents the semantic difference loss between the -th element in the sign language semantic feature vector and the -th element in the speech prosody feature vector.
[0009] In a possible implementation, the synchronizing the sign language semantic feature vector and the speech prosody feature vector in time by combining the DTW algorithm includes: Construct a cumulative distance matrix between the sign language semantic feature vector and the speech prosody feature vector based on the distance metric between them; Trace back from the end point of the cumulative distance matrix along the direction of the minimum cumulative distance to obtain the optimal time alignment path; Synchronize the sign language semantic feature vector and the speech prosody feature vector in time based on the optimal time alignment path.
[0010] In a possible implementation, the generating sign language actions, facial expressions, and lip shapes based on the multi-modal feature sequence includes: Input the multi-modal feature sequence into a sign language action generation model, a facial expression generation model, and a lip shape generation model respectively to generate sign language actions, facial expressions, and lip shapes; The sign language action generation model is trained by a generative adversarial network with sample speech semantic features, sample emotion features, and sample sign language action pairs; The facial expression generation model is trained by a generative adversarial network with sample emotion features and sample facial features; The lip shape generation model is trained on a Wav2Lip model using a sample Chinese face video and speech dataset, and a multi-head self-attention layer is arranged between the encoder and the decoder of the Wav2Lip model.
[0011] In a possible implementation manner, generating sign language actions, facial expressions, and lip shapes based on the multi-modal feature sequence, and controlling the digital human to perform a display, includes: Performing weighted fusion on the generated sign language actions, facial expressions, and lip shapes based on a fusion attention mechanism to obtain fused features, and controlling the digital human to perform a display based on the fused features.
[0012] In a possible implementation manner, controlling the digital human to perform a display based on the fused features includes: Extracting the hand and limb movement parameters in the fused features, determining the positioning of the end position of the digital human's hand based on inverse kinematics, and determining the limb movement trajectory of the digital human based on forward kinematics; Extracting the facial muscle movement parameters in the fused features, and determining the facial expression of the digital human based on blend shapes; Extracting the lip shape control parameters in the fused features, and determining the lip shape of the digital human.
[0013] On the other hand, the present invention further provides a multi-modal interaction device for synchronously displaying speech and sign language, including: A time synchronization module, configured to determine a distance metric based on the semantic difference loss between the sign language semantic feature vector and the speech prosody feature vector, and perform time synchronization on the sign language semantic feature vector and the speech prosody feature vector based on the distance metric and the DTW algorithm; A generation module, configured to fuse the emotion feature vector with the time-synchronized sign language semantic feature vector and speech prosody feature vector to generate a multi-modal feature sequence; A control module, configured to generate sign language actions, facial expressions, and lip shapes based on the multi-modal feature sequence, and control the digital human to perform a display; The sign language semantic feature vector is extracted from sign language action data, the speech prosody feature vector and the emotion feature vector are extracted from speech signal data, the sign language action data includes the hand and limb action data of the signer, and the speech signal data includes the corresponding speech content read by the signer during sign language performance.
[0014] In a second aspect, the present invention further provides an interaction device, including a memory and a processor, wherein, The memory is used to store programs; The processor is coupled to the memory and is used to execute the program stored in the memory to implement the steps in the multi-modal interaction method of synchronizing speech and sign language display described in any of the above implementation manners.
[0015] In a third aspect, the present invention also provides a computer-readable storage medium for storing computer-readable programs or instructions, and when the programs or instructions are executed by a processor, the steps in the multi-modal interaction method of synchronizing speech and sign language display described in any of the above implementation manners can be implemented.
[0016] The beneficial effects of the present invention are as follows: The multi-modal interaction method and device for synchronizing speech and sign language display provided by the present invention first increase the semantic difference loss between the sign language semantic feature vector and the speech prosody feature vector on the basis of the traditional DTW algorithm to improve the accuracy of time alignment. Then, an emotion feature vector is fused into the sign language semantic feature vector and the speech prosody feature vector after time synchronization to obtain a multi-modal feature sequence to enrich the feature representation. Finally, sign language actions, facial expressions, and lip shapes are generated through the multi-modal feature sequence to control the digital human for display, realizing the synchronous speech and sign language display of the digital human. The present invention expresses emotion information while ensuring the consistency between the sign language actions of the digital human and the speech output, improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a schematic flowchart of an embodiment of the multi-modal interaction method for synchronizing speech and sign language display provided by the present invention; Figure 2 It is a schematic flowchart of an embodiment of the multi-modal interaction process provided by the present invention; Figure 3 It is a schematic flowchart of an embodiment of the multi-modal time synchronization and alignment process provided by the present invention; Figure 4 It is a schematic flowchart of an embodiment of the training process of the sign language action generation model provided by the present invention; Figure 5 It is a schematic flowchart of an embodiment of the synchronization information generation process of the optimized Wav2Lip model provided by the present invention; Figure 6 It is a schematic structural diagram of an embodiment of the multi-modal interaction device for synchronizing speech and sign language display provided by the present invention; Figure 7 It is a schematic structural diagram of an embodiment of the interaction device provided by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the protection scope of the present invention.
[0019] In the description of the embodiments of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0020] The descriptions such as "first" and "second" involved in the embodiments of the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Therefore, the technical features defined with "first" and "second" may explicitly or implicitly include at least one such feature.
[0021] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment may be included in at least one embodiment of the present invention. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.
[0022] The present invention provides a multi-modal interaction method and device for synchronizing speech and sign language display, which will be described separately below.
[0023] Figure 1 It is a schematic flowchart of an embodiment of the multi-modal interaction method for synchronizing speech and sign language display provided by the present invention. As Figure 1 shown, the multi-modal interaction method for synchronizing speech and sign language display includes: S101. Determine a distance metric based on the semantic difference loss between the sign language semantic feature vector and the speech prosody feature vector, and perform time synchronization on the sign language semantic feature vector and the speech prosody feature vector based on the distance metric and the DTW algorithm.
[0024] It should be noted that: in order to improve the accuracy of time alignment, the present invention combines the sign language semantic feature and the speech prosody feature on the basis of the traditional Dynamic Time Warping (DTW) algorithm, and adds the semantic difference loss between the sign language semantic feature vector and the speech prosody feature vector.
[0025] S102. Integrate the emotional feature vector with the sign language semantic feature vector and the speech prosody feature vector after time synchronization to generate a multi-modal feature sequence.
[0026] It should be noted that: By integrating the emotional feature vector into the sign language semantic feature vector and the speech prosody feature vector after time synchronization to obtain a multi-modal feature sequence, the feature representation can be enriched.
[0027] S103. Generate sign language actions, facial expressions, and lip shapes based on the multi-modal feature sequence, and control the digital human for display.
[0028] It should be noted that: Through the multi-modal feature sequence, sign language actions, facial expressions, and lip shapes can be generated. And according to the sign language actions, facial expressions, and lip shapes, the digital human can be controlled to synchronize the speech and sign language display, expressing emotional information while ensuring the consistency of the sign language actions and the speech output.
[0029] The sign language semantic feature vector is extracted from the sign language action data, and the speech prosody feature vector and the emotional feature vector are extracted from the speech signal data. The sign language action data includes the hand and limb action data of the signer, and the speech signal data includes the corresponding speech content read by the signer during the sign language performance.
[0030] It should be noted that: When extracting the sign language semantic feature vector, a deep learning-based sign language recognition model (Sign Language Transformer, SL-Transformer) can be used for the preprocessed sign language action data to encode the sign language action sequence into a semantic feature vector. The model uses the self-attention mechanism to capture the temporal and spatial features in the sign language action sequence. Through the encoder of SL-Transformer, the sign language action sequence can be converted into a fixed-length semantic feature vector, representing the semantic content of the sign language action. This feature vector contains the semantic information of the sign language vocabulary and the temporal features of the action.
[0031] When extracting the speech prosody feature vector, OpenSMILE can be used for the preprocessed speech signal to extract the prosody features of the speech, including the fundamental frequency (F0), energy (Energy), duration (Duration), etc. These prosody features can reflect the rhythm, stress, and emotional information of the speech. The extracted prosody features are combined into a fixed-length feature vector as the prosody feature representation of the speech signal. To capture the temporal information, a sliding window method can be used to calculate the prosody features in each time period.
[0032] When extracting the emotional feature vector, a hybrid model based on a convolutional neural network and a long short-term memory network can be used to classify the speech emotion, identify the emotion category and intensity expressed by the speech, such as happy, sad, etc., and then obtain the emotional feature vector.
[0033] In summary, the multi-modal interaction method for synchronous speech and sign language display provided by the embodiments of the present invention first adds the semantic difference loss between the sign language semantic feature vector and the speech prosody feature vector on the basis of the traditional DTW algorithm to improve the accuracy of time alignment. Then, an emotional feature vector is fused into the sign language semantic feature vector and the speech prosody feature vector after time synchronization to obtain a multi-modal feature sequence to enrich the feature representation. Finally, sign language actions, facial expressions, and lip shapes are generated through the multi-modal feature sequence to control the digital human for display, realizing the synchronous speech and sign language display of the digital human. The present invention expresses emotional information while ensuring the consistency between the sign language actions of the digital human and the speech output, improving the user experience.
[0034] Combined Figure 2 with the above, the multi-modal interaction process provided by the present invention includes the following specific steps: I. Data collection and processing.
[0035] To ensure the effectiveness and accuracy of the model in the Chinese environment, the present invention independently collects and constructs a high-quality Chinese sign language and speech multi-modal dataset. The construction of the dataset includes three main links: sign language action data collection, speech signal data collection, and data synchronization and annotation.
[0036] 1. Sign language action data collection.
[0037] The present invention uses a high-precision motion capture device Azure Kinect DK based on depth camera technology to obtain the hand and limb action data of sign language users. The specific method is as follows: 1) Device selection and configuration: Use Azure Kinect DK as the motion capture device. This device has a high-resolution depth sensor and a wide-angle RGB camera, and can capture the three-dimensional bone data of sign language users at a speed of 30 frames per second.
[0038] 2) Collection environment preparation: Conduct data collection in an indoor environment with uniform light and a simple background. The sign language user wears dark clothes, with no obstruction on the hands and face, ensuring the accuracy and integrity of motion capture.
[0039] 3) Sign language content design: Develop a sign language script covering rich Chinese sign language vocabulary and sentences, including daily expressions, professional terms, and emotional expressions, etc. Invite sign language users with rich sign language experience to perform to ensure the accuracy and naturalness of sign language actions.
[0040] 4) Data acquisition process: The sign language user performs sign language according to the script content, and the Azure Kinect DK records the three-dimensional skeletal data of the hands and limbs in real time. To ensure data quality, each sign language action is collected multiple times, and the action with the best performance is selected as valid data.
[0041] 5) Preliminary data check: During the acquisition process, monitor the data quality in real time to check for data loss, occlusion, or anomalies. For unqualified data, re-acquire it immediately to ensure the integrity and accuracy of the data.
[0042] 2. Speech signal data acquisition.
[0043] Synchronously acquire the speech signals of the sign language user to ensure the clarity and integrity of the speech data. The specific methods are as follows: 1) Microphone selection and configuration: Use a professional condenser microphone, the Shure SM7B, which has a wide frequency response range and a flat frequency response curve. Set the sampling rate to 48 kHz and the quantization precision to 24 bits to ensure the high fidelity of the speech signal.
[0044] 2) Audio acquisition environment: Conduct speech acquisition in a recording studio with good acoustic conditions. The walls are covered with sound-absorbing materials to reduce environmental noise and reverberation. Keep the microphone about 20 centimeters away from the sign language user's mouth and use a pop filter to avoid plosive sounds.
[0045] 3) Synchronization of speech content and sign language: While performing sign language, the sign language user clearly reads the corresponding speech content. To achieve audio-video synchronization, use a professional timecode synchronizer, the Tentacle Sync, to unify the timecodes of the motion capture device and the audio recording device, ensuring the precise alignment of the data in time.
[0046] 4) Audio monitoring and adjustment: Monitor the level and waveform of the speech signal in real time to prevent overloading and distortion. Adjust the gain of the microphone and the speech volume of the sign language user as needed to ensure that the recorded speech signal is clear and noise-free.
[0047] 3. Data synchronization and annotation.
[0048] To achieve the precise alignment of sign language action data and speech signal data, data synchronization and annotation are required.
[0049] 1) Time synchronization: Use the Tentacle Sync timecode synchronizer to unify the timecodes of the audio devices recorded by the Azure Kinect DK and the Shure SM7B, ensuring that the sign language action data and the speech signal data have the same time reference.
[0050] 2) Data annotation: Use the professional annotation tool ELAN to perform fine annotation on the collected data. The annotation content includes: Sign language action annotation: Mark the start time and end time of each sign language action, annotate the corresponding sign language words or sentences, as well as the semantic and emotional information of the action.
[0051] Speech signal annotation: Mark the speech segments and silence segments in the speech signal, annotate the corresponding speech content, pronunciation features and emotional expressions.
[0052] 4. Data preprocessing.
[0053] The preprocessing of sign language action data includes data format conversion, data cleaning and denoising, coordinate correction and normalization, action segmentation and alignment, data augmentation, etc.
[0054] The preprocessing of speech signal data includes noise reduction, pre-emphasis, framing and windowing, feature extraction, speech normalization, speech segmentation and alignment, etc.
[0055] II. Multimodal time synchronization and alignment.
[0056] Combined Figure 3 Viewed, the multimodal time synchronization and alignment process provided by the present invention specifically includes: 1. Definition of distance metric.
[0057] In some embodiments of the present invention, determining the distance metric based on the semantic difference loss between the sign language semantic feature vector and the speech prosody feature vector includes: Based on the semantic difference loss and Euclidean distance between the sign language semantic feature vector and the speech prosody feature vector, determine the distance metric between the sign language semantic feature vector and the speech prosody feature vector.
[0058] In some embodiments of the present invention, determining the distance metric between the sign language semantic feature vector and the speech prosody feature vector based on the semantic difference loss and Euclidean distance between the sign language semantic feature vector and the speech prosody feature vector includes: Determine the distance metric between the sign language semantic feature vector and the speech prosody feature vector based on the following formula:
[0059] Among them, represents the -th element in the sign language semantic feature vector, represents the -th element in the speech prosody feature vector, represents the -th element in the sign language semantic feature vector and the The distance metric between elements and is the weight coefficient indicating Euclidean distance calculation indicating the th element in the sign language semantic feature vector and the th element in the speech prosody feature vector, representing the semantic difference loss between them.
[0060] 2. Time synchronization and alignment.
[0061] In some embodiments of the present invention, the time synchronization of the sign language semantic feature vector and the speech prosody feature vector by combining with the DTW algorithm includes: Based on the distance metric between the sign language semantic feature vector and the speech prosody feature vector, constructing an accumulated distance matrix between the sign language semantic feature vector and the speech prosody feature vector; Starting from the end point of the accumulated distance matrix and backtracking along the direction of the minimum accumulated distance to obtain the optimal time alignment path; Based on the optimal time alignment path, performing time synchronization on the sign language semantic feature vector and the speech prosody feature vector.
[0062] 3. Generation of multi-modal feature sequences.
[0063] 1) Construction of synchronized feature sequences: According to the time mapping relationship, aligning the sign language semantic features and the speech prosody features in time to generate synchronized multi-modal feature sequences.
[0064] 2) Emotional feature fusion: Extracting emotional features from the speech signal, classifying the speech emotions to obtain the emotional categories and intensities. Fusing the emotional features with the synchronized multi-modal feature sequences to enrich the feature representation.
[0065] 3) Feature normalization and storage: Performing normalization processing on the synchronized multi-modal feature sequences to reduce the differences between different feature scales. Saving the processed feature sequences for subsequent use.
[0066] In some embodiments of the present invention, the generation of sign language actions, facial expressions, and lip shapes based on the multi-modal feature sequences includes: Inputting the multi-modal feature sequences into a sign language action generation model, a facial expression generation model, and a lip shape generation model respectively to generate sign language actions, facial expressions, and lip shapes; The sign language action generation model is trained by a generative adversarial network with sample speech semantic features, sample emotional features, and sample sign language actions; The facial expression generation model is trained by a generative adversarial network with sample emotional features and sample facial features; The lip shape generation model is obtained by training the Wav2Lip model with a sample Chinese face video and speech dataset, and a multi-head self-attention layer is set between the encoder and the decoder of the Wav2Lip model.
[0067] III. Sign language action generation.
[0068] Combined Figure 4 seen, the training process of the sign language action generation model provided by the present invention includes: 1. Design of the sign language action generation model based on the generative adversarial network.
[0069] 1.1. Model structure design.
[0070] 1) Generator design.
[0071] The generator aims to convert the input speech semantic features and emotional features into a sign language action sequence. The generator adopts an encoder-decoder structure: Encoder part: Input layer: Accept the preprocessed speech semantic feature vector and emotional feature vector , and splice the two to form a joint feature vector .
[0072] Feature extraction layer: Use a multi-layer fully connected neural network to perform a non-linear transformation on the joint feature vector to extract high-level semantic representations.
[0073] Decoder part: Sign language action generation layer: Input the high-level semantic representation output by the encoder into a multi-layer bidirectional long short-term memory network (Bi-LSTM) to generate the corresponding sign language action sequence .
[0074] Output layer: Use a fully connected layer to map the output of the Bi-LSTM to the joint coordinate space of the sign language action to generate the predicted value of the sign language action sequence.
[0075] 2) Discriminator design.
[0076] Adopt a one-dimensional convolutional neural network (1D-CNN) structure: Input layer: Accept the sign language action sequence , including the sign language action sequence generated by the generator and the real sign language action sequence.
[0077] Convolutional layer: Through multi-layer one-dimensional convolution, extract features from the input sign language action sequence to capture the temporal and spatial features of the action sequence.
[0078] Fully connected layer: The output of the convolutional layer passes through the fully connected layer to further extract high-level features.
[0079] Output layer: Using the sigmoid activation function, output a scalar , representing the probability that the input sign language action sequence is real data.
[0080] 1.2 Functions of the generator and discriminator.
[0081] 1) Generator function.
[0082] Semantic to action mapping: Convert speech semantic features and emotional features into sign language action sequences to achieve the mapping from semantics to actions.
[0083] Action sequence generation: Generate natural and fluent sign language action sequences by learning the temporal patterns and spatial features of sign language actions.
[0084] 2) Discriminator function.
[0085] Real and fake sample discrimination: Evaluate whether the input sign language action sequence is real data to guide the generator to improve the quality of generated sign language actions.
[0086] Adversarial training: Through adversarial training, force the sign language action sequence generated by the generator to approximate the distribution of real sign language actions.
[0087] 2. Model training and optimization.
[0088] Design a multi-task loss function and adopt an effective model training strategy to optimize the model.
[0089] 2.1 Design of the multi-task loss function.
[0090] The loss function consists of three parts: adversarial loss, action reconstruction loss, and semantic consistency loss.
[0091] 1) Adversarial loss: Used to guide the adversarial training of the generator and discriminator, defined as follows:
[0092] Among them, is the real sign language action sequence, is the sign language action sequence generated by the generator, is the probability output by the discriminator.
[0093] Action reconstruction loss: Used to measure the difference between the generated sign language action sequence and the real sign language action sequence in the action space, defined as the mean squared error (MSE):
[0094] 3) Semantic consistency loss: It is used to ensure the semantic consistency between the generated sign language action sequence and the input speech semantic features. The pre-trained sign language recognition model F(⋅) is used to perform semantic prediction on the generated sign language action sequence, and the loss function is the cross-entropy loss:
[0095] where is the input speech semantic label.
[0096] 4) Total loss function:
[0097] where is the weight coefficient of the loss term, which is adjusted according to the experiment.
[0098] 2.2 Model training strategy.
[0099] During the model training process, first use the preprocessed multi-modal dataset, including synchronized speech semantic features, emotional features, and sign language action sequences. The dataset is divided into a training set, a validation set, and a test set, with a ratio of 80%, 10%, and 10%. The training process is divided into three stages. In the first stage, the discriminator is fixed and the generator is pre-trained so that it can initially generate sign language action sequences that conform to the semantics. The optimization objective is to minimize the action reconstruction loss and the semantic consistency loss. In the second stage, the generator is fixed and the discriminator is trained to improve its ability to distinguish real and generated samples. The optimization objective is to minimize the adversarial loss. In the third stage, both the generator and the discriminator are trained and multiple rounds of iteration are performed. The optimization objective is to minimize the total loss function.
[0100] 2.3 Model performance evaluation.
[0101] The quantitative evaluation metrics include the mean squared error (MSE) and the semantic accuracy rate. On the test set, the MSE of the model is lower than the set threshold, and the semantic accuracy rate reaches 92%.
[0102] In terms of qualitative evaluation, professional sign language experts are invited to subjectively evaluate the generated sign language actions. The scoring criteria include the fluency, naturalness, and semantic accuracy of the actions. The average score for the naturalness of the actions is 4.7 points (out of 5), indicating that the quality of the generated sign language actions is relatively high.
[0103] IV. Speech-driven facial expression and lip shape generation.
[0104] Combined with Figure 5 it can be seen that the synchronization information generation process of the optimized Wav2Lip model provided by the present invention includes: 1. The Wav2Lip model optimized for Chinese speech.
[0105] Wav2Lip is a deep learning-based lip-sync model that can generate a video of lip movements synchronized with the speech according to the input speech signal and static face image. However, the original Wav2Lip model was mainly trained based on English corpora and has poor performance when directly applied to the Chinese environment. To address this issue, the present invention specifically optimizes the Wav2Lip model.
[0106] First, a large-scale Chinese face video and speech dataset is constructed, covering a variety of speakers, mouth shape changes, and emotional expressions. This dataset is used to retrain the Wav2Lip model to adapt to the tones and pronunciation characteristics of Chinese speech.
[0107] In terms of the model structure, an attention mechanism is introduced to enhance the model's ability to capture key speech features and corresponding lip movement. Specifically, a multi-head self-attention layer is added between the encoder and the decoder, enabling the model to better align the correspondence between the speech signal and the lip movement.
[0108] In addition, for the polyphones, liaisons, and tone changes in Chinese speech, the present invention adds acoustic features such as tone information and pronunciation position features during the model training process. These features are input into the model together with the original speech features to help the model generate corresponding lip movements more accurately.
[0109] In the design of the loss function, in addition to the original reconstruction loss and adversarial loss, a lip-sync loss and a perceptual loss are added. The lip-sync loss is used to measure the consistency of the generated lip movements and the real lip movements in time and space; the perceptual loss ensures that the generated lip movements are visually similar to the real lip movements through a pre-trained face recognition network.
[0110] 2. Facial expression generation method.
[0111] To enable the digital human's facial expressions to reflect the speech content and emotional information, the present invention combines speech emotion recognition technology and an expression generation model to achieve context-compliant facial expression generation.
[0112] First, emotional features are extracted from the speech signal, and a hybrid model based on a convolutional neural network and a long short-term memory network is used to classify the speech emotion, identifying the emotional category and intensity expressed by the speech, such as happy, sad, etc.
[0113] Then, using the generative adversarial network framework, corresponding facial expression parameters are generated. The generator inputs the emotional features and basic facial features and outputs the corresponding facial expression changes. The discriminator is used to evaluate whether the generated expression is real and natural. Through adversarial training, the generated facial expressions have a high sense of reality and naturalness.
[0114] During the generation process, ensure the coordination between facial expressions and lip movements. To this end, in the expression generation model, a lip shape constraint condition is added, that is, when generating expressions, the lip area generated by the Wav2Lip model is retained, and only the expression changes are made to other facial areas. Ensure the accuracy of lip synchronization while enriching the expressiveness of facial expressions.
[0115] Finally, fuse the generated facial expression parameters and lip movement sequences and map them to the facial model of the digital human.
[0116] V. Multimodal data fusion and digital human driving.
[0117] Effectively fuse sign language action features, speech features, and emotional features and map them to the bones and controllers of the digital human model to achieve synchronous display of the sign language actions, lip movements, and facial expressions of the digital human.
[0118] 1. Multimodal data fusion with fusion attention mechanism.
[0119] In some embodiments of the present invention, generating sign language actions, facial expressions, and lip shapes based on multimodal feature sequences and controlling a digital human for display includes: Weightedly fuse the generated sign language actions, facial expressions, and lip shapes based on the fusion attention mechanism to obtain the fused features, and control the digital human for display based on the fused features.
[0120] First, obtain sign language action features, speech features, and emotional features respectively. The sign language action features are output by the sign language action generation model and include the motion parameters of the hands and limbs; the speech features are obtained after being processed by the optimized Wav2Lip model and include lip movement parameters and facial muscle movement information; the emotional features are obtained through the speech emotion recognition model and represent the emotional category and intensity of the current speech.
[0121] Then, preliminarily process the features of each modality to unify the feature dimensions and scales. Embed each feature using a fully connected neural network and map it to the same feature space. Reduce the differences between features of different modalities to facilitate subsequent fusion. Introduce the fusion attention mechanism to weightedly fuse the features of different modalities. Specifically, calculate the attention weights of the sign language action features, speech features, and emotional features for each time step. The calculation of the attention weights is based on the correlation between the features of each modality and the current task, and the formula is as follows:
[0122] Where represents the attention weight of the th modality feature, is the feature vector of the th modality, and is a trainable parameter matrix and a bias vector, is an activation function, is the number of modalities.
[0123] Dynamically adjust the contribution degree of each modality feature in the fusion process by calculating the attention weights. Add the weighted modality features to obtain a unified feature representation after fusion:
[0124] 2. Digital human model driving.
[0125] In some embodiments of the present invention, controlling the digital human for display based on the fused features includes: Extract the hand and limb movement parameters in the fused features, determine the positioning of the end position of the digital human's hand based on inverse kinematics, and determine the limb movement trajectory of the digital human based on forward kinematics; Extract the facial muscle movement parameters in the fused features, and determine the facial expression of the digital human based on blend shape; Extract the lip shape control parameters in the fused features, and determine the lip shape of the digital human.
[0126] After obtaining the fused feature representation, map it to the skeleton and controller of the digital human model to drive the digital human to synchronously display sign language actions, lip shape actions and facial expressions.
[0127] First, construct a high-fidelity digital human model. The skeletal structure adopts the general human bone naming and hierarchy, which is convenient for mapping with action data. Then, decompose the fused feature representation according to different parts and map it to the hand and limb skeletons, facial expression controller and lip shape controller of the digital human respectively.
[0128] For the sign language action part, extract the hand and limb movement parameters in the fused features, and use a method combining inverse kinematics (IK) and forward kinematics (FK) to calculate the rotation and displacement of each bone joint. Specifically, use IK to solve the precise positioning of the end position of the hand, and use FK to maintain the natural movement trajectory of the limb.
[0129] For the facial expression and lip shape action parts, extract the facial muscle movement parameters and lip shape control parameters in the fused features. The facial expression of the digital human is controlled based on blend shape. Map the facial expression parameters to the corresponding blend shape weights to drive the change of the digital human's facial expression. The lip shape action is realized by a special lip shape controller, which adjusts the shape and position of the lip part according to the lip shape parameters to achieve lip shape changes synchronized with the speech.
[0130] During the action mapping process, action smoothing and transition processing are adopted to filter and smooth the parameters of sign language actions, facial expressions, and lip movements, eliminating possible jitters and discontinuities. For the switching between different actions, an easing function is used to generate a natural transition effect.
[0131] Finally, the processed action parameters are applied to the digital human model to render the sign language, expression, and lip animation of the digital human in real time.
[0132] VI. System Architecture Optimization and Emotional Computing Interaction.
[0133] To enable the digital human to understand and express rich emotional information, emotional features are extracted from speech signals and sign language content, and a hybrid model integrating a convolutional neural network and a long short-term memory network is adopted to accurately identify the user's emotional state. In terms of emotional expression, the digital human presents an emotional performance that conforms to the context by adjusting facial expressions, speech intonations, and gesture actions. Using a pre-constructed emotional expression library and a dynamic expression generation model, the digital human can generate corresponding facial expression changes in real time according to the results of emotional recognition. The sign language actions and speech output will also be appropriately adjusted according to the emotional state, enhancing the naturalness and affinity of the interaction.
[0134] The multi-modal time synchronization method based on Chinese sign language semantics and speech prosody features proposed by the present invention effectively solves the problem of accurate synchronization of sign language actions and speech signals in terms of time and content. Compared with traditional time synchronization methods, by using the fusion of sign language semantics and speech prosody features, the accuracy of time alignment is improved, making the sign language actions and speech output of the digital human highly consistent during the display process, and significantly enhancing the user experience.
[0135] The present invention adopts a sign language action generation model for Chinese sign language. The generated sign language action sequence is natural, fluent, and semantically accurate, and the model has obvious improvements in both the naturalness and accuracy of sign language actions.
[0136] The present invention optimizes the Wav2Lip model for Chinese speech. By introducing an attention mechanism and acoustic feature assistance, the accuracy and naturalness of lip synchronization are improved, making the lip movements of the digital human highly match the Chinese speech, and enhancing the lip expression ability of the digital human.
[0137] Through a multi-modal data fusion framework integrating an attention mechanism, the present invention effectively integrates sign language action features, speech features, and emotional features, enabling the digital human to simultaneously present sign language actions, facial expressions, and lip movements, and achieving the coordinated unity of multi-modal information.
[0138] The present invention enables the digital human to recognize the user's emotional state based on the speech and sign language content, and make appropriate emotional responses through facial expressions, speech intonation, and gesture actions, thereby enhancing the affinity and intelligence of human-computer interaction.
[0139] The present invention has broad application prospects in the fields of human-computer interaction, virtual reality, public services, etc., and has significant technical value and social benefits.
[0140] To better implement the multi-modal interaction method for synchronous speech and sign language display in the embodiments of the present invention, correspondingly, based on the multi-modal interaction method for synchronous speech and sign language display, as Figure 6 shown, the embodiments of the present invention further provide a multi-modal interaction device for synchronous speech and sign language display. The multi-modal interaction device 600 for synchronous speech and sign language display includes: A time synchronization module 601, configured to determine a distance metric based on the semantic difference loss between the sign language semantic feature vector and the speech prosody feature vector, and perform time synchronization on the sign language semantic feature vector and the speech prosody feature vector based on the distance metric and the DTW algorithm; A generation module 602, configured to fuse the emotional feature vector with the time-synchronized sign language semantic feature vector and speech prosody feature vector to generate a multi-modal feature sequence; A control module 603, configured to generate sign language actions, facial expressions, and lip shapes based on the multi-modal feature sequence, and control the digital human to perform a display; The sign language semantic feature vector is extracted from the sign language action data, the speech prosody feature vector and the emotional feature vector are extracted from the speech signal data. The sign language action data includes the hand and limb action data of the signer, and the speech signal data includes the corresponding speech content read by the signer during the sign language performance.
[0141] The multi-modal interaction device 600 for synchronous speech and sign language display provided in the above embodiments can implement the technical solutions described in the embodiments of the multi-modal interaction method for synchronous speech and sign language display. For the specific implementation principles of the above modules or units, reference can be made to the corresponding content in the embodiments of the multi-modal interaction method for synchronous speech and sign language display, which will not be elaborated here.
[0142] As Figure 7 shown, the present invention also correspondingly provides an interaction device 700. The interaction device 700 includes a processor 701, a memory 702, and a display 703. Figure 7 Only some components of the interaction device 700 are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0143] In some embodiments, the processor 701 may be a central processing unit (CPU), a microprocessor, or other data processing chips, which are used to run the program code stored in the memory 702 or process data, such as the magnetic resonance image optimization method in the present invention.
[0144] In some embodiments, the processor 701 may be a single server or a server group. The server group may be centralized or distributed. In some embodiments, the processor 701 may be local or remote. In some embodiments, the processor 701 may be implemented on a cloud platform. In one embodiment, the cloud platform may include a private cloud, a public cloud, a hybrid cloud, a community cloud, a distributed cloud, an internal cloud, a multi-cloud, etc., or any combination thereof.
[0145] In some embodiments, the memory 702 may be an internal storage unit of the interaction device 700, such as the hard disk or memory of the interaction device 700. In some other embodiments, the memory 702 may also be an external storage device of the interaction device 700, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the interaction device 700.
[0146] Furthermore, the memory 702 may also include both the internal storage unit and the external storage device of the interaction device 700. The memory 702 is used to store the application software installed in the interaction device 700 and various types of data.
[0147] In some embodiments, the display 703 may be an LED display, a liquid crystal display, a touch liquid crystal display, an Organic Light-Emitting Diode (OLED) toucher, etc. The display 703 is used to display the information in the interaction device 700 and to display a visual user interface. The components 701-703 of the interaction device 700 communicate with each other through a system bus.
[0148] In one embodiment, when the processor 701 executes the multi-modal interaction program for synchronizing speech and sign language display in the memory 702, the following steps may be implemented: Determine a distance metric based on the semantic difference loss between the sign language semantic feature vector and the speech prosody feature vector, and perform time synchronization on the sign language semantic feature vector and the speech prosody feature vector based on the distance metric and the DTW algorithm; Fuse the emotion feature vector with the sign language semantic feature vector and the speech prosody feature vector after time synchronization to generate a multi-modal feature sequence; Generate sign language actions, facial expressions, and lip shapes based on multi-modal feature sequences to control the digital human for display; The sign language semantic feature vector is extracted from sign language action data, and the speech prosody feature vector and the emotion feature vector are extracted from speech signal data. The sign language action data includes the hand and limb action data of the sign language user, and the speech signal data includes the corresponding speech content read by the sign language user during sign language performance.
[0149] It should be understood that when the processor 701 executes the multi-modal interaction program for synchronizing speech and sign language display in the memory 702, in addition to the above functions, other functions can also be realized. For specific details, reference can be made to the description of the corresponding method embodiments above.
[0150] Furthermore, the type of the interaction device 700 mentioned in the embodiments of the present invention is not specifically limited. The interaction device 700 can be a portable electronic device such as a mobile phone, a tablet computer, a personal digital assistant (PDA), a wearable device, a laptop computer, etc. Exemplary embodiments of the portable electronic device include, but are not limited to, portable electronic devices equipped with IOS, android, microsoft, or other operating systems. The above portable electronic devices can also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of the present invention, the interaction device 700 may not be a portable electronic device, but a desktop computer with a touch-sensitive surface (e.g., a touch panel).
[0151] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium, which is used to store computer-readable programs or instructions. When the programs or instructions are executed by a processor, the steps or functions in the multi-modal interaction method for synchronizing speech and sign language display provided by the above method embodiments can be realized.
[0152] Those skilled in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program. The computer program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.
[0153] The above has introduced in detail the multi-modal interaction and device for synchronous voice and sign language display provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A multimodal interaction method for synchronously displaying speech and sign language, characterized in that, Including: Determine a distance metric based on the semantic difference loss between the sign language semantic feature vector and the speech prosody feature vector, and perform time synchronization on the sign language semantic feature vector and the speech prosody feature vector based on the distance metric and the DTW algorithm; Fuse the emotion feature vector with the sign language semantic feature vector and the speech prosody feature vector after time synchronization to generate a multi-modal feature sequence; Generate sign language actions, facial expressions, and lip shapes based on the multi-modal feature sequence to control the digital human for display; The sign language semantic feature vector is extracted from the sign language action data, the speech prosody feature vector and the emotion feature vector are extracted from the speech signal data, the sign language action data includes the hand and limb action data of the signer, and the speech signal data includes the corresponding speech content read by the signer during sign language performance.
2. The multimodal interaction method for synchronous voice and sign language display according to claim 1, wherein The determining the distance metric based on the semantic difference loss between the sign language semantic feature vector and the speech prosody feature vector includes: Determine the distance metric between the sign language semantic feature vector and the speech prosody feature vector based on the semantic difference loss and the Euclidean distance between the sign language semantic feature vector and the speech prosody feature vector.
3. The multimodal interaction method for synchronous speech and sign language display according to claim 2, wherein The determining the distance metric between the sign language semantic feature vector and the speech prosody feature vector based on the semantic difference loss and the Euclidean distance between the sign language semantic feature vector and the speech prosody feature vector includes: Determine the distance metric between the sign language semantic feature vector and the speech prosody feature vector based on the following formula: Among them, represents the th element in the sign language semantic feature vector, represents the th element in the speech prosody feature vector, represents the distance metric between the th element in the sign language semantic feature vector and the th element in the speech prosody feature vector, and are weight coefficients, represents Euclidean distance calculation, represents the semantic difference loss between the th element in the sign language semantic feature vector and the th element in the speech prosody feature vector.
4. The multimodal interaction method for synchronous voice and sign language display according to claim 1, wherein The performing time synchronization on the sign language semantic feature vector and the speech prosody feature vector in combination with the DTW algorithm includes: Construct a cumulative distance matrix between the sign language semantic feature vector and the speech prosody feature vector based on the distance metric between the sign language semantic feature vector and the speech prosody feature vector; Trace back from the end point of the cumulative distance matrix along the direction of the minimum cumulative distance to obtain the optimal time alignment path; Perform time synchronization on the sign language semantic feature vector and the speech prosody feature vector based on the optimal time alignment path.
5. The multimodal interaction method for synchronous speech and sign language display according to claim 1, wherein The generating sign language actions, facial expressions, and lip shapes based on the multi-modal feature sequence includes: Input the multi-modal feature sequence into a sign language action generation model, a facial expression generation model, and a lip shape generation model respectively to generate sign language actions, facial expressions, and lip shapes; The sign language action generation model is trained by a generative adversarial network with sample speech semantic features, sample emotion features, and sample sign language actions; The facial expression generation model is trained by a generative adversarial network with sample emotion features and sample facial features; The lip shape generation model is trained by a Wav2Lip model with a sample Chinese face video and a speech dataset, and a multi-head self-attention layer is arranged between the encoder and the decoder of the Wav2Lip model.
6. The multimodal interaction method for synchronous speech and sign language display according to claim 1, wherein The generating sign language actions, facial expressions, and lip shapes based on the multi-modal feature sequence to control the digital human for display includes: Perform weighted fusion on the generated sign language actions, facial expressions, and lip shapes based on a fusion attention mechanism to obtain the fused features, and control the digital human for display based on the fused features.
7. The multimodal interaction method for synchronous speech and sign language display according to claim 6, wherein The controlling the digital human for display based on the fused features includes: Extract the hand and limb movement parameters from the fused features, determine the positioning of the end position of the digital human's hand based on inverse kinematics, and determine the limb movement trajectory of the digital human based on forward kinematics; Extract the facial muscle movement parameters from the fused features and determine the facial expressions of the digital human based on blend shapes; Extract the lip control parameters from the fused features and determine the lip shapes of the digital human.
8. A multimodal interaction device for synchronously displaying speech and sign language, characterized in that, It includes: A time synchronization module, which is used to determine the distance metric between the sign language semantic feature vector and the speech prosody feature vector based on the semantic difference loss between the sign language semantic feature vector and the speech prosody feature vector, and synchronize the sign language semantic feature vector and the speech prosody feature vector in time by combining the DTW algorithm; A generation module, which is used to fuse the emotion feature vector with the sign language semantic feature vector and the speech prosody feature vector after time synchronization to generate a multi-modal feature sequence; A control module, which is used to generate sign language actions, facial expressions and lip shapes based on the multi-modal feature sequence to control the digital human for display; The sign language semantic feature vector is extracted from the sign language action data, the speech prosody feature vector and the emotion feature vector are extracted from the speech signal data, the sign language action data includes the hand and limb action data of the signer, and the speech signal data includes the corresponding speech content read by the signer during the sign language performance.
9. An interactive device, characterized in that, It includes a memory and a processor, where the memory is used to store programs; the processor is coupled to the memory and is used to execute the program stored in the memory to implement the steps in the multi-modal interaction method for synchronizing speech and sign language display described in any one of claims 1 to 7 above.
10. A computer-readable storage medium, characterized in that, For storing computer-readable programs or instructions, when the programs or instructions are executed by a processor, they can implement the steps in the multi-modal interaction method for synchronizing speech and sign language display described in any one of claims 1 to 7 above.
Citation Information
Cited By
Voice-driven three-dimensional human body movement method based on local style encoder
CN120894473A