Multi-language and multi-mode short play virtual face image generation method
By adopting the multilingual multimodal short drama virtual face image generation method in the field of image generation, and using the deep fusion technology of semantic extraction and image generation modules, the problems of difficulty in fusion of multimodal information and insufficient generation quality in the prior art are solved, and high-quality multilingual and multimodal image generation is achieved.
Patent Information
- Application Number
- CN202411871074.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When processing multimodal inputs, it is difficult for the prior art to achieve efficient fusion of information of different modalities, resulting in the generative model failing to make full use of the semantic information of each modal, and the generalization ability is weak, the language processing ability is limited, and the generation quality needs to be improved.
A multilingual multimodal short drama virtual face image generation method is adopted. Semantic features and time-step features are extracted through the semantic extraction module and input them into the image generation module. Technical means such as Text Embedding module, speech encoder and multi-head self-attention mechanism are used to achieve deep fusion and image generation of information in different modes.
The effective fusion of text and pronunciation in different languages is achieved, the impact of low-quality data on the model effect is reduced, and the performance and image generation quality of the model in multilingual and multimodal scenarios is improved.
Smart Images

Figure CN119941885A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of skit virtual face image generation, and in particular to a multi-language and multi-modal skit virtual face image generation method. Background Art
[0002] At present, many deep learning-based methods have emerged in the field of image generation, among which the more classic ones are generative adversarial networks (GANs) and variational autoencoders (VAEs). These methods have shown quite good results when processing single-modal inputs (such as text or images) and can generate high-quality images. However, when it is necessary to process multi-modal inputs (such as text and audio at the same time), existing methods usually face some significant technical challenges, which are specifically manifested in the following aspects:
[0003] 1. Lack of effective cross-modal fusion mechanism:
[0004] Existing image generation methods, especially when dealing with multimodal inputs, often have difficulty achieving efficient fusion of information from different modalities. For example, in scenarios where text and audio are input simultaneously, the information between the modalities cannot fully interact, resulting in the generative model failing to fully utilize the rich semantic information contained in each modality. The lack of this fusion mechanism makes it difficult for the model to achieve results comparable to those of a single modality in multimodal scenarios.
[0005] 2. Weak generalization ability:
[0006] Existing generative models, especially those designed for a single modality, often exhibit weak transfer capabilities when faced with new modal inputs. When the input data form is inconsistent with the modality used during training, the performance of the model often drops significantly, indicating that these models lack sufficient robustness and have difficulty achieving good generalization capabilities in multimodal scenarios.
[0007] 3. Limited language processing capabilities:
[0008] Most existing image generation methods can only process text input in a single language, usually limited to English or other mainstream languages. This limitation makes these methods unable to cope with the needs of multilingual input, especially in the context of globalization, where cross-language text processing capabilities become increasingly important. Existing methods do not perform well in multilingual scenarios and are difficult to meet the needs of users of different languages.
[0009] 4. The quality of generation needs to be improved:
[0010] Although methods such as GAN and VAE have made significant progress in image generation, they still have certain shortcomings in generating high-quality, realistic images, especially when dealing with complex scenes. These methods sometimes fail to accurately capture the details in complex scenes, and the quality of the generated images is still far behind the real images. In addition, when multimodal inputs are involved, the quality of the generated images often deteriorates further and cannot meet the requirements of practical applications. Summary of the invention
[0011] The purpose of the present invention is to provide a multi-language and multi-modal method for generating virtual face images for skits, thereby solving the aforementioned problems existing in the prior art.
[0012] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0013] A multi-language and multi-modal method for generating virtual face images for skits, comprising the following steps:
[0014] S100, inputting the audio and text to be processed into a semantic extraction module to extract semantic features;
[0015] S200, input the time step into the embedding layer to obtain the time step feature;
[0016] S300, inputting the extracted semantic features and time step features together into an image generation module to obtain a final image.
[0017] In some specific embodiments, the semantic extraction module includes:
[0018] The Text Embedding module is used to fuse two sets of embeddings from different encoders to generate the final unified text embedding;
[0019] Speech encoder: The speech encoder uses the Whisper encoder architecture; it can efficiently process speech input and extract deep speech features;
[0020] A speech encoder converts an audio signal into a two-dimensional spectral representation.
[0021] In some specific embodiments, the step of generating text embedding by the Text Embedding module includes:
[0022] Norm+Linear processing: Normalize and linearly transform the output embedding of Chinese-CLIP and the output embedding of mT5+MLP respectively;
[0023] Introducing the cross-attention mechanism: used to capture the association between semantic information from different encoders; semantic information is specifically the similarities and differences between different languages and modalities;
[0024] A multi-head self-attention mechanism is introduced to further process the fused features; the self-attention mechanism can capture the long-distance dependencies and contextual information in the text, so that the final text embedding can contain local information and reflect the overall context;
[0025] Introducing residual connections: The initial input information is retained by introducing residual connections and added to the output of the attention mechanism; residual connections are used to avoid information loss and improve gradient propagation problems during training.
[0026] In some specific embodiments, it also includes: an embedding fusion module, which is used to align the embeddings of the text encoder and the audio encoder.
[0027] In some specific embodiments, Chinese-CLIP is used to process Chinese texts. Relying on its advantages in joint learning of images and texts, the model can more accurately capture the semantic relationship between texts and images when processing Chinese texts containing visual elements.
[0028] The mT5 encoder has the ability to process multilingual texts. Through its multi-task learning mechanism, it can effectively handle input in multiple languages including English, French, and Spanish, ensuring the compatibility and adaptability of the system in multilingual scenarios.
[0029] In some specific embodiments, in Cross Attention, the output embedding of Chinese-CLIP is used as a query, and the output embedding of mT5 participates in the attention calculation as a key and value.
[0030] The beneficial effects of the present invention are:
[0031] The present invention discloses a multi-language and multi-modal method for generating virtual face images for skits, comprising the following steps: inputting the audio and text to be processed into a semantic extraction module to extract semantic features; inputting the time step into an embedding layer to obtain the time step features; and inputting the extracted semantic features and the time step features together into an image generation module to obtain the final image. The present invention can not only realize the use of text and speech in different languages as prompts to control the generated images, but also effectively reduce the impact of low-quality data on the model effect. At the same time, the online data screening method is adopted, which also plays a positive role in the model effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 It is a flow chart of a multi-language and multi-modal skit virtual face image generation method of the present invention;
[0033] Figure 2 It is the analysis flow chart of the Text Embedding module of the present invention;
[0034] Figure 3 It is a schematic diagram of the structure of the speech encoder of the present invention;
[0035] Figure 4 It is a schematic structural diagram of another embodiment of the present invention. DETAILED DESCRIPTION
[0036] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation methods described herein are only used to explain the present invention and are not used to limit the present invention.
[0037] The present invention first designs an embedding fusion module to align the embeddings of the text encoder and the audio encoder.
[0038] The design of the text encoder part of the present invention is the key to realize multilingual input processing of the present invention. In order to take into account the text understanding and generation capabilities in a multilingual environment, two text encoders are used in the system, namely Chinese-CLIP and mT5 text encoders.
[0039] The output of mT5 passes through a multi-layer perceptron (MLP) to further deeply fuse and map the features. MLP not only enhances the semantic representation from the outputs of different encoders, but also plays a role in unifying the feature dimensions. This processing step ensures that the representation of texts in different languages in the feature space is more consistent, allowing the model to more smoothly capture its inherent semantic relevance and contextual dependencies when processing multilingual texts. This unification of features greatly improves the accuracy and performance of the model in multilingual translation, generation and other tasks, especially in complex cross-language scenarios, enabling more accurate text generation and translation.
[0040] After obtaining two sets of embeddings with the same dimensions, the present invention designs a fusion module Textembedding, which aims to fuse the two sets of embeddings from different encoders to generate the final unified text embedding. The design of this module is not just a simple splicing of features from different sources, but through a series of carefully designed steps, it ensures the deep fusion of multi-source information, so that the model can more comprehensively understand the semantics and contextual information of multilingual texts.
[0041] Reference Figure 1 , Figure 2 , Figure 3 and Figure 4 A multi-language and multi-modal method for generating virtual face images for skits is provided, comprising the following steps:
[0042] S100, inputting the audio and text to be processed into a semantic extraction module to extract semantic features;
[0043] S200, input the time step into the embedding layer to obtain the time step feature;
[0044] S300, inputting the extracted semantic features and time step features together into an image generation module to obtain a final image.
[0045] In some specific embodiments, the semantic extraction module includes:
[0046] The Text Embedding module is used to fuse two sets of embeddings from different encoders to generate the final unified text embedding.
[0047] Speech encoder: The speech encoder uses the Whisper encoder architecture; it can efficiently process speech input and extract deep speech features;
[0048] A speech encoder converts an audio signal into a two-dimensional spectral representation.
[0049] Text Embedding module analysis:
[0050] The text embedding module consists of several key steps, each of which plays a vital role in the entire system:
[0051] In some specific embodiments, the step of generating text embedding by the Text Embedding module includes:
[0052] Norm+Linear processing: Normalize and linearly transform the output embedding of Chinese-CLIP and the output embedding of mT5+MLP respectively.
[0053] In this embodiment, first, the output embeddings of Chinese-CLIP and mT5+MLP are normalized (Norm) and linearly transformed (Linear). The normalization process can reduce the distribution differences of embeddings from different sources in the feature space, ensuring that they have similar numerical scales, thus laying the foundation for subsequent fusion. The linear transformation further adjusts the representation of the embedding through a mapping operation, so that it can be more effectively compared and processed in the same feature space. This step not only improves the comparability between different embeddings, but also provides a more coordinated input for the cross-attention module.
[0054] Introduce the cross-attention mechanism: used to capture the association between semantic information from different encoders; semantic information specifically refers to the similarities and differences between different languages and modalities.
[0055] Cross Attention:
[0056] The cross attention mechanism is introduced to establish a strong interactive connection between the two sets of embeddings. It can capture the association between the semantic information from different encoders, especially the similarities and differences between different languages and modalities. In Cross Attention, the output embedding of Chinese-CLIP will be used as a query, while the output embedding of mT5 will participate in the attention calculation as a key and value. Through this mechanism, the model can better understand the common patterns and unique semantic features in different languages, thereby improving the accuracy of cross-language tasks.
[0057] A multi-head self-attention mechanism is introduced to further process the fused features; the self-attention mechanism can capture the long-distance dependencies and contextual information in the text, so that the final text embedding can contain local information and reflect the overall context.
[0058] Multi-head Self-Attention:
[0059] After the cross-attention processing, the model uses the multi-head self-attention mechanism to further process the fused features. The purpose of introducing the self-attention mechanism is to capture the long-distance dependencies and contextual information in the text, so that the final text embedding can not only contain local information but also reflect the overall context. This module processes the input in parallel through a multi-head mechanism, mining the information in the features from different angles, and then generating richer and more diverse semantic representations. This step significantly improves the system's language understanding ability, making the model perform better in multilingual text tasks.
[0060] Introducing residual connections: The initial input information is retained by introducing residual connections and added to the output of the attention mechanism; residual connections are used to avoid information loss and improve gradient propagation problems during training.
[0061] Residual Connection:
[0062] In the figure, the residual connection is introduced to avoid information loss and improve the gradient propagation problem during training. By retaining the initial input information and adding it to the output of the attention mechanism, the residual connection ensures the stability of the model and enhances its expressiveness. This mechanism allows the model to process deep semantic features without losing the original shallow information, thereby improving the accuracy of the model in multilingual semantic representation.
[0063] In some specific embodiments, it also includes: an embedding fusion module, which is used to align the embeddings of the text encoder and the audio encoder.
[0064] In some specific embodiments, Chinese-CLIP is used to process Chinese texts. Relying on its advantages in joint learning of images and texts, the model can more accurately capture the semantic relationship between texts and images when processing Chinese texts containing visual elements.
[0065] The mT5 encoder has the ability to process multilingual texts. Through its multi-task learning mechanism, it can effectively handle input in multiple languages including English, French, and Spanish, ensuring the compatibility and adaptability of the system in multilingual scenarios.
[0066] In some specific embodiments, in Cross Attention, the output embedding of Chinese-CLIP is used as a query, and the output embedding of mT5 participates in the attention calculation as a key and value.
[0067] Whisper Encoder
[0068] In the present invention, the speech encoder uses the Whisper encoder. The Whisper model is based on the classic Transformer Encoder-Decoder architecture, which can efficiently process speech input and extract deep speech features. The input is 30 seconds of audio, which is processed and converted into a log-Mel spectrogram (using a sliding window size of 20ms). This preprocessing step converts the audio signal into a two-dimensional spectrum representation for subsequent feature extraction and time compression.
[0069] The specific processing steps are as follows:
[0070] 1. One-dimensional convolution layer: The input log-Mel spectrum graph first passes through two layers of one-dimensional convolution network (1DConvolution), which is used to extract the underlying features of the audio and perform time-series compression on the audio. The convolution layer can reduce the length of the data in the time dimension and extract a more compact audio feature representation, so that the subsequent Transformer Encoder can process shorter feature sequences and improve the computational efficiency of the model.
[0071] 2. Transformer Encoder: After being processed by the convolutional layer, the audio features will enter several layers of Transformer Encoder. Transformer Encoder captures the long-term and short-term dependencies and complex timing patterns in the audio signal through the multi-head self-attention mechanism. The multi-layer Transformer Encoder can gradually extract high-level feature representations in the audio and generate audio embeddings. These embeddings represent important semantic information in the input audio and can provide rich feature support for subsequent tasks.
[0072] 3. Fusion of text and audio embeddings: While the Whisper encoder generates audio embeddings, the text encoder generates text embeddings. Next, these two embeddings are fused through a module called PromptFuseBlock. This module is designed to integrate features from different modalities (text and audio) to generate the final prompt embedding.
[0073] οNorm+Linear: The audio and text embeddings are first normalized (Norm) and linearly transformed (Linear) to ensure the consistency of the two in the feature space and are ready for further fusion.
[0074] οCross Attention: In PromptFuseBlock, the interaction between audio and text features is first realized through the cross-attention mechanism. This mechanism can capture the cross-modal association between audio and text, so that the model can better understand the intrinsic connection between them when fusing multimodal information.
[0075] οMulti-head Self-Attention: Subsequently, the fused features are further processed through a multi-head self-attention mechanism. This module ensures that the fused features not only retain the uniqueness of each modality, but also combine the contextual information in the audio and text at a deeper level, thereby generating richer prompt embeddings.
[0076] 4. Prompt Embedding output: Finally, after the fusion processing of PromptFuseBlock, the generated prompt embedding contains a comprehensive representation of audio and text information. This fused representation has a wide range of application scenarios in multimodal tasks, and is particularly suitable for complex tasks that require simultaneous processing of voice and text input, such as multi-language translation, voice question and answer, and multimodal generation tasks.
[0077] Whisper Encoder architecture advantages:
[0078] Multimodal processing capabilities: Through convolutional layers, Transformer Encoder, and PromptFuseBlock, Whisper encoder can process audio and text inputs simultaneously and effectively fuse the features of the two, giving the system powerful semantic understanding and generation capabilities in multimodal scenarios.
[0079] Efficient feature extraction: One-dimensional convolution is used for time series compression, which enables the subsequent Transformer Encoder to improve computational efficiency without sacrificing feature expression capabilities, making it suitable for processing long-duration audio data.
[0080] Cross-modal fusion: Through cross-attention and self-attention mechanisms, PromptFuseBlock can deeply fuse features of different modalities and improve the performance of the system when processing multimodal tasks.
[0081] This architectural design can not only process voice input, but also deeply integrate with text input to generate high-quality prompt embeddings, providing strong support for multimodal natural language processing tasks.
[0082] Image generator: image generation model. The present invention adopts the same image generation structure as HunyuanDit. Different from the original structure, in the present invention, Figure 4 The Text Prompt in is the multimodal multilingual prompt in the present invention, and the overall structure is shown in the structural diagram of the speech encoder part.
[0083] The present invention makes full use of open source datasets to train models, but in practice, it is found that some datasets have quality problems, especially the mismatch between text and image descriptions or the improper use of ambiguous words, such as the words "apple" and "Xiaomi" have different meanings in different scenarios. In order to solve these problems, the present invention designs a set of data processing procedures to improve data quality and optimize model performance.
[0084] Data processing flow:
[0085] 1. Image-text similarity screening:
[0086] First, for the English dataset, the CLIP model is used to extract the embeddings of the image and text; for the Chinese dataset, the Chinese CLIP model is used to calculate the similarity. By calculating the embedding similarity between the image and the corresponding text, the mismatched image-text pairs are roughly screened out and the samples with low similarity are removed. This step can effectively eliminate the image-text pairs that obviously do not conform to the expected semantics, thereby improving the quality of the dataset.
[0087] 2. Automatic image annotation:
[0088] Next, the GPT-4o model is used to automatically annotate the image to generate more accurate description text. The newly generated text labels will form new image-text pairs with the images, further improving the accuracy of the data. Then, the similarity calculation process in the first step is repeated to remove the image-text pairs with low similarity again to ensure that the generated description has a high semantic match with the image. This step solves the problem of inaccurate image-text descriptions through an automated annotation system, especially in the case of polysemous words.
[0089] 3. Image stitching and new description generation:
[0090] In order to further expand the size of the training data set, an innovative method was adopted: randomly splicing two pictures together and using GPT-4o to generate new text describing these spliced pictures. This not only increases the diversity of training samples, but also generates more complex and scene-rich picture-text pairs, allowing the model to better learn semantic relationships in different scenarios. This step improves the model's ability to cope with complex scenarios by increasing the complexity and diversity of the data.
[0091] 4. Generate similarity screening of pictures:
[0092] Use the trained model to generate corresponding images based on the text in the training set and extract the embedding of the generated image. Then, calculate the similarity with the embedding of the original image and remove the image-text pairs with too high similarity. This step reduces samples that are too similar, prevents data redundancy, and ensures that sufficient diversity is retained in the training set. After similarity screening, continue to fine-tune the model to further improve the robustness and generalization ability of the model.
[0093] Overall advantages:
[0094] Through such a data processing flow, the present invention not only solves the problems of inaccurate image-text matching and semantic ambiguity in open source datasets, but also significantly improves the quality and diversity of datasets through automatic annotation and data expansion. This process can ensure that the model has higher accuracy and robustness in complex multi-language and multi-modal scenarios, and significantly improves the training effect of the model.
[0095] By adopting the above technical solution disclosed in the present invention, the following beneficial effects are obtained:
[0096] The present invention discloses a multi-language and multi-modal method for generating virtual face images for skits, comprising the following steps: inputting the audio and text to be processed into a semantic extraction module to extract semantic features; inputting the time step into an embedding layer to obtain the time step features; and inputting the extracted semantic features and the time step features together into an image generation module to obtain the final image. The present invention can not only realize the use of text and speech in different languages as prompts to control the generated images, but also effectively reduce the impact of low-quality data on the model effect. At the same time, the online data screening method is adopted, which also plays a positive role in the model effect.
[0097] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be considered as the scope of protection of the present invention.
Claims
1. A multi-language and multi-modal method for generating virtual face images for skits, characterized in that: The following steps are involved: S100, inputting the audio and text to be processed into a semantic extraction module to extract semantic features; S200, input the time step into the embedding layer to obtain the time step feature; S300: Input the extracted semantic features and the time step features together into an image generation module to obtain a final image.
2. The multi-language and multi-modal skit virtual face image generation method according to claim 1, characterized in that: The semantic extraction module comprises: The Text Embedding module is used to fuse two sets of embeddings from different encoders to generate the final unified text embedding; A speech encoder, wherein the speech encoder adopts a Whisper encoder architecture; the speech encoder can efficiently process speech input and extract deep speech features; The speech encoder is capable of converting an audio signal into a two-dimensional spectral representation.
3. The multi-language and multi-modal skit virtual face image generation method according to claim 2, characterized in that: The steps of generating the text embedding by the Text Embedding module include: Norm+Linear processing: Normalize and linearly transform the output embedding of Chinese-CLIP and the output embedding of mT5+MLP respectively; Introducing the cross-attention mechanism: used to capture the association between semantic information from different encoders; the semantic information is specifically the similarities and differences between different languages and modalities; A multi-head self-attention mechanism is introduced to further process the fused features; the self-attention mechanism can capture the long-distance dependencies and contextual information in the text, so that the final text embedding can contain local information and reflect the overall context; Introducing residual connections: The initial input information is retained by introducing the residual connections and added to the output of the attention mechanism; the residual connections are used to avoid information loss and improve the gradient propagation problem during training.
4. The multi-language and multi-modal skit virtual face image generation method according to claim 3, characterized in that: It also includes: an embedding fusion module, which is used to align the embeddings of the text encoder and the audio encoder.
5. The multi-language and multi-modal skit virtual face image generation method according to claim 4, characterized in that: The Chinese-CLIP is used to process Chinese text. Relying on its advantages in joint learning of images and text, the model can more accurately capture the semantic relationship between text and images when processing Chinese text containing visual elements. The mT5 encoder has the ability to process multilingual texts. Through its multi-task learning mechanism, it can effectively handle input in multiple languages including English, French, and Spanish, ensuring the compatibility and adaptability of the system in multilingual scenarios.
6. The multi-language and multi-modal skit virtual face image generation method according to claim 5, characterized in that: In the Cross Attention, the output embedding of the Chinese-CLIP is used as a query, and the output embedding of the mT5 is used as a key and value to participate in the attention calculation.
Citation Information
Patent Citations
Traffic scene video description generation method and device based on multi-modal feature fusion
CN115496134A
Text image cross-modal pedestrian retrieval method and system based on implicit relation reasoning alignment
CN116383671A
Mongolian multi-modal sentiment analysis method based on cross-modal transformer
CN118364427A
Multi-modal sentiment analysis method combining pre-training model and self-attention block
CN118898046A
Cited By
Digital human video generation method and system based on audio driving
CN120602740A