A user prompt-based audio-driven digital human generation system and method

CN118968579BActive Publication Date: 2026-09-08HARBIN INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410903224.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-08
Publication Date
2026-09-08
Estimated Expiration
2044-07-08

AI Technical Summary

Technical Problem

[0006]本发明为解决传统的数字人生成方式在语音表现方面往往存在不足且对输入人脸的要求高的技术问题,进而提出一种基于用户提示的音频驱动数字人生成系统及方法

Benefits of technology

[0053] 1. In the process of digital human generation, this invention ensures a high degree of correlation between the generated digital human image and the input audio content. Through advanced audio feature extraction technology, the system can capture semantic and emotional information in the audio signal and map it onto the digital human's facial expressions and features, thereby achieving precise matching between audio and digital human behavior. Furthermore, the system of this invention is highly personalized. Users can choose their preferred digital human style according to their own preferences and needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968579B_ABST
    Figure CN118968579B_ABST
Patent Text Reader

Abstract

The application provides an audio-driven digital human generation system and method based on user prompts, wherein the system comprises a database module, an audio feature extraction module, an AIGC generated face image module, an Audioface module, an audio-driven face image module and an audio-driven digital human action generation module.The application realizes audio-driven digital human generation based on user prompts, generates content according to user input prompts, gives digital humans highly personalized features and natural behavior performance, and has important application value and prospects, and with the continuous development and improvement of related fields, the application can bring richer experiences and application scenarios to the fields of digital entertainment, virtual reality, human-computer interaction and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an audio-driven digital human generation system and method based on user prompts, belonging to the field of artificial intelligence technology. Background Technology

[0002] With the rapid advancement of technology and the explosive growth of artificial intelligence, the digital human industry is ushering in unprecedented development opportunities. Digital humans, in their diverse forms, have deeply penetrated various fields such as film and entertainment, marketing, e-commerce live streaming, and financial services. They are not only reshaping the business ecosystem and user experience, but also becoming important tools for enterprises to enhance service innovation and improve efficiency. In the process of driving the digital transformation of industries, digital humans have demonstrated enormous application value and development potential.

[0003] The generation of digital humans is a hot research topic in the field of artificial intelligence. Its core task is to generate highly realistic and natural digital human models that can achieve near-human realism in terms of morphology, movement, facial expressions, and speech.

[0004] Currently, there are two main approaches in the field of digital human generation. The first is based on static images. This method typically utilizes deep learning and image processing techniques to analyze and extract key data such as facial features, texture information, and posture from static images, thereby generating digital human models with similar appearances. While this method can reproduce facial details well in static states, it may be relatively limited in handling dynamic expressions and movements. The other approach is based on dynamic video generation. This method captures and analyzes information such as the movements, facial expressions, and voice synchronization of characters in video sequences to generate more vivid and realistic digital human characters. Video-based digital human generation technology can not only simulate continuous human movement but also achieve high synchronization with voice, allowing the generated digital human model to exhibit more natural movements and facial expressions during interaction.

[0005] However, traditional digital human generation methods perform poorly in terms of voice-action consistency and have high requirements for the input face. To address this issue, this embodiment proposes an audio-driven digital human generation system and method based on user prompts, i.e., audio-based digital human generation. This method mines key information such as speech features, voiceprint information, and semantic expressions from audio signals, and constructs a digital human character that highly matches the timbre, semantics, and phonemes in the audio based on this information. Audio-based digital human generation not only enables voice interaction of the digital human character but also endows the digital human with unique personality and emotional expression based on voice characteristics. This new method is expected to open up broader possibilities for the application of digital human technology and enhance the interactive experience and realism of digital human characters. Summary of the Invention

[0006] To address the technical problems of traditional digital human generation methods, which often suffer from deficiencies in voice performance and have high requirements for input facial features, this invention proposes an audio-driven digital human generation system and method based on user prompts.

[0007] The technical solution adopted by this invention to solve the above problems is as follows: This invention proposes an audio-driven digital human generation system based on user prompts, comprising:

[0008] The system includes a database module, an audio feature extraction module, an AIGC facial portrait generation module, an Audioface module, an audio-driven facial image generation module, and an audio-driven digital human motion generation module.

[0009] The database module is used to accurately associate each person's voice sample with the corresponding image information, and to match real-world speaking humans with their corresponding language content and emotional information.

[0010] The audio feature extraction module is used to extract speech feature information from audio data, convert speech information into text information and analyze it, and transmit the audio features to the AIGC face portrait generation module.

[0011] The AIGC facial profile generation module receives audio features from a specific user and prompts input by the user to generate a personalized facial image.

[0012] The Audioface module is used to generate face images that match human voice characteristics;

[0013] The audio-driven face image module is used to align audio and lip movements, generating digital human videos where lip movements are consistent with the audio.

[0014] The audio-driven digital human motion generation module is used to generate audio-based digital human videos with natural human movements and lip movements that are highly consistent with the audio.

[0015] Optionally, the AIGC facial image generation module includes a first generator and a first discriminator;

[0016] The first generator is used to deeply integrate semantic extraction results, user prompts, and user-selected styles to generate a facial profile;

[0017] The first discriminator is used to judge the facial image to determine whether the facial image is real and whether it is consistent with the distribution of the input data.

[0018] Optionally, the Audioface module includes a second generator and a second discriminator;

[0019] The second generator is used to cross-fuse the face image generated by the AIGC face portrait generation module and the speech feature information of the audio feature extraction module to generate a new face image.

[0020] The second discriminator is used to determine whether the new face images obtained by the generator are real or fake, and returns the discrimination result to the generator in the form of negative feedback.

[0021] An audio-driven digital human generation method based on user prompts, comprising:

[0022] Step 1: Construct a database module that matches human voices with facial images and matches semantic and emotional information in the audio.

[0023] Step 2: Extract speech feature information from the audio data based on the audio feature extraction module, convert the speech information into text information and analyze it to obtain audio analysis results, and transmit the audio analysis results to the AIGC face portrait generation module;

[0024] Step 3: Based on the deep integration of semantic extraction results, user prompts, and user-selected styles by the AIGC-generated face portrait module, a face image is generated;

[0025] Step 4: Based on the Audioface module, combined with the face images generated by the AIGC face portrait generation module and the speech extraction results of the audio feature extraction module, the database module is used as the dataset for the Audioface module, and face images that match human voice features are selected from the dataset.

[0026] Step 5: Use the face image with human voice characteristics in the Audioface module and the speech extraction result of the audio feature extraction module as input to the audio-driven face image module, identify and reconstruct the feature points of the lips, and output a digital human video with lip movements that match the audio.

[0027] Step 6: Input the driving video and a face image that matches human voice characteristics into the audio-driven digital human motion generation module. Use First Order Motion technology to perform digital human motion transfer. Input the image frame sequence from the motion-transfer processed video and the speech extraction results extracted by the audio feature extraction module into the generative adversarial network. Train the network using a library in the database module that matches the semantic and emotional information in the audio to generate a digital human that matches the semantic and emotional information. Then, perform masking reconstruction on the digital human video output by the audio-driven face image module where the lip movements match the audio to generate a digital human video with natural human movements and lip shapes that are highly consistent with the audio.

[0028] Optionally, the step of constructing a database module in step 1 that matches human speech with facial images and matches semantic and emotional information in the audio includes:

[0029] Step 1.1: Using speech recognition and image processing technologies, the sound samples are accurately extracted and compared, and the face images are accurately identified and located to obtain a library of matching human voice and face images. Based on the library of matching human voice and face images, the comprehensive information of sound and image is accurately matched.

[0030] Step 1.2: Construct a library that matches semantic and emotional information in audio. By using the speech, semantic, emotional, and nonverbal information in the library that matches semantic and emotional information in audio, capture the speaker's real communication state and generate a vivid and natural digital human.

[0031] Optionally, step 2, which involves obtaining the audio analysis results, includes:

[0032] Step 2.1: Divide the continuous speech signal into several frames, and perform a comprehensive analysis of the speech signal in the time domain and frequency domain for each frame;

[0033] Step 2.2: Based on time domain analysis, obtain the correspondence between waveform changes, amplitude fluctuations and time changes of the speech signal. Based on frequency domain analysis, use Fourier transform to convert the speech signal in the time domain to the frequency domain, explore and extract the spectral characteristics of the speech signal, and complete the speech extraction.

[0034] Step 2.3: Convert the speech signal to text using natural language processing technology to complete semantic extraction;

[0035] Step 2.4: Combine the speech extraction results and semantic extraction results to obtain the audio analysis results.

[0036] Optionally, step 3, which generates the face image, includes:

[0037] Step 3.1: Based on the AIGC facial profile generation module, perform image data learning and analysis to extract facial features and patterns;

[0038] Step 3.2: Deeply integrate the semantic extraction results, user prompts, and user-selected styles, and input them into the first generator to generate a facial profile by combining facial features and patterns;

[0039] Step 3.3: Based on the first discriminator, judge whether the face image is real and whether it is consistent with the distribution of the input data, and feed the judgment result back to the first generator;

[0040] Step 3.4: The first generator adjusts its internal parameters and structure based on feedback to generate an optimized facial image, which is then transmitted to the first discriminator for judgment.

[0041] Step 3.5: Repeat steps 3.3-3.4 until the accuracy of the generated facial image reaches the preset value.

[0042] Optionally, step 4, which involves filtering facial images that match human voice characteristics, includes:

[0043] Step 4.1: Input the face image and speech extraction results into the second generator in the Audioface module. The generator processes and analyzes the face image and speech extraction results through a neural network, cross-merges the face image and speech extraction results, and generates a new face image.

[0044] Step 4.2: The second discriminator is used to determine whether the new face images obtained by the generator are real or fake, and returns the determination result to the second generator in the form of negative feedback;

[0045] Step 4.3: The second generator adjusts its internal parameters and structure based on feedback, optimizes the face image, and transmits the optimized face image to the second discriminator for judgment;

[0046] Step 4.4: Repeat steps 4.2-4.3 until the second discriminator can no longer determine the authenticity of the face image generated by the second generator and outputs the corresponding face image.

[0047] Optionally, step 5, which involves outputting a digital human video whose lip movements match the audio, includes:

[0048] Step 5.1: Input the face image and speech extraction results into the generative adversarial network in the audio-driven face image module;

[0049] Step 5.2: Use gradient constraints to limit the gradient variation between adjacent pixels, and introduce the L-softmax loss function and the cross-entropy loss function;

[0050] Step 5.3: The L-softmax loss function adjusts the decision boundary between lip movements and other movements by introducing the parameter m, enabling the audio-driven face image module to learn the mapping relationship between audio and lip movements;

[0051] Step 5.4: Based on the cross-entropy loss function, calculate the error between the digital human video output by the audio-driven face image module that matches the audio lip movements and the real video, and adjust the output digital human video that matches the audio lip movements until the error is less than a preset threshold.

[0052] The beneficial effects of this invention are:

[0053] 1. In the process of digital human generation, this invention ensures a high degree of correlation between the generated digital human image and the input audio content. Through advanced audio feature extraction technology, the system can capture semantic and emotional information in the audio signal and map it onto the digital human's facial expressions and features, thereby achieving precise matching between audio and digital human behavior. Furthermore, the system of this invention is highly personalized. Users can choose their preferred digital human style according to their own preferences and needs.

[0054] 2. Building upon the high accuracy of lip-reading, this invention further introduces a motion module to enhance the expressiveness and naturalness of the digital human. Through deep learning technology, the system proposed in this invention can learn and simulate real human facial expressions and movements, such as blinking and smiling, as well as head movements like nodding and tilting, and even hand gestures. These features make the generated digital human more natural and vivid in its behavior, better conveying the emotional and semantic information in the audio.

[0055] 3. This invention realizes audio-based digital human generation and endows the digital human with highly personalized features and natural behavioral performance, making this invention of significant application value and prospects. With the continuous development and improvement of related fields, this invention can bring richer experiences and application scenarios to the fields of digital entertainment, virtual reality, and human-computer interaction. Attached Figure Description

[0056] Figure 1 A structural framework diagram of an audio-driven digital human generation system based on user prompts provided by the present invention;

[0057] Figure 2 A flowchart of an audio-driven digital human generation method based on user prompts provided by the present invention;

[0058] Figure 3 This is a structural diagram of the database module provided by the present invention;

[0059] Figure 4 This is a structural diagram of the audio feature extraction module provided by the present invention;

[0060] Figure 5 This is a structural diagram of the AIGC facial image generation module provided by the present invention;

[0061] Figure 6 This is a framework diagram of the Audioface module provided by the present invention;

[0062] Figure 7 A framework diagram of the audio-driven face image module provided by the present invention;

[0063] Figure 8 This is a framework diagram of the audio-based digital human motion generation module provided by the present invention. Example

[0064] Example 1

[0065] like Figure 1 As shown in the figure, the structure of the user-prompt-based audio-driven digital human generation system described in this embodiment includes:

[0066] The system includes a database module, an audio feature extraction module, an AIGC facial portrait generation module, an Audioface model module, an audio-driven facial image module, and an audio-based digital human motion generation module.

[0067] The database module is used to accurately associate each person's voice sample with the corresponding image information, and to match real-world speaking humans with their corresponding language content and emotional information.

[0068] The audio feature extraction module is used to extract speech feature information from audio data, convert speech information into text information and analyze it, and transmit the audio analysis results to the AIGC face portrait generation module.

[0069] The AIGC facial image generation module is used to generate facial images.

[0070] The Audioface model module is used to generate face images that match human voice characteristics;

[0071] An audio-driven face image module is used to generate digital human videos where lip movements match the audio.

[0072] The audio-based digital human motion generation module is used to generate audio-based digital human videos with natural human movements and lip movements that are highly consistent with the audio.

[0073] The AIGC facial image generation module includes a first generator and a first discriminator;

[0074] The first generator is used to deeply integrate semantic extraction results, user prompts, and user-selected styles to generate a facial profile;

[0075] The first discriminator is used to judge the facial image to determine whether the facial image is real and whether it is consistent with the distribution of the input data.

[0076] The AIGC facial image generation module includes a first generator and a first discriminator;

[0077] The first generator is used to deeply integrate semantic extraction results, user prompts, and user-selected styles to generate a facial profile;

[0078] The first discriminator is used to judge the facial image to determine whether the facial image is real and whether it is consistent with the distribution of the input data.

[0079] Example 2

[0080] Combination Figure 2-8 This embodiment will be described as follows: Figure 2 As shown in the figure, the process of the audio-driven digital human generation method based on user prompts described in this embodiment includes:

[0081] S1: Construct a database module that matches human voices with facial images and matches semantic and emotional information in the audio.

[0082] like Figure 3 As shown, this embodiment establishes a precise and comprehensive correspondence between sound and image information, ensuring that each person's sound sample perfectly matches their corresponding facial image information. To this end, this embodiment employs advanced speech recognition and image processing technologies to accurately extract and compare features from sound samples, while simultaneously performing high-precision recognition and localization of facial images. Through these steps, it is ensured that the sound samples and facial image information of each individual in the database are accurate and mutually corresponding. The construction of this database not only facilitates subsequent identification of individuals but also ensures that subtle differences in each individual's short-term spectrum, sound source, temporal dynamics, prosody, and linguistic features can be accurately identified during recognition and comparison processes. This information is crucial for understanding and analyzing an individual's voice characteristics and speech patterns. Through such a matching database, the combined information of sound and images can be used for matching more comprehensively and deeply. This not only helps this embodiment further understand the correspondence between speech and facial profiles but also enhances the technical level of this embodiment in fields such as facial recognition and speech recognition.

[0083] Furthermore, this embodiment constructs a database that matches real-world human speakers with the semantic and emotional information in their audio content. This database will focus more on people's natural facial expressions and movements during speech, such as nodding, tilting their heads, and gestures, which play a crucial role in real communication. However, existing research often limits itself to the analysis of speech and human lip movements, neglecting this rich nonverbal information. Therefore, this embodiment constructs a database that integrates speech, semantics, emotion, and nonverbal information, covering a wider range of movement and facial expression data to capture a more realistic communication state of the speaker. Through such a database, it will be possible to generate more vivid and natural digital humans, allowing them to exhibit expressions and reactions closer to humans during communication, thereby significantly improving the interactive experience and application value of digital humans.

[0084] S2: Extract speech feature information from audio data based on the audio feature extraction module, convert speech information into text information and analyze it to obtain audio analysis results;

[0085] like Figure 4 As shown, audio feature extraction includes speech extraction and semantic extraction.

[0086] S201: Speech extraction;

[0087] This embodiment divides the continuous speech signal into several short time periods, which are called frames. The purpose of this division is to ensure that the features presented by the speech signal can maintain relative stability within a relatively short time range. This stability provides great convenience for subsequent feature extraction and analysis operations. For each frame of the speech signal, various methods can be flexibly selected for feature extraction according to actual needs. This embodiment uses a combination of time-domain and frequency-domain analysis, filter bank technology, and deep learning-based feature extraction to extract audio features.

[0088] S20101: Comprehensive analysis based on time and frequency domains.

[0089] In time-domain analysis, the focus is primarily on the waveform changes and amplitude fluctuations of speech signals over time. These characteristics provide rich information for understanding the intuitive representation of speech. In frequency-domain analysis, the powerful capabilities of mathematical tools such as the Fourier transform are used to cleverly convert speech signals, originally in the time domain, to the frequency domain, enabling in-depth exploration and extraction of the spectral characteristics inherent in speech. These characteristics are closely related to the individual characteristics of the speaker, such as vocal tract structure and pronunciation habits, and can reflect the uniqueness and differences in their speech to a certain extent.

[0090] S20102: Extracting audio features based on filter bank technology to obtain unique audio features for each individual.

[0091] Filter banks are an effective signal processing method that decomposes complex audio signals into multiple sub-bands, each corresponding to a different frequency range. This method allows for more detailed analysis of the features contained within the audio signal. Filter banks play a crucial role in audio feature extraction. They operate on specific frequency regions to extract features closely related to an individual's vocal characteristics. These features include pitch, timbre, and formants, which together constitute each person's unique audio "fingerprint." Using the audio features extracted by filter banks, the distinctive characteristics of each person's voice can be further recorded. By comparing the audio features of different individuals, the speaker's vocal characteristics can be accurately identified, and the speech signal can be categorized into the appropriate class.

[0092] S20103: Feature extraction based on deep learning.

[0093] Deep learning-based speech feature extraction is an important technique in speech recognition and speech processing. It utilizes deep neural network models to automatically extract key features from raw speech signals, providing effective input for subsequent speech recognition, speech synthesis, or other speech processing tasks. First, raw speech signals are typically presented in the form of waveforms, amplitudes, and frequencies, which contain rich speech information. Deep learning-based speech feature extraction methods use deep neural network models to learn and extract useful features from these raw speech signals. These deep neural network models include, but are not limited to, convolutional neural networks (CNNs), recurrent neural networks (RNNs), and long short-term memory networks (LSTMs). Through layer-by-layer learning and abstraction, these network structures can automatically discover complex patterns and structures in speech signals and extract feature representations useful for subsequent tasks.

[0094] In deep learning-based speech feature extraction, the raw speech signal first needs to be preprocessed, such as through framing, windowing, and feature calculation. Then, the processed speech data is input into a deep neural network model. The network model learns how to extract key features from the speech data through training. These features can be low-level acoustic features, such as spectrum and energy, or higher-level semantic features, such as phonemes, word, or sentence-level information. Compared with traditional feature extraction methods, deep learning-based speech feature extraction has many advantages. First, it allows for automatic feature learning; that is, deep neural networks can automatically learn and extract features useful for speech recognition without the need for manual feature design and selection. Furthermore, it possesses strong feature representation capabilities and robustness.

[0095] S202: Semantic extraction;

[0096] This embodiment extracts specific content information from complex audio data through meticulous processing and analysis. This information plays an indispensable role in comprehensively and accurately understanding and judging a person's identity. The semantic extraction process primarily focuses on the rich linguistic features contained in the audio, including but not limited to precise word selection, rigorous grammatical structure, and personalized pronunciation habits. These linguistic features can largely reflect the speaker's educational background, professional experience, regional culture, and other deep-seated social attributes, thus providing strong clues and evidence for revealing their true identity in this embodiment. To achieve efficient semantic extraction from audio, this embodiment can utilize the currently popular Natural Language Processing (NLP) technology. With its powerful text processing capabilities, NLP technology can easily and accurately transcribe audio, rapidly transforming the originally elusive speech signal into clear and understandable text data. Based on this, this embodiment can further utilize NLP technology to deeply analyze the rich semantic information contained in this text data, thereby providing a deeper and more comprehensive perspective for fully interpreting the speaker's true intentions and identity characteristics.

[0097] S3: Based on the deep integration of semantic extraction results, user prompts and user-selected styles in the AIGC-generated face portrait module, a face image is generated;

[0098] like Figure 5As shown, the AIGC facial portrait generation module learns and analyzes a large amount of image data to extract facial features and patterns. This data may include facial images of various styles, angles, and expressions, providing rich materials and references for the subsequent generation process. In the generation phase, the AIGC facial portrait generation module utilizes the framework of Generative Adversarial Networks (GANs). A GAN consists of two main parts: a generator and a discriminator. The generator's task is to generate new facial portraits based on input information (such as a user's photo or specific parameters). The discriminator's task is to determine whether the generated facial portraits are realistic and consistent with the original data distribution. Initially, the generator produces some relatively rough facial portraits, which the discriminator then evaluates to determine their approximation of realistic facial portraits. Based on the discriminator's feedback, the generator adjusts its internal parameters and structure to generate more realistic facial portraits. This process is iterative; with each iteration, the facial portraits generated by the generator become increasingly realistic and detailed. During this process, the discriminator also continuously learns and improves its discrimination ability. It attempts to identify flaws and shortcomings in the paintings generated by the generator, allowing the generator to improve in subsequent iterations. Ultimately, through continuous competition and optimization between the generator and discriminator, the AIGC facial portrait generation module can produce high-quality facial paintings with diverse styles. These works can be realistic, abstract, or even innovative, incorporating multiple elements. In this paper, the AIGC painting process for generating portraits has been pre-trained and demonstrates good generation results; therefore, it is not used in subsequent model training.

[0099] The system employs highly sophisticated algorithms, first deeply analyzing and utilizing the core semantic text information extracted from the speech signal. Simultaneously, user-provided descriptive words and their chosen style serve as key inputs, deeply integrating with the core semantic text information. Users can vividly and meticulously depict their desired facial features through descriptive words, including the shape of the eyes, the contour of the nose, and the curve of the mouth—each detail reflecting the user's personalized needs. The user's chosen style provides a more diverse range of representations for the facial image, whether classical or modern, realistic or abstract, all can be expressed through style selection. Through this combination, the system can generate facial portraits that not only meet the user's personalized needs but also possess a high degree of creativity. This creativity is not only reflected in the appearance of the portrait but also in the artistry and emotional expression it embodies.

[0100] Even if the user does not provide descriptive terms or select a style, this embodiment still demonstrates powerful intelligent processing capabilities. AIGC technology automatically evaluates whether the semantic information extracted from the speech signal is sufficient to support the generation of a high-quality facial profile. If the system determines that the information is insufficient, this embodiment intelligently filters popular and context-appropriate descriptive terms from a vast database to supplement the profile. This intelligent filtering and random combination method ensures the completeness of the profile while increasing its diversity and interest.

[0101] The AIGC facial portrait generation module, leveraging the powerful capabilities of AIGC technology, has successfully generated new, realistic, and artistic facial portraits. These portraits not only bring users an unprecedented creative experience but also inject new vitality and creativity into the field of facial image generation.

[0102] S4: Based on the Audioface model module, combining the face image and speech extraction results, the database module is used as the dataset for the Audioface model module, and face images that match human voice characteristics are selected from the dataset.

[0103] like Figure 6 As shown, the Audioface model is a generative adversarial network module. Its core task is to use the AIGC to generate face images generated by the face portrait module, and combine them with the speech features of human voice to further generate face images that are more consistent with human voice features.

[0104] There is a subtle yet close connection between human vocal qualities and physical appearance. This connection is not only evident in acoustics and morphology but also widely applied in technologies such as speech generation and facial reconstruction. One of the most intuitive and significant connections is the substantial impact of gender and age on vocal frequencies. For example, male voices are typically deep and powerful, while female voices are often higher-pitched and more melodious; as people age, their voices may also become more mature and stable or slightly hoarse. Furthermore, an individual's body shape, particularly body type and facial features, plays a crucial role in the production and transmission of sound. This is because the body's morphological structure directly affects the vibration patterns and frequencies of air within the body, thus influencing the timbre and quality of the voice. For instance, differences in facial bone structure and muscle distribution can subtly affect the resonance and reflection of sound, making each person's voice unique. In addition, the degree of importance and attention people place on their image also influences their pronunciation. Generally, people who pay more attention to their appearance may tend to use clearer and more standard pronunciation to showcase their confidence and charm; while those who are more casual or less concerned about their appearance may pronounce words more naturally and casually. Based on the above analysis, we can conclude that generating corresponding facial images based on the input human voice features has a certain degree of rationality and scientific basis. This is because there is an inseparable intrinsic connection between voice and face, and this connection can be extracted and reconstructed through technical means to achieve more accurate and vivid facial image generation. The implementation of the Audioface model module relies on a key dataset, namely the database module's library for matching human voice and images. This library provides the module with rich voice and image samples, enabling the system to learn the intrinsic connections and patterns between voice and images.

[0105] In the Audioface module, the first step is to input facial images and human speech features into the generator. The generator processes and analyzes the two types of input data through a neural network, effectively combining their information. A key aspect of this process is the cross-fusion of speech features and facial images. This means the generator needs to find a way to effectively map speech features onto facial images, so that the generated image better reflects the characteristics of the speech.

[0106] After the generator completes the fusion process and generates new face images, these images are sent to the discriminator for verification. The discriminator's main task is to distinguish between real and fake generated face images. It performs binary classification on the real sample data matched in the database and the data input from the generator to determine whether the data belongs to real samples. The discriminator returns the discrimination result to the generator in the form of negative feedback. The generator adjusts its parameters and structure based on this feedback to improve its generation results.

[0107] As training progresses, the generator and discriminator continuously compete and optimize. The generator constantly improves the quality and diversity of its generated face images to deceive the discriminator as much as possible; while the discriminator continuously improves its discrimination ability to more accurately identify the differences between generated and real images. When the discriminator can no longer accurately distinguish between real and fake samples, the entire system reaches an equilibrium state, at which point the training process ends.

[0108] Ultimately, through processing and optimization by the Audioface module, this embodiment can obtain facial images that more closely resemble human voice characteristics. These images not only possess high realism and detail but also fully demonstrate the mapping and representation of speech features on facial images. The successful implementation of this module not only improves the level of facial image generation technology but also provides new ideas and methods for cross-modal generation between speech and images.

[0109] S5: The face image with human voice features in the Audioface model module and the speech extraction results of the audio feature extraction module are used as inputs to the audio-driven face image module to identify and reconstruct the feature points of the lips, and output a digital human video with lip movements that match the audio.

[0110] like Figure 7 As shown, lip movements are a crucial part of speech expression and are essential for achieving realism and naturalness in digital human videos. Therefore, by accurately capturing and reconstructing key lip points, the system can generate digital human videos with lip movements highly consistent with the audio.

[0111] To achieve consistency between lip features and audio, this module employs further pixel-based constraints. The core objective of these constraints is to ensure that the generated image maintains pixel-level consistency with the original image or target features. This means that the color, brightness, saturation, and other attributes of each pixel need to be strictly controlled. Achieving this goal relies primarily on precise pixel-level comparison and adjustment. This embodiment first uses gradient constraints to focus on the rate of change between adjacent pixels. In image generation or editing tasks, by limiting the gradient of pixel value changes, it can be ensured that the generated image remains smooth in local areas, avoiding abrupt jumps. This guarantees that the lip shape of the generated facial video image does not change drastically in a short period, resulting in a more coherent and natural lip shape in the generated digital human.

[0112] Furthermore, by introducing a more complex generative adversarial network structure and more accurate loss functions such as L-softmax loss function and cross-entropy loss function, the system can better learn the mapping relationship between audio and lip movements, enabling the generated digital human video to achieve a high degree of consistency between lip movements and audio.

[0113] The L-Softmax loss function is a loss function used for training deep learning models, especially in classification tasks such as face recognition and object recognition. It's an improvement on the traditional Softmax loss function, aiming to increase the compactness of intra-class features and the separability of inter-class features, thereby enhancing the model's classification performance. The traditional Softmax loss function primarily maximizes the probability of the correct class while minimizing the probability of the incorrect class. However, its control over inter-class distance and intra-class compactness is relatively weak, sometimes leading to samples from different classes being too tightly distributed in the feature space, making them difficult to distinguish effectively. The L-Softmax loss function introduces a parameter m (often called the margin) to adjust the decision boundary between different classes. This margin parameter makes the distribution of samples from different classes more dispersed in the feature space while maintaining the compactness of intra-class samples. In this way, the L-Softmax loss function can improve the model's discriminative ability, enabling the model to more accurately determine the class of new, unknown samples. Specifically, the L-Softmax loss function achieves these goals by optimizing an objective function related to sample features and class labels. During training, the model adjusts its parameters based on the output of the loss function to minimize its value. This allows the model to learn more meaningful feature representations, thereby improving classification performance. Furthermore, the L-Softmax loss function does not directly change the model's structure or the number of parameters; rather, it influences the training process by altering how the loss function is calculated. Therefore, it can be combined with various deep learning models (such as convolutional neural networks and recurrent neural networks) to enhance classification performance.

[0114] The cross-entropy loss function is a widely used loss function in machine learning and deep learning, especially when dealing with classification problems. It measures the difference between the probability distribution predicted by the model and the true probability distribution. In classification problems, this typically involves a sample's true label (usually one-hot encoded) and the model's output (usually the output processed by a softmax function, representing the predicted probability for each class). The cross-entropy loss function measures the difference between these two.

[0115] Furthermore, the audio-driven face image module draws upon numerous existing research cases and advanced technologies to ensure its effectiveness and reliability in practical applications. These research cases provide valuable references and experience for the module's design and implementation, enabling this embodiment to better handle various complex scenarios and changes. Through the processing of the audio-driven face image module, this embodiment ultimately obtains a digital human video that is highly consistent with the audio, where lip movements are natural and smooth, perfectly matching the speech content. The successful implementation of this module enhances the realism and naturalness of the digital human video. S6: Input the generation results of S1-S5 into the audio-based digital human motion generation module to obtain the digital human video;

[0116] like Figure 8 As shown, this embodiment combines advanced computer vision technology and deep learning algorithms, innovatively employing First Order Motion technology to achieve motion transfer in digital humans. This enables the digital human to exhibit actions and behaviors similar to those of real humans. To achieve this, this embodiment first imports the face image processed by the driving video and Audioface model module as input data into the module. First Order Motion technology utilizes motion information from the video and images to extract keyframes and their corresponding motion patterns, then maps them onto the digital human model to achieve motion transfer. In this way, the digital human can imitate various actions of real humans, including postures and gestures, thereby enhancing its expressiveness and realism.

[0117] To further enhance the digital human's motion performance and emotional expression, this embodiment further fuses the image frame sequence from the motion transfer processed video with the human voice audio content information extracted by the second module. This step is achieved through a carefully designed generative adversarial network, which can effectively combine image and audio data to generate digital human actions that match the audio semantic information and emotional expression. In this way, the digital human's actions not only remain consistent with the audio content but also reflect corresponding emotional changes, thereby enhancing the realism and expressiveness of the video. Furthermore, this embodiment fully utilizes the real-world speaking human and language content matching library constructed in the first module. This library contains a large amount of real human language and motion data. By using this data for training, the system in this embodiment can better understand human language habits and action expressions. This makes the digital human more natural and realistic when imitating human actions, and better adaptable to the needs of various scenarios.

[0118] After generating the digital human motion, to ensure a high degree of consistency between lip movements and audio, this embodiment further integrates the lip movement data obtained from the audio-driven face image module. By performing precise occlusion and reconstruction processing on the lip movements, this embodiment ensures that the final generated digital human video perfectly matches the audio in terms of lip movements, further enhancing the realism and naturalness of the video.

[0119] Through this series of complex processing steps, this embodiment ultimately generates audio-generated digital human videos with natural human movements and perfect lip-sync to audio. The successful implementation of this module not only significantly enhances the expressiveness of digital humans but also brings new breakthroughs and application prospects to audio-based digital human video generation technology. Whether in entertainment, education, or other fields, this digital human technology can provide users with more realistic and engaging experiences, driving the development and application of related technologies.

[0120] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent substitutions, and improvements made to the above embodiments without departing from the scope of the present invention, based on the technical essence of the present invention and within the spirit and principles of the present invention, shall still fall within the protection scope of the present invention.

Claims

1. A user-prompt-based audio-driven digital human generation system, characterized in that, The structure of the user-prompt-based audio-driven digital human generation system includes: The system includes a database module, an audio feature extraction module, an AIGC facial portrait generation module, an Audioface module, an audio-driven facial image generation module, and an audio-driven digital human motion generation module. The database module is used to accurately associate each person's voice sample with the corresponding image information, and to match real-world speaking humans with their corresponding language content and emotional information. The audio feature extraction module is used to extract speech feature information from audio data, convert speech information into text information and analyze it, and transmit the audio features to the AIGC face portrait generation module. The AIGC facial profile generation module receives audio features from a specific user and prompts input by the user to generate a personalized facial image. The AIGC-based facial profiling module deeply integrates semantic extraction results, user prompts, and user-selected styles to generate facial images. The Audioface module is used to generate face images that match human voice characteristics; Based on the Audioface module, combined with the face images generated by the AIGC face portrait generation module and the speech extraction results of the audio feature extraction module, the database module is used as the dataset of the Audioface module, and face images that match human voice features are selected from the dataset. The audio-driven face image module is used to align audio and lip movements, generating digital human videos where lip movements are consistent with the audio. The facial image that conforms to human voice characteristics in the Audioface module and the speech extraction result of the audio feature extraction module are used as inputs to the audio-driven facial image module to identify and reconstruct the feature points of the lips, and output a digital human video in which the lip movements match the audio. The audio-driven digital human motion generation module is used to generate audio-generated digital human videos with natural human movements and lip movements that are highly consistent with the audio. The driving video and a face image that matches human voice features are input into the audio-driven digital human motion generation module. First Order Motion technology is used for digital human motion transfer. The image frame sequence from the motion-transferred video and the speech extraction results extracted by the audio feature extraction module are input into the generative adversarial network. The network is trained using a library in the database module that matches the semantic and emotional information in the audio. A digital human that matches the semantic and emotional information is generated. The digital human video with lip movements that match the audio output by the audio-driven face image module is then masked and reconstructed to generate a digital human video with natural human movements and lip shapes that are highly consistent with the audio.

2. The audio-driven digital human generation system based on user prompts according to claim 1, characterized in that, The AIGC facial image generation module includes a first generator and a first discriminator; The first generator is used to deeply integrate semantic extraction results, user prompts, and user-selected styles to generate a facial profile; The first discriminator is used to judge the facial image to determine whether the facial image is real and whether it is consistent with the distribution of the input data.

3. The audio-driven digital human generation system based on user prompts according to claim 1, characterized in that, The Audioface module includes a second generator and a second discriminator; The second generator is used to cross-fuse the face image generated by the AIGC face portrait generation module and the speech feature information of the audio feature extraction module to generate a new face image; The second discriminator is used to determine whether the new face image obtained by the generator is real or fake, and returns the determination result to the generator in the form of negative feedback.

4. A user-prompt-based audio-driven digital human generation method, applied to the user-prompt-based audio-driven digital human generation system according to any one of claims 1-3, characterized in that, include: Step 1: Construct a database module that matches human voices with facial images and matches semantic and emotional information in the audio. Step 2: Extract speech feature information from the audio data based on the audio feature extraction module, convert the speech information into text information and analyze it to obtain audio analysis results, and transmit the audio analysis results to the AIGC face portrait generation module; Step 3: The AIGC-based facial profile generation module deeply integrates semantic extraction results, user prompts, and the user-selected style to generate a facial image; Step 4: Based on the Audioface module, combined with the face images generated by the AIGC face portrait generation module and the speech extraction results of the audio feature extraction module, the database module is used as the dataset of the Audioface module, and face images that match human voice features are selected from the dataset. Step 5: Use the face image that matches human voice characteristics in the Audioface module and the speech extraction result of the audio feature extraction module as input to the audio-driven face image module, identify and reconstruct the feature points of the lips, and output a digital human video whose lip movements match the audio. Step 6: Input the driving video and a face image that matches human voice features into the audio-driven digital human motion generation module. Use First Order Motion technology to perform digital human motion transfer. Input the image frame sequence from the motion-transfer processed video and the speech extraction results extracted by the audio feature extraction module into the generative adversarial network. Train the network using a library in the database module that matches the semantic and emotional information in the audio to generate a digital human that matches the semantic and emotional information. Then, perform masking reconstruction on the digital human video output by the audio-driven face image module where the lip movements match the audio to generate a digital human video with natural human movements and lip shapes that are highly consistent with the audio.

5. The method for generating an audio-driven digital human based on user prompts according to claim 4, characterized in that, Step 1 involves constructing a database module that matches human voices with facial images and also matches semantic and emotional information in the audio. Step 1.1: Using speech recognition and image processing technologies, the sound samples are accurately extracted and compared, and the face images are accurately identified and located to obtain a library of matching human voice and face images. Based on the library of matching human voice and face images, the comprehensive information of sound and image is accurately matched. Step 1.2: Construct a library that matches semantic and emotional information in audio. By using the speech, semantic, emotional, and nonverbal information in the library that matches semantic and emotional information in audio, capture the speaker's real communication state and generate a vivid and natural digital human.

6. The method for generating an audio-driven digital human based on user prompts according to claim 4, characterized in that, Step 2, which involves obtaining the audio analysis results, includes the following steps: Step 2.1: Divide the continuous speech signal into several frames, and perform a comprehensive analysis of the speech signal in the time domain and frequency domain for each frame; Step 2.2: Based on time domain analysis, obtain the correspondence between waveform changes, amplitude fluctuations and time changes of the speech signal. Based on frequency domain analysis, use Fourier transform to convert the speech signal in the time domain to the frequency domain, explore and extract the spectral characteristics of the speech signal, and complete the speech extraction. Step 2.3: Convert the speech signal to text using natural language processing technology to complete semantic extraction; Step 2.4: Combine the speech extraction results and semantic extraction results to obtain the audio analysis results.

7. The method for generating an audio-driven digital human based on user prompts according to claim 4, characterized in that, Step 3, which involves generating a face image, includes: Step 3.1: Based on the AIGC facial profile generation module, perform image data learning and analysis to extract facial features and patterns; Step 3.2: Deeply integrate the semantic extraction results, user prompts, and user-selected styles, and input them into the first generator to generate a facial profile by combining facial features and patterns; Step 3.3: Based on the first discriminator, judge whether the face image is real and whether it is consistent with the distribution of the input data, and feed the judgment result back to the first generator; Step 3.4: The first generator adjusts its internal parameters and structure based on feedback to generate an optimized facial image, which is then transmitted to the first discriminator for judgment. Step 3.5: Repeat steps 3.3-3.4 until the accuracy of the generated facial image reaches the preset value.

8. The method for generating an audio-driven digital human based on user prompts according to claim 4, characterized in that, Step 4, which involves selecting facial images that match human voice characteristics, includes the following steps: Step 4.1: Input the face image and speech extraction results into the second generator in the Audioface module. The generator processes and analyzes the face image and speech extraction results through a neural network, and cross-merges the face image and speech extraction results to generate a new face image. Step 4.2: The second discriminator is used to determine whether the new face images obtained by the generator are real or fake, and returns the determination result to the second generator in the form of negative feedback; Step 4.3: The second generator adjusts its internal parameters and structure based on feedback, optimizes the face image, and transmits the optimized face image to the second discriminator for judgment; Step 4.4: Repeat steps 4.2-4.3 until the second discriminator can no longer determine the authenticity of the face image generated by the second generator and outputs the corresponding face image.

9. The method for generating an audio-driven digital human based on user prompts according to claim 4, characterized in that, Step 5, which involves outputting a digital human video where the lip movements match the audio, includes the following steps: Step 5.1: Input the face image and speech extraction results into the generative adversarial network in the audio-driven face image module; Step 5.2: Use gradient constraints to limit the gradient variation between adjacent pixels, and introduce the L-softmax loss function and the cross-entropy loss function; Step 5.3: The L-softmax loss function adjusts the decision boundary between lip movements and other movements by introducing a parameter m, enabling the audio-driven face image module to learn the mapping relationship between audio and lip movements; Step 5.4: Based on the cross-entropy loss function, calculate the error between the digital human video output by the audio-driven face image module that matches the audio and the real video, and adjust the output digital human video that matches the audio until the error is less than a preset threshold.