Virtual human animation generation method and device, electronic equipment and storage medium
By combining large language models, image generation models, voice-face generation models, and text-to-speech models, the problem of poor matching between virtual human voice and image has been solved, enabling rapid generation of personalized virtual human images and efficient animation synthesis, thereby improving the generation efficiency and immersiveness of virtual human animations.
Patent Information
- Application Number
- CN202610025030.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-09
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2046-01-09
AI Technical Summary
In existing virtual human animation generation technologies, the adaptability of virtual human voice timbre is poor. The voice timbre provided by the platform is mostly bound to the preset virtual human model and cannot be flexibly changed. Moreover, it is difficult to match the generated virtual human image in terms of age, occupation, temperament, and other dimensions, which affects the overall coordination and immersion of the virtual human animation.
By analyzing user needs and generating detailed features using a pre-set large language model, generating candidate virtual human images using a pre-trained image generation model, performing feature analysis and timbre matching using a pre-trained voice-face generation model, performing speech synthesis using a text-to-speech model, and performing animation synthesis using an audio-driven image model, the virtual human image and timbre are automatically matched.
It enables the rapid generation of personalized virtual human images based on simple needs, reduces the complexity and cost of image customization, ensures the coordination between the generated text and audio and the target virtual human image, and improves the generation efficiency and overall immersion of virtual human animation.
Smart Images

Figure CN121482229A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of computers, in particular to a virtual person animation generation method and device, electronic equipment and a storage medium. BACKGROUND
[0002] Virtual person animation generation technology, as an important application direction of the integration of artificial intelligence and multimedia technology, has been widely used in live streaming, intelligent customer service, media broadcasting and other scenarios. At present, related Internet and AI (Artificial Intelligence) companies have launched virtual person related products, and the core technology logic of which is to combine text-to-speech (Text To Speech, TTS) technology, image generation technology and voice-driven image generation animation technology to realize the rapid generation of virtual person speaking video. Specifically, the user inputs the dialogue text, selects the virtual person image provided by the platform, and then generates the corresponding voice through the platform. After synthesis processing, a virtual person animation video is obtained, and some platforms also support limited selection of voice tone and timbre.
[0003] However, the related virtual person animation generation technology still has defects to be solved. For example, the adaptability of the virtual person's timbre is poor, and the timbre provided by the platform is mostly bound to the preset virtual person model and cannot be flexibly changed. Even in the scene of uploading pictures to generate an image, the number of optional timbres is small, and it is difficult to match the generated virtual person image in the dimensions of age, occupation, temperament, etc., affecting the overall coordination and immersion of the virtual person animation.
[0004] Therefore, how to provide a virtual person animation generation scheme that can quickly generate personalized virtual person images based on simple user requirements and automatically match the timbre consistent with the virtual person image has become a key problem to be solved in the current technical field. SUMMARY
[0005] The embodiments of the present application provide a virtual person animation generation method, device, electronic equipment, computer program and storage medium, which can improve the adaptability of virtual person timbre and image.
[0006] The embodiments of the present application provide a virtual person animation generation method, comprising: obtaining an input requirement for generating a virtual person image, analyzing and processing the requirement through a preset large language model, and generating detailed features corresponding to the virtual person image; inputting the detailed features as a generation prompt word into a pre-trained image generation model, and generating at least one candidate virtual person image based on the detailed features through the image generation model; The pre-trained voice-face generation model is used for feature analysis and timbre matching analysis processing on a target virtual human image selected from the at least one candidate virtual human image, to obtain a target timbre matched with the target virtual human image; the voice-face generation model is trained based on a pre-constructed speech data set and a paired face data set; The text-to-speech model is used for speech synthesis processing on the target timbre and input scripts associated with the target virtual human image, to obtain script audio of the target timbre. The audio-driven image model is used for animation synthesis processing on the target virtual human image and the script audio, to obtain corresponding virtual human animation.
[0007] The embodiment of the application further provides a virtual human animation generation device, comprising: A feature generation unit is configured to obtain input requirements for generating a virtual human image, analyze and process the requirements by using a pre-set large language model, and generate detailed features corresponding to the virtual human image. An image generation unit is configured to input the detailed features as a generation prompt to a pre-trained image generation model, and generate at least one candidate virtual human image based on the detailed features by using the image generation model. A timbre matching unit is configured to perform feature analysis and timbre matching analysis processing on a target virtual human image selected from the at least one candidate virtual human image by using a pre-trained voice-face generation model, to obtain a target timbre matched with the target virtual human image; the voice-face generation model is trained based on a pre-constructed speech data set and a paired face data set. A speech synthesis unit is configured to perform speech synthesis processing on the target timbre and input scripts associated with the target virtual human image by using a text-to-speech model, to obtain script audio of the target timbre. An animation synthesis unit is configured to perform animation synthesis processing on the target virtual human image and the script audio by using an audio-driven image model, to obtain corresponding virtual human animation.
[0008] The embodiment of the application further provides an electronic device comprising a processor and a memory, wherein the memory stores a plurality of instructions; the processor loads the instructions from the memory to perform steps in any virtual human animation generation method provided by the embodiment of the application.
[0009] The embodiment of the application further provides a computer-readable storage medium storing a plurality of instructions, wherein the instructions are adapted to be loaded by a processor to perform steps in any virtual human animation generation method provided by the embodiment of the application.
[0010] In addition, the embodiment of the present application also provides a computer program product comprising a computer program or instructions, which, when executed by a processor, implements the steps of the virtual human animation generation method provided by the embodiment of the present application.
[0011] In the present application, by presetting a large language model to analyze and process the simple image demand input by the user, the detailed features corresponding to the virtual human image can be automatically generated without the user providing a complex description or uploading a picture with privacy and copyright risks, and at least one candidate virtual human image can be generated by the image generation model, realizing the rapid generation of personalized virtual human images based on simple demand, greatly reducing the complexity and cost of image customization. Further, by using the voice-face generation model trained based on the pre-constructed voice data set and the paired face data set, the feature analysis and tone matching can be automatically performed for the target virtual human image selected by the user, and the target tone matching the target image can be accurately obtained, effectively solving the problem of mismatch between tone and virtual human image in age, occupation, temperament and other dimensions in related technologies, and ensuring the coordination of the subsequent generated script audio and target virtual human image. Finally, the script audio with the corresponding tone is generated by the text-to-speech model, and the virtual human animation is synthesized by the audio driving image model, realizing the automation and intelligent processing from the virtual human image demand to the animation output, significantly improving the generation efficiency and overall immersion of the virtual human animation, and meeting the user's demand for personalized and highly adaptive virtual human animation. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0013] Figure 1a is a scene schematic diagram of the virtual human animation generation method provided by the embodiment of the present application; Figure 1b is a flowchart of the virtual human animation generation method provided by the embodiment of the present application; Figure 2 is a flowchart of another virtual human animation generation method provided by the embodiment of the present application; Figure 3 is a structural schematic diagram of the virtual human animation generation device provided by the embodiment of the present application; Figure 4 is a structural schematic diagram of the electronic device provided by the embodiment of the present application. DETAILED DESCRIPTION
[0014] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0015] The related virtual human animation generation technology has significant technical defects in actual application, and it is difficult to meet the needs of users for personalized and highly adaptive virtual human content. Firstly, the customization ability of virtual human image is insufficient, and the related technology mainly relies on a small number of virtual human image libraries pre-prepared by the platform, and users can only make limited choices. If personalized image customization is needed, either personal photos or network pictures need to be uploaded as the basis of the image, which may cause compliance risks such as user privacy leakage (such as misuse of facial biometric information) and copyright infringement (such as unauthorized use of other people's images), or detailed image descriptions need to be provided to the platform, which not only has a complicated process and high time cost, but also the customization effect is limited by the efficiency of human communication, and it is difficult to guarantee the user's expectations. Secondly, the adaptability of the voice tone and the image of the virtual human is poor. In the related technology, the voice tone is usually bound to the pre-prepared virtual human image, and cannot be changed flexibly. Even in the scenario of uploading pictures to generate images, the number of optional voice tones is also very limited, and there is no association mechanism between voice tone and image characteristics (such as age, occupation, temperament, and gender), resulting in a serious gap between the generated voice tone and the virtual human image in the sensory dimension, which greatly reduces the immersion and credibility of the virtual human animation. Thirdly, there is a lack of coordination between image generation and speech synthesis. The related technology does not form a linkage logic of "image characteristics-speech characteristics", and the image generation and speech selection are independent of each other, which needs to be manually matched by the user, not only increasing the operation complexity, but also easily leading to inconsistent final animation effect due to human judgment deviation.
[0016] To solve the above technical problems, the embodiments of the present application provide a virtual human animation generation method, device, electronic equipment and storage medium.
[0017] The virtual human animation generation device can be integrated in an electronic equipment, which can be a terminal, a server or the like.
[0018] The server can be a standalone physical server, a server cluster or a distributed system formed by multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery network (CDN), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, or a personal computer (PC), but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application.
[0019] In some embodiments, the terminal can also be used as a server to implement part or all of the functions of the server.
[0020] For example, referring to Figure 1a , Figure 1a is a scene schematic diagram of a virtual human animation generation method provided by an embodiment of the present application. Taking the case that a virtual human animation generation apparatus is integrated in Figure 1a the electronic device shown in the figure, the electronic device can be a terminal, a server, or the like as shown in Figure 1a . The electronic device can obtain an input requirement for generating a virtual human image, analyze and process the requirement through a pre-set large language model, and generate detailed features corresponding to the virtual human image; input the detailed features as a generation prompt into a pre-trained image generation model, and generate at least one candidate virtual human image based on the detailed features through the image generation model; analyze and process the features and the timbre matching of a target virtual human image selected from the at least one candidate virtual human image through a pre-trained voice-face generation model, to obtain a target timbre matched with the target virtual human image; the voice-face generation model is trained based on a pre-constructed voice data set and a paired face data set; perform voice synthesis processing on the target timbre and input scripts associated with the target virtual human image through a text-to-speech model, to obtain script audio of the target timbre; perform animation synthesis processing on the target virtual human image and the script audio through an audio-driven image model, to obtain a corresponding virtual human animation. The adaptability of the virtual human timbre and image can be improved.
[0021] The following will be described in detail. It should be noted that the serial numbers of the following embodiments do not limit the preferred order of the embodiments.
[0022] A virtual human animation generation method provided by an embodiment of the present application is applied to an electronic device as shown in Figure 1a , and the specific process of the virtual human animation generation method can be as follows as shown in Figure 1b . 101. Obtain the input requirements for generating a virtual human image, analyze and process the requirements through a preset large language model, and generate detailed features corresponding to the virtual human image.
[0023] Virtual human images refer to the digital human images ultimately used to generate animations. They have visually recognizable appearance attributes, such as facial features, hairstyle, clothing, and expression. They are the core visual carriers of virtual human animations, such as "subway service guide images" and "children's program host images".
[0024] The requirements for generating virtual avatars refer to the initial instructions input by users that define the core attributes of the virtual avatar. These requirements do not need to include detailed descriptions of the avatar; they only need to clearly define the core scene or role positioning. They are concise requirements with low information density, such as "a 25-year-old female hospital guide" or "a cartoon-style elementary school science teacher." Furthermore, these requirements can be in text or image format, further enhancing the flexibility of input and adapting to the usage habits of different users.
[0025] Large language models refer to artificial intelligence models with natural language understanding and feature analysis capabilities. They can perform semantic parsing and attribute expansion based on input information and support input requirements in the form of text or images. For example, the LLM (Large Language Model) model with image description analysis capabilities has the core function of transforming the user's simplified requirements into structured features required for image generation.
[0026] The pre-deployed large language model refers to a large language model that has been pre-deployed and whose parameters have been tuned. This model has optimized the feature output logic through customized question templates to ensure that the generated features meet the prompt word requirements of the subsequent image generation model, rather than an unadapted general large language model, such as a pre-tuned GPT4 (Generative Pre-trained Transformer 4) model that can output standard descriptive words for AI painting.
[0027] Detailed features refer to the set of all-dimensional attributes generated after the large language model analyzes user needs, which are used to accurately define the virtual human image. They cover the virtual human image's age, gender, occupation, hairstyle, clothing, expression, scene adaptation style, etc., and are the core input basis for the image generation model. For example, the detailed features generated for the need of "25-year-old female hospital guide" are "25-year-old female, short hair, wearing light blue medical guide uniform, wearing a name tag, with a gentle smile, standing upright, background adapted to the hospital guide desk scene, and overall style realistic".
[0028] In an exemplary embodiment, the user needs to generate a virtual human image for a hospital guidance scenario. The user inputs the requirement "a 25-year-old female hospital guide" into the system. The system then calls a pre-deployed and debugged large language model, such as the GPT4 model adapted for AI painting descriptive word output. This model performs semantic parsing and attribute expansion on the requirement "25-year-old female hospital guide" based on a built-in question template, automatically supplementing the typical image features of a guide in the scenario. Finally, the large language model generates detailed features corresponding to the virtual human image: "25-year-old female, short black hair, no bangs, wearing a light blue short-sleeved medical guide uniform, a white name tag with the word 'Guide' printed on her left chest, light makeup, a natural smile, upright posture, hands naturally folded in front of her body, realistic image style, suitable for the bright and clean guidance environment of a hospital," providing accurate prompts for the subsequent image generation model.
[0029] In the above embodiments, the user's input requirements for generating a virtual human image are analyzed and processed using a preset large language model, generating detailed features. This effectively solves the technical shortcomings of existing technologies where users need to provide complex image descriptions or upload images to customize a virtual human image. On the one hand, users only need to input concise requirements containing core attributes (such as age, occupation, and scene), without manually adding details or uploading images that pose risks of privacy leaks or copyright infringements. This significantly reduces the user's operational threshold and usage costs, while avoiding the risks of privacy leaks and copyright infringements caused by image uploads. On the other hand, the preset large language model can automatically expand and generate detailed features covering multi-dimensional attributes based on customized logic, ensuring the completeness and accuracy of feature information. It can directly adapt to the prompts of the subsequent image generation model without manual intervention, avoiding the inefficiency and deviation problems caused by manual feature addition. In addition, this step supports input of requirements in the form of text or images, further improving the flexibility of requirement input and adapting to the usage habits of different users.
[0030] 102. Input detailed features as generation prompts into a pre-trained image generation model, and generate at least one candidate virtual human image based on the detailed features through the image generation model.
[0031] Among them, image generation models refer to artificial intelligence models that have the ability to generate visual images based on text prompts. Their core function is to transform structured text descriptions into digital images that conform to the description features. They especially support the accurate restoration of character appearance and scene style, and are the core execution carrier for virtual human image generation. Examples include the stable diffusion (SD) model and the Midjourney model. Among them, the stable diffusion model is more suitable for the refined generation needs of virtual human images because it supports multi-dimensional input control.
[0032] The pre-trained image generation model refers to an image generation model that has been trained through a large-scale image-text paired data set and integrated with an auxiliary control module to optimize the generation effect, and is not a basic model without adaptation. The model has completed parameter optimization in advance and can stably analyze the "detailed feature" prompt word, and improve the controllability of image generation (such as face key point precision, style consistency) through the auxiliary module. For example, the pre-trained stable diffusion model integrated with the ControlNet auxiliary model has preset functions such as canny edge detection and face key point control, which can avoid problems such as face distortion and feature deviation in generated images.
[0033] The candidate virtual human image refers to multiple virtual human image versions generated by the image generation model based on the detailed features, which are selected by the user. Each version meets the core attribute requirements of the detailed features, but there are differences in the details of the style (such as hair curl degree, clothing texture, and subtle differences in expression), providing personalized selection space for the user. For example, based on the detailed features of a "25-year-old female hospital guide", three candidate images are generated: version 1 is straight hair with ear length, light blue long-sleeved guide clothes, version 2 is short hair with slight curl, light blue short-sleeved guide clothes, and version 3 is straight hair with shoulder length, light blue guide clothes with stripe decoration. The three meet the core attributes of "guide", only the details of the style are different.
[0034] In an exemplary embodiment, suppose the generated detailed features are: "25-year-old female, short black hair, no bangs, wearing a light blue short-sleeved medical guide uniform, a white name tag with the word 'Guide' printed on her left chest, light makeup, a natural smile, upright posture, hands naturally folded in front of her body, realistic image style, suitable for the bright and clean guide environment of a hospital." The system uses these detailed features as generation prompts and inputs them into a pre-trained image generation model, specifically a stable model integrated with a ControlNet auxiliary model. The diffusion model, pre-trained for facial keypoint control and realistic style adaptation, first analyzes the core features in the prompts, including age, occupation, clothing, and expression. It then uses ControlNet's facial keypoint control module to locate key positions such as the eyes, corners of the mouth, and shoulders, avoiding the generation of facial proportion imbalances or expressions that deviate from a "smiling" look. Simultaneously, it optimizes the background color (primarily white and light blue) and lighting effects (bright and soft) based on the description of the "hospital guidance environment." Finally, the model generates four candidate virtual human images, all meeting the core requirements of the detailed features. The differences are: Image 1 has straight black hair to the ears and an undecorated guide uniform; Image 2 has slightly wavy black hair to the ears and white stripes on the cuffs of the guide uniform; Image 3 has straight black hair to the ears and light blue lace trim on the collar of the guide uniform; and Image 4 has slightly wavy black hair to the ears and a red border on the name tag on the left chest of the guide uniform. All candidate virtual human images maintain a natural smile and upright posture, with a bright hospital guidance desk scene as the background, allowing users to further select their target virtual human image.
[0035] In the above embodiments, detailed features are input as prompt words into the pre-trained image generation model to generate at least one candidate virtual human image. This effectively solves the technical defects of existing technologies, such as virtual human image customization relying on pre-made libraries or manual design, and the generation effect being uncontrollable. In addition, the high efficiency of the pre-trained model enables the image generation process to be completed quickly (usually a few seconds to tens of seconds), which greatly shortens the customization cycle of virtual human images. Compared with traditional manual design, it saves a lot of time and costs, lays a high-quality visual foundation for subsequent voice matching and animation synthesis, and further improves the overall efficiency of virtual human animation generation and user experience.
[0036] 103. Using a pre-trained voice-face generation model, feature analysis and timbre matching analysis are performed on the target virtual human image selected from at least one candidate virtual human image to obtain the target timbre that matches the target virtual human image; the voice-face generation model is trained based on a pre-constructed speech dataset and its paired face dataset.
[0037] Among them, the voice-face generation model refers to an artificial intelligence model with the ability to establish the correlation between "human face visual features and voice timbre features". The core function is to output a voice timbre that matches the virtual image in the sensory dimension (such as age, gender, and temperament) based on the input virtual image (face features). It is the core carrier of "image-timbre" intelligent adaptation, including two types of matching models based on voice-face relationship (matching the optimal timbre from the timbre library) and generation models based on encoder-decoder (directly generating matching timbre). For example, a matching model can be established by measuring the correlation between face and timbre features, or a generation model can be established by aligning face and voice features with an encoder and generating voice with a decoder.
[0038] The pre-trained voice-face generation model refers to a voice-face generation model that has been trained and optimized using a pre-built voice dataset and its paired face dataset. It is not a basic model that has not been trained. This model has learned the mapping rules of face features and timbre features through training and can stably output a timbre that matches the input image. For example, a voice-face generation model trained based on paired data containing "25-year-old female medical staff face-soft female voice" and "50-year-old male teacher face-steady male voice" can be directly used for timbre matching of different virtual images.
[0039] The target virtual image refers to the final virtual image selected by the user from the multiple candidate virtual images generated in step 102 for subsequent animation generation. This image has clear and fixed visual features (such as facial features, age, and professional attire) and is the core input object for feature analysis and timbre matching by the voice-face generation model. For example, the user selects a virtual image of "black short hair with a slight curl, light blue short sleeve guide uniform, and a gentle smile" from four candidate images of hospital guides.
[0040] The target timbre refers to the voice timbre output by the voice-face generation model after analyzing and processing the target virtual image, which is highly adapted to the image in the sensory dimension. It needs to meet the characteristics of age, gender, professional temperament, and other characteristics conveyed by the target virtual image. It is the core basis for the subsequent text-to-speech model to generate script audio. For example, the "25-year-old female hospital guide" image matches the "25-year-old female voice with a soft tone, moderate speed, and no obvious accent" in standard Chinese.
[0041] The voice dataset refers to a collection of a large number of different voice timbre data used to train the voice-face generation model. The data needs to be standardized (such as cropping to the same length and uniform sampling rate). The source can be obtained by intercepting voice segments from publicly available videos on the Internet. It needs to ensure that each voice sample has a corresponding face sample. For example, it contains "daily conversation voice of different age groups, occupations, and gender groups", "broadcasting voice", "service voice", etc., and each voice segment is 3 seconds long.
[0042] The face dataset refers to a set of face images paired with the voice dataset for training the voice-face generation model, each face sample comes from the video of the corresponding voice segment source, and needs to be consistent with the "personality" of the voice sample, that is, the face and voice of the same person are paired samples, for example, the "service voice of a 25-year-old female medical staff" in the voice dataset corresponds to the face dataset of the female medical staff in the video of the same period of the voice segment.
[0043] In an exemplary embodiment, the user selects the image of "black short curly hair, light blue short sleeve guide dress, white work card on the left chest, gentle smile, and overall 25-year-old female temperament" from the four hospital guide candidate virtual human images as the target virtual human image; the system calls the pre-trained voice-face generation model (the model has been trained based on a dataset containing 100,000 "face-voice" paired data, which covers a large number of medical staff, service personnel, and various professional face and voice data); the voice-face generation model first analyzes the features of the target virtual human image, extracts its face contour (soft lines), age (about 25 years old), and professional related features (guide dress reflects service attributes), and other visual information, and then matches the optimal voice based on the learned "face-voice" mapping rule in the preset voice feature library; the final model outputs the target voice that matches the target image - "25-year-old female voice, gentle and friendly tone, about 120 words per minute, no obvious regional accent, high voice clarity, and service communication temperament in the hospital guide scene", which provides the voice standard for the subsequent text-to-speech step.
[0044] In the above embodiment, the pre-trained voice-face generation model is used to analyze the features of the target virtual human image and match the voice, effectively solving the technical defects of poor virtual human voice and image adaptation, manual selection by the user, and limited selection range in the prior art.
[0045] 104、Through the text-to-speech model, the input script associated with the target voice and the target virtual human image is processed for voice synthesis, and the script audio of the target voice is obtained.
[0046] The text-to-speech model refers to an artificial intelligence model that can convert text information into natural speech signals. Its core function is to synthesize speech audio that meets semantic expression and timbre requirements based on input text content and specified timbre characteristics. It supports adjusting parameters such as speech speed, intonation, and clarity. It is the core carrier connecting "text" and "audio". For example, the mockingbird model and the Tacotron 2 model. The mockingbird model can more accurately reproduce the target timbre and is more suitable for the "audio synthesis based on target timbre" requirement in this scheme.
[0047] It should be noted that the mockingbird model is an open-source speech cloning and synthesis model. It can quickly extract the core features such as timbre and intonation of a target speech sample in just 5 seconds, and synthesize natural speech that is highly consistent with the target timbre based on input text. It supports Chinese adaptation and can adjust parameters such as speech speed and intonation. The Tacotron 2 model is an end-to-end TTS benchmark model. Its core is to directly map text to speech waveform through deep learning, without the complex intermediate steps of traditional TTS (such as phoneme conversion and spectrum generation separation design). The synthesized speech has excellent naturalness and prosody performance.
[0048] The input text associated with the target virtual human image refers to the text content that matches the application scenario and professional attributes of the target virtual human image. It is the semantic basis for the text-to-speech model to synthesize speech and needs to conform to the role positioning of the virtual human image (such as service, broadcast, and teaching). It avoids situations where the text content is inconsistent with the image attributes. This input text can be user input or generated based on related models. For example, the user input text associated with the "25-year-old female hospital guide" image is "Hello, welcome to AB Hospital. Please go to the first floor lobby for registration, and please sign in at the corresponding department for treatment. If you have any questions, you can consult the guide desk staff at any time."
[0049] The target timbre audio refers to the speech audio synthesized by the TTS model based on the target timbre and the input text, which has both semantic accuracy and timbre adaptation. This audio needs to fully convey the text semantics and accurately reproduce the characteristics of the target timbre (such as age, intonation style, and professional temperament). It is the core speech material for the subsequent audio-driven image model to generate virtual human animations. For example, the "standard Chinese female audio with gentle intonation, moderate speed, and clear delivery of guide information" that matches the "25-year-old female hospital guide" image and the corresponding guide text.
[0050] In an exemplary embodiment, assuming that the target virtual human image has been determined to be a "25-year-old female hospital guide", the corresponding target voice tone is "a standard Chinese female voice of about 25 years old, with a gentle tone, a speech rate of about 120 words per minute, and no obvious accent", and the input script associated with the target virtual human image is obtained, that is, "Hello everyone, please show your health code and appointment certificate after entering the hospital. There are registration windows and self-service registration machines in the first floor lobby. The internal medicine clinic is located on the east side of the third floor, and the surgical clinic is located on the west side of the third floor. If you need help, please call the guide personnel around you. I wish you a smooth medical treatment." The system calls a text-to-speech model, for example, a mockingbird model supporting voice cloning function, and inputs the feature parameters of the target voice tone (such as the fundamental frequency range, the intonation fluctuation interval, and the speech rate value) and the input script into the mockingbird model simultaneously. The mockingbird model first analyzes the script semantics to determine the pause position and emotional tone of the sentence (the guide scene needs to be smooth and friendly), and then adjusts the acoustic features of the speech synthesis based on the target voice tone parameters to ensure that the synthesized speech not only conforms to the script semantic expression but also accurately matches the target voice tone. Finally, the script audio of the target voice tone is generated, which has a duration of about 25 seconds, a gentle tone without harshness, and a stable speech rate of 120 words per minute. The audio clearly conveys all the information of the guide script, and the voice tone is highly adapted to the "25-year-old female hospital guide" image, which can be directly used for subsequent animation synthesis.
[0051] In the above embodiment, the text-to-speech model is used to synthesize the voice of the input script associated with the target voice tone and the target virtual human image, effectively solving the problem of poor adaptation of virtual human voice and image in related technologies. The synthesized script audio accurately matches the target voice tone (such as the gentle female voice tone corresponding to the guide image), which is consistent with the age, occupation, and temperament of the virtual human image, and can adjust the pause and tone of the voice based on the script semantics to ensure accurate semantic transmission and natural voice without manual parameter intervention or voice recording. This greatly reduces the operation cost, and the generated audio can be directly used for subsequent animation synthesis, providing a guarantee for the voice quality and overall coordination of virtual human animation.
[0052] 105. The audio-driven image model is used to synthesize the animation of the target virtual human image and the script audio, and the corresponding virtual human animation is obtained.
[0053] The audio-driven image model refers to an artificial intelligence model that can combine speech audio with a static virtual human image to generate dynamic speaking effects. The core function is to analyze the acoustic characteristics of the audio (such as speech speed, tone, and pauses) to drive the facial movements (such as lip opening and closing, facial muscle micro-movements) and body posture (such as slight nodding) of the virtual human image, achieving "voice-action" synchronization. For example, the Sadtalker plugin of Stable Diffusion, the Wav2Lip model, etc. The Sadtalker plugin is a high-fidelity virtual human lip shape synchronization and animation generation plugin under the Stable Diffusion ecosystem, supporting high-definition image driving and natural motion, and is more suitable for virtual human animation generation needs. The Wav2Lip model is a classic cross-modal lip shape synchronization generation model that establishes a mapping relationship between speech audio and lip movements through deep learning. By inputting any face image (real person / virtual human) and speech segment, it can generate a video with precise alignment of lip shape and speech.
[0054] The animation synthesis process refers to the process of coordinating the target virtual human image and the script audio by the audio-driven image model. It specifically includes analyzing the acoustic characteristics of the audio, mapping the virtual human facial movement parameters, generating frame sequences and splicing them into dynamic videos. The core is to ensure the precise alignment of the virtual human motion and the audio timeline, avoiding the problem of "lip shape not matching the speech". For example, each syllable of the guide script audio is matched with the opening and closing amplitude and frequency of the virtual human's lips, and slight head micro-movements are added to enhance naturalness.
[0055] Virtual human animation refers to the video file containing dynamic virtual human images and synchronized speech output after the animation synthesis process. This animation needs to fully present the semantic delivery of the script audio, and the virtual human facial movements, posture, and audio are completely synchronized. It is the final application carrier of virtual human technology. For example, a 30-second dynamic video of a "25-year-old female hospital guide image" making lip opening and closing speaking movements while slightly nodding, accompanied by guide script audio, with a hospital guide desk scene in the background.
[0056] In an exemplary embodiment, it is assumed that a target virtual human image (a 25-year-old female hospital guide) and corresponding script audio have been obtained; the system calls an audio-driven image model (specifically, the Sadtalker plug-in of stable diffusion), inputs the static image of the target virtual human image and the script audio into the audio-driven image model synchronously; the audio-driven image model first analyzes the acoustic characteristics of the audio, extracts the speech rate and fundamental frequency change at each time node, maps them to the opening and closing degree of the virtual human's lips (such as the lips being wide open when pronouncing "a" and the lips being closed when pronouncing "b") and the subtle movements of the facial muscles (such as the degree of upward turning of the corners of the mouth when smiling), and adds slight head rotation and nodding movements based on the guide scene; then the audio-driven image model generates a dynamic image sequence of 24 frames per second, splices all the frame sequences and aligns them with the script audio on the time axis, and finally outputs the corresponding virtual human animation - the animation is 25 seconds long, the speaking lip movements of the virtual human image are completely synchronized with the audio, the posture is natural and not stiff, the background remains the hospital guide desk scene, and the animation can be directly used for actual application in the hospital guide scene.
[0057] In the above embodiment, the audio-driven image model is used to animate the target virtual human image and the script audio, effectively solving the problems of "speech-motion" asynchronization and stiff dynamic effects in the prior art: it can accurately realize the time axis alignment of the virtual human facial movements (such as lip movements) and the script audio, ensure the realism and immersion of the animation, adapt natural body postures based on the scene, avoid stiff dynamic effects, and does not require manual adjustment of action parameters, greatly improving the animation generation efficiency, and the final output of the virtual human animation can be directly applied to the target scene, providing users with high-quality virtual human interaction experience.
[0058] From the above, the embodiment of the present application adopts the virtual human animation generation method, which can effectively solve the technical problems of inconvenience in self-defining virtual human image, poor adaptability of tone and image in the prior art, and has the following specific beneficial technical effects: through the preset large language model, the simple image demand input by the user is analyzed and processed, the detailed features corresponding to the virtual human image can be automatically generated, the user does not need to provide complex description or upload pictures with privacy and copyright risk, at least one candidate virtual human image can be generated by the image generation model, realizing the rapid generation of personalized virtual human image based on simple demand, greatly reducing the complexity and cost of image customization; further, through the voice-face generation model trained based on the pre-constructed voice data set and the paired face data set, the feature analysis and tone matching can be automatically performed for the target virtual human image selected by the user, the target tone matched with the target image is accurately obtained, effectively solving the problem of mismatch between tone and virtual human image in age, occupation, temperament and other dimensions in related technology, ensuring the coordination of the subsequent generated script audio and target virtual human image; finally, the script audio with corresponding tone is generated by the text-to-speech model, and the virtual human animation is synthesized by the audio-driven image model, realizing the automation and intelligent processing from the virtual human image demand to the animation output, significantly improving the generation efficiency and overall immersion of the virtual human animation, meeting the user's demand for personalized and highly adaptive virtual human animation In one exemplary embodiment, the voice-face generation model in step 103 includes two types of voice-face relationship-based matching models and encoder-decoder-based generation models, which will be described in detail below.
[0059] In one embodiment, the training method of the voice-face relationship-based matching model includes the following steps: Step 1: Construct a voice data set and a paired face data set.
[0060] The voice data set is obtained by intercepting the voice in the target video and cutting it into audio files of a preset length, and the face data set is obtained by intercepting the corresponding face image from the video of the voice source.
[0061] The target video refers to the original video material used to extract voice data and corresponding face data to construct the paired data set, which needs to contain clear human voice and contemporaneous human face pictures, and the voice and face of the person in the video need to be one-to-one corresponding, i.e. the voice comes from the person in the picture, which can include public interview videos, service scene recording videos, education teaching videos, etc., to ensure that the human voice and corresponding face image can be synchronously obtained.
[0062] The preset length refers to a uniform length set when the voice segment extracted from the target video is standardized. The purpose is to ensure the format consistency of all audio files in the voice data set, facilitate feature extraction and parameter calculation during subsequent model training, and determine the preset length according to the integrity of the voice features and the training efficiency. Usually, it is 1-5 seconds, for example, the preset length is set to 3 seconds, that is, all voice segments extracted from the target video need to be cut to 3 seconds long. If the original extracted segment is less than 3 seconds, it is processed by zero padding, and if it exceeds 3 seconds, the core voice segment is extracted.
[0063] Exemplarily, in order to construct a voice data set and its paired face data set for training a matching model based on the relationship between voice and face, a target video is first selected. The target video is specifically a public scene recording video containing characters of different professions (such as medical staff, teachers, and bank clerks), different age groups (such as young people aged 20-30 and middle-aged people aged 40-50), and the video needs to meet the requirements that the character's voice is clear and free of noise, the character's face in the same period is unobstructed and the face contour and expression can be clearly identified. For example, the hospital's public guide service recording video, the school's public teacher's teaching video and the bank's public customer service video are selected. Then, the voice and corresponding face data are extracted from each target video: taking the hospital guide service recording video as an example, the voice segment of the guide personnel is extracted during the communication period between the guide personnel and the patient in the video, and the extracted voice segment is standardized according to the preset length (the preset length is set to 3 seconds). If the length of the original voice segment is 3.8 seconds, the middle 3 seconds of the segment with coherent semantics are retained, and if the length of the original voice segment is 2.2 seconds, the length is supplemented to 3 seconds by audio zero padding technology. At the same time, the same period face image (selecting 3 frames of face image with consistent angle and no blur to ensure the completeness of the face features) corresponding to the 3-second voice segment is extracted from the video, and a unique pairing relationship between the 3-second standardized voice segment and the corresponding 3 frames of face image is established. According to the same extraction and processing method as described above, the rest of the target videos such as the teacher's teaching video and the bank's customer service video are processed. Finally, all audio files in the voice data set are 3 seconds long, and each face image in the face data set is paired with an audio file in the voice data set.
[0064] Step two: convert the face images in the face data set into face feature vectors, and convert the audio files in the voice data set into voice feature vectors.
[0065] Exemplarily, in the feature extraction of the constructed face dataset and voice dataset, for each face image in the face dataset, the encoder is used to convert it into a feature vector representation, or a preset image feature extraction algorithm (such as a convolutional neural network-based feature extraction method) is used to analyze the facial key regions (such as facial contour, facial texture, and expression features) in the image, and the analyzed facial feature information is mapped into a fixed-dimension vector form, i.e., a facial feature vector of the corresponding face image is obtained; at the same time, for each audio file in the voice dataset, the acoustic feature extraction technology (such as the Mel frequency cepstral coefficient extraction, fundamental frequency analysis, etc.) is used to quantitatively process the acoustic features such as tone, intonation, and speech rate in the audio, and the processed acoustic feature information is converted into a vector form that is adapted to the dimension of the facial feature vector, and then a voice feature vector of the corresponding audio file is obtained, so as to ensure that the two types of feature vectors can be used for subsequent correlation analysis in model training.
[0066] Step three: taking the voice feature vector as the label and the facial feature vector as the input, calculating the distance loss of the voice feature vector and the facial feature vector in the feature space by the metric learning function, training the initial matching model based on the feature correlation model method of metric learning, and optimizing the initial matching model parameters through multiple rounds of iterative training until the initial matching model can output the tone corresponding to the voice feature vector that matches the facial feature vector corresponding to the input face image to a preset threshold.
[0067] The initial matching model refers to a model prototype that has not been trained and optimized, has a basic “input-output” framework, but the parameters are not adapted to the “face-voice” feature correlation rule. The core architecture includes a feature input layer, a metric learning calculation layer, and a result output layer, which can receive facial feature vector input and try to output corresponding voice feature vector prediction results, but the output results have large deviations from the true voice feature vector in the initial state, and the matching accuracy needs to be improved through training and optimization of parameters. For example, a “face-voice” feature matching model based on a fully connected neural network is built, which only completes the initialization of the network structure but does not perform parameter training.
[0068] The preset threshold refers to a preset determination standard for determining whether the matching accuracy of the voice feature vector output by the model and the true voice feature vector meets the standard. It is usually expressed in terms of distance value (such as Euclidean distance, cosine distance) in the feature space or matching similarity percentage, and needs to be set in combination with the model training target and actual application requirements to ensure that the tone output by the model can be adapted to the sensory dimension of the input face image. For example, the preset threshold is set to “cosine similarity in feature space ≥ 0.85”, i.e., when the cosine similarity of the voice feature vector output by the model and the true voice feature vector reaches or exceeds 0.85, it is determined that the output result meets the standard, and the corresponding tone matches the face image.
[0069] The feature correlation model method based on metric learning is a model training technology that optimizes the correlation between different modal data (face feature vectors and speech feature vectors) in a high-dimensional feature space by constructing a specific metric criterion (i.e., a metric learning function). The core logic is to quantify the distance or similarity between the "input feature (face feature vector)" and the "label feature (speech feature vector)" in the feature space through the metric learning function, to "minimize the positive sample pair distance and maximize the negative sample pair distance" as the optimization goal, to iteratively train the parameters of the initial matching model, and finally to enable the model to have the ability to "input one type of modal feature and output another type of modal feature with high correlation". The essence is to realize accurate feature mapping and correlation matching of cross-modal data. For example, the Learnable Pins model method is a core technical solution for face-speech cross-modal feature correlation learning. The core logic is to construct an "identity-aware cross-modal feature space" to realize the aggregation of face and speech features of the same identity (or high correlation) and the separation of different identity features, while accurately solving the problem of false negative samples in the training data.
[0070] Exemplarily, when training the matching model based on the audio-face relationship, the face feature vector obtained in step two is input as input data, and the corresponding speech feature vector is input as label data into the pre-built initial matching model. Through a preset metric learning function, such as a triplet loss or a contrastive loss, the distance loss between the predicted value of the speech feature vector output by the model and the true speech feature vector label in the feature space is calculated. The feature correlation model method based on metric learning (such as the Learnable Pins model method) iteratively trains the initial matching model based on the loss value as the optimization goal. In each training process, the weights, biases, and other parameters of the model are dynamically adjusted to gradually reduce the distance loss, while the matching accuracy of the model output is periodically verified. After several iterations, the face feature vector corresponding to any face image in the input face data set is input into the initial matching model, and the speech feature vector corresponding to the face feature vector is output, and the matching degree between the speech feature vector and the true speech feature vector reaches a preset threshold (such as a feature space Euclidean distance ≤ 0.2 or a similarity ≥ 85%). At this time, the training is stopped, and the model can output the speech feature vector with a matching degree that meets the standard based on the face feature vector corresponding to the input face image, and then obtain the tone that matches the input face image.
[0071] In one example, the feature correlation model method based on metric learning also includes a false negative sample exclusion step when training the initial matching model, including: During model training, the similarity between the speech feature vector to be screened and the known paired speech feature vectors is calculated. If the similarity is greater than or equal to the preset similarity threshold, the speech feature vector to be screened is determined as a positive sample and excluded. Only speech feature vectors with similarity lower than the preset similarity threshold are retained as negative samples to participate in model training, so as to optimize the loss calculation accuracy of the metric learning function.
[0072] For example, in the process of training the initial matching model using a feature association model based on metric learning, a false negative sample exclusion operation is performed to address potential false negative samples (i.e., speech feature vectors that should be positive samples but are mistakenly labeled as negative samples) in the training data. In each iteration of training, during the negative sample selection phase, the speech feature vector to be selected for input into the current model is extracted. Simultaneously, known paired speech feature vectors that have already established a pairing relationship with the facial feature vectors used in the current training are retrieved. A preset similarity calculation algorithm (such as cosine similarity algorithm) is used to calculate the similarity between the speech feature vector to be selected and the known paired speech feature vectors. Then, the calculated similarity is compared with a pre-set preset similarity... A similarity threshold (e.g., 0.9) is used for comparison. If the similarity between the voice feature vector to be screened and the known paired voice feature vector is greater than or equal to the preset similarity threshold, it is determined that the voice feature vector to be screened and the current facial feature vector belong to the same identity-associated positive sample category, and it is excluded from the negative sample candidate set to avoid it being used as a negative sample in training and causing bias in loss calculation. Only voice feature vectors to be screened with a similarity lower than the preset similarity threshold are determined as real negative samples and input into the model to participate in the loss calculation and parameter optimization of the subsequent metric learning function. This ensures the authenticity of the negative samples used for training, further improves the accuracy of the loss calculation of the metric learning function, and ensures the training effect of the initial matching model.
[0073] In one exemplary embodiment, the training of a generative model based on an encoder-decoder includes the following steps: Step 1: Construct a training dataset containing face images and their corresponding voice data.
[0074] Step 2: Convert the face images in the training dataset into facial feature vectors, and convert the speech data corresponding to the face images into Mel spectrograms.
[0075] Exemplarily, in the feature preprocessing of the training data set, for each face image in the data set, a preset image feature extraction algorithm (such as a feature extraction framework based on a convolutional neural network) is used to analyze and quantify the key facial regions in the image, and the analyzed facial visual information is mapped into a fixed-dimension vector form, that is, the facial feature vector of the corresponding face image is obtained; at the same time, for the speech data corresponding to the face image, the acoustic features such as frequency and amplitude of the speech are converted through acoustic signal processing technology (such as short-time Fourier transform combined with mel filter bank), and the time-domain speech signal is converted into frequency-domain mel spectrum. The mel spectrum can simulate the perception characteristics of human ears to different frequency sounds, and more accurately represent the timbre characteristics of the speech, thereby providing adaptive input data for subsequent encoding processing of the speech encoder.
[0076] Step three: encoding processing of the mel spectrum by the speech encoder to generate a speech feature vector; encoding processing of the facial feature vector by the face encoder to generate an aligned feature vector consistent with the length and format of the speech feature vector.
[0077] Exemplarily, in the model feature encoding stage, the mel spectrum obtained in step two is input into a pre-constructed speech encoder (which can adopt a time convolutional network or a Transformer architecture, and has the ability to capture the time sequence features of the speech frequency domain), and through multi-layer feature extraction and dimension compression processing of the encoder, the speech acoustic features represented by the mel spectrum are converted into a fixed-dimension vector form, that is, a speech feature vector for subsequent matching is generated; at the same time, the facial feature vector obtained in step two is input into the corresponding face encoder (which can adopt a convolutional neural network or an improved architecture thereof, and can enhance the representation ability of the facial key features), and in the feature conversion process of the face encoder, the dimension parameters of the network output layer are adjusted to make the finally generated facial encoding vector consistent with the speech feature vector in length (dimension number) and data format, forming an aligned feature vector that can be used for cross-modal feature comparison, and laying a foundation for the difference calculation and spatial alignment of the two types of features in the subsequent training process.
[0078] Step four: taking the speech feature vector as the target reference, training the face encoder, calculating the difference between the aligned feature vector and the speech feature vector through a preset loss function, and controlling the training process to make the aligned feature vector consistent with the speech feature vector in the feature space.
[0079] Exemplarily, in the process of training the face encoder, the speech feature vector generated in step three is taken as the target reference for feature matching, and it is included in the training cycle with the alignment feature vector generated at the same time (from the face encoder); the difference value of the two types of vectors in the feature space is calculated through the preset loss function (such as mean square error loss function, cosine similarity loss function, etc.), which directly reflects the matching degree of the alignment feature vector and the speech feature vector; taking “minimizing the difference value” as the training target, the network weight, bias and other parameters of the face encoder are dynamically adjusted through gradient descent and other optimization algorithms, and the optimization direction is fed back according to the new difference value after each training, and the feature distance of the two types of vectors is gradually reduced through continuous iterative adjustment, so that the distribution of the alignment feature vector and the speech feature vector in the feature space tends to be consistent, and the face encoder can learn the correlation mapping rule of face features and speech features.
[0080] Step five: when the loss function is less than the specified threshold, the training is completed, and the target face encoder is obtained. The target face encoder and the preset speech decoder are combined to form a generation model. The preset speech decoder is used to convert the alignment feature vector output by the target face encoder into a corresponding speech signal, and a target timbre matching the input face image is obtained.
[0081] The specified threshold refers to a preset determination standard for determining whether the face encoder training reaches the expected effect, and its value is determined based on the model training accuracy requirement and the adaptability requirement of the actual application scene, and is presented in the form of the output value of the loss function. The core function is to quantify the matching degree of the alignment feature vector and the speech feature vector. When the loss function value is less than the specified threshold, it indicates that the two types of features have been sufficiently converged in the feature space, and the face encoder has learned a stable “face-speech” feature mapping rule, and the training can be stopped. For example, combined with the timbre matching accuracy requirement of the voice-face generation model, the specified threshold is set to 0.05 (corresponding to the mean square error loss function) or 0.1 (corresponding to the cosine similarity loss function), to ensure that the trained model can output a timbre highly adapted to the input face image.
[0082] Exemplarily, during the model training process, the output value of the loss function is continuously monitored and compared with a predetermined specified threshold. When the loss function output value first stabilizes to be less than the specified threshold, it is determined that the face encoder training achieves the desired effect, the training is stopped, and the face encoder at this time is determined as the target face encoder. Then, the target face encoder is integrated with a pre-set voice decoder (the decoder uses a network architecture adapted to the generated voice signal and has the ability to convert high-dimensional feature vectors into time-domain voice signals) to form a complete audio face generation model. In the actual application of the generation model, the input face image is processed by the target face encoder to obtain an aligned feature vector. After the aligned feature vector is input into the pre-set voice decoder, the decoder converts it into an audible voice signal through acoustic feature reconstruction and signal conversion. The timbre of the voice signal is highly adapted to the input face image in terms of age, temperament and other dimensions, that is, the target timbre matched with the input face image is obtained.
[0083] In the above embodiment, by constructing a paired training data set and through feature conversion, cross-modal encoding alignment and targeted training, the generation model can accurately learn the feature mapping rule of face and voice. Without user input of voice files, the model can directly generate an adapted target timbre based on a face image, effectively improving the adaptability of timbre and image, simplifying the training process, ensuring the stability of the model output, and providing efficient and reliable technical support for intelligent matching of virtual human image and timbre.
[0084] In an exemplary embodiment, the image generation model in step 102 is a stable diffusion model, and is configured with a ControlNet auxiliary model to enhance the controllability and accuracy of the image generation process; wherein the ControlNet auxiliary model is modularly coupled with the backbone network of the stable diffusion model, which can expand the input control dimension of the stable diffusion model, and has multiple input control condition modules pre-built, including but not limited to a canny edge detection module, a semantic segmentation graph analysis module, a face / human key point positioning module, a scribble contour recognition module, etc., each of which can solve the generation accuracy problem in different dimensions. For example, the image contour information described in the detailed features is extracted by the canny edge detection module and converted into an edge graph, which can ensure that the generated virtual human image contour is regular and distortion-free; the core facial positions such as eyes and mouth corners are locked by the face key point positioning module, which can accurately restore the expressions and appearances such as "gentle smile" and "light makeup" in the detailed features; the subject and background (such as the image of the receptionist and the hospital environment) are distinguished by the semantic segmentation graph analysis module, which can avoid the confusion between the subject and the background. By using the ControlNet auxiliary model to constrain the generation process of the stable diffusion model in multiple dimensions, the capability boundary of the stable diffusion model is effectively expanded, and compared with the basic stable diffusion model without auxiliary model, the controllability of AI painting and the accuracy of the generated results are greatly improved, ensuring that the generated candidate virtual human image can strictly match the description requirements of age, occupation, dressing, expression, scene style and other multi-dimensional attributes in the detailed features, reducing the deviation between the generated results and the requirements.
[0085] In an embodiment, the detailed features are input into the pre-trained image generation model as generation prompt words, and at least one candidate virtual human image is generated based on the detailed features by the image generation model. The detailed features are input into the stable diffusion model as generation prompt words, and at least one candidate virtual human image is generated based on the detailed features by the stable diffusion model under the control of the ControlNet auxiliary model, and the face key point generation effect of the at least one candidate virtual human image is optimized by the ControlNet auxiliary model.
[0086] For example, the detailed features obtained in step 101 are first input into the pre-trained stable diffusion model as a generated prompt; at the same time, the ControlNet auxiliary model coupled with the stable diffusion model is started, and in the control flow of the ControlNet auxiliary model, the facial key point positioning module is preferentially called, which generates corresponding facial key point constraint conditions (such as eye opening degree, mouth corner lifting angle, and facial contour curve) based on the description of the virtual person's expression (such as “gentle smile” and “calm eye contact”) and facial appearance (such as “no obvious makeup” and “soft facial contour”) in the detailed features; after receiving the generated prompt, the stable diffusion model combines the facial key point constraint conditions output by the ControlNet auxiliary model to generate an image, and the ControlNet auxiliary model continuously optimizes the facial key point generation effect of the virtual person image during the generation process to avoid problems such as facial proportion imbalance, expression deviation from the description (such as a smile turning into a serious expression), and mispositioning of facial features; finally, at least one candidate virtual person image is output by the stable diffusion model, all candidate images meet the core description of the detailed features, and the facial key point generation effect is accurate, which can truly restore the requirements for the virtual person's facial expression and appearance in the detailed features.
[0087] In an exemplary embodiment, the application can also receive auxiliary audio input through a text-to-speech model to adjust the timbre characteristics of the generated script audio; or, through a text-to-speech model, the timbre characteristics of the received auxiliary audio input are generated.
[0088] For example, in the application process of the TTS model, the user can input auxiliary audio (which contains the user's desired timbre style features, such as tone softness and timbre brightness) to the TTS model according to actual timbre optimization needs; after receiving the auxiliary audio, the TTS model can integrate the target timbre features in the auxiliary audio into the generation process of the script audio through a timbre feature extraction and adaptation adjustment mechanism, which can not only adjust the timbre of the originally generated script audio based on the timbre features of the auxiliary audio (such as enhancing the softness of the audio and adjusting the rhythm), but also directly generate script audio with the same timbre features as the auxiliary audio, making the final output script audio timbre more in line with the user's individual needs while maintaining the adaptability to the target virtual person image.
[0089] In the above embodiments, the user can flexibly adjust or define the timbre characteristics of the script audio by inputting auxiliary audio, which not only meets the user's individual needs for timbre, but also further optimizes the adaptability of the target timbre to the target virtual person image, improves the naturalness and practicality of the script audio, and enhances the user experience of the overall virtual person animation.
[0090] In one exemplary embodiment, such as Figure 2 The diagram shown is a flowchart of another virtual human animation generation method provided in this application embodiment. When the electronic device executes the virtual human speaking audio generation process, the user first inputs a simple requirement about the virtual human, such as "generate a 25-year-old female hospital guide image to explain the registration process." After this simple requirement is input into the large language model, the large language model analyzes the requirement to generate detailed features including dimensions such as the virtual human's age, occupation, clothing, expression, scene, and writing style. These detailed features are then input into an image generation model, such as a stable model configured with a ControlNet auxiliary model. The diffusion model generates at least one candidate virtual human image that matches detailed features from the image generation model. After the user selects the target virtual human image from the candidate images, the target image is input into the voice-face generation model to obtain the corresponding target voice tone. At the same time, the user inputs text that matches the virtual human image, such as "Hello, please go to the service desk on the first floor to register." Then, the matched target voice tone and the input text are input into the text-to-speech model to generate the corresponding text audio. Finally, the text audio and the target virtual human image are input into the audio-driven image model, which completes the animation synthesis process of the virtual human image and the text audio, and finally outputs the virtual human speaking audio. In the virtual human speaking audio, the virtual human image's lip movements and expressions are completely synchronized with the text audio.
[0091] It is understood that, in the specific embodiments of this application, voice data, facial data and other related data are involved. When the following embodiments of this application are applied to specific products or technologies, permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0092] To better implement the above methods, this application also provides a virtual human animation generation device, which can be integrated into an electronic device, such as a terminal or server. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, or personal computer; the server can be a single server or a server cluster composed of multiple servers.
[0093] For example, in this embodiment, the method of this application embodiment will be described in detail by taking the virtual human animation generation device specifically integrated into the server as an example.
[0094] For example, such as Figure 3As shown, the virtual human animation generation apparatus can include a feature generation unit 310, an image generation unit 320, a timbre matching unit 330, a speech synthesis unit 340, and an animation synthesis unit 350, as follows: (I) The feature generation unit 310; The feature generation unit 310 is configured to obtain an input requirement for generating a virtual human image, analyze and process the requirement through a pre-set large language model, and generate detailed features corresponding to the virtual human image.
[0095] (II) The image generation unit 320; The image generation unit 320 is configured to input the detailed features as a generation prompt into a pre-trained image generation model, and generate at least one candidate virtual human image based on the detailed features through the image generation model.
[0096] (III) The timbre matching unit 330; The timbre matching unit 330 is configured to perform feature analysis and timbre matching analysis on a target virtual human image selected from the at least one candidate virtual human image through a pre-trained voice-face generation model, to obtain a target timbre matched with the target virtual human image; the voice-face generation model is trained based on a pre-constructed speech data set and a paired face data set.
[0097] (IV) The speech synthesis unit 340; The speech synthesis unit 340 is configured to perform speech synthesis processing on the target timbre and input scripts associated with the target virtual human image through a text-to-speech model, to obtain script audio of the target timbre.
[0098] (V) The animation synthesis unit 350; The animation synthesis unit 350 is configured to perform animation synthesis processing on the target virtual human image and the script audio through an audio-driven image model, to obtain a corresponding virtual human animation.
[0099] In some embodiments, the virtual human animation generation apparatus further includes a matching model training unit configured to construct the speech data set and the paired face data set; the speech data set is obtained by intercepting speech in a target video and cutting it into audio files of a pre-set length, and the face data set is obtained by intercepting corresponding face images from the video where the speech comes from; convert the face images in the face data set into face feature vectors, and convert the audio files in the speech data set into speech feature vectors; The voice feature vector is taken as a label, and the face feature vector is taken as an input. A distance loss of the voice feature vector and the face feature vector in a feature space is calculated through a metric learning function. An initial matching model is trained based on a feature correlation model method of metric learning. The initial matching model parameters are optimized through multiple rounds of iterative training until the initial matching model can output a voice feature vector with a matching degree reaching a preset threshold based on a face feature vector corresponding to an input face image.
[0100] The matching model training unit is further configured to calculate a similarity between the to-be-screened voice feature vector and the known paired voice feature vector during the model training process. If the similarity is greater than or equal to a preset similarity threshold, the to-be-screened voice feature vector is determined as a positive sample and excluded. Only voice feature vectors with a similarity lower than the preset similarity threshold are reserved as negative samples to participate in the model training, so as to optimize the loss calculation accuracy of the metric learning function.
[0101] In some embodiments, the virtual human animation generation apparatus further includes a generation model training unit configured to construct a training data set containing face images and corresponding voice data; The face images in the training data set are converted into face feature vectors, and the voice data corresponding to the face images are converted into mel spectra. The mel spectra are processed through a voice encoder to generate voice feature vectors. The face feature vectors are processed through a face encoder to generate aligned feature vectors with the same length and format as the voice feature vectors. The face encoder is trained with the voice feature vectors as a target reference. A preset loss function is used to calculate the difference between the aligned feature vectors and the voice feature vectors, and the training process is controlled to make the aligned feature vectors consistent with the voice feature vectors in the feature space. The training is completed when the loss function is less than a specified threshold, and a target face encoder is obtained. The target face encoder and a preset voice decoder are combined to form the generation model. The preset voice decoder is used to convert the aligned feature vectors output by the target face encoder into corresponding voice signals, so as to obtain a target voice color matching the input face image.
[0102] In some embodiments, the image generation model is a stable diffusion model, and is configured with a control net auxiliary model; the control net auxiliary model is used to expand the input control dimension of the stable diffusion model; the image generation unit 320 is further configured to input the detailed features as a generation prompt word into the stable diffusion model, generate at least one candidate virtual human image based on the detailed features through the stable diffusion model under the control of the ControlNet auxiliary model, and optimize the face key point generation effect of the at least one candidate virtual human image through the ControlNet auxiliary model.
[0103] In some embodiments, the speech synthesis unit 340 is further configured to receive auxiliary audio input through the text-to-speech model to adjust the timbre features of the generated script audio; Or, The timbre features of the script audio generated by the timbre features of the received auxiliary audio input through the text-to-speech model.
[0104] In implementation, each of the above units can be implemented as an independent entity, or can be combined as the same or several entities, and the specific implementation of each of the above units can refer to the method embodiments in the foregoing, which will not be repeated here.
[0105] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory) or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an overall module or unit that includes the functions of the module or unit.
[0106] As can be seen from the above, the feature generation unit 310 in this embodiment analyzes and processes the user's simple image requirements through a preset large language model, and can automatically generate detailed features corresponding to the virtual human image. Without requiring the user to provide complex descriptions or upload images with privacy and copyright risks, at least one candidate virtual human image can be generated by the image generation model of the image generation unit 320. This achieves rapid generation of personalized virtual human images based on simple requirements, significantly reducing the complexity and cost of image customization. Furthermore, through a voice-face generation model trained on a pre-constructed voice dataset and its paired face dataset, the voice matching unit 330 can match the user's selected target virtual human. The system automatically performs feature analysis and timbre matching to accurately obtain the target timbre that matches the target image. This effectively solves the problem of timbre mismatch between virtual human images and their age, occupation, temperament, etc., in related technologies, ensuring the coordination between the subsequently generated text audio and the target virtual human image. Finally, the text audio with the corresponding timbre is generated by the text-to-speech model of the speech synthesis unit 340, and the virtual human animation is synthesized by the audio-driven image model of the animation synthesis unit 350. This realizes automated and intelligent processing from virtual human image requirements to animation output, significantly improving the generation efficiency and overall immersion of virtual human animation, and meeting users' needs for personalized and highly adaptable virtual human animation.
[0107] This application also provides an electronic device, which can be a terminal, a server, or other similar device. The terminal can be a mobile phone, tablet computer, smart Bluetooth device, laptop computer, personal computer, etc.; the server can be a single server or a server cluster composed of multiple servers, etc.
[0108] In some embodiments, the virtual human animation generation device can also be integrated into multiple electronic devices. For example, the virtual human animation generation device can be integrated into multiple servers, and the virtual human animation generation method of this application can be implemented by multiple servers.
[0109] In this embodiment, a server will be used as an example for detailed description. For example, ... Figure 4 As shown, it illustrates a structural schematic diagram of the electronic device involved in the embodiments of this application, specifically: The electronic device may include components such as a processor 410 with one or more processing cores, a memory 420 with one or more computer-readable storage media, a power supply 430, an input module 440, and a communication module 450. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 410 is the control center of the electronic device, connects various parts of the entire electronic device through various interfaces and lines, and performs various functions of the electronic device and processes data by running or executing software programs and / or modules stored in the memory 420 and calling data stored in the memory 420, thereby overall detecting the electronic device. In some embodiments, the processor 410 can include one or more processing cores; in some embodiments, the processor 410 can integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application program, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 410.
[0110] The memory 420 can be used to store software programs and modules, and the processor 410 executes various function applications and data processing by running the software programs and modules stored in the memory 420. The memory 420 can mainly include a storage program area and a storage data area, wherein the storage program area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.) and the like; the storage data area can store data created according to the use of the electronic device and the like. In addition, the memory 420 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device. Accordingly, the memory 420 can also include a memory controller to provide the processor 410 with access to the memory 420.
[0111] The electronic device also includes a power supply 430 for powering various components, and in some embodiments, the power supply 430 can be logically connected to the processor 410 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 430 can also include one or more direct current or alternating current power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power state indicator, and any other components.
[0112] The electronic device can also include an input module 440, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0113] The electronic device can also include a communication module 450, which in some embodiments can include a wireless module, and the electronic device can perform short-range wireless transmission through the wireless module of the communication module 450, thereby providing the user with wireless broadband Internet access. For example, the communication module 450 can be used to help the user send and receive emails, browse web pages, and access streaming media, etc.
[0114] Although not shown, the electronic device can further include a display unit, etc., which will not be described herein. Specifically, in the present embodiment, the processor 410 in the electronic device will load the executable file corresponding to the process of one or more application programs into the memory 420 according to the following instructions, and run the application program stored in the memory 420 by the processor 410, thereby realizing various functions, as follows: obtain an input requirement for generating a virtual human image, analyze and process the requirement through a preset large language model, and generate detailed features corresponding to the virtual human image; input the detailed features as a generation prompt into a pre-trained image generation model, and generate at least one candidate virtual human image based on the detailed features through the image generation model; perform feature analysis and timbre matching analysis processing on a target virtual human image selected from the at least one candidate virtual human image through a pre-trained audio-face generation model, to obtain a target timbre matched with the target virtual human image; the audio-face generation model is trained based on a pre-constructed voice data set and a paired face data set; perform voice synthesis processing on the target timbre and input scripts associated with the target virtual human image through a text-to-speech model, to obtain script audio of the target timbre; perform animation synthesis processing on the target virtual human image and the script audio through an audio-driven image model, to obtain a corresponding virtual human animation.
[0115] The specific implementation of each operation can be referred to the foregoing embodiments, which will not be described herein.
[0116] As can be seen from the above, the embodiments of the present application can improve the adaptability of virtual human timbre and image.
[0117] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by instructions controlling related hardware, which can be stored in a computer readable storage medium and loaded and executed by a processor.
[0118] To this end, the present application provides a computer readable storage medium, which stores a plurality of instructions capable of being loaded by a processor to execute the steps in any virtual human animation generation method provided by the embodiments of the present application. For example, the instructions can perform the following steps: obtain an input requirement for generating a virtual human image, analyze and process the requirement through a preset large language model, and generate detailed features corresponding to the virtual human image; input the detailed features into a pre-trained image generation model as a generation prompt word, and generate at least one candidate virtual human image based on the detailed features through the image generation model; perform feature analysis and timbre matching analysis processing on a target virtual human image selected from the at least one candidate virtual human image through a pre-trained voice-face generation model to obtain a target timbre matched with the target virtual human image; the voice-face generation model is trained based on a pre-constructed voice data set and a paired face data set; perform voice synthesis processing on the target timbre and input scripts associated with the target virtual human image through a text-to-speech model to obtain script audio of the target timbre; perform animation synthesis processing on the target virtual human image and the script audio through an audio-driven image model to obtain a corresponding virtual human animation.
[0119] The storage medium can include a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, or the like.
[0120] According to an aspect of the present application, a computer program product or computer program is provided, which includes computer instructions stored in a computer readable storage medium. The processor of the electronic device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the electronic device to perform the virtual human animation generation method provided in the above embodiments.
[0121] Due to the instructions stored in the storage medium, the steps in any of the virtual human animation generation methods provided in the embodiments of the present application can be performed, and thus the beneficial effects of any of the virtual human animation generation methods provided in the embodiments of the present application can be achieved. Details are described in the foregoing embodiments, which will not be repeated here.
[0122] The above describes in detail a virtual human animation generation method, device, electronic device, and computer readable storage medium provided in the embodiments of the present application. The principles and implementation manners of the present application are described by applying specific examples in this paper. The above embodiment descriptions are only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, the specific implementation manners and application ranges will be changed according to the idea of the present application. In summary, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A method for generating virtual human animation, characterized in that, include: The system obtains the input requirements for generating a virtual human character, analyzes and processes the requirements using a pre-defined large language model, and generates detailed features corresponding to the virtual human character. The detailed features are used as input to a pre-trained image generation model to generate prompt words, and the image generation model generates at least one candidate virtual human image based on the detailed features. The pre-trained voice-face generation model performs feature analysis and timbre matching analysis on the target virtual human image selected from the at least one candidate virtual human image to obtain the target timbre that matches the target virtual human image; the voice-face generation model is trained based on a pre-constructed speech dataset and its paired face dataset. Using a text-to-speech model, speech synthesis processing is performed on the input text associated with the target timbre and the target virtual human image to obtain the audio text of the target timbre. By using an audio-driven image model, animation synthesis is performed on the target virtual human image and the audio of the text to obtain the corresponding virtual human animation.
2. The method as described in claim 1, characterized in that, The voice-face generation model is a matching model based on voice-face relationships; the training methods for the matching model include: Construct the aforementioned voice dataset and the paired face dataset; the voice dataset is obtained by extracting voice from the target video and cropping it into an audio file of a preset length, and the face dataset is obtained by extracting the corresponding face image from the video from which the voice originates; The face images in the face dataset are converted into facial feature vectors, and the audio files in the speech dataset are converted into speech feature vectors. Using the speech feature vector as a label and the facial feature vector as input, the distance loss between the speech feature vector and the facial feature vector in the feature space is calculated through a metric learning function. An initial matching model is trained based on the feature association model method of metric learning. The parameters of the initial matching model are optimized through multiple rounds of iterative training until the initial matching model can output the timbre corresponding to the speech feature vector whose matching degree reaches a preset threshold based on the facial feature vector corresponding to the input face image.
3. The method as described in claim 2, characterized in that, The feature association model method based on metric learning, when training the initial matching model, also includes a false negative sample exclusion step, including: During model training, the similarity between the speech feature vector to be screened and the known paired speech feature vectors is calculated. If the similarity is greater than or equal to a preset similarity threshold, the speech feature vector to be screened is determined as a positive sample and excluded. Only speech feature vectors with a similarity lower than the preset similarity threshold are retained as negative samples to participate in model training, so as to optimize the loss calculation accuracy of the metric learning function.
4. The method as described in claim 1, characterized in that, The voice-face generation model is a generator model based on an encoder-decoder; the training methods for the generator model include: Construct a training dataset containing face images and their corresponding speech data; The face images in the training dataset are converted into facial feature vectors, and the speech data corresponding to the face images are converted into Mel spectrograms. The Mel spectrum is encoded by a speech encoder to generate a speech feature vector; the facial feature vector is encoded by a face encoder to generate an aligned feature vector with the same length format as the speech feature vector. Using the speech feature vector as the target reference, the face encoder is trained. The difference between the alignment feature vector and the speech feature vector is calculated through a preset loss function. The training process is controlled to make the alignment feature vector and the speech feature vector tend to be consistent in the feature space. Training is completed when the loss function is less than a specified threshold, resulting in a target face encoder. The target face encoder is then combined with a preset speech decoder to form the generative model. The preset speech decoder is used to convert the alignment feature vector output by the target face encoder into a corresponding speech signal, thereby obtaining a target timbre that matches the input face image.
5. The method as described in claim 1, characterized in that, The image generation model is a stable diffusion model and is configured with a control network auxiliary model; the control network auxiliary model is used to expand the input control dimension of the stable diffusion model; the image generation model, which uses the detailed features as input to generate prompt words, generates at least one candidate virtual human image based on the detailed features, including: The detailed features are input as generation prompts into the stable diffusion model. Under the control of the control network-assisted model, at least one candidate virtual human image is generated based on the detailed features by the stable diffusion model, and the facial key point generation effect of the at least one candidate virtual human image is optimized by the control network-assisted model.
6. The method as described in claim 1, characterized in that, Also includes: The text-to-speech model receives auxiliary audio input to adjust the timbre characteristics of the generated text audio. or, The text-to-speech model generates the timbre features of the text audio from the timbre features of the received auxiliary audio input.
7. A virtual human animation generation device, characterized in that, include: The feature generation unit is used to obtain the input requirements for generating a virtual human image, analyze and process the requirements through a preset large language model, and generate detailed features corresponding to the virtual human image. An image generation unit is used to input the detailed features as generation prompts into a pre-trained image generation model, and generate at least one candidate virtual human image based on the detailed features through the image generation model. The timbre matching unit is used to perform feature analysis and timbre matching analysis on a target virtual human image selected from at least one candidate virtual human image using a pre-trained voice-face generation model to obtain a target timbre that matches the target virtual human image; the voice-face generation model is trained based on a pre-constructed speech dataset and its paired face dataset; The speech synthesis unit is used to perform speech synthesis processing on the target timbre and the input text associated with the target virtual human image through a text-to-speech model to obtain the text audio of the target timbre. An animation synthesis unit is used to perform animation synthesis processing on the target virtual human image and the text audio through an audio-driven image model to obtain the corresponding virtual human animation.
8. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing multiple instructions; the processor loads instructions from the memory to perform the steps in the virtual human animation generation method as described in any one of claims 1 to 6.
9. A computer program product, characterized in that, It includes a computer program / instruction that, when executed by a processor, implements the steps in the virtual human animation generation method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the virtual human animation generation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method for generating virtual character and interacting with user
CN120495484A
Cited By
Speech synthesis data acquisition method and device, electronic equipment and storage medium
CN117577091A