Method and device for automatically producing whole-body digital human data and recording medium

Through the dual-stage training strategy, including data augmentation and model fine-tuning, the time-consuming and labor-intensive problem of digital people's speech data acquisition is solved, and efficient and automated data production and model accuracy are achieved.

CN120472260APending Publication Date: 2025-08-12BEIJING YUNSIZHIXUE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510543906.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art costs huge and time-consuming in the process of collecting digital human speech data, and the existing data cleaning models are insufficiently accurate, resulting in low lip accuracy and requires a lot of manual screening.

Method used

A two-stage training strategy is adopted, firstly, the lip-aligned model and digital human-generated model are trained through the initial training data set, and data enhancement and expansion are carried out until the preset accuracy is reached; then when the data volume reaches the threshold, the description text data is generated for fine-tuning of the model, replacing the old model, and iterating cycle until the training goal is met.

Benefits of technology

It significantly reduces the cost of digital human data acquisition, improves model accuracy and robustness, reduces dependence on manual annotation, enhances the model's adaptability to new scenarios and new speakers, and realizes a highly automated data production process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472260A_ABST
    Figure CN120472260A_ABST
Patent Text Reader

Abstract

The invention discloses a method for automatically producing whole-body digital human data, which comprises a first model training stage: collecting a real human speaking data set and acquiring an open source data set as an initial training data set, training a lip shape alignment model and a digital human generation model, and performing data enhancement and data expansion processing based on the initial training data set to obtain a second model training stage; continuing to train the lip shape alignment model and the digital human generation model until the precision of the lip shape alignment model and the digital human generation model meets a preset model parameter threshold value; and a second model training stage: when the data volume of the training data set in the first model training stage reaches a preset data volume threshold value, generating new description text data based on the video data in the training data set, and adding the description text data into the video data set, performing fine tuning on the digital human generation model and replacing the corresponding model in the first model training stage; and executing loop iteration of the first model training stage and the second model training stage at the same time until a preset training target is met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital human technology, and specifically provides a method, device and recording medium for automatically generating full-body digital human data. Background Art

[0002] A digital human (or metahuman) is a digitized humanoid created through digital technology that closely resembles a human form. It is the product of the fusion of information science and life science, utilizing information science methods to simulate the human body at various levels of form and function. The development of digital human speech has been a key area of application in recent years. Training a digital human speech model presupposes a wealth of high-precision speech data. However, collecting this data has always been a tedious, time-consuming, labor-intensive, and costly undertaking. To date, the typical approach to collecting digital human speech data in academia and industry has been:

[0003] 1. Speech data is collected manually and strictly in accordance with high standards.

[0004] 2. Collect a bunch of speech data, and then filter out high-precision speech data according to a series of data cleaning logic.

[0005] The first option was extremely expensive, time-consuming, and labor-intensive. The second, automated approach was what we were considering, but we found that the cleaning logic and models used in these steps were not very accurate. Ultimately, the cleaned data still had issues due to low lip shape accuracy, requiring extensive manual screening.

[0006] In view of this, the present invention patent is proposed. Summary of the Invention

[0007] To address the above technical issues, the present invention proposes a method, device, and recording medium for automatically generating full-body digital human data. Unlike simply removing unqualified data from a large amount of dirty data, the present invention can generate more speech data than the data to be cleaned. Specifically, the following technical solutions are adopted:

[0008] In a first aspect, the present invention provides a method for automatically generating full-body digital human data, comprising a first model training phase and a second model training phase:

[0009] The first model training stage: the collected real-person speech dataset and the obtained open-source dataset are used as the initial training dataset, the lip alignment model and the digital human generation model are trained using the initial training dataset, data enhancement and data expansion processing are performed based on the initial training dataset, and the lip alignment model and the digital human generation model are continuously trained until the accuracy of the lip alignment model and the digital human generation model meets the preset model parameter threshold;

[0010] The second model training stage: when the data volume of the training data set in the first model training stage reaches a preset data volume threshold, new descriptive text data is generated based on the video data in the training data set, the descriptive text data is added to the video data set, the digital human generation model is fine-tuned, and the corresponding model in the first model training stage is replaced;

[0011] The first model training phase and the second model training phase are executed in a loop iteration at the same time until the preset training target is met.

[0012] As an optional embodiment of the present invention, in a method for automatically generating full-body digital human data of the present invention, the first model training stage includes:

[0013] Step S101: The collected real-person speech dataset and the obtained open-source dataset are used as the initial training dataset, and the lip alignment model and the digital human generation model are trained using the initial training dataset;

[0014] Step S102, performing data augmentation processing on the data in the initial training dataset to obtain an enhanced training dataset with more data, and using the enhanced training dataset to continue training the lip alignment model and the digital human generation model;

[0015] Step S103: Add a new speech document and generate voice data corresponding to the speech document. Based on the enhanced training data set in step S102, generate new digital human speech data. Filter the new digital human speech data to obtain an extended training data set. Use the extended training data set to continue training the lip alignment model and the digital human generation model.

[0016] Step S104, iterating steps S102 to S103 in a loop until the accuracy of the lip alignment model and the digital human generated model meets a preset threshold.

[0017] As an optional embodiment of the present invention, in a method for automatically generating full-body digital human data of the present invention, the step S102 performs data enhancement processing on the data in the initial training dataset to obtain an enhanced training dataset with more data, including:

[0018] Data enhancement processing is performed on the video data and image data in the initial training data set by changing faces and backgrounds to obtain a first enhanced training data set.

[0019] As an optional embodiment of the present invention, in a method for automatically generating full-body digital human data of the present invention, the step S102 performs data enhancement processing on the data in the initial training dataset to obtain an enhanced training dataset with more data, including:

[0020] Performing data enhancement processing on the video data and the speech data in the first enhanced training data set by timbre conversion to obtain a second enhanced training data set;

[0021] The data in the initial training data set, the first enhanced training data set and the second enhanced training data set together constitute an enhanced training data set.

[0022] As an optional embodiment of the present invention, in a method for automatically generating full-body digital human data of the present invention, the step S103 of filtering the new digital human speech data to obtain an extended training data set includes:

[0023] The new digital human speech data is input into the lip alignment model trained and generated in step S102 for screening and filtering, noise data in the new digital human speech data is removed, and the data is added to the enhanced training data set to form an extended training data set.

[0024] As an optional embodiment of the present invention, in a method for automatically generating full-body digital human data of the present invention, the second model training stage includes:

[0025] Step S201, when the data volume of the training data set in the first model training phase reaches a preset data volume threshold, the second model training phase begins;

[0026] Step S202: Generate new description text data based on the video data in the training data set, add the description text data to the video data set, and perform fine-tuning training on the parameters of the text2video model in the digital human generation model;

[0027] Step S203: Replace the corresponding models in the first model training stage with the lip synchronization model and texture generation model in the text2video model obtained through fine-tuning training.

[0028] As an optional embodiment of the present invention, in a method for automatically generating full-body digital human data of the present invention, generating new descriptive text data based on the video data in the training data set in step S202 includes:

[0029] Extract speech from video data in the training dataset and convert it into text data text1, and extract frames from video data in the training dataset and convert them into text data text2;

[0030] Assemble text data text1 and text data text2 into new description text data.

[0031] In a second aspect, the present invention provides an apparatus for automatically generating full-body digital human data, comprising a first model training phase module and a first model phase process module:

[0032] The first model training phase module: uses the collected real-person speech dataset and the obtained open-source dataset as the initial training dataset, uses the initial training dataset to train the lip alignment model and the digital human generation model, performs data enhancement and data expansion processing based on the initial training dataset, and continues to train the lip alignment model and the digital human generation model until the accuracy of the lip alignment model and the digital human generation model meets the preset model parameter threshold;

[0033] The second model training phase module: when the data volume of the training data set in the first model training phase reaches a preset data volume threshold, generates new descriptive text data based on the video data in the training data set, adds the descriptive text data to the video data set, fine-tunes the digital human generation model, and replaces the corresponding model in the first model training phase;

[0034] The first model training phase module and the second model training phase module are executed in a loop iteration at the same time until the preset training target is met.

[0035] In a third aspect, the present invention provides an electronic device comprising a processor and a memory, wherein the memory is used to store a computer executable program. When the computer program is executed by the processor, the processor executes the method for automatically generating full-body digital human data.

[0036] In a fourth aspect, the present invention provides a computer-readable recording medium storing a computer-executable program, wherein when the computer-executable program is executed, the method for automatically generating full-body digital human data is implemented.

[0037] Compared with the prior art, the present invention has the following beneficial effects:

[0038] The method of automatically generating full-body digital human data in the present invention mainly includes two core stages: the first model training stage and the second model training stage.

[0039] The first model training phase: The lip alignment model and the digital human generation model are trained based on the initial dataset. The initial dataset includes a dataset of real-life speech collected and an open-source dataset. This avoids the costly, time-consuming, and labor-intensive problem of manually collecting speech data strictly according to high standards. In addition, more speech data is generated through data augmentation and expansion, and the lip alignment model and the digital human generation model are continuously trained to gradually improve the model accuracy.

[0040] Second model training phase: When the data volume reaches a threshold, description text is generated based on the video data, the digital human generation model is fine-tuned, and the old model components are replaced.

[0041] The first model training phase and the second model training phase are iterated repeatedly until the training objectives are met.

[0042] The training data in the first model training phase mainly targets close-up video data. The scope of training data is expanded in the first model training phase to include half-body and full-body videos and their text descriptions.

[0043] Therefore, the present invention's method for automatically generating full-body digital human data uses a portion of a real-person speech dataset combined with an open-source dataset as the initial dataset to train the lip alignment model and digital human generation model. This solves the enormous, time-consuming, and labor-intensive problem of relying solely on manual speech data collection to strict, high standards. By enhancing and expanding data to generate more speech data, compared to simply eliminating unqualified data from a large amount of dirty data, more speech data can be generated, gradually improving model accuracy. The present invention's method for automatically generating full-body digital human data significantly reduces the cost of digital human data collection and can produce more data at a very low cost.

[0044] In addition, influenced by some solutions in the multimodal field, such as the BLIP series, the present invention provides a method for automatically generating full-body digital human data, which uses a method of cleaning, training, and generating at the same time to generate digital human speech data.

[0045] The method of automatically generating full-body digital human data of the present invention has the following beneficial effects:

[0046] Reduce manual annotation costs: Through automated data augmentation, data expansion, and descriptive text generation technologies, the reliance on manually labeled data is significantly reduced.

[0047] Improve generation quality: The two-stage loop training strategy effectively improves the accuracy and robustness of the model.

[0048] Enhanced generalization capability: Through continuous data expansion and model fine-tuning, the model's adaptability to new scenarios and new speakers is improved.

[0049] High degree of automation: The entire process from data preparation to model training is highly automated.

[0050] Flexible and scalable: Supports continuous data expansion and model updates to adapt to changing application needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 A flow chart of a method for automatically generating full-body digital human data according to an embodiment of the present invention;

[0052] Figure 2 A schematic structural diagram of an electronic device according to an embodiment of the present invention;

[0053] Figure 3 Schematic diagram of a computer-readable recording medium according to an embodiment of the present invention. DETAILED DESCRIPTION

[0054] To make the purpose, technical solutions and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them.

[0055] Therefore, the following detailed description of the embodiments of the present invention is not intended to limit the scope of the claimed invention, but merely represents some embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work are within the scope of protection of the present invention.

[0056] It should be noted that, in the absence of conflict, the embodiments of the present invention and the features and technical solutions therein may be combined with each other.

[0057] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0058] In the description of the present invention, it should be noted that the terms "upper" and "lower" and the like indicate orientations or positional relationships based on the orientations or positional relationships shown in the accompanying drawings, or the orientations or positional relationships in which the inventive product is typically placed when in use, or the orientations or positional relationships commonly understood by those skilled in the art. Such terms are intended solely to facilitate the description of the present invention and simplify the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, be constructed, or operate in a specific orientation. Therefore, they should not be construed as limitations on the present invention. Furthermore, the terms "first" and "second" and the like are used solely for distinction and should not be construed as indicating or implying relative importance.

[0059] See also Figure 1 As shown, this embodiment provides a method for automatically generating full-body digital human data, including a first model training phase and a second model training phase:

[0060] The first model training stage: the collected real-person speech dataset and the obtained open-source dataset are used as the initial training dataset, the lip alignment model and the digital human generation model are trained using the initial training dataset, data enhancement and data expansion processing are performed based on the initial training dataset, and the lip alignment model and the digital human generation model are continuously trained until the accuracy of the lip alignment model and the digital human generation model meets the preset model parameter threshold;

[0061] The second model training stage: when the data volume of the training data set in the first model training stage reaches a preset data volume threshold, new descriptive text data is generated based on the video data in the training data set, the descriptive text data is added to the video data set, the digital human generation model is fine-tuned, and the corresponding model in the first model training stage is replaced;

[0062] The first model training phase and the second model training phase are executed in a loop iteration at the same time until the preset training target is met.

[0063] The method of automatically generating full-body digital human data in this embodiment mainly includes two core stages: the first model training stage and the second model training stage.

[0064] The first model training phase: The lip alignment model and the digital human generation model are trained based on the initial dataset. The initial dataset includes a dataset of real-life speech collected and an open-source dataset. This avoids the costly, time-consuming, and labor-intensive problem of manually collecting speech data strictly according to high standards. In addition, more speech data is generated through data augmentation and expansion, and the lip alignment model and the digital human generation model are continuously trained to gradually improve the model accuracy.

[0065] Second model training phase: When the data volume reaches a threshold, description text is generated based on the video data, the digital human generation model is fine-tuned, and the old model components are replaced.

[0066] The first model training phase and the second model training phase are iterated repeatedly until the training objectives are met.

[0067] The training data in the first model training phase mainly targets close-up video data. The scope of training data is expanded in the first model training phase to include half-body and full-body videos and their text descriptions.

[0068] Therefore, the method for automatically generating full-body digital human data in this embodiment uses a portion of a real-person speech dataset combined with an open-source dataset as the initial dataset to train the lip alignment model and the digital human generation model. This solves the problem of being extremely expensive, time-consuming, and labor-intensive compared to relying solely on manual speech data collection to strict high standards. By generating more speech data through data augmentation and expansion, more speech data can be generated, gradually improving model accuracy, compared to simply eliminating unqualified data from a large amount of dirty data. This method for automatically generating full-body digital human data in this embodiment significantly reduces the cost of digital human data collection and can produce more data at a very low cost.

[0069] In addition, influenced by some solutions in the multimodal field, such as the BLIP series, a method for automatically generating full-body digital human data in this embodiment uses a method of cleaning, training, and generating at the same time to generate digital human speech data.

[0070] The method of automatically generating full-body digital human data in this embodiment has the following beneficial effects:

[0071] Reduce manual annotation costs: Through automated data augmentation, data expansion, and descriptive text generation technologies, the reliance on manually labeled data is significantly reduced.

[0072] Improve generation quality: The two-stage loop training strategy effectively improves the accuracy and robustness of the model.

[0073] Enhanced generalization capability: Through continuous data expansion and model fine-tuning, the model's adaptability to new scenarios and new speakers is improved.

[0074] High degree of automation: The entire process from data preparation to model training is highly automated.

[0075] Flexible and scalable: Supports continuous data expansion and model updates to adapt to changing application needs.

[0076] As an optional implementation of this embodiment, in a method for automatically generating full-body digital human data of this embodiment, the first model training stage includes:

[0077] Step S101: The collected real-person speech dataset and the obtained open-source dataset are used as the initial training dataset, and the lip alignment model and the digital human generation model are trained using the initial training dataset;

[0078] Step S102, performing data augmentation processing on the data in the initial training dataset to obtain an enhanced training dataset with more data, and using the enhanced training dataset to continue training the lip alignment model and the digital human generation model;

[0079] Step S103: Add a new speech document and generate voice data corresponding to the speech document. Based on the enhanced training data set in step S102, generate new digital human speech data. Filter the new digital human speech data to obtain an extended training data set. Use the extended training data set to continue training the lip alignment model and the digital human generation model.

[0080] Step S104, iterating steps S102 to S103 in a loop until the accuracy of the lip alignment model and the digital human generated model meets a preset threshold.

[0081] In step S101 of this embodiment, a small amount of real-person speech data is collected as the beginning of a cold start. The lip alignment model and the digital human generation model are trained using the real-person speech data and the open source dataset. The current model is greatly influenced by the open source dataset.

[0082] As a specific embodiment of the initial training data set, the initial training data set is composed as follows:

[0083] Self-collected data: 50 people × 200 videos (1080p / 60fps);

[0084] Open source datasets: LRS3, VoxCeleb2.

[0085] The speech data used for model training in this embodiment mainly include the following types:

[0086] Video Data: Training a digital human model requires a large amount of video data, typically with a frame rate of 25fps, a resolution of 512x512, and a recommended duration of 3-5 minutes. These videos must contain high-quality facial images so that the model can accurately recognize and track faces.

[0087] Image data: Image data includes various expressions, gestures, etc. This data helps the model understand the facial expressions and body language of the characters, so that it can interact with users more naturally.

[0088] Speech data: Speech data contains information such as different intonations, speaking speeds, and accents. This data helps the model understand speech input and generate natural speech output.

[0089] Text data: Text data such as news, novels, and encyclopedias is used to train models for natural language processing tasks, such as understanding user questions and providing accurate answers.

[0090] The data preprocessing steps include:

[0091] Video frame extraction: Split the video into frames and detect and separate the face area.

[0092] Facial feature extraction: Use image processing tools and libraries (such as OpenCV, dlib) to extract facial features.

[0093] Text cleaning: remove noise and erroneous characters from text data and perform normalization.

[0094] Speech feature extraction: Extract the spectral features of speech to facilitate model calculation and reasoning.

[0095] As an optional implementation of this embodiment, the lip alignment model architecture of this embodiment adopts an improved SyncNet architecture: video branch: 3D ResNet-18, audio branch: 1DCNN+BiLSTM, loss function: Among them, δ=0.2 is the boundary threshold.

[0096] The digital human generation model of this embodiment includes: Body movement generator: Transformer-based diffusion model Texture generator: StyleGAN2-ADA architecture.

[0097] Since the initial training dataset is used to train the lip alignment model and the digital human generation model, it is greatly affected by the open source dataset. Therefore, this embodiment requires a more targeted expansion of the initial training dataset to improve the training accuracy requirements of the model. Therefore, in step S102 of this embodiment, data enhancement processing is performed on the data in the initial training dataset to obtain an enhanced training dataset with more data, including:

[0098] Data enhancement processing is performed on the video data and image data in the initial training data set by changing faces and backgrounds to obtain a first enhanced training data set.

[0099] Among them, the generative adversarial network (GAN) is used to replace the background and face of the video frame. The formula is: G face (I src , I tar )→I output , where I src is the source image, I tar is the target background / face.

[0100] In addition, in step S102 of this embodiment, data enhancement processing is performed on the data in the initial training data set to obtain an enhanced training data set with more data, including:

[0101] Performing data enhancement processing on the video data and the speech data in the first enhanced training data set by timbre conversion to obtain a second enhanced training data set;

[0102] The data in the initial training data set, the first enhanced training data set and the second enhanced training data set together constitute an enhanced training data set.

[0103] Among them, the timbre is modified by the voice feature transfer model, and the formula is: new =VoiceConv(A orig , F target ), where F target is the target timbre feature vector.

[0104] Furthermore, in the method for automatically generating full-body digital human data of this embodiment, the step S103 of filtering the new digital human speech data to obtain an extended training data set includes:

[0105] The new digital human speech data is input into the lip alignment model trained and generated in step S102 for screening and filtering, noise data in the new digital human speech data is removed, and the data is added to the enhanced training data set to form an extended training data set.

[0106] The specific implementation process of step S103 in this embodiment includes:

[0107] 1. Input the new text to generate voice data Anew.

[0108] 2. Generate digital human video V new =Text2Video(A new , Model gen ).

[0109] 3. Use the lip alignment model to filter data and filter out noise data: Data with scores above the threshold are retained.

[0110] In this embodiment, a method for automatically generating full-body digital human data includes: after the first model training phase is executed for a certain period of time and the amount of speech data reaches a certain level, the second model training phase includes:

[0111] Step S201, when the data volume of the training data set in the first model training phase reaches a preset data volume threshold, the second model training phase begins;

[0112] Step S202: Generate new description text data based on the video data in the training data set, add the description text data to the video data set, and perform fine-tuning training on the parameters of the text2video model in the digital human generation model;

[0113] Step S203: Replace the corresponding models in the first model training stage with the lip synchronization model and texture generation model in the text2video model obtained through fine-tuning training.

[0114] The BLIP-2 model is used to generate the description text data of the video data: Where Q(·) is the Query Transformer.

[0115] Specifically, in step S202 of this embodiment, generating new description text data based on the video data in the training data set includes:

[0116] Extract speech from video data in the training dataset and convert it into text data text1, and extract frames from video data in the training dataset and convert them into text data text2;

[0117] Assemble text data text1 and text data text2 into new description text data.

[0118] Specifically, in step S202 of this embodiment, fine-tuning the parameters of the text2video model in the digital human generation model includes:

[0119] Fine-tuning loss function combination:

[0120] Reconstruction loss:

[0121] Optical flow consistency loss:

[0122] In step S203 of this embodiment, replacing the corresponding model in the first model training phase includes:

[0123] Smooth transition strategy between old and new models: W new =αW old +(1-α)W fine-tuned , α=0.3;

[0124] Dynamic threshold adjustment:

[0125] This embodiment also provides a device for automatically generating full-body digital human data, including a first model training phase module and a first model phase process module:

[0126] The first model training phase module: uses the collected real-person speech dataset and the obtained open-source dataset as the initial training dataset, uses the initial training dataset to train the lip alignment model and the digital human generation model, performs data enhancement and data expansion processing based on the initial training dataset, and continues to train the lip alignment model and the digital human generation model until the accuracy of the lip alignment model and the digital human generation model meets the preset model parameter threshold;

[0127] The second model training phase module: when the data volume of the training data set in the first model training phase reaches a preset data volume threshold, generates new descriptive text data based on the video data in the training data set, adds the descriptive text data to the video data set, fine-tunes the digital human generation model, and replaces the corresponding model in the first model training phase;

[0128] The first model training phase module and the second model training phase module are executed in a loop iteration at the same time until the preset training target is met.

[0129] The device of this embodiment for automatically generating full-body digital human data mainly includes two core modules: a first model training phase module and a second model training phase module.

[0130] The first model training phase module: trains the lip alignment model and digital human generation model based on the initial dataset. The initial dataset includes a dataset of real-life speech collected and an open-source dataset. This avoids the costly, time-consuming, and labor-intensive problem of manually collecting speech data strictly according to high standards. In addition, by data augmentation and expansion, more speech data is produced to continue training the lip alignment model and digital human generation model, gradually improving model accuracy.

[0131] Second model training phase module: When the data volume reaches the threshold, description text is generated based on the video data, the digital human generation model is fine-tuned and the old model components are replaced.

[0132] The first model training phase module and the second model training phase module are iterated cyclically until the training objectives are met.

[0133] Therefore, in the present embodiment, a device for automatically generating full-body digital human data, the first model training phase module collects a portion of a real-person speech data set and combines it with an acquired open-source data set as the initial data set to train the lip alignment model and the digital human generation model. Compared with relying solely on manual and strict high-standard speech data collection, this solves the problem of being extremely costly, time-consuming, and labor-intensive. By enhancing and expanding data to generate more speech data, more speech data can be generated compared to simply eliminating unqualified data from a large amount of dirty data, thereby gradually improving the model accuracy.

[0134] In addition, influenced by some solutions in the multimodal field, such as the BLIP series, the first model training phase module and the second model training phase module of this embodiment use a method of cleaning, training and generating at the same time to produce digital human speech data.

[0135] The device of this embodiment for automatically generating full-body digital human data has the following beneficial effects:

[0136] Reduce manual annotation costs: Through automated data augmentation, data expansion, and descriptive text generation technologies, the reliance on manually labeled data is significantly reduced.

[0137] Improve generation quality: The two-stage loop training strategy effectively improves the accuracy and robustness of the model.

[0138] Enhanced generalization capability: Through continuous data expansion and model fine-tuning, the model's adaptability to new scenarios and new speakers is improved.

[0139] High degree of automation: The entire process from data preparation to model training is highly automated.

[0140] Flexible and scalable: Supports continuous data expansion and model updates to adapt to changing application needs.

[0141] The following describes an electronic device embodiment of the present invention, which can be considered a specific physical implementation of the method and apparatus embodiments of the present invention described above. Details described in the electronic device embodiment of the present invention should be considered supplementary to the above-mentioned method or apparatus embodiments; details not disclosed in the electronic device embodiment of the present invention can be implemented with reference to the above-mentioned method or apparatus embodiments.

[0142] Figure 2 This is a schematic structural diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a processor and a memory. The memory is used to store a computer executable program. When the computer program is executed by the processor, the processor executes a method for automatically generating full-body digital human data according to embodiment one, two, three, or four.

[0143] like Figure 2 As shown, the electronic device is implemented as a general-purpose computing device. The processor may be one or multiple processors working in concert. The present invention also does not exclude distributed processing, meaning that the processors may be dispersed across different physical devices. The electronic device of the present invention is not limited to a single entity but may also be the sum of multiple physical devices.

[0144] The memory stores a computer executable program, typically a machine-readable code, which can be executed by the processor to enable the electronic device to perform the method of the present invention, or at least some of the steps in the method.

[0145] The memory includes a volatile memory, such as a random access memory unit (RAM) and / or a cache memory unit, and may also be a non-volatile memory, such as a read-only memory unit (ROM).

[0146] Optionally, in this embodiment, the electronic device further includes an I / O interface for exchanging data with an external device. The I / O interface may represent one or more of several types of bus structures, including a storage unit bus or storage unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0147] It should be understood that Figure 2 The electronic device shown is merely an example of the present invention. The electronic device of the present invention may also include elements or components not shown in the above examples. For example, some electronic devices also include display units such as screens, and some electronic devices also include human-computer interaction elements such as buttons and keyboards. As long as the electronic device can execute a computer-readable program stored in its memory to implement the method of the present invention or at least some of the steps of the method, it is considered an electronic device covered by the present invention.

[0148] Figure 3 FIG is a schematic diagram of a computer readable recording medium according to an embodiment of the present invention. Figure 3 As shown, a computer-readable recording medium stores a computer-executable program. When executed, the computer-executable program implements a method for automatically generating full-body digital human data according to Embodiment 1, 2, 3, or 4 of the present invention. The computer-readable recording medium may include a data signal propagated in baseband or as part of a carrier wave, carrying readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The readable recording medium may also be any readable medium other than a readable recording medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, device, or component. The program code contained on the readable recording medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0149] The program code for performing the operations of the present invention may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0150] Through the above description of the implementation mode, it is easy for those skilled in the art to understand that the present invention can be implemented by hardware capable of executing a specific computer program, such as the system of the present invention, and the electronic processing unit, server, client, mobile phone, control unit, processor, etc. contained in the system. The present invention can also be implemented by computer software that executes the method of the present invention, such as control software executed by a microprocessor, an electronic control unit, a client, a server, etc. However, it should be noted that the computer software that executes the method of the present invention is not limited to being executed by one or a specific hardware entity, and it can also be implemented in a distributed manner by unspecified specific hardware. For computer software, the software product can be stored in a computer-readable recording medium (which can be a CD-ROM, a USB flash drive, a mobile disk, etc.), or it can be distributed and stored on a network, as long as it enables an electronic device to execute the method according to the present invention.

[0151] The above embodiments are only used to illustrate the present invention and are not intended to limit the technical solutions described in the present invention. Although this specification has described the present invention in detail with reference to the above embodiments, the present invention is not limited to the above specific implementation methods. Therefore, any modification or equivalent replacement of the present invention; and all technical solutions and improvements thereof that do not depart from the spirit and scope of the invention are included in the scope of the claims of the present invention.

Claims

1. A method for automatically generating full-body digital human data, characterized in that: It includes the first model training phase and the second model phase: The first model training stage: the collected real-person speech dataset and the obtained open-source dataset are used as the initial training dataset, the lip alignment model and the digital human generation model are trained using the initial training dataset, data enhancement and data expansion processing are performed based on the initial training dataset, and the lip alignment model and the digital human generation model are continuously trained until the accuracy of the lip alignment model and the digital human generation model meets the preset model parameter threshold; The second model training stage: when the data volume of the training data set in the first model training stage reaches a preset data volume threshold, new descriptive text data is generated based on the video data in the training data set, the descriptive text data is added to the video data set, the digital human generation model is fine-tuned, and the corresponding model in the first model training stage is replaced; The first model training phase and the second model training phase are executed in a loop iteration at the same time until the preset training target is met.

2. The method for automatically generating full-body digital human data according to claim 1, characterized in that: The first model training phase includes: Step S101: The collected real-person speech dataset and the obtained open-source dataset are used as the initial training dataset, and the lip alignment model and the digital human generation model are trained using the initial training dataset; Step S102, performing data augmentation processing on the data in the initial training dataset to obtain an enhanced training dataset with more data, and using the enhanced training dataset to continue training the lip alignment model and the digital human generation model; Step S103: Add a new speech document and generate voice data corresponding to the speech document. Based on the enhanced training data set in step S102, generate new digital human speech data. Filter the new digital human speech data to obtain an extended training data set. Use the extended training data set to continue training the lip alignment model and the digital human generation model. Step S104, iterating steps S102 to S103 in a loop until the accuracy of the lip alignment model and the digital human generated model meets a preset threshold.

3. The method for automatically generating full-body digital human data according to claim 2, characterized in that: In step S102, data enhancement processing is performed on the data in the initial training data set to obtain an enhanced training data set with more data, including: Data enhancement processing is performed on the video data and image data in the initial training data set by changing faces and backgrounds to obtain a first enhanced training data set.

4. The method for automatically generating full-body digital human data according to claim 3, characterized in that: In step S102, data enhancement processing is performed on the data in the initial training data set to obtain an enhanced training data set with more data, including: Performing data enhancement processing on the video data and the speech data in the first enhanced training data set by timbre conversion to obtain a second enhanced training data set; The data in the initial training data set, the first enhanced training data set and the second enhanced training data set together constitute an enhanced training data set.

5. The method for automatically generating full-body digital human data according to claim 4, characterized in that: The expanded training data set obtained by screening and filtering the new digital human speech data in step S103 includes: The new digital human speech data is input into the lip alignment model trained and generated in step S102 for screening and filtering, noise data in the new digital human speech data is removed, and the data is added to the enhanced training data set to form an extended training data set.

6. The method for automatically generating full-body digital human data according to any one of claims 1 to 5, characterized in that: The second model training phase includes: Step S201, when the data volume of the training data set in the first model training phase reaches a preset data volume threshold, the second model training phase begins; Step S202: Generate new description text data based on the video data in the training data set, add the description text data to the video data set, and perform fine-tuning training on the parameters of the text2video model in the digital human generation model; Step S203: Replace the corresponding models in the first model training stage with the lip synchronization model and texture generation model in the text2video model obtained through fine-tuning training.

7. The method for automatically generating full-body digital human data according to claim 6, characterized in that: The step S202 of generating new description text data based on the video data in the training data set includes: Extract speech from video data in the training dataset and convert it into text data text1, and extract frames from video data in the training dataset and convert them into text data text2; Assemble text data text1 and text data text2 into new description text data.

8. A device for automatically generating full-body digital human data, characterized in that: Including the first model training stage module and the first model stage program module: The first model training phase module: uses the collected real-person speech dataset and the obtained open-source dataset as the initial training dataset, uses the initial training dataset to train the lip alignment model and the digital human generation model, performs data enhancement and data expansion processing based on the initial training dataset, and continues to train the lip alignment model and the digital human generation model until the accuracy of the lip alignment model and the digital human generation model meets the preset model parameter threshold; The second model training phase module: when the data volume of the training data set in the first model training phase reaches a preset data volume threshold, generates new descriptive text data based on the video data in the training data set, adds the descriptive text data to the video data set, fine-tunes the digital human generation model, and replaces the corresponding model in the first model training phase; The first model training phase module and the second model training phase module are executed in a loop iteration at the same time until the preset training target is met.

9. An electronic device comprising a processor and a memory, wherein the memory is used to store a computer executable program, wherein: When the computer program is executed by the processor, the processor executes the method for automatically generating full-body digital human data as described in any one of claims 1 to 7.

10. A computer-readable recording medium storing a computer-executable program, characterized in that: When the computer executable program is executed, the method for automatically generating full-body digital human data as described in any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Digital human image generation method combining cold start driving and active learning mechanism

    CN120833401A