A method, system, device and storage medium for quickly generating a talking digital human
By applying NeRF and multimodal attention mechanisms in digital human generation, combining adaptive pose coding and octree decomposition, the problems of poor rendering effect and slow convergence speed in traditional methods are solved, and high-quality and fast-generated conversational digital people are achieved.
Patent Information
- Application Number
- CN202410571446.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-09
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-05-09
AI Technical Summary
Traditional digital human production methods are difficult to achieve high-quality rendering effects, and the convergence speed is slow, making it difficult to adapt to complex facial movements and posture changes.
A NeRF-based method is adopted to combine audio features with spatial regions through a multimodal attention mechanism to realize facial motion modeling; adaptive pose coding is used to map complex pose information to spatial coordinates, and combined with octree to decompose 3D space, NeRF model is trained to generate high-quality conversational digital people.
The conversational digital person with fast convergence and high-quality rendering can meet the requirements in real-time applications and improve the computing efficiency and rendering quality of digital person generation.
Smart Images

Figure CN118379407B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computer vision technology, and particularly to a method, system, device, and storage medium for quickly generating a talking digital human. Background Art
[0002] With the continuous development of computer graphics and human-computer interaction technologies, digital humans have been widely used in various fields. However, traditional digital human production methods often have certain limitations, making it difficult to achieve high-quality rendering effects and prone to problems such as slow convergence speed.
[0003] NeRF (Neural Radiance Field) is a deep learning-based method for three-dimensional scene representation, which can generate realistic three-dimensional objects and scenes. However, applying it to the field of digital human synthesis still faces some challenges. First, the head and body of a digital human have complex spatial relationships, and it is difficult for traditional three-dimensional reconstruction methods to accurately capture these relationships. Second, the expression and motion control of a digital human require accurate simulation of audio signals, which requires an efficient model to achieve. In addition, issues such as real-time performance and computational efficiency need to be considered during the generation process of a digital human. Summary of the Invention
[0004] Therefore, embodiments of the present invention provide a method, system, device, and storage medium for quickly generating a talking digital human to solve the technical problems in the prior art that it is difficult to achieve high-quality rendering effects and prone to slow convergence speed when synthesizing digital humans through traditional methods.
[0005] To achieve the above object, embodiments of the present invention provide the following technical solutions:
[0006] According to the first aspect of embodiments of the present invention, a method for quickly generating a talking digital human is provided, and the method includes:
[0007] S1. Collect audio data and related face image data for synthesis and perform preprocessing to generate NeRF format audio data and NeRF format face image data applicable to the NeRF model;
[0008] S2. Use an octree to decompose the 3D space into multiple orthogonal planes and subdivide each spatial cube into entities in the space;
[0009] S3. Use a multimodal attention mechanism to combine audio features with specific spatial regions to achieve facial motion modeling;
[0010] S4. Use adaptive pose encoding to map complex pose information to spatial coordinates to generate pose space coordinate information, providing clear position relationship data for implicit pose learning of the NeRF of the body part;
[0011] S5. Train a preset NeRF model using NeRF format audio data and NeRF format face image data, and adjust the parameters using a loss function and an optimization algorithm to obtain a trained NeRF model.
[0012] S6. Obtain a NeRF format audio data and a NeRF format face image data, use the trained NeRF model, the NeRF format audio data and the NeRF format face image data to generate a digital human to be rendered for a conversation, and render the digital human to be rendered for a conversation to generate a rendered digital human for a conversation.
[0013] Further, collect audio data and related face image data to be used for synthesis and perform preprocessing to generate NeRF format audio data and NeRF format face image data suitable for the NeRF model, including:
[0014] Extract features from the audio data to obtain audio features.
[0015] Perform face detection and alignment on the related face image data to generate corresponding face image data.
[0016] Convert the audio features and the corresponding face image data into NeRF format audio data and NeRF format face image data respectively.
[0017] Further, decompose the 3D space into multiple orthogonal planes using an octree and subdivide each spatial cube into entities in the space, including:
[0018] Efficiently store and traverse voxel data through a recursive structure and an index expander.
[0019] Perform factorization using NeRF-based tri-plane decomposition to reduce the number of hash collisions.
[0020] Further, combine audio features with specific spatial regions using a multimodal attention mechanism to achieve facial motion modeling, including:
[0021] The multimodal attention mechanism is a relational attention network.
[0022] Further, map complex pose information to spatial coordinates using adaptive pose encoding to generate pose space coordinate information, providing clear position relationship data for the NeRF learning of implicit poses of body parts, including:
[0023] Map the complex change information of the head pose to spatial coordinates using adaptive pose encoding.
[0024] Further, train a preset NeRF model using NeRF-format audio data and NeRF-format face image data, and adjust the parameters using a loss function and an optimization algorithm to obtain a trained NeRF model, including:
[0025] The loss function here is the loss function of the generative adversarial network.
[0026] Further, obtain a NeRF-format audio data and a NeRF-format face image data, generate a digital talking person to be rendered using the trained NeRF model and the NeRF-format audio data and NeRF-format face image data, and render the digital talking person to be rendered to generate a rendered digital talking person, including:
[0027] Implement efficient inference and rendering by optimizing the model structure and corresponding parameters;
[0028] Among them, a neural network compression mechanism is used during the rendering process.
[0029] According to the second aspect of the embodiments of the present invention, a system for quickly generating a digital talking person is provided, and the system includes:
[0030] A preprocessing module, configured to collect audio data and related face image data required for synthesis and perform preprocessing to generate NeRF-format audio data and NeRF-format face image data suitable for the NeRF model;
[0031] A 3D space decomposition module, configured to decompose the 3D space into multiple orthogonal planes using an octree and subdivide each space cube into entities in the space;
[0032] A region attention module, configured to combine audio features with specific spatial regions using a multimodal attention mechanism to implement facial motion modeling;
[0033] An adaptive pose encoding module, configured to map complex pose information to spatial coordinates using adaptive pose encoding to generate pose space coordinate information, providing clear position relationship data for the NeRF learning implicit pose of the body part;
[0034] A training module, configured to train a preset NeRF model using NeRF-format audio data and NeRF-format face image data, and adjust the parameters using a loss function and an optimization algorithm to obtain a trained NeRF model;
[0035] A rendering module, configured to obtain an audio data in NeRF format and a face image data in NeRF format, generate a digital human for conversation to be rendered by using a trained NeRF model and the audio data in NeRF format and the face image data in NeRF format, and render the digital human for conversation to be rendered to generate a rendered digital human for conversation.
[0036] According to a third aspect of an embodiment of the present invention, there is provided a device for quickly generating a digital human for conversation, the device including: a processor and a memory;
[0037] The memory is configured to store one or more program instructions;
[0038] The processor is configured to run one or more program instructions to execute the steps of a method for quickly generating a digital human for conversation as described in any one of the above.
[0039] According to a fourth aspect of an embodiment of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of a method for quickly generating a digital human for conversation as described in any one of the above are implemented.
[0040] Embodiments of the present invention have the following advantages:
[0041] 1) High-quality rendering: The present invention adopts a method for generating 2D face modeling videos based on NERF, which can achieve fast convergence in the case of a small model size, so as to obtain a high-quality rendering effect;
[0042] 2) Real-time inference and rendering: By optimizing the model structure and parameters, the present invention realizes an efficient inference and rendering process, so that the synthesized avatar can meet the requirements in real-time applications;
[0043] 3) Dynamic head reconstruction: By using an octree-based representation method to decompose the 3D space into multiple orthogonal planes and further subdivide or represent each spatial cube as an entity in the space, the present invention can achieve dynamic head reconstruction, improve the rendering quality and the convergence speed;
[0044] 4) Adaptive pose encoding: The present invention introduces adaptive pose encoding, maps complex pose transformations to spatial coordinates, and provides a clear positional relationship for implicit pose learning of NeRF for body parts, thereby solving the separation problem between the head and the body;
[0045] 5) Audio Feature and Spatial Region Association Capture: The RegionAttentionModule is proposed, which combines audio features with specific spatial regions through a multimodal attention mechanism to achieve more accurate facial motion modeling. Such a module can better adapt to the needs of different people or scenarios and improve the accuracy and realism of the generated results;
[0046] 6) Reducing the Number of Hash Collisions: Factorization using NeRF-based tri-plane decomposition can effectively reduce the number of hash collisions and further improve the rendering quality;
[0047] 7) Real-time Applications: Due to the efficient inference and rendering process of the present invention, the synthesized avatars can meet the requirements in real-time applications and are widely used in various scenarios such as video conferencing, online education, game characters, etc. Description of the Drawings
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings described below are only exemplary. For those of ordinary skill in the art, without creative efforts, other implementation drawings can be obtained based on the provided drawings.
[0049] The structures, ratios, sizes, etc. shown in this specification are only used to cooperate with the content disclosed in the specification for those familiar with this technology to understand and read, and are not used to limit the limited conditions under which the present invention can be implemented. Therefore, they do not have technical substance. Any modification of the structure, change in the proportional relationship, or adjustment of the size should still fall within the scope that can be covered by the technical content disclosed in the present invention without affecting the effects that the present invention can produce and the purposes that can be achieved.
[0050] Figure 1 It is a schematic logical structure diagram of a system for quickly generating talking digital humans provided by an embodiment of the present invention;
[0051] Figure 2 It is a schematic application principle diagram of a system for quickly generating talking digital humans provided by an embodiment of the present invention. Detailed Embodiments
[0052] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0053] With the continuous development of computer graphics and human-computer interaction technologies, digital humans have been widely used in various fields. However, traditional digital human production methods often have certain limitations, making it difficult to achieve high-quality rendering effects and prone to problems such as slow convergence speed.
[0054] NeRF (Neural Radiance Field) is a deep learning-based 3D scene representation method that can generate realistic 3D objects and scenes. However, applying it to the field of digital human synthesis still faces some challenges. First, the head and body of a digital human have complex spatial relationships, and it is difficult for traditional 3D reconstruction methods to accurately capture these relationships. Second, the expression and motion control of a digital human require precise simulation of audio signals, which requires an efficient model to achieve. In addition, issues such as real-time performance and computational efficiency also need to be considered during the generation process of a digital human.
[0055] To solve the technical problems that it is difficult to achieve high-quality rendering effects and prone to slow convergence speed when synthesizing digital humans through traditional methods.
[0056] Reference Figure 1 , an embodiment of the present invention discloses a system for quickly generating a talking digital human, which includes: a preprocessing module 1; a 3D space decomposition module 2; a regional attention module 3; an adaptive pose encoding module 4; a training module 5; a rendering module 6.
[0057] Corresponding to the above-disclosed system for quickly generating a talking digital human, an embodiment of the present invention also discloses a method for quickly generating a talking digital human. The following details the method for quickly generating a talking digital human disclosed in the embodiments of the present invention in combination with the above-described system for quickly generating a talking digital human.
[0058] Reference Figure 2, the present invention discloses a method for quickly generating a talking digital human. S1. Collect the audio data and related face image data to be used for synthesis and perform preprocessing to generate NeRF format audio data and NeRF format face image data suitable for the NeRF model; S2. Use an octree to decompose the 3D space into multiple orthogonal planes and subdivide each space cube into entities in the space; S3. Use a multimodal attention mechanism to combine audio features with specific spatial regions to achieve facial motion modeling; S4. Use adaptive pose encoding to map complex pose information to spatial coordinates to generate pose space coordinate information, providing clear position relationship data for the NeRF learning of implicit poses of the body part; S5. Use the NeRF format audio data and NeRF format face image data to train a preset NeRF model and use a loss function and an optimization algorithm to adjust the parameters to obtain a trained NeRF model; S6. Obtain a NeRF format audio data and NeRF format face image data, use the trained NeRF model and the NeRF format audio data and NeRF format face image data to generate a to-be-rendered talking digital human, and render the to-be-rendered talking digital human to generate a rendered talking digital human.
[0059] Further, collecting the audio data and related face image data to be used for synthesis and performing preprocessing to generate NeRF format audio data and NeRF format face image data suitable for the NeRF model includes: extracting features from the audio data to obtain audio features; performing face detection and alignment on the related face image data to generate corresponding face image data; and respectively converting the audio features and the corresponding face image data into NeRF format audio data and NeRF format face image data.
[0060] Further, using an octree to decompose the 3D space into multiple orthogonal planes and subdivide each space cube into entities in the space includes: efficiently storing and traversing voxel data through a recursive structure and an index expander; and performing factorization using NeRF-based tri-plane decomposition to reduce the number of hash collisions.
[0061] Further, using a multimodal attention mechanism to combine audio features with specific spatial regions to achieve facial motion modeling includes: the multimodal attention mechanism is a relational attention network.
[0062] Further, using adaptive pose encoding to map complex pose information to spatial coordinates to generate pose space coordinate information, providing clear position relationship data for the NeRF learning of implicit poses of the body part includes: using adaptive pose encoding to map the complex change information of the head pose to spatial coordinates.
[0063] Further, train a preset NeRF model using NeRF-format audio data and NeRF-format face image data, and adjust the parameters using a loss function and an optimization algorithm to obtain a trained NeRF model, including: The loss function here is the loss function of a generative adversarial network.
[0064] Further, obtain a NeRF-format audio data and a NeRF-format face image data, generate a digital human for a conversation to be rendered using the trained NeRF model and the NeRF-format audio data and the NeRF-format face image data, and render the digital human for a conversation to be rendered to generate a rendered digital human for a conversation, including: By optimizing the model structure and corresponding parameters, efficient inference and rendering are achieved;
[0065] Among them, a neural network compression mechanism is used during the rendering process.
[0066] In addition, an embodiment of the present invention further provides a device for quickly generating a digital human for a conversation. The device includes: a processor and a memory; the memory is used to store one or more program instructions; the processor is used to run one or more program instructions to execute the steps of a method for quickly generating a digital human for a conversation as described in any one of the above.
[0067] In addition, an embodiment of the present invention further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of a method for quickly generating a digital human for a conversation as described in any one of the above are implemented.
[0068] In an embodiment of the present invention, the processor may be an integrated circuit chip with signal processing capabilities. The processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0069] The various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention can be implemented or executed. The general-purpose processor can be a microprocessor, or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or can be executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The processor reads the information in the storage medium and combines its hardware to complete the steps of the above method.
[0070] The storage medium can be a memory, for example, it can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories.
[0071] Among them, the non-volatile memory can be a read-only memory (ROM for short), a programmable read-only memory (PROM for short), an erasable programmable read-only memory (EPROM for short), an electrically erasable programmable read-only memory (EEPROM for short), or a flash memory.
[0072] The volatile memory can be a random access memory (RAM for short), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM for short), dynamic random access memory (DRAM for short), synchronous dynamic random access memory (SDRAM for short), double data rate synchronous dynamic random access memory (DDR SDRAM for short), enhanced synchronous dynamic random access memory (ESDRAM for short), synchronous link dynamic random access memory (SLDRAM for short), and direct rambus random access memory (DRRAM for short).
[0073] The storage medium described in the embodiments of the present invention is intended to include but not limited to these and any other suitable types of memories.
[0074] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the present invention can be implemented by a combination of hardware and software. When applying software, the corresponding functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transmission of a computer program from one place to another. The storage media can be any available medium accessible by a general-purpose or special-purpose computer.
[0075] Although the present invention has been described in detail above with general descriptions and specific examples, on the basis of the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.
Claims
1. A method for quickly generating a talking digital person, characterized in that: The method comprises: S1. Collect audio data and related facial image data required for synthesis and perform preprocessing to generate NeRF format audio data and NeRF format facial image data suitable for the NeRF model; S2, using octree to decompose the 3D space into multiple orthogonal planes and subdivide each space cube into entities in the space; S3, using multimodal attention mechanism to combine audio features with specific spatial regions to achieve facial motion modeling; S4, using adaptive posture encoding to map complex posture information to spatial coordinates, generate posture spatial coordinate information, and provide clear position relationship data for NeRF learning implicit posture of body parts; S5. Use the NeRF format audio data and the NeRF format face image data to train a preset NeRF model and use a loss function and an optimization algorithm to adjust parameters to obtain a trained NeRF model; S6, obtaining a NeRF format audio data and a NeRF format facial image data, generating a to-be-rendered talking digital human by using the trained NeRF model and the NeRF format audio data and the NeRF format facial image data, rendering the to-be-rendered talking digital human to generate a rendered talking digital human; The octree is used to decompose the 3D space into multiple orthogonal planes and subdivide each space cube into entities in the space, including: Efficiently store and traverse voxel data through recursive structures and index expressors; NeRF-based tri-plane decomposition is used for factorization to reduce the number of hash collisions.
2. A method for quickly generating a talking digital human as claimed in claim 1, characterized in that: Collect the audio data and related facial image data required for synthesis and perform preprocessing to generate NeRF format audio data and NeRF format facial image data suitable for the NeRF model, including: Performing feature extraction on the audio data to obtain audio features; Perform face detection and alignment on relevant face image data to generate corresponding face image data; The audio features and the corresponding facial image data are converted into NeRF format audio data and NeRF format facial image data respectively.
3. A method for quickly generating a talking digital human as claimed in claim 2, characterized in that: Multimodal attention mechanism is used to combine audio features with specific spatial regions to achieve facial motion modeling, including: The multimodal attention mechanism is a relational attention network.
4. A method for quickly generating a talking digital human as claimed in claim 3, characterized in that: Adaptive posture encoding is used to map complex posture information to spatial coordinates, generate posture space coordinate information, and provide clear position relationship data for NeRF learning implicit posture of body parts, including: Adaptive pose encoding is used to map the complex changes in head posture information into spatial coordinates.
5. A method for quickly generating a talking digital human as claimed in claim 4, characterized in that: Use NeRF format audio data and NeRF format face image data to train the preset NeRF model and use the loss function and optimization algorithm to adjust the parameters to obtain the trained NeRF model, including: The loss function here is the loss function of the generative adversarial network.
6. A method for quickly generating a talking digital human as claimed in claim 5, characterized in that: Obtaining a NeRF format audio data and a NeRF format face image data, generating a to-be-rendered talking digital human by using the trained NeRF model and the NeRF format audio data and the NeRF format face image data, rendering the to-be-rendered talking digital human to generate a rendered talking digital human, including: By optimizing the model structure and corresponding parameters, efficient inference and rendering can be achieved; Among them, a neural network compression mechanism is used in the rendering process.
7. A system for quickly generating a talking digital person, characterized in that: The system comprises: A preprocessing module is used to collect audio data and related facial image data required for synthesis and perform preprocessing to generate NeRF format audio data and NeRF format facial image data suitable for the NeRF model; 3D space decomposition module, used to decompose 3D space into multiple orthogonal planes using octree and subdivide each space cube into entities in space; A region attention module for combining audio features with specific spatial regions using a multimodal attention mechanism for facial motion modeling; An adaptive posture encoding module, used for mapping complex posture information to spatial coordinates using adaptive posture encoding, generating posture spatial coordinate information, and providing clear position relationship data for NeRF learning implicit postures of body parts; A training module is used to train a preset NeRF model using NeRF format audio data and NeRF format face image data and adjust parameters using a loss function and an optimization algorithm to obtain a trained NeRF model; A rendering module is used to obtain audio data in NeRF format and facial image data in NeRF format, generate a talking digital human to be rendered by using the trained NeRF model and the audio data in NeRF format and the facial image data in NeRF format, and render the talking digital human to be rendered to generate a rendered talking digital human; The octree is used to decompose the 3D space into multiple orthogonal planes and subdivide each space cube into entities in the space, including: Efficiently store and traverse voxel data through recursive structures and index expressors; NeRF-based tri-plane decomposition is used for factorization to reduce the number of hash collisions.
8. A device for quickly generating a talking digital person, characterized in that: The device comprises: a processor and a memory; The memory is used to store one or more program instructions; The processor is used to run one or more program instructions to execute the steps of a method for quickly generating a talking digital human as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of a method for quickly generating a talking digital human as claimed in any one of claims 1 to 6.
Citation Information
Patent Citations
Three-dimensional digital human generation and interaction method and system
CN117496072A
Hash coding-based digital human head reconstruction method, system, equipment and medium
CN117893650A