Virtual digital human generation algorithm system based on radiation light field
By adopting a virtual digital life generation algorithm system based on the radiation light field in 3D speaking life generation, a dynamic implicit neural light segment field network is built and the teacher radiation field network is introduced, and the problem of poor quality and efficiency of 3D speaking life generation in the existing technology is solved, and high-quality, fast rendering and low resource dependence virtual digital life generation is achieved.
Patent Information
- Application Number
- CN202411721873.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-28
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to achieve a balance of high quality and efficiency in 3D speaking life, especially in terms of facial expressions and lip simultaneous synchronization, and lacks flexibility and adaptability, so it is impossible to quickly generate virtual digital people that meet the requirements of specific scenarios.
Using a virtual digital life generation algorithm system based on radiation light field, a dynamic implicit neural light segment field network (NLDF) is constructed and the teacher radiation field network is introduced, combined with knowledge distillation technology, high-quality 3D digital people are generated, and the rendering is optimized through active beam strategies to improve the generation efficiency.
It realizes high visual quality 3D speaking life generation, with facial expressions and lip shapes that are highly similar to real speakers, and has greatly improved rendering speed, which can quickly meet application scenarios with high real-time requirements, such as virtual live broadcasts, and reduces dependence on large-scale data sets and high-computing computing resources.
Smart Images

Figure CN119941935A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to a virtual digital human generation algorithm system based on radiation light field. Background Art
[0002] In the fields of virtual digital humans, film and television production, human-computer interaction, etc., high-quality and efficient 3D speaker generation technology is of great significance. Traditional generation methods often find it difficult to strike a balance between generation quality and efficiency. For example, some early methods did not produce enough details when generating the appearance of virtual digital humans, their facial expressions were stiff, their movements were not natural and smooth enough, and they could not bring a realistic experience to users. At the same time, these methods usually require a lot of computing resources and time costs, and it is difficult to meet application scenarios with high real-time requirements, such as real-time virtual live broadcasts.
[0003] In terms of synchronization with audio, many existing technologies also have shortcomings. The lip movements of virtual digital humans cannot be accurately matched with the voice content, resulting in an obvious sense of incoordination during the voice interaction process, which seriously affects the user's immersion and interactive experience. In addition, the existing virtual digital human generation system lacks sufficient flexibility and adaptability when facing the diverse needs of different application scenarios, and cannot quickly and conveniently generate virtual digital humans that meet specific scenario requirements (such as different styles, different resolutions, different interactive functions, etc.).
[0004] Therefore, a virtual digital human generation algorithm system based on radiation light field is proposed. Summary of the invention
[0005] In view of the above problems, the present invention provides a virtual digital human generation algorithm system based on radiation light field to solve the problems raised in the above background technology.
[0006] The technical solution of the present invention is:
[0007] A virtual digital human generation algorithm system based on radiation light field, comprising:
[0008] The data set acquisition unit is used to acquire and preprocess video data, including:
[0009] Video data acquisition module, which acquires several minutes of video data from external devices. This module supports multiple video formats to ensure a wide range of data sources;
[0010] The data preprocessing module divides the video data into 80% training set and 20% test set. p en g l or open source view extraction model to extract the camera view, extract the audio sequence from the video and save it in WAV format, and perform pre-processing operations such as cropping and normalization on the background image;
[0011] Model building unit for building audio-tuned dynamic implicit neural light segment field network (NLDF), including:
[0012] NLDF network building module, which builds a network that splices beam sampling points into a long sequence and outputs the color value of the light segment. The network depth can be selected from 28-layer, 50-layer or 88-layer ResMLP networks;
[0013] The knowledge distillation module introduces the teacher radiation field network and constrains the color value of the light segment through knowledge distillation.
[0014] Its network function is:
[0015] F φ :
[0016] F′ θ :(a,d,p)→(c,σ)(Teacher)
[0017] The rendering formula is:
[0018]
[0019] The loss is set as:
[0020]
[0021] The temperature parameter is set to 0.5;
[0022] A generating unit, for generating 3D digital content according to an audio signal, comprising:
[0023] Audio receiving module, receiving audio signals of various formats and sampling rates;
[0024] The 3D digital content generation module inputs the audio signal into the trained NLDF network and combines the data set information to generate dynamically controllable 3D digital content. The resolution can be adjusted between 360p and 1080p, and the video sequence can be displayed or saved in real time.
[0025] The neural light dynamic field optimization module reduces the amount of NeRF calculations by rendering pixels in a single light segment, speeding up the generation of 3D talking heads. Its rendering follows specific rules.
[0026] Dynamic implicit neural light segment field network enhancement module, which splices beam sampling points into long sequences, predicts facial movements based on audio, improves quality using knowledge distillation, and constrains light segment color values through a teacher network;
[0027] Hyperparameter setting module, which sets common hyperparameters for the above-mentioned related networks, including learning rate (which can be initially set to 0.001), etc., to improve the applicability and stability of the model and enable adaptive optimization;
[0028] The model training module uses a few minutes of video data to train an audio-driven, dynamically controllable 3D virtual digital human generation model. The facial expressions and pronunciation details of the digital human generated by the model are high-quality and synchronized with the voice, which is applicable to multiple fields.
[0029] The Active Beam Strategy application module uses the Active Beam Strategy in the Neural Light Dynamic Field module to optimize rendering, improve speed and quality, and its strategy is based on a specific algorithm.
[0030] In a further technical solution, the data preprocessing module can accurately calculate parameters such as position, direction, focal length, etc. when extracting the camera perspective, providing accurate perspective information for 3D reconstruction.
[0031] In a further technical solution, the hyperparameter setting module can dynamically adjust hyperparameters, such as learning rate, temperature parameters, etc., according to the system operation status and generation effect, to improve the training and generation quality.
[0032] In a further technical solution, the model training module can perform real-time enhancement on the data during training, such as random cropping, rotation, flipping, etc., to enhance the generalization ability of the model.
[0033] In a further technical solution, when generating 3D digital content, the generation unit can adjust facial expressions and movement amplitudes according to audio emotional characteristics to enhance emotional expression.
[0034] In further technical solutions, the system can switch between different style templates according to application scenarios and customize the appearance features of digital humans, such as clothing, hairstyle, etc.
[0035] In further technical solutions, the system has multi-language support capabilities to ensure that the pronunciation and lip shape of the digital human generated by audio in different languages are accurate.
[0036] In a further technical solution, the system can monitor resource usage while running, automatically adjust algorithm parameters or strategies when resources are tight, and provide logs and reports to facilitate maintenance and optimization.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] 1. Through the unique NLDF network architecture and knowledge distillation technology, the present invention can generate 3D speakers with high visual quality. The facial expressions and lip shapes generated are highly similar to those of real speakers, and the details are more realistic, such as accurate blinking motion capture.
[0039] 2. Compared with traditional 3D speaker generation methods (such as NeRF-based methods), the rendering speed of this system is greatly improved. For example, it takes about 7 hours to generate a 30-second video with a resolution of 512×512 using a V100 GPU with ADNeRF. The NLDF method of this system renders 30 times faster than ADNeRF and 70 times faster than DFRF, which can more quickly meet application scenarios with high real-time requirements, such as virtual live broadcast.
[0040] 3. By using a few minutes of short video as a data set and a few hours of single-card training, high-quality 3D speaker generation can be achieved, which greatly reduces the dependence on large-scale data sets and high-computing power computing resources (such as multi-card parallel computing, etc.), reduces the application cost, and enables efficient 3D speaker generation on ordinary computing devices.
[0041] 4. The dynamically controllable 3D speakers generated by this system can be widely used in virtual anchors, film and television production, virtual digital humans, human-computer interaction and other fields, providing more vivid and realistic digital characters for these fields and improving the quality of user experience and content creation. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 It is a schematic diagram of the overall operation flow of the present invention;
[0043] Figure 2 It is a schematic diagram of the data preprocessing flow chart of the present invention;
[0044] Figure 3 It is a schematic diagram of the model construction and training flow chart of the present invention;
[0045] Figure 4 It is a schematic diagram of the generation, interaction and optimization flow chart of the present invention. DETAILED DESCRIPTION
[0046] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the present invention is further explained below in conjunction with specific implementation methods.
[0047] In the description of the present invention, it should be noted that the terms "upper", "lower", "inner", "outer", "front end", "rear end", "two ends", "one end", "the other end" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention. In addition, the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance.
[0048] In the description of the present invention, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "provided with", "connected", etc. should be understood in a broad sense. For example, "connected" can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be indirectly connected through an intermediate medium, or it can be the internal communication of two components. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0049] Example:
[0050] See also Figure 1-Figure 4 , a virtual digital human generation algorithm system based on radiation light field, including:
[0051] The data set acquisition unit is used to acquire and preprocess video data, including:
[0052] Video data acquisition module, which acquires several minutes of video data from external devices. This module supports multiple video formats to ensure a wide range of data sources;
[0053] The data preprocessing module divides the video data into 80% training set and 20% test set. p en g l or open source view extraction model to extract the camera view, extract the audio sequence from the video and save it in WAV format, and perform pre-processing operations such as cropping and normalization on the background image;
[0054] Model building unit for building audio-tuned dynamic implicit neural light segment field network (NLDF), including:
[0055] NLDF network building module, which builds a network that splices beam sampling points into a long sequence and outputs the color value of the light segment. The network depth can be selected from 28-layer, 50-layer or 88-layer ResMLP networks;
[0056] The knowledge distillation module introduces the teacher radiation field network and constrains the color value of the light segment through knowledge distillation.
[0057] Its network function is:
[0058] F φ :
[0059] F′ θ :(a,d,p)→(c,σ)(Teacher)
[0060] The rendering formula is:
[0061]
[0062] The loss is set as:
[0063]
[0064] The temperature parameter is set to 0.5;
[0065] A generating unit, for generating 3D digital content according to an audio signal, comprising:
[0066] Audio receiving module, receiving audio signals of various formats and sampling rates;
[0067] The 3D digital content generation module inputs the audio signal into the trained NLDF network and combines the data set information to generate dynamically controllable 3D digital content with a resolution of 360 p -1080 p Time adjustment, real-time display or saving of video sequences;
[0068] The neural light dynamic field optimization module reduces the amount of NeRF calculations by rendering pixels in a single light segment, speeding up the generation of 3D talking heads. Its rendering follows specific rules.
[0069] Dynamic implicit neural light segment field network enhancement module, which splices beam sampling points into long sequences, predicts facial movements based on audio, improves quality using knowledge distillation, and constrains light segment color values through a teacher network;
[0070] Hyperparameter setting module, which sets common hyperparameters for the above-mentioned related networks, including learning rate (which can be initially set to 0.001), etc., to improve the applicability and stability of the model and enable adaptive optimization;
[0071] The model training module uses a few minutes of video data to train an audio-driven, dynamically controllable 3D virtual digital human generation model. The facial expressions and pronunciation details of the digital human generated by the model are high-quality and synchronized with the voice, which is applicable to multiple fields.
[0072] The Active Beam Strategy application module uses the Active Beam Strategy in the Neural Light Dynamic Field module to optimize rendering, improve speed and quality, and its strategy is based on a specific algorithm.
[0073] The working principle of the above technical solution is as follows:
[0074] The system first obtains several minutes of video data in various formats from external devices through the video data acquisition module to ensure that the data source is extensive. Then, the data preprocessing module divides the video data into 80% training set and 20% test set, and uses opengl or open source perspective extraction model to extract camera perspective information, laying the foundation for building accurate 3D scenes. At the same time, the audio sequence is extracted from the video and saved in WAV format to facilitate subsequent audio feature extraction and processing, and the background image is preprocessed by cropping and normalizing to better integrate the background with the generated virtual digital human;
[0075] NLDF network construction: The NLDF network construction module constructs a dynamic implicit neural light segment field network (NLDF), splices the sampling points on the light beam into a long sequence, feeds it into a large network for processing at one time, and outputs the color value of each light segment. The network depth can be selected from 28-layer, 50-layer or 88-layer ResMLP networks to adapt to different computing resources and generation quality requirements;
[0076] Knowledge distillation: The knowledge distillation module introduces the teacher radiation field network. Through the knowledge distillation technology, the teacher network is used to constrain the light segment color value output by the NLDF network. Its network function, rendering formula and loss setting follow specific rules, and the temperature parameter is set to 0.5 to improve the generation quality;
[0077] Model training: The model training module uses several minutes of preprocessed video data for training. It uses advanced training algorithms and optimization strategies, and sets hyperparameters such as the initial learning rate to 0.001. During the training process, the data is enhanced in real time to increase data diversity and improve the generalization ability of the model. The training obtains an audio-driven, dynamically controllable 3D virtual digital human generation model, so that the generated digital human's facial expressions and pronunciation details are of high quality and synchronized with the speech;
[0078] The audio receiving module receives audio signals of various formats and sampling rates from external input. The 3D digital content generation module inputs the audio signal into the trained NLDF network and generates dynamically controllable 3D digital content in combination with the data set information. The resolution can be adjusted between 360p-1080p to meet the needs of different application scenarios. The generated video sequence can be displayed in real time or saved as a video file;
[0079] Neural-optical dynamic field optimization: The Neural-optical dynamic field optimization module renders pixels in a single pass through light segments, reducing the extensive calculations required by traditional NeRF, speeding up the process of generating 3D talking heads, and following specific rendering rules;
[0080] Dynamic Implicit Neural Light Segment Field Network Enhancement: The Dynamic Implicit Neural Light Segment Field Network Enhancement module splices the beam sampling points into long sequences and predicts facial movements based on audio signals. It uses knowledge distillation to improve quality and further constrains the light segment color values through the teacher network;
[0081] Hyperparameter setting and adaptive optimization: The hyperparameter setting module sets common hyperparameters for related networks to improve model applicability and stability, and can perform adaptive optimization based on system operation status and generation effects;
[0082] Active Beam Strategy Application: The Active Beam Strategy Application module uses the active beam strategy in the neural light dynamic field module to optimize rendering based on the activity of the beam, improving rendering speed and quality.
[0083] See also Figure 1 and Figure 4 ,The data preprocessing module can accurately calculate the position, direction, focal length and other parameters when extracting the camera ,viewing angle, providing accurate viewing angle information for 3D reconstruction.
[0084] In the data preprocessing stage, when extracting the camera perspective, the system uses tools such as OpenGL or open source perspective extraction models to analyze each frame of the video data. By identifying specific feature points and object edges in the image, and combining the geometric relationship of the image and the physical characteristics of the shooting scene, a series of mathematical algorithms are used to infer the camera's position, direction, focal length and other parameters.
[0085] For example, to calculate the position of a camera, we can analyze the position changes of the same object in multiple frames and the movement patterns of the camera, and use triangulation and other methods to determine the coordinates of the camera in three-dimensional space. To determine the direction of the camera, we may infer the shooting direction of the camera based on clues such as the perspective relationship of objects in the image and the convergence point of lines. The focal length can be calculated by analyzing the size ratio, blur level and other features of the objects in the image, combined with known information such as the actual size of the object, to estimate the focal length of the camera.
[0086] See also Figure 1 and Figure 4 ,The hyperparameter setting module can dynamically adjust the hyperparameters such as learning rate and ,temperature parameters according to the system operation status and generation effect, ,to improve the training and generation quality.
[0087] By dynamically adjusting hyperparameters such as the learning rate, we can avoid slow convergence or non-convergence during training. A reasonable learning rate can enable the model to quickly learn the general characteristics of the data in the early stages of training, and finely adjust the model parameters in the later stages, thereby accelerating the convergence of the model and reducing training time.
[0088] See also Figure 1 and Figure 4 ,During training, the model training module can perform real-time data enhancement, such as random cropping, rotation, and flipping, to enhance the model generalization ,ability.
[0089] Through random cropping, rotation, and flipping operations, the diversity of training data is effectively increased. After these transformations, the originally limited data set can generate a large number of different samples, allowing the model to be exposed to more different scenes and situations. For example, after different cropping, rotation, and flipping, the video data of a person can present different postures, angles, and backgrounds, allowing the model to learn a wider range of feature representations.
[0090] See also Figure 1 and Figure 4 When the generation unit generates 3D digital content, it can adjust the expression and movement amplitude according to the audio emotional characteristics to enhance the emotional expression.
[0091] By adjusting the virtual human's expression and movement amplitude according to the emotional characteristics of the audio, the virtual human can convey the emotional information in the audio more vividly. The audience can feel the emotional state of the virtual human more intuitively, enhancing the effect of emotional expression. For example, when the virtual anchor tells a touching story, the virtual human's sad expression and slow movements can help the audience better understand the emotional atmosphere of the story and enhance the audience's immersion.
[0092] See also Figure 1 and Figure 4 The system can switch between different style templates according to the application scenario and customize the appearance features of the digital human, such as clothing, hairstyle, etc.
[0093] Different application scenarios have different requirements for the style and appearance of virtual digital people. By switching style templates and customizing appearance features according to application scenarios, the system can meet the diverse needs of various scenarios. For example, in the virtual anchor scenario, the virtual digital person needs to have a unique personality and charm to attract the audience's attention; in the film and television production scenario, the virtual digital person needs to be able to integrate with the plot and scene to enhance the visual effect of the film. By providing a variety of style templates and appearance feature customization options, the system can meet the needs of different users and application scenarios and improve the applicability and flexibility of the system.
[0094] See also Figure 1 and Figure 4 ,The system has multi-language support capabilities, ensuring that the pronunciation and mouth shape of the digital human generated by audio in different languages are accurate.
[0095] Having multilingual support can break down language barriers and enable virtual digital humans to be widely used around the world. Users from different countries and regions can interact with virtual digital humans in their own languages without having to worry about language barriers. This is of great significance to multinational companies, international exchanges, multilingual education and other fields. For example, in virtual meetings of multinational companies, virtual digital humans can speak and communicate in different languages to improve communication efficiency; in multilingual education, virtual digital humans can teach in different languages to help students learn different languages.
[0096] See also Figure 1 and Figure 4 ,The system can monitor resource usage when running, automatically adjust algorithm parameters or strategies when resources are tight, and provide logs and reports to facilitate maintenance and optimization.
[0097] By monitoring resource usage in real time and automatically adjusting algorithm parameters or strategies when resources are tight, the system can avoid crashes or performance degradation caused by insufficient resources. This improves the stability and reliability of the system, ensuring that the system can continue to run stably, especially when running for a long time or facing high load;
[0098] For example, in a virtual live broadcast scenario, if the system can adjust resource allocation in a timely manner to avoid freezes and degradation of picture quality, it can provide a better user experience and ensure the smooth progress of the live broadcast.
[0099] When working, after the system is started, it first actively collects video data of several minutes in length from external devices through the video data acquisition module. The multiple video formats it supports ensure the wide range of data sources. Subsequently, the data preprocessing module quickly starts working and divides the video data according to the established ratio (80% training set, 20% test set) to provide a reasonable data distribution for subsequent model training and verification. In this process, opengl or open source perspective extraction models are used to accurately extract the camera perspective information of each frame of video, including parameters such as the camera's position, direction, and focal length, providing an accurate perspective basis for 3D reconstruction. At the same time, the audio sequence is extracted from the video and saved in WAV format for subsequent audio feature extraction and processing, and the background image is preprocessed by cropping and normalization, so that the background image can be better integrated with the subsequently generated virtual digital human, presenting a natural and harmonious visual effect;
[0100] The NLDF network building module starts to build a dynamic implicit neural light segment field network (NLDF), splicing the sampling points on the light beam into a long sequence, and then sending it to the large network for processing at one time, and finally outputting the color value of each light segment. According to actual needs, the network depth can be flexibly selected in the 28-layer, 50-layer or 88-layer ResMLP network to adapt to different computing resources and generation quality requirements;
[0101] The knowledge distillation module introduces the teacher radiation field network. Through the knowledge distillation technology, the teacher network is used to constrain the color value of the light segment output by the NLDF network. It operates according to specific network functions, rendering formulas and loss settings, and the temperature parameter is set to 0.5 to improve the generation quality;
[0102] The model training module uses several minutes of preprocessed video data for training, and adopts advanced training algorithms and optimization strategies, such as setting hyperparameters such as the initial learning rate to 0.001. During the training process, real-time data enhancement operations are performed, including random cropping, rotation, flipping, etc., to increase data diversity and improve the generalization ability of the model. After multiple iterative training, an audio-driven, dynamically controllable 3D virtual digital human generation model is obtained, so that the generated digital human's facial expressions and pronunciation details are high-quality and synchronized with the voice;
[0103] The audio receiving module is always ready to receive external audio signals of various formats and sampling rates. Once the audio signal is received, the 3D digital content generation module immediately inputs the audio signal into the trained NLDF network, combines the data set information, and generates dynamically controllable 3D digital content through complex calculation and processing. The resolution can be adjusted between 360p-1080p to meet the needs of different application scenarios. The generated content can be displayed on the screen in real time, so that users can view the effect immediately, or it can be saved as a video file for subsequent editing and use;
[0104] Neural-optical dynamic field optimization: The Neural-optical dynamic field optimization module renders pixels in a single pass through light segments, reducing the extensive calculations required by traditional NeRF, accelerating the generation process of 3D talking heads, following specific rendering rules, and improving generation efficiency;
[0105] The dynamic implicit neural light segment field network enhancement module splices the beam sampling points into long sequences for processing, predicts facial movements based on audio signals, uses knowledge distillation to improve quality, and further constrains the light segment color values through the teacher network, making the generated virtual digital human more realistic and natural;
[0106] The hyperparameter setting module continuously monitors the system operation status and generation effect, and dynamically adjusts hyperparameters such as learning rate and temperature parameters as needed to improve the model applicability and stability, and improve the training and generation quality;
[0107] The Active Beam Strategy Application Module uses the active beam strategy in the Neural Light Dynamic Field Module to optimize rendering based on the activity of the beam, thereby improving rendering speed and quality.
[0108] The resource monitoring function is always active throughout the system operation. It continuously monitors the usage of various computing resources, including CPU usage, GPU memory usage, memory usage, network bandwidth, etc. When it is judged that resources are tight, it automatically adjusts algorithm parameters or strategies, such as reducing the learning rate, reducing the resolution of virtual digital people, adjusting light segment rendering parameters, etc., to optimize resource usage and ensure system stability and performance. At the same time, the system generates detailed log files and regular reports to provide system administrators and developers with a basis for maintenance and optimization.
[0109] In summary, the various modules of the virtual digital human generation algorithm system work together to efficiently generate high-quality, dynamically controllable 3D digital content, and can be adaptively adjusted according to different application scenarios and resource conditions, providing users with powerful virtual digital human creation tools.
[0110] The above-mentioned embodiments only express the specific implementation of the present invention, and the description thereof is relatively specific and detailed, but it cannot be understood as limiting the scope of the present invention. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.
Claims
1. A virtual digital human generation algorithm system based on radiation light field, characterized in that: include: The data set acquisition unit is used to acquire and preprocess video data, including: Video data acquisition module, which acquires several minutes of video data from external devices. This module supports multiple video formats to ensure a wide range of data sources; The data preprocessing module divides the video data into 80% training set and 20% test set. p en g l or open source view extraction model to extract the camera view, extract the audio sequence from the video and save it in WAV format, and perform pre-processing operations such as cropping and normalization on the background image; Model building unit for building audio-tuned dynamic implicit neural light segment field network (NLDF), including: NLDF network building module, which builds a network that splices beam sampling points into a long sequence and outputs the color value of the light segment. The network depth can be selected from 28-layer, 50-layer or 88-layer ResMLP networks; The knowledge distillation module introduces the teacher radiation field network and constrains the color value of the light segment through knowledge distillation. Its network function is: Fθ:(a,d,p)→(c,σ)(Teacher) The rendering formula is: The loss is set as: The temperature parameter is set to 0.5; A generating unit, for generating 3D digital content according to an audio signal, comprising: Audio receiving module, receiving audio signals of various formats and sampling rates; The 3D digital content generation module inputs the audio signal into the trained NLDF network and combines the data set information to generate dynamically controllable 3D digital content. The resolution can be adjusted between 360p and 1080p, and the video sequence can be displayed or saved in real time. The neural light dynamic field optimization module reduces the amount of NeRF calculations by rendering pixels in a single light segment, speeding up the generation of 3D talking heads. Its rendering follows specific rules. Dynamic implicit neural light segment field network enhancement module, which splices beam sampling points into long sequences, predicts facial movements based on audio, improves quality using knowledge distillation, and constrains light segment color values through a teacher network; Hyperparameter setting module, which sets common hyperparameters for the above-mentioned related networks, including learning rate (which can be initially set to 0.001), etc., to improve the applicability and stability of the model and enable adaptive optimization; The model training module uses a few minutes of video data to train an audio-driven, dynamically controllable 3D virtual digital human generation model. The facial expressions and pronunciation details of the digital human generated by the model are high-quality and synchronized with the voice, which is applicable to multiple fields. The Active Beam Strategy application module uses the Active Beam Strategy in the Neural Light Dynamic Field module to optimize rendering, improve speed and quality, and its strategy is based on a specific algorithm.
2. The system according to claim 1, characterized in that: The data preprocessing module can accurately calculate parameters such as position, direction, focal length, etc. when extracting the camera perspective, providing accurate perspective information for 3D reconstruction.
3. The system according to claim 1, characterized in that: The hyperparameter setting module can dynamically adjust hyperparameters, such as learning rate, temperature parameters, etc., according to the system operation status and generation effect to improve the training and generation quality.
4. The system according to claim 1, characterized in that: The model training module can perform real-time data enhancement during training, such as random cropping, rotation, flipping, etc., to enhance the model's generalization ability.
5. The system according to claim 1, characterized in that: When generating 3D digital content, the generation unit can adjust facial expressions and movement amplitudes according to audio emotional characteristics to enhance emotional expression.
6. The system according to claim 1, characterized in that: The system can switch between different style templates according to the application scenario and customize the appearance features of the digital human, such as clothing, hairstyle, etc.
7. The system according to claim 1, characterized in that: The system has multi-language support capabilities to ensure that the pronunciation and lip shape of digital humans generated by audio in different languages are accurate.
8. The system according to claim 1, characterized in that: The system can monitor resource usage while running, automatically adjust algorithm parameters or strategies when resources are tight, and provide logs and reports to facilitate maintenance and optimization.