A digital human rendering method and related apparatus
By generating multi-level Gaussian primitive LOD digital humans, and optimizing the rendering detail levels by utilizing the distance between the user's viewpoint and the Gaussian primitive positions, the problem of high computational resources in digital human rendering technology is solved, achieving efficient rendering effects.
Patent Information
- Application Number
- CN202511445282.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-10-10
AI Technical Summary
Existing digital human rendering technologies have high computational resource requirements and long rendering time per frame, making it difficult to meet the needs of large-scale digital human rendering scenarios.
By generating hierarchical detail LOD digital humans based on multi-level Gaussian primitives, the rendering detail hierarchy is optimized by utilizing the distance between the user's real-time viewpoint position and the Gaussian primitive position, thereby reducing the computational load.
While ensuring the visual effects of digital humans, the computational load has been reduced and the rendering efficiency has been improved to meet the needs of large-scale digital human rendering scenarios.
Smart Images

Figure CN120912792B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a digital human rendering method and related apparatus. Background Technology
[0002] Digital human rendering technology aims to use computer graphics algorithms and hardware resources to transform digital human models into realistic and vivid two-dimensional or three-dimensional visual images.
[0003] Currently, digital human rendering can be achieved based on Neural Radiation Field (NeRF) technology. However, this method has high computational resource requirements, and rendering a single frame takes several seconds. This not only increases the computational load on terminal devices but also reduces rendering efficiency, making it difficult to meet the needs of large-scale digital human rendering scenarios.
[0004] Therefore, there is an urgent need for a solution to address the aforementioned technical problems. Summary of the Invention
[0005] Based on the above problems, this application provides a digital human rendering method and related apparatus, which aims to reduce the computational load and meet the needs of large-scale digital human rendering scenarios while ensuring the visual effect of digital humans.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] Firstly, this application provides a digital human rendering method, which generates a 3DGS digital human based on user images from multiple perspectives; determines the level of each Gaussian unit in the 3DGS digital human based on the mean square error between the rendered image and the user image, and generates a hierarchical level of detail (LOD) digital human composed of multiple levels of Gaussian units; the rendered image is generated by rendering the 3DGS digital human; and the LOD digital human is rendered based on the distance between the user's real-time viewpoint position and the positions of each Gaussian unit, generating the rendered digital human.
[0008] This application embodiment generates LOD digital humans, which can optimize the level of detail of LOD digital humans based on the distance between the user's real-time viewpoint position and each Gaussian unit position. While ensuring the visual effect of the digital human, it reduces the computational load and can meet the needs of large-scale digital human rendering scenarios.
[0009] In one possible implementation, determining the hierarchy of each Gaussian unit in the 3DGS digital human based on the mean square error between the rendered image and the user image, and generating a hierarchical level of detail (LOD) digital human composed of multiple levels of Gaussian units, includes:
[0010] Based on the mean square error between the rendered image and the user image, the gradient of each Gaussian unit in the 3DGS digital human is determined; based on the gradient, the level of each Gaussian unit in the 3DGS digital human is determined, and a hierarchical LOD digital human composed of multiple levels of Gaussian units is generated.
[0011] In this embodiment, the level of each Gaussian element in the 3DGS digital human can be determined based on the gradient of Gaussian elements, thereby automatically generating the LOD digital human. This can improve the generation efficiency of the rendered digital human while ensuring the visual effect of the digital human and reducing the computational load.
[0012] In one possible implementation, determining the hierarchy of each Gaussian unit in the 3DGS digital human based on the gradient, and generating a hierarchical level-of-detail (LOD) digital human composed of multiple levels of Gaussian units, includes:
[0013] The Gaussian elements are sorted according to the magnitude of the gradient to obtain the sorting result; based on the sorting result, the level of each Gaussian element in the 3DGS digital human is determined, and a hierarchical LOD digital human composed of multi-level Gaussian elements is generated.
[0014] In this embodiment, the magnitude of the gradient can indicate the level of Gaussian units. The smaller the level, the less detailed information the digital human has, while the larger the level, the more detailed information the digital human has. Based on this, this embodiment can automatically generate LOD digital humans based on gradients, which can improve the generation efficiency of rendered digital humans while ensuring the visual effect of the digital human and reducing the computational load.
[0015] In one possible implementation, determining the hierarchy of each Gaussian unit in the 3DGS digital human based on the sorting result, and generating a hierarchical level of detail (LOD) digital human composed of multiple levels of Gaussian units, includes:
[0016] Based on the total number of Gaussian elements and the attenuation coefficient, the number of Gaussian elements contained in each level is determined; according to the sorting result and the number of Gaussian elements contained in each level, the level of each Gaussian element in the 3DGS digital human is determined, and a hierarchical detail (LOD) digital human composed of multi-level Gaussian elements is generated.
[0017] In this embodiment, the LOD digital human can be automatically generated based on the total number of Gaussian primitives, the attenuation coefficient, and the sorting result, which improves the generation efficiency of LOD digital humans and thus improves the generation efficiency of rendered digital humans.
[0018] In one possible implementation, rendering the LOD digital human based on the distance between the user's real-time viewpoint position and each Gaussian pixel position to generate the rendered digital human includes:
[0019] Based on the distance between the user's real-time viewpoint position and the positions of each Gaussian primitive, the rendering layer corresponding to each Gaussian primitive is determined; each Gaussian primitive is rendered to the corresponding rendering layer to generate the rendered digital human.
[0020] In this application embodiment, different Gaussian primitives can correspond to different rendering levels. For example, Gaussian primitives that are close to the user's real-time viewpoint can correspond to a higher rendering level, while Gaussian primitives that are far from the user's real-time viewpoint can correspond to a lower rendering level. This reduces the computational load while ensuring the visual effect of the digital human, thus meeting the needs of large-scale digital human rendering scenarios.
[0021] In one possible implementation, determining the rendering level corresponding to each hexagonal element based on the distance between the user's real-time viewpoint position and the positions of each hexagonal element includes:
[0022] Based on the ratio between the distance between the user's real-time viewpoint position and the position of each Gaussian primitive and the baseline distance threshold, the distance ratio corresponding to each Gaussian primitive is determined. Taking the attenuation coefficient as the base, the logarithm of each distance ratio is rounded down to obtain the rendering level corresponding to each Gaussian primitive. This facilitates the subsequent rendering of each Gaussian primitive to the corresponding rendering level to generate the rendered digital human. This reduces the computational load while ensuring the visual effect of the digital human and meets the needs of large-scale digital human rendering scenarios.
[0023] In one possible implementation, rendering the LOD digital human based on the distance between the user's real-time viewpoint position and each Gaussian pixel position to generate the rendered digital human includes:
[0024] The LOD digital human is divided into multiple regions. Based on the distance between the user's real-time viewpoint position and the multiple Gaussian positions included in each region, the dominant viewing distance of the Gaussian units corresponding to each region is determined. Based on the dominant viewing distance of the Gaussian units corresponding to each region, the LOD digital human is rendered to generate the rendered digital human.
[0025] In this embodiment of the application, by dividing the LOD digital human into multiple regions, multiple regions of the LOD digital human can be rendered simultaneously, which can effectively improve rendering efficiency.
[0026] In one possible implementation, after rendering the LOD digital man based on the dominant view distance of each region's Gaussian metameter and generating the rendered digital man, the method further includes: performing pixel-level fusion on the rendering results of the transition regions between adjacent regions in the rendered digital man to obtain the fused digital man.
[0027] In this embodiment of the application, by performing pixel-level fusion on the rendering results of the transition area, it can be ensured that the transition area can be as natural as possible in terms of geometric edges and color presentation, so that the fused digital human can present a visually coherent display effect.
[0028] Second aspect: Embodiments of this application provide a digital human rendering apparatus, including:
[0029] The system comprises a first generation unit, a second generation unit, and a rendering unit;
[0030] The first generation unit is used to generate a 3DGS digital human based on user images from multiple perspectives;
[0031] The second generation unit is used to determine the level of each Gaussian unit in the 3DGS digital human based on the mean square error between the rendered image and the user image, and to generate a hierarchical level of detail (LOD) digital human composed of multiple levels of Gaussian units; the rendered image is generated by rendering the 3DGS digital human.
[0032] The rendering unit is used to render the LOD digital human based on the distance between the user's real-time viewpoint position and each Gaussian unit position, and generate the rendered digital human.
[0033] Thirdly: This application provides a terminal device, which includes a processor and a memory;
[0034] The memory is used to store program code and transmit the program code to the processor;
[0035] The processor is used to execute the steps of a digital human rendering method as described above, according to the instructions in the program code.
[0036] Fourth aspect: Embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a digital human rendering method as described in the first aspect above.
[0037] Fifth aspect: This application provides a computer program product, which, when run on a computer, executes the steps of a digital human rendering method as described above.
[0038] Sixth aspect: This application provides a chip including a processor coupled to a memory for executing computer programs or instructions stored in the memory, such that the chip implements a digital human rendering method as described in the first aspect above. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 A scene illustration showing the application of a digital human rendering method provided in this application embodiment;
[0041] Figure 2 A flowchart illustrating a digital human rendering method provided in this application embodiment;
[0042] Figure 3 A flowchart illustrating another digital human rendering method provided in this application embodiment;
[0043] Figure 4 This is a schematic diagram of the structure of a digital human rendering device provided in an embodiment of this application;
[0044] Figure 5 This is a schematic diagram of the hardware structure of a terminal device provided in an embodiment of this application. Detailed Implementation
[0045] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0046] Digital human rendering technology aims to transform digital human models into realistic and vivid two-dimensional or three-dimensional visual images using computer graphics algorithms and hardware resources. The core challenge facing digital human rendering technology lies in the mismatch between the limited computing power of terminal devices and the computing resources required for high-quality rendering.
[0047] Currently, traditional polygon mesh rendering technology faces the problem of rigid topological structure when representing details. While the emerging Neural Radiation Field (NeRF) technology has better rendering effects, it has high requirements for computing resources and each frame takes several seconds to render. This not only increases the computing load on terminal devices but also reduces rendering efficiency, making it difficult to meet the needs of large-scale digital human rendering scenarios.
[0048] Based on this, this application provides a digital human rendering method. It generates a 3DGS digital human from user images based on multiple viewpoints; determines the level of each Hexagonal element in the 3DGS digital human based on the mean square error between the rendered image and the user image, and generates a Hierarchical Detail (LOD) digital human composed of multiple levels of Hexagonal elements; the rendered image is generated by rendering the 3DGS digital human; and the LOD digital human is rendered based on the distance between the user's real-time viewpoint position and the positions of each Hexagonal element, generating the rendered digital human. This application, by generating an LOD digital human, can optimize the level of detail of the LOD digital human based on the distance between the user's real-time viewpoint position and the positions of each Hexagonal element, reducing the computational load while ensuring the visual effect of the digital human, and meeting the needs of large-scale digital human rendering scenarios.
[0049] like Figure 1 As shown, the terminal device 101 can interact with the server 102 to implement a digital human rendering method provided in this application. Figure 1 This is a schematic diagram illustrating a scene in which a digital human rendering method provided in this application embodiment can be applied.
[0050] Terminal device 101 can acquire user images from multiple perspectives and send these images to server 102. Server 102 generates a 3DGS digital human based on the user images from multiple perspectives; it determines the level of each Gaussian unit in the 3DGS digital human based on the mean square error between the rendered image and the user images, generating a hierarchical level of detail (LOD) digital human composed of multiple levels of Gaussian units; the rendered image is generated by rendering the 3DGS digital human; based on the distance between the user's real-time viewpoint position and the positions of each Gaussian unit, the LOD digital human is rendered to generate the rendered digital human. Server 102 can send the rendered digital human to terminal device 101 so that terminal device 101 can display the rendered digital human to the user.
[0051] It is understood that in this embodiment, the server 102 can be integrated into the terminal device 101 or deployed independently, and this embodiment does not specifically limit this.
[0052] In one example, the method provided in this application embodiment can be applied to a scenario of large-scale digital human interaction on the same screen. In this scenario, a digital human at a distance can be generated based on a lower rendering layer, and a digital human at a closer distance can be generated based on a higher rendering layer to display more detailed information.
[0053] In another example, the method provided in this application can be applied to industrial simulation and digital twin scenarios. For example, a scenario where a high-quality digital human is displayed on a regular workstation or terminal device, and the digital human interacts collaboratively with a 3D mechanical model.
[0054] In another example, the method provided in this application can be applied to social media and content creation scenarios. For instance, it determines the corresponding rendering layer based on the user's network conditions. That is, the better the network conditions, the higher the rendering layer. Based on this rendering layer, a rendered digital human corresponding to the broadcaster is generated.
[0055] The following is combined with Figure 2 This application introduces a digital human rendering method provided by an embodiment, such as... Figure 2 As shown, this figure is a flowchart of a digital human rendering method provided in an embodiment of this application, including S201-S203.
[0056] S201. Generate a 3DGS digital human based on user images from multiple perspectives.
[0057] Among them, 3D Gaussian Splatting (3DGS) technology is a real-time 3D scene reconstruction and rendering technology based on explicit Gaussian point cloud representation.
[0058] In this embodiment of the application, a multi-camera array can be used to capture user images from multiple perspectives while ensuring full coverage of the user's face.
[0059] In one possible implementation, to ensure the quality of the generated 3DGS digital human, user images from multiple perspectives can be acquired under stable lighting conditions and without interference from moving objects. After obtaining user images from multiple perspectives, blurry and low-quality user images are removed to obtain preprocessed user images from multiple perspectives.
[0060] For example, in this embodiment of the application, the sharpness of the user image can be determined by calculating the Laplacian gradient of the user image; a deep learning algorithm can be used to determine whether the user image contains a face, and if the confidence of detecting a face is less than a threshold, the corresponding user image is determined to be a low-quality image.
[0061] It is understood that the method for filtering blurry or low-quality user images in this application embodiment is not specifically limited, and the above is only an example.
[0062] After obtaining preprocessed user images from multiple perspectives, semantic segmentation can be performed on the preprocessed user images from multiple perspectives based on the face parsing algorithm to obtain the corresponding segmentation results.
[0063] The segmentation result can be used to indicate the facial features corresponding to each pixel in the user image. Based on the segmentation result, feature matching can be performed on user images from multiple perspectives to obtain the feature-matched user image. Incremental reconstruction of the feature-matched user image can generate sparse point clouds and camera pose.
[0064] Based on user images and camera poses from multiple perspectives, sparse point clouds can be converted into Gaussian elements, and 3DGS digital humans can be generated based on these Gaussian elements.
[0065] Each Gaussian element contains data such as Gaussian element location, covariance matrix (including scaling and rotation parameters), opacity (α), and spherical harmonic coefficients (SH).
[0066] In one example, to improve the quality of the 3DGS digital human, after generating the sparse point cloud and camera pose, global bundle adjustment (BA) can be used to optimize the camera pose and improve the geometric consistency of the sparse point cloud. Based on this, using user images from multiple viewpoints and the optimized camera pose, the sparse point cloud is converted into Gaussian units, and a 3DGS digital human is generated based on these Gaussian units.
[0067] S202. Based on the mean square error between the rendered image and the user image, determine the level of each Gaussian unit in the 3DGS digital human, and generate a hierarchical detailed digital human composed of multiple levels of Gaussian units.
[0068] The rendered image is generated by rendering the 3DGS digital human.
[0069] The 3DGS digital human is composed of a set of Gaussian elements. In this embodiment, the hierarchical level of the Gaussian elements that make up the 3DGS digital human is determined by classifying them.
[0070] In one possible implementation, each level corresponds to a level label. During the classification process, a level label can be assigned to each Gaussian element, such as L0, L1, ..., Ln-1.
[0071] For example, lower-level Gaussian units may include basic contour information of the digital human, while higher-level Gaussian units may include more detailed information about the digital human.
[0072] In one possible implementation, to determine the hierarchy of Gaussian elements, the gradient of each Gaussian element in the 3DGS digital human can be determined based on the mean square error between the rendered image and the user image; based on the gradient, the hierarchy of each Gaussian element in the 3DGS digital human is determined, generating a hierarchical detail (LOD) digital human composed of multiple levels of Gaussian elements.
[0073] For example, in this embodiment of the application, the mean squared error between the rendered image and the user image can be used as the loss function, and the gradient of each Gaussian unit can be determined through gradient backpropagation. Based on this, the Gaussian units can be sorted according to the magnitude of the gradient to obtain a sorting result; according to the sorting result, the level of each Gaussian unit in the 3DGS digital human can be determined to generate a Level of Dimension (LOD) digital human composed of multiple levels of Gaussian units.
[0074] In the process of determining the hierarchy of each Gaussian element in the 3DGS digital human based on the sorting results and generating a LOD digital human composed of multiple levels of Gaussian elements, the number of Gaussian elements contained in each level can be determined based on the total number of Gaussian elements and the attenuation coefficient; based on the sorting results and the number of Gaussian elements contained in each level, the hierarchy of each Gaussian element in the 3DGS digital human is determined, and a LOD digital human composed of multiple levels of Gaussian elements is generated.
[0075] The smaller the gradient, the smaller the level label corresponding to the Gaussian element. That is, the gradient of the Gaussian element with level label L0 is less than the gradient L1 of the Gaussian element with level label L1, the gradient of the Gaussian element with level label L1 is less than the gradient of the Gaussian element with level label L2, and so on.
[0076] In this embodiment of the application, the number of Gaussian elements with hierarchical label L0 can be expressed by equation (1) as follows:
[0077] (1)
[0078] Where m0 represents the number of Gaussian cells with the hierarchical label L0; M represents the total number of Gaussian cells; α represents the attenuation coefficient; i represents the hierarchical level; and n-1 represents the total number of hierarchical levels.
[0079] After determining m0, the number of Gaussian elements contained in each level can be determined based on m0 and the attenuation coefficient, as shown in equation (2):
[0080] (2)
[0081] Based on the sorting results and the number of Gaussian elements contained in each level, the level of each Gaussian element in the 3DGS digital human can be determined, and a LOD digital human composed of multiple levels of Gaussian elements can be generated. The LOD digital human corresponding to each level can be represented by equation (3) as follows:
[0082] (3)
[0083] Among them, LOD m This represents the m-th level of the 3DGS digital human.
[0084] S203. Based on the distance between the user's real-time viewpoint position and each Gaussian unit position, render the LOD digital human to generate the rendered digital human.
[0085] In this embodiment, by acquiring the user's viewpoint position in real time, the rendering layer corresponding to each Hexagonal pixel can be determined based on the distance between the user's real-time viewpoint position and the positions of each Hexagonal pixel. Based on this, by rendering each Hexagonal pixel to its corresponding rendering layer, a rendered digital human can be generated. The viewpoint position can be determined based on parameters such as camera position, camera orientation, and frustum range.
[0086] For example, based on the ratio between the distance between the user's real-time viewpoint position and each Gaussian primitive position and the reference distance threshold, multiple ratios can be determined; by rounding down the logarithm of each ratio with the attenuation coefficient as the base, the rendering level corresponding to each Gaussian primitive can be obtained, as shown in Equation (4) below:
[0087] (4)
[0088] Where k represents the rendering level; d represents the distance between the viewpoint and each high-order primitive; α represents the attenuation coefficient, used to control the LOD switching sensitivity; and β represents the baseline distance threshold.
[0089] In this embodiment, the greater the distance between the viewpoint and each high-order primitive position, the lower the corresponding rendering level; the closer the distance between the viewpoint and each high-order primitive position, the higher the corresponding rendering level, so as to display more detailed information. This can effectively reduce the requirements on the computing power of the terminal device, reduce the computing load while ensuring the visual effect of the digital human, and meet the needs of large-scale digital human rendering scenarios.
[0090] In one possible implementation, to improve the display effect of the rendered digital human, this embodiment of the application renders the LOD digital human based on the distance between the user's real-time viewpoint position and each Gaussian pixel position. During the process of generating the rendered digital human, the LOD digital human can be divided into multiple regions. Based on the distance between the user's real-time viewpoint position and the multiple Gaussian pixel positions included in each region, the dominant viewing distance of each region's Gaussian pixels is determined. Based on the dominant viewing distance of each region's Gaussian pixels, the LOD digital human is rendered to generate the rendered digital human.
[0091] For example, in the embodiments of this application, the LOD digital human can be evenly divided into a 16×16 pixel tile structure region. By independently analyzing the spatial relationship between the multiple Gaussian pixel positions contained in each region and the camera position, the dominant viewing distance of the Gaussian pixels corresponding to each region can be determined.
[0092] In one example, the dominant line-of-sight distance is the weighted average of the distances from multiple Gaussian cell locations within the region to the viewpoint location.
[0093] In another example, the dominant viewing distance is the distance from the viewpoint to the location of one of the multiple Gaussian elements contained within the region. This Gaussian element can be selected according to actual needs.
[0094] After determining the dominant view distance of the Gaussian pixels for each region, the rendering level for each region can be determined based on the dominant view distance of the Gaussian pixels for each region; based on the rendering level for each region, the LOD digital human can be rendered to generate the rendered digital human.
[0095] This process fully utilizes the parallel computing power of the GPU, enabling simultaneous rendering of LOD digital humans at different rendering levels, thus improving rendering efficiency.
[0096] In one possible implementation, since different regions may correspond to different rendering levels, there may be layer switching traces between adjacent regions due to the different rendering levels. Based on this, in this embodiment, for the transition regions between adjacent regions with different rendering levels, an alpha blending algorithm or a weighted superposition algorithm based on view distance difference is used to perform pixel-level fusion of the rendering results of different rendering levels to obtain the fused digital human.
[0097] For example, in the process of pixel-level fusion of rendering results based on weighted superposition algorithm, a smooth interpolation function (such as linear interpolation or smooth step function) can be used for calculation to ensure that the transition area is as natural as possible in terms of geometric edges and color presentation, so that the fused digital human presents a visually coherent display effect.
[0098] To facilitate understanding, the following will be combined with... Figure 3 The methods provided in the embodiments of this application will be described in general. Figure 3 A flowchart of another digital human rendering method provided in an embodiment of this application.
[0099] A multi-camera array is used to capture user images from multiple perspectives, covering as much of the user's face as possible. A 3D GDS digital human is then generated based on these images. Rendering the 3D GDS digital human yields a rendered image. The mean squared error between the rendered image and the user image is used as a loss function, and the gradient of each Gaussian primitive (GRP) is determined through gradient backpropagation. Based on the magnitude of these gradients, the GRPs are sorted. A Level of Distance (LOD) model is then constructed using these subsequent GRPs to generate the LOD digital human.
[0100] To generate a Level-of-Depth (LOD) digitized human, the rendering level corresponding to each hexagonal pixel in the LOD digitized human is determined based on the distance between the user's real-time viewpoint position and the distance between hexagonal pixels. For example, the farther the hexagonal pixel is, the lower the rendering level it corresponds to, and the closer the hexagonal pixel is, the higher the rendering level it corresponds to, in order to display more detail. Based on the rendering level corresponding to each hexagonal pixel, the LOD digitized human is rendered, generating the rendered digitized human.
[0101] In summary, the embodiments of this application, by generating LOD (Level of Detail) digital humans, can optimize the level of detail of the LOD digital human based on the distance between the user's real-time viewpoint position and each Gaussian unit position, thereby reducing the computational load while ensuring the visual effect of the digital human. While maintaining the clarity of the digital human's main features, the rendering frame rate can be increased to over 60 FPS, which can meet the needs of VR / AR interaction.
[0102] Meanwhile, in the case of rendering multiple people in the same scene, the method provided by the embodiments of this application only requires each user to independently maintain their rendered digital human, and the total computing load increases linearly with the number of people rather than exponentially.
[0103] This application provides a digital human rendering device, see [link]. Figure 4 The figure is a schematic diagram of the structure of a digital human rendering device provided in an embodiment of this application. Its specific implementation method is consistent with the implementation method and the technical effect achieved in the embodiments of the above method, and some contents will not be repeated.
[0104] This application provides a digital human rendering device 4100, including:
[0105] The first generation unit 4101, the second generation unit 4102, and the rendering unit 4103;
[0106] The first generation unit 4101 is used to generate a 3DGS digital human based on user images from multiple perspectives;
[0107] The second generation unit 4102 is used to determine the level of each Gaussian unit in the 3DGS digital human based on the mean square error between the rendered image and the user image, and generate a hierarchical detail (LOD) digital human composed of multiple levels of Gaussian units; the rendered image is generated by rendering the 3DGS digital human.
[0108] The rendering unit 4103 is used to render the LOD digital human based on the distance between the user's real-time viewpoint position and each Gaussian unit position, and generate the rendered digital human.
[0109] In one possible implementation, the second generating unit is specifically used for:
[0110] The gradient of each Gaussian unit in the 3DGS digital human is determined based on the mean square error between the rendered image and the user image.
[0111] Based on the gradient, the level of each Gaussian unit in the 3DGS digital human is determined, and a hierarchical LOD digital human composed of multi-level Gaussian units is generated.
[0112] In one possible implementation, the second generating unit is specifically used for:
[0113] The Gaussian units are sorted according to the magnitude of the gradient to obtain the sorting result;
[0114] Based on the sorting results, the hierarchy of each Gaussian unit in the 3DGS digital human is determined, and a hierarchical LOD digital human composed of multiple levels of Gaussian units is generated.
[0115] In one possible implementation, the second generating unit is specifically used for:
[0116] Based on the total number of Gaussian elements and the attenuation coefficient, determine the number of Gaussian elements contained in each level;
[0117] Based on the sorting results and the number of Gaussian elements contained in each level, the level of each Gaussian element in the 3DGS digital human is determined, and a hierarchical LOD digital human composed of multiple levels of Gaussian elements is generated.
[0118] In one possible implementation, the rendering unit is specifically used for:
[0119] Based on the distance between the user's real-time viewpoint position and the positions of each high-order primitive, the rendering level corresponding to each high-order primitive is determined.
[0120] Each Gaussian primitive is rendered to its corresponding rendering layer to generate the rendered digital human.
[0121] In one possible implementation, the rendering unit is specifically used for:
[0122] Based on the ratio between the distance between the user's real-time viewpoint position and the position of each Gaussian element and the baseline distance threshold, the distance ratio corresponding to each Gaussian element is determined.
[0123] Using the attenuation coefficient as the base, the logarithm of each distance ratio is rounded down to obtain the rendering level corresponding to each Gaussian unit.
[0124] In one possible implementation, the rendering unit is specifically used for:
[0125] The LOD digital human is divided into multiple regions. Based on the distance between the user's real-time viewpoint position and the multiple Gaussian positions included in each region, the dominant viewing distance of each region is determined.
[0126] Based on the dominant viewing distance of Gaussian elements corresponding to each region, the LOD digital human is rendered to generate the rendered digital human.
[0127] In one possible implementation, the device further includes: a fusion unit;
[0128] The fusion unit is specifically used to perform pixel-level fusion of the rendering results of the transition areas between adjacent areas in the rendered digital human, so as to obtain the fused digital human.
[0129] In summary, the embodiments of this application, by generating LOD digital humans, can optimize the level of detail of LOD digital humans based on the distance between the user's real-time viewpoint position and each Gaussian meta-position. While ensuring the visual effect of the digital human, it reduces the computational load and can meet the needs of large-scale digital human rendering scenarios.
[0130] This application provides a terminal device, such as... Figure 5 As shown in the figure, this figure is a schematic diagram of the hardware structure of a terminal device provided in an embodiment of this application. The terminal device 510 includes: a memory 511 and a processor 512.
[0131] The memory 511 is used to store program code and transmit the program code to the processor 512; the processor 512 is used to execute the steps of a digital human rendering method as described above according to the instructions in the program code.
[0132] This application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a digital human rendering method as described above.
[0133] This application provides a computer program product that, when run on a computer, executes the steps of a digital human rendering method as described above.
[0134] This application provides a chip including a processor coupled to a memory for executing computer programs or instructions stored in the memory, thereby enabling the chip to implement a digital human rendering method as described above.
[0135] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiments. The device and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components indicated as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the solution in this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0136] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A digital human rendering method, characterized in that, include: Generate 3DGS digital humans based on user images from multiple perspectives; Based on the mean square error between the rendered image and the user image, the level of each Gaussian unit in the 3DGS digital human is determined, and a hierarchical level of detail (LOD) digital human composed of multiple levels of Gaussian units is generated; the rendered image is generated by rendering the 3DGS digital human. The LOD digital human is rendered based on the distance between the user's real-time viewpoint position and each Gaussian unit position, generating the rendered digital human. The process of determining the hierarchy of each Gaussian unit in the 3DGS digital human based on the mean square error between the rendered image and the user image, and generating a hierarchical level of detail (LOD) digital human composed of multiple levels of Gaussian units, includes: The gradient of each Gaussian unit in the 3DGS digital human is determined based on the mean square error between the rendered image and the user image. Based on the gradient, the level of each Gaussian unit in the 3DGS digital human is determined, and a hierarchical LOD digital human composed of multi-level Gaussian units is generated.
2. The method according to claim 1, characterized in that, The step of determining the hierarchy of each Gaussian unit in the 3DGS digital human based on the gradient, and generating a hierarchical level of detail (LOD) digital human composed of multiple levels of Gaussian units, includes: The Gaussian units are sorted according to the magnitude of the gradient to obtain the sorting result; Based on the sorting results, the hierarchy of each Gaussian unit in the 3DGS digital human is determined, and a hierarchical LOD digital human composed of multiple levels of Gaussian units is generated.
3. The method according to claim 2, characterized in that, The step of determining the hierarchy of each Gaussian unit in the 3DGS digital human based on the sorting result, and generating a hierarchical level-of-detail (LOD) digital human composed of multiple levels of Gaussian units, includes: Based on the total number of Gaussian elements and the attenuation coefficient, determine the number of Gaussian elements contained in each level; Based on the sorting results and the number of Gaussian elements contained in each level, the level of each Gaussian element in the 3DGS digital human is determined, and a hierarchical LOD digital human composed of multiple levels of Gaussian elements is generated.
4. The method according to claim 1, characterized in that, The rendering of the LOD digital human based on the distance between the user's real-time viewpoint position and each Gaussian unit position, to generate the rendered digital human, includes: Based on the distance between the user's real-time viewpoint position and the positions of each high-order primitive, the rendering level corresponding to each high-order primitive is determined. Each Gaussian primitive is rendered to its corresponding rendering layer to generate the rendered digital human.
5. The method according to claim 4, characterized in that, The process of determining the rendering level corresponding to each hexagonal element based on the distance between the user's real-time viewpoint position and the positions of each hexagonal element includes: Based on the ratio between the distance between the user's real-time viewpoint position and the position of each Gaussian element and the baseline distance threshold, the distance ratio corresponding to each Gaussian element is determined. Using the attenuation coefficient as the base, the logarithm of each distance ratio is rounded down to obtain the rendering level corresponding to each Gaussian unit.
6. The method according to claim 1, characterized in that, The rendering of the LOD digital human based on the distance between the user's real-time viewpoint position and each Gaussian unit position, to generate the rendered digital human, includes: The LOD digital human is divided into multiple regions. Based on the distance between the user's real-time viewpoint position and the multiple Gaussian positions included in each region, the dominant viewing distance of each region is determined. Based on the dominant viewing distance of Gaussian elements corresponding to each region, the LOD digital human is rendered to generate the rendered digital human.
7. The method according to claim 6, characterized in that, After rendering the LOD digital human based on the dominant view distance of Gaussian elements corresponding to each region, and generating the rendered digital human, the process further includes: For the transition areas between adjacent areas in the rendered digital human, the rendering results of the transition areas are fused at the pixel level to obtain the fused digital human.
8. A digital human rendering device, characterized in that, include: The first generation unit is used to generate 3DGS digital humans based on user images from multiple perspectives; The second generation unit is used to determine the level of each Gaussian unit in the 3DGS digital human based on the mean square error between the rendered image and the user image, and to generate a hierarchical level of detail (LOD) digital human composed of multiple levels of Gaussian units; the rendered image is generated by rendering the 3DGS digital human. The rendering unit is used to render the LOD digital human based on the distance between the user's real-time viewpoint position and each Gaussian unit position, and generate the rendered digital human. The second generation unit is specifically used for: The gradient of each Gaussian unit in the 3DGS digital human is determined based on the mean square error between the rendered image and the user image. Based on the gradient, the level of each Gaussian unit in the 3DGS digital human is determined, and a hierarchical LOD digital human composed of multi-level Gaussian units is generated.
9. A terminal device, characterized in that, The terminal device includes: a processor and a memory; The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the steps of a digital human rendering method as described in any one of claims 1-7 according to instructions in the program code.
Citation Information
Patent Citations
Three-dimensional tree modeling method and device based on 3DGS
CN120689517A