Digital human rendering method and related device

By generating hierarchical LOD digital humans with multiple layers of high-level primitives, the rendering detail levels are optimized, solving the problem of high computational resources in digital human rendering technology and achieving efficient rendering effects.

CN120912792AActive Publication Date: 2025-11-07HANGZHOU QIUGUOJIHUA TECHNOLOGY CO LTD
View PDF 14 Cites 0 Cited by

Patent Information

Application Number
CN202511445282.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-11-07
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

Existing digital human rendering technologies have high computational resource requirements and long rendering time per frame, making it difficult to meet the needs of large-scale digital human rendering scenarios.

Method used

By generating hierarchical detail LOD digital humans based on multi-level Gaussian primitives, the rendering detail hierarchy is optimized by utilizing the distance between the user's real-time viewpoint position and the Gaussian primitive position, thereby reducing the computational load.

Benefits of technology

While ensuring the visual effects of digital humans, the computational load has been reduced and the rendering efficiency has been improved to meet the needs of large-scale digital human rendering scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120912792A_ABST
    Figure CN120912792A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a digital human rendering method and a related device, and the method comprises the steps: generating a 3DGS digital human through a user image based on a plurality of visual angles; based on the mean square error between the rendered image and the user image, determining the level of each Gaussian primitive in the 3DGS digital human, and generating an LOD digital human composed of multiple levels of Gaussian primitives; the rendered image is generated by rendering the 3DGS digital human; and rendering the LOD digital human based on the distance between the real-time viewpoint position of the user and each Gaussian primitive position, and generating a rendered digital human. According to the embodiment of the invention, by generating the LOD digital human, the detail level of the LOD digital human can be optimized based on the distance between the real-time viewpoint position of the user and each Gaussian primitive position, the calculation load is reduced on the premise of ensuring the visual effect of the digital human, and the demand of rendering a scene by a large-scale digital human can be met.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a digital human rendering method and related device. BACKGROUND

[0002] The digital human rendering technology aims to convert a digital human model into a realistic and vivid two-dimensional or three-dimensional visual image by using computer graphics algorithms and hardware resources.

[0003] At present, the digital human rendering can be realized based on the neural radiation field (NeRF) technology, however, this method has high demand for computing resources, and a single frame rendering needs to consume several seconds of time, which not only increases the computing load of the terminal device, but also reduces the rendering efficiency, and thus it is difficult to meet the demand of large-scale digital human rendering scene.

[0004] Therefore, there is an urgent need for a solution to solve the above technical problems. SUMMARY

[0005] In view of the above problems, the embodiments of the present application provide a digital human rendering method and related device, which aims to reduce the computing load and meet the demand of large-scale digital human rendering scene while ensuring the visual effect of the digital human.

[0006] The embodiments of the present application disclose the following technical solutions: In a first aspect, the embodiments of the present application provide a digital human rendering method, which generates a 3DGS digital human based on user images of multiple perspectives; determines the level of each Gaussian primitive in the 3DGS digital human based on the mean square error between a rendered image and the user images, generates a hierarchical detail LOD digital human composed of multiple levels of Gaussian primitives; the rendered image is generated by rendering the 3DGS digital human; and renders the LOD digital human based on the distance between the real-time viewpoint position of the user and the position of each Gaussian primitive, to generate a rendered digital human.

[0007] The embodiments of the present application can optimize the detail level of the LOD digital human based on the distance between the real-time viewpoint position of the user and the position of each Gaussian primitive by generating the LOD digital human, reduce the computing load while ensuring the visual effect of the digital human, and meet the demand of large-scale digital human rendering scene.

[0008] In a possible implementation manner, the determining the level of each Gaussian primitive in the 3DGS digital human based on the mean square error between the rendered image and the user images, and generating the hierarchical detail LOD digital human composed of multiple levels of Gaussian primitives comprises: determine a gradient of each Gaussian cell in the 3D GS digital human based on a mean square error between the rendered image and the user image; determine a level of each Gaussian cell in the 3D GS digital human based on the gradient, and generate a hierarchical LOD digital human composed of Gaussian cells of multiple levels.

[0009] In the embodiments of the present application, the gradient of the Gaussian cell can be used to determine the level of each Gaussian cell in the 3D GS digital human, and thus to automatically generate the LOD digital human. This can improve the generation efficiency of the rendered digital human while ensuring the visual effect of the digital human and reducing the computational load.

[0010] In a possible implementation, the determining the level of each Gaussian cell in the 3D GS digital human based on the gradient, and generating the hierarchical LOD digital human composed of Gaussian cells of multiple levels includes: sorting the Gaussian cells according to the size of the gradient to obtain a sorting result; determining the level of each Gaussian cell in the 3D GS digital human based on the sorting result, and generating the hierarchical LOD digital human composed of Gaussian cells of multiple levels.

[0011] In the embodiments of the present application, the size of the gradient can indicate the level of the Gaussian cell. The smaller the level, the less the detail information of the digital human, and the larger the level, the more the detail information of the digital human. Based on this, the embodiments of the present application can automatically generate the LOD digital human based on the gradient, which can improve the generation efficiency of the rendered digital human while ensuring the visual effect of the digital human and reducing the computational load.

[0012] In a possible implementation, the determining the level of each Gaussian cell in the 3D GS digital human based on the sorting result, and generating the hierarchical LOD digital human composed of Gaussian cells of multiple levels includes: determining the number of Gaussian cells contained in each level based on the total number of Gaussian cells and the attenuation coefficient; determining the level of each Gaussian cell in the 3D GS digital human based on the sorting result and the number of Gaussian cells contained in each level, and generating the hierarchical LOD digital human composed of Gaussian cells of multiple levels.

[0013] In the embodiments of the present application, the total number of Gaussian cells, the attenuation coefficient, and the sorting result can be used to automatically generate the LOD digital human, which improves the generation efficiency of the LOD digital human, and thus improves the generation efficiency of the rendered digital human.

[0014] In a possible implementation, the rendering the LOD digital human based on the distance between the real-time viewpoint position of the user and the position of each Gaussian cell, and generating the rendered digital human includes: The distance between the real-time viewpoint position of the user and each Gaussian cell position is determined to determine a rendering level corresponding to each Gaussian cell.

[0015] In the embodiments of the present application, different Gaussian cells can correspond to different rendering degrees. For example, a Gaussian cell close to the real-time viewpoint position of the user can correspond to a higher rendering degree, and a Gaussian cell far from the real-time viewpoint position of the user can correspond to a lower rendering degree, so that the visual effect of the digital person can be ensured, the computing load is reduced, and the demand of large-scale digital person rendering scene can be met.

[0016] In a possible implementation, the distance between the real-time viewpoint position of the user and each Gaussian cell position is determined to determine a rendering level corresponding to each Gaussian cell, including: The distance between the real-time viewpoint position of the user and each Gaussian cell position is determined to determine a distance ratio corresponding to each Gaussian cell based on the distance ratio between the reference distance threshold, and the logarithm of each distance ratio is taken with the attenuation coefficient as the base to obtain a rendering level corresponding to each Gaussian cell, so that each Gaussian cell can be rendered to the corresponding rendering level to generate a rendered digital person, the visual effect of the digital person can be ensured, the computing load is reduced, and the demand of large-scale digital person rendering scene can be met.

[0017] In a possible implementation, the distance between the real-time viewpoint position of the user and each Gaussian cell position is determined to render the LOD digital person to generate a rendered digital person, including: The LOD digital person is divided into multiple regions, the distance between the real-time viewpoint position of the user and the multiple Gaussian cell positions included in each region is determined to determine a Gaussian cell dominant view distance corresponding to each region, and the LOD digital person is rendered based on the Gaussian cell dominant view distance corresponding to each region to generate a rendered digital person.

[0018] In the embodiments of the present application, the LOD digital person is divided into multiple regions, so that multiple regions of the LOD digital person can be rendered at the same time, and the rendering efficiency can be effectively improved.

[0019] In a possible implementation, after the distance between the real-time viewpoint position of the user and each Gaussian cell position is determined to render the LOD digital person based on the Gaussian cell dominant view distance corresponding to each region to generate a rendered digital person, the method further includes: performing pixel-level fusion on the rendering result of a transition region between adjacent regions in the rendered digital person to obtain a fused digital person.

[0020] In the embodiments of the present application, pixel-level fusion is performed on the rendering result of the transition area, so that the transition area can be as natural as possible in terms of geometric edges and color presentation, and the fused digital person can present a visually coherent display effect.

[0021] In a second aspect, the embodiments of the present application provide a digital person rendering device, comprising: a first generation unit, a second generation unit, and a rendering unit; The first generation unit is configured to generate a 3D GS digital person based on user images of multiple perspectives. The second generation unit is configured to determine levels of each Gaussian primitive in the 3D GS digital person based on a mean square error between a rendering image and the user images, and generate a hierarchical detail LOD digital person composed of multiple-level Gaussian primitives; the rendering image is generated by rendering the 3D GS digital person. The rendering unit is configured to render the LOD digital person based on distances between a real-time viewpoint position of a user and positions of each Gaussian primitive, and generate a rendered digital person.

[0022] In a third aspect, the embodiments of the present application provide a terminal device, comprising a processor and a memory. The memory is configured to store program codes and transmit the program codes to the processor. The processor is configured to execute steps of a digital person rendering method according to instructions in the program codes.

[0023] In a fourth aspect, the embodiments of the present application provide a computer readable storage medium, which stores a computer program. When the computer program is executed by a processor, steps of a digital person rendering method according to the first aspect are implemented.

[0024] In a fifth aspect, the embodiments of the present application provide a computer program product. When the computer program product is run on a computer, the computer executes steps of a digital person rendering method.

[0025] In a sixth aspect, the embodiments of the present application provide a chip, comprising a processor and a memory coupled to the processor. The processor is configured to execute a computer program or instructions stored in the memory, so that the chip implements a digital person rendering method according to the first aspect. BRIEF DESCRIPTION OF DRAWINGS

[0026] In order to make the technical solutions in the embodiments of the present application or the prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and other accompanying drawings can be obtained by those skilled in the art without creative labor.

[0027] Figure 1 A digital human rendering method provided by an embodiment of the present application can be applied to a scene schematic diagram in which the method is used. Figure 2 A flowchart of a digital human rendering method provided by an embodiment of the present application. Figure 3 A flowchart of another digital human rendering method provided by an embodiment of the present application. Figure 4 A structural schematic diagram of a digital human rendering device provided by an embodiment of the present application. Figure 5 A hardware structural schematic diagram of a terminal device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0028] In order for those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0029] Digital human rendering technology aims to use computer graphics algorithms and hardware resources to convert digital human models into realistic and vivid two-dimensional or three-dimensional visual images. The core contradiction faced by digital human rendering technology is that the limited computing power of terminal devices does not match the computing resources required for high-quality rendering.

[0030] At present, traditional polygon mesh rendering technology faces the problem of rigid topology when expressing details, while the emerging NeRF technology has good rendering effect, but it requires high computing resources, and single-frame rendering takes several seconds, which not only increases the computing load of terminal devices, but also reduces the rendering efficiency, and thus it is difficult to meet the needs of large-scale digital human rendering scenarios.

[0031] Based on this, the embodiment of the application provides a digital person rendering method, which generates a 3D GS digital person based on user images of multiple perspectives; determines the level of each Gaussian cell in the 3D GS digital person based on the mean square error between a rendered image and the user images, generates a level of detail (LOD) digital person composed of multiple-level Gaussian cells; the rendered image is generated by rendering the 3D GS digital person; and renders the LOD digital person based on the distance between the real-time viewpoint position of the user and the position of each Gaussian cell, to generate a rendered digital person. The embodiment of the application can optimize the detail level of the LOD digital person based on the distance between the real-time viewpoint position of the user and the position of each Gaussian cell by generating the LOD digital person, reduces the calculation load under the premise of guaranteeing the visual effect of the digital person, and can meet the demand of large-scale digital person rendering scenarios.

[0032] As shown in the figure, the terminal device 101 can interact with the server 102 to implement a digital person rendering method provided by the embodiment of the application. Figure 1 The digital person rendering method provided by the embodiment of the application can be applied to the scene diagram shown in the figure. Figure 1 The digital person rendering method provided by the embodiment of the application can be applied to the scene diagram shown in the figure.

[0033] The terminal device 101 can collect user images of multiple perspectives and send the user images of multiple perspectives to the server 102. The server 102 generates a 3D GS digital person based on the user images of multiple perspectives; determines the level of each Gaussian cell in the 3D GS digital person based on the mean square error between a rendered image and the user images, generates a level of detail (LOD) digital person composed of multiple-level Gaussian cells; the rendered image is generated by rendering the 3D GS digital person; and renders the LOD digital person based on the distance between the real-time viewpoint position of the user and the position of each Gaussian cell, to generate a rendered digital person. The server 102 can send the rendered digital person to the terminal device 101, so that the terminal device 101 shows the rendered digital person to the user.

[0034] It can be understood that the server 102 in the embodiment of the application can be integrated in the terminal device 101 or can be independently deployed, and the embodiment of the application does not make specific limitation on this.

[0035] In an example, the method provided by the embodiment of the application can be applied to the scene of large-scale digital person on-screen interaction. In this scene, a digital person at a long distance can be generated based on a lower rendering level; and a digital person at a short distance can be generated based on a higher rendering level, to show more detailed information.

[0036] In another example, the method provided by the embodiments of the present application can be applied to the scene of industrial simulation and digital twin. For example, a high-quality digital human is displayed on a common workstation or terminal device, and the digital human interacts with a 3D mechanical model.

[0037] In another example, the method provided by the embodiments of the present application can be applied to the scene of social media and content creation. For example, according to the network condition of the user, a corresponding rendering level to be rendered is determined. That is, the better the network condition, the higher the rendering level to be rendered. Based on the rendering level to be rendered, a rendered digital human corresponding to the anchor is generated.

[0038] The following describes the method for rendering a digital human provided by the embodiments of the present application in combination with Figure 2 The method for rendering a digital human provided by the embodiments of the present application is described below. Figure 2 As shown in the figure, the figure is a flowchart of the method for rendering a digital human provided by the embodiments of the present application, including S201-S203.

[0039] S201, generating a 3DGS digital human based on user images of multiple perspectives.

[0040] Among them, the three-dimensional Gaussian splatting (3D Gaussian Splatting, 3DGS) technology is a real-time three-dimensional scene reconstruction and rendering technology based on an explicit Gaussian point cloud representation.

[0041] In the embodiments of the present application, a multi-camera array can be used to collect user images of multiple perspectives while ensuring a full range of perspectives covering the user's face.

[0042] In a possible implementation, to ensure the quality of the generated 3DGS digital human, the user images of multiple perspectives can be collected under the condition of stable lighting and avoiding interference of dynamic objects. After obtaining the user images of multiple perspectives, the user images that are blurred or of low quality are removed to obtain preprocessed user images of multiple perspectives.

[0043] For example, in the embodiments of the present application, the sharpness of the user image can be determined by calculating the Laplacian gradient of the user image, and whether the user image includes a face can be determined by a deep learning algorithm. In a case where the confidence of detecting a face is less than a threshold, the corresponding user image is determined as a low-quality image.

[0044] It can be understood that the method for screening blurred and low-quality user images in the embodiments of the present application is not specifically limited, and the above is only an example.

[0045] After obtaining the preprocessed user images of multiple perspectives, the preprocessed user images of multiple perspectives can be subjected to semantic segmentation based on a face parsing algorithm to obtain a corresponding segmentation result.

[0046] The segmentation result can be used to indicate the face part corresponding to each pixel point in the user image. Feature matching can be performed on the user images of multiple views based on the segmentation result, and the user images after feature matching can be obtained. Incremental reconstruction can be performed on the user images after feature matching, and sparse point cloud and camera pose can be generated.

[0047] Based on the user images of multiple views and the camera pose, the sparse point cloud can be converted into Gaussian primitives, and a 3D GS digital human can be generated based on the Gaussian primitives.

[0048] Each Gaussian primitive contains data such as Gaussian primitive position, covariance matrix (including scaling parameter and rotation parameter), opacity (a), and spherical harmonic coefficient (SH).

[0049] In an example, to improve the quality of the 3D GS digital human, after generating the sparse point cloud and the camera pose, the camera pose can be optimized by global bundle adjustment (BA), and the geometric consistency of the sparse point cloud can be improved. On this basis, the sparse point cloud is converted into Gaussian primitives based on the user images of multiple views and the optimized camera pose, and a 3D GS digital human is generated based on the Gaussian primitives.

[0050] S202, based on the mean square error between the rendered image and the user image, determine the level of each Gaussian primitive in the 3D GS digital human, and generate a hierarchical detailed digital human composed of multiple levels of Gaussian primitives.

[0051] The rendered image is generated by rendering the 3D GS digital human.

[0052] The 3D GS digital human is composed of a group of Gaussian primitives. In the embodiments of the present application, the levels of the Gaussian primitives are determined by classifying the Gaussian primitives that make up the 3D GS digital human.

[0053] In a possible implementation, each level corresponds to a level label. In the classification process, each Gaussian primitive can be assigned a level label, such as L0, L1, …, Ln-1.

[0054] For example, low-level Gaussian primitives can include basic contour information of the digital human, and high-level Gaussian primitives can include more detailed information of the digital human.

[0055] In a possible implementation, to determine the level of the Gaussian primitive, the gradient of each Gaussian primitive in the 3D GS digital human can be determined based on the mean square error between the rendered image and the user image; based on the gradient, the level of each Gaussian primitive in the 3D GS digital human is determined, and a hierarchical detailed (LOD) digital human composed of multiple levels of Gaussian primitives is generated.

[0056] For example, in the embodiments of the present application, the mean square error between the rendered image and the user image can be taken as a loss function (Loss), and the gradient of each Gaussian cell can be determined by gradient back propagation. On this basis, the Gaussian cells can be sorted based on the size of the gradient to obtain a sorting result; and the level of each Gaussian cell in the 3D GS digital person can be determined according to the sorting result, and the LOD digital person composed of multiple levels of Gaussian cells can be generated.

[0057] In the process of determining the level of each Gaussian cell in the 3D GS digital person according to the sorting result, and generating the LOD digital person composed of multiple levels of Gaussian cells, the number of Gaussian cells contained in each level can be determined based on the total number of Gaussian cells and the attenuation coefficient; and the level of each Gaussian cell in the 3D GS digital person can be determined according to the sorting result and the number of Gaussian cells contained in each level, and the LOD digital person composed of multiple levels of Gaussian cells can be generated.

[0058] For example, the smaller the gradient, the smaller the level label of the corresponding Gaussian cell. For example, the gradient of the Gaussian cell with a level label of L0 is smaller than the gradient L1 of the Gaussian cell with a level label of L1, the gradient of the Gaussian cell with a level label of L1 is smaller than the gradient of the Gaussian cell with a level label of L2, and so on.

[0059] In the embodiments of the present application, the number of Gaussian cells with a level label of L0 can be represented by formula (1) as follows: (1) Wherein, m0 represents the number of Gaussian cells with a level label of L0; M represents the total number of Gaussian cells; a represents the attenuation coefficient; i represents the level; and n-1 represents the total number of levels.

[0060] After determining m0, the number of Gaussian cells contained in each level can be determined based on m0 and the attenuation coefficient, which can be represented by formula (2) as follows: (2) According to the sorting result and the number of Gaussian cells contained in each level, the level of each Gaussian cell in the 3D GS digital person can be determined, and the LOD digital person composed of multiple levels of Gaussian cells can be generated. The LOD digital person corresponding to each level can be represented by formula (3) as follows: (3) Wherein, LOD m represents the 3D GS digital person of the mth level.

[0061] S203, based on the distance between the real-time viewpoint position of the user and the position of each Gaussian cell, the LOD digital person is rendered to generate a rendered digital person.

[0062] In the embodiments of the present application, the viewpoint position of the user is obtained in real time, and the distance between the real-time viewpoint position of the user and each Gaussian cell position is determined to determine the rendering level corresponding to each Gaussian cell. On this basis, the Gaussian cells are rendered to the corresponding rendering level to generate the rendered digital person. The viewpoint position can be determined based on the camera position, camera orientation, and the range of the view frustum.

[0063] For example, based on the distance between the real-time viewpoint position of the user and each Gaussian cell position and the ratio between the reference distance threshold, a plurality of ratios can be determined. The logarithm of each ratio with the attenuation coefficient as the base is taken to obtain the rendering level corresponding to each Gaussian cell, as shown in formula (4) as follows: (4) In formula (4), k represents the rendering level; d represents the distance between the viewpoint position and each Gaussian cell position; a represents the attenuation coefficient, which is used to control the LOD switching sensitivity; and b represents the reference distance threshold.

[0064] In the embodiments of the present application, the farther the distance between the viewpoint position and each Gaussian cell position, the lower the corresponding rendering level; the closer the distance between the viewpoint position and each Gaussian cell position, the higher the corresponding rendering level, so as to display more detailed information, which can effectively reduce the requirement for the computing capacity of the terminal device, reduce the computing load under the premise of ensuring the visual effect of the digital person, and meet the demand of large-scale digital person rendering scene.

[0065] In a possible implementation, to improve the display effect of the rendered digital person, in the process of rendering the LOD digital person based on the distance between the real-time viewpoint position of the user and each Gaussian cell position to generate the rendered digital person, the LOD digital person can be divided into a plurality of regions, the distance between the real-time viewpoint position of the user and each Gaussian cell position included in each region is determined to determine the dominant view distance of the Gaussian cell corresponding to each region, and the LOD digital person is rendered based on the dominant view distance of the Gaussian cell corresponding to each region to generate the rendered digital person.

[0066] For example, in the embodiments of the present application, the LOD digital person can be uniformly divided into regions of a tile structure of 16x16 pixels, and the spatial relationship between the plurality of Gaussian cell positions included in each region and the camera position is independently analyzed to determine the dominant view distance of the Gaussian cell corresponding to each region.

[0067] In an example, the dominant view distance is the weighted average of the distances from the plurality of Gaussian cell positions included in the region to the viewpoint position.

[0068] In another example, the dominant view distance is the distance from a position of a certain Gaussian cell included in the region to the position of the viewpoint. The Gaussian cell can be selected according to actual needs.

[0069] After determining the dominant view distance of the Gaussian cell corresponding to each region, the level to be rendered corresponding to each region can be determined based on the dominant view distance of the Gaussian cell corresponding to each region; and the LOD digital person is rendered based on the level to be rendered corresponding to each region, to generate the rendered digital person.

[0070] This process fully utilizes the parallel computing capability of the GPU, and can simultaneously render the LOD digital person of different levels to be rendered, thereby improving the rendering efficiency.

[0071] In a possible implementation, since different regions can correspond to different levels to be rendered, there can be level switching marks between adjacent regions due to the different levels to be rendered. Based on this, in the embodiments of the present application, for the transition region between adjacent regions with different levels to be rendered, an Alpha Blending algorithm or a weight superposition algorithm based on the distance difference is used to perform pixel-level fusion on the rendering results of different levels to be rendered, to obtain the fused digital person.

[0072] For example, in the process of performing pixel-level fusion on the rendering results based on the weight superposition algorithm, a smooth interpolation function (such as a linear interpolation or a smooth step function) can be used for calculation to ensure that the transition region is as natural as possible in terms of geometric edges and color presentation, so that the fused digital person presents a visually coherent display effect.

[0073] To facilitate understanding, the following will be described in conjunction with Figure 3 The method provided by the embodiments of the present application is introduced as a whole. Figure 3 The flowchart of another digital person rendering method provided by the embodiments of the present application.

[0074] Through the multi-camera array, multi-view user images are collected in a full range of angles covering the user's face as much as possible, and a 3DGDS digital person is generated based on the multi-view user images. The rendering image can be obtained by rendering the 3DGDS digital person. The mean square error between the rendering image and the user image is taken as a loss function, and the gradient of each Gaussian cell is determined through gradient back propagation. On this basis, the Gaussian cells are sorted based on the size of the gradient. The subsequent Gaussian cells can be used to construct an LOD model to generate an LOD digital person.

[0075] The LOD digital human is generated, and based on distances between a real-time viewpoint position of a user and each Gaussian cell in the LOD digital human, a level to be rendered corresponding to each Gaussian cell can be determined. For example, the farther the distance of a Gaussian cell, the lower the level to be rendered corresponding to the Gaussian cell, and the closer the distance of a Gaussian cell, the higher the level to be rendered corresponding to the Gaussian cell, to display more detailed information. Based on the level to be rendered corresponding to each Gaussian cell, the LOD digital human is rendered, and a rendered digital human can be generated.

[0076] In conclusion, by generating the LOD digital human, the distance between the real-time viewpoint position of the user and each Gaussian cell can be determined, the level of detail of the LOD digital human is optimized, the calculation load is reduced under the premise of ensuring the visual effect of the digital human, the frame rate of the digital human can be increased to more than 60 FPS under the premise of keeping the main features of the digital human clear, and the VR / AR interaction demand can be met.

[0077] Meanwhile, in the case of rendering in the same scene by multiple people, the method provided in the embodiment of the application only needs to maintain the rendered digital human of each user independently, and the total calculation load increases linearly with the number of people rather than exponentially.

[0078] The application provides a digital human rendering device, referring to Figure 4 The figure is a structural schematic diagram of a digital human rendering device provided in the embodiment of the application, and the specific implementation manner and the achieved technical effects are consistent with those described in the embodiment of the method, and part of the content will not be described again.

[0079] The embodiment of the application provides a digital human rendering device 4100, which comprises: a first generation unit 4101, a second generation unit 4102, and a rendering unit 4103; The first generation unit 4101 is configured to generate a 3D GS digital human based on user images of multiple viewpoints. The second generation unit 4102 is configured to determine levels of each Gaussian cell in the 3D GS digital human based on a mean square error between a rendering image and the user images, and generate a hierarchical detail LOD digital human composed of multiple levels of Gaussian cells; the rendering image is generated by rendering the 3D GS digital human. The rendering unit 4103 is configured to render the LOD digital human based on distances between a real-time viewpoint position of a user and each Gaussian cell position, and generate a rendered digital human.

[0080] In a possible implementation manner, the second generation unit is specifically configured to: determine a gradient of each Gaussian cell in the 3D GS digital human based on a mean square error between a rendering image and the user images. Based on the gradient, a level of each Gaussian cell in the 3D GS digital human is determined, and a hierarchical detail LOD digital human composed of multiple levels of Gaussian cells is generated.

[0081] In a possible implementation, the second generation unit is specifically configured to: sort the Gaussian cells according to the size of the gradient, to obtain a sorting result; determine a level of each Gaussian cell in the 3D GS digital human according to the sorting result, and generate a hierarchical detail LOD digital human composed of multiple levels of Gaussian cells.

[0082] In a possible implementation, the second generation unit is specifically configured to: determine a number of Gaussian cells contained in each level based on the total number of Gaussian cells and the attenuation coefficient; determine a level of each Gaussian cell in the 3D GS digital human according to the sorting result and the number of Gaussian cells contained in each level, and generate a hierarchical detail LOD digital human composed of multiple levels of Gaussian cells.

[0083] In a possible implementation, the rendering unit is specifically configured to: determine a corresponding rendering level of each Gaussian cell based on a distance between a real-time viewpoint position of a user and a position of each Gaussian cell; render each Gaussian cell to the corresponding rendering level, to generate a rendered digital human.

[0084] In a possible implementation, the rendering unit is specifically configured to: determine a distance ratio of each Gaussian cell based on a ratio between a distance between a real-time viewpoint position of a user and a position of each Gaussian cell and a reference distance threshold; take a logarithm of each distance ratio with the attenuation coefficient as a base, and down-round to obtain a corresponding rendering level of each Gaussian cell.

[0085] In a possible implementation, the rendering unit is specifically configured to: divide the LOD digital human into multiple regions, and determine a Gaussian cell dominant view distance corresponding to each region based on a distance between a real-time viewpoint position of a user and a position of each Gaussian cell included in each region; render the LOD digital human based on the Gaussian cell dominant view distance corresponding to each region, to generate a rendered digital human.

[0086] In a possible implementation, the apparatus further includes a fusion unit. The fusion unit is specifically configured to perform pixel-level fusion on the rendering result of the transition area between adjacent areas in the rendered digital person, to obtain a fused digital person.

[0087] To sum up, the embodiments of the present application can optimize the detail level of the LOD digital person based on the distance between the real-time viewpoint position of the user and each Gaussian cell position by generating the LOD digital person, reduce the calculation load under the premise of guaranteeing the visual effect of the digital person, and meet the demand of large-scale digital person rendering scene.

[0088] The embodiments of the present application provide a terminal device, as shown in the figure, which is a hardware structure schematic diagram of a terminal device provided by the embodiments of the present application. Figure 5 The terminal device 510 includes a memory 511 and a processor 512.

[0089] The memory 511 is configured to store program code and transmit the program code to the processor 512; and the processor 512 is configured to execute the steps of the digital person rendering method according to the instructions in the program code.

[0090] The embodiments of the present application provide a computer readable storage medium, and the computer readable storage medium stores a computer program.

[0091] The embodiments of the present application provide a computer program product, and when the computer program product runs on a computer, the computer executes the steps of the digital person rendering method.

[0092] The embodiments of the present application provide a chip, which includes a processor coupled with a memory, and is configured to execute a computer program or instructions stored in the memory, so that the chip implements the digital person rendering method.

[0093] It should be noted that each of the embodiments described in the specification of the present application adopts a progressive mode for description, and the same or similar parts between the embodiments can be mutually referred to. Each of the embodiments focuses on the differences from other embodiments. In particular, the device and system embodiments are described more simply because they are basically similar to the method embodiments, and the relevant parts can be referred to the part of the method embodiments. The above-described device and system embodiments are only illustrative, and the units described as separate components can or can not be physically separated, and the components indicated as units can or can not be physical units, i.e., they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiments according to actual needs. Those skilled in the art can understand and implement it without creative labor.

[0094] The above describes only one specific embodiment of the present application, but the protection scope of the present application is not limited to this. Any skilled person in the art can easily think of changes or replacements within the technical range disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method of digital human rendering, the method comprising: The method comprises the steps of: generating a 3D GS digital person based on a plurality of perspective images of a user; determining the level of each Gaussian cell in the 3D GS digital person based on the mean square error between a rendered image and the user image, and generating a hierarchical detail LOD digital person composed of multiple levels of Gaussian cells; the rendered image is generated by rendering the 3D GS digital person; rendering the LOD digital person based on the distance between the real-time viewpoint position of the user and the position of each Gaussian cell, and generating a rendered digital person.

2. The method of claim 1, wherein, The method for determining the level of each Gaussian cell in the 3D GS digital person based on the mean square error between the rendered image and the user image, and generating a hierarchical detail LOD digital person composed of multiple levels of Gaussian cells, comprises the steps of: determining the gradient of each Gaussian cell in the 3D GS digital person based on the mean square error between the rendered image and the user image; determining the level of each Gaussian cell in the 3D GS digital person based on the gradient, and generating a hierarchical detail LOD digital person composed of multiple levels of Gaussian cells.

3. The method of claim 2, wherein, The method for determining the level of each Gaussian cell in the 3D GS digital person based on the gradient, and generating a hierarchical detail LOD digital person composed of multiple levels of Gaussian cells, comprises the steps of: sorting the Gaussian cells according to the size of the gradient to obtain a sorting result; determining the level of each Gaussian cell in the 3D GS digital person based on the sorting result, and generating a hierarchical detail LOD digital person composed of multiple levels of Gaussian cells.

4. The method of claim 3, wherein, The method for determining the level of each Gaussian cell in the 3D GS digital person based on the sorting result, and generating a hierarchical detail LOD digital person composed of multiple levels of Gaussian cells, comprises the steps of: determining the number of Gaussian cells contained in each level based on the total number of Gaussian cells and an attenuation coefficient; determining the level of each Gaussian cell in the 3D GS digital person based on the sorting result and the number of Gaussian cells contained in each level, and generating a hierarchical detail LOD digital person composed of multiple levels of Gaussian cells.

5. The method of claim 1, wherein, The method for rendering the LOD digital person based on the distance between the real-time viewpoint position of the user and the position of each Gaussian cell, and generating a rendered digital person, comprises the steps of: determining the corresponding rendering level of each Gaussian cell based on the distance between the real-time viewpoint position of the user and the position of each Gaussian cell; rendering each Gaussian cell to the corresponding rendering level to generate a rendered digital person.

6. The method of claim 5, wherein, The method for determining the corresponding rendering level of each Gaussian cell based on the distance between the real-time viewpoint position of the user and the position of each Gaussian cell, comprises the steps of: determining the distance ratio of each Gaussian cell based on the ratio between the distance between the real-time viewpoint position of the user and the position of each Gaussian cell and a reference distance threshold; taking the attenuation coefficient as the base, and taking the logarithm of each distance ratio to the minus first power to obtain the corresponding rendering level of each Gaussian cell.

7. The method of claim 1, wherein, The method for rendering the LOD digital person based on the distance between the real-time viewpoint position of the user and the position of each Gaussian cell, and generating a rendered digital person, comprises the steps of: The LOD digital person is divided into multiple regions, and a corresponding Gaussian cell dominant view distance of each region is determined based on a distance between a real-time viewpoint position of a user and a plurality of Gaussian cell positions included in each region; Based on the corresponding Gaussian cell dominant view distance of each region, the LOD digital person is rendered to generate a rendered digital person.

8. The method of claim 7, wherein, After the LOD digital person is rendered based on the corresponding Gaussian cell dominant view distance of each region to generate a rendered digital person, the method further includes: For a transition region between adjacent regions in the rendered digital person, pixel-level fusion is performed on a rendering result of the transition region to obtain a fused digital person.

9. A digital human rendering apparatus, characterized by, It includes: A first generation unit is configured to generate a 3D GS digital person based on user images of multiple perspectives; A second generation unit is configured to determine levels of each Gaussian cell in the 3D GS digital person based on a mean square error between a rendered image and the user images, and generate a hierarchical detail LOD digital person composed of multiple-level Gaussian cells; the rendered image is generated by rendering the 3D GS digital person; A rendering unit is configured to render the LOD digital person based on a distance between a real-time viewpoint position of a user and positions of each Gaussian cell, and generate a rendered digital person.

10. A terminal device, comprising: The terminal device includes a processor and a memory; The memory is configured to store program code and transmit the program code to the processor; The processor is configured to execute the steps of the digital person rendering method according to the instructions in the program code.

Citation Information

Patent Citations

  • Real-time rendering method and device based on multi-level Gaussian sputtering

    CN118096972A

  • Three-dimensional dynamic scene rendering method and device, equipment, storage medium and program product

    CN119048662A

  • Large-scale scene parallel processing method based on 3D Gaussian

    CN119444956A

  • Novel interested target three-dimensional reconstruction method in complex environment

    CN120318420A

  • Dynamic scene reconstruction method and system based on spatial decomposition and Gaussian splashing

    CN120355851A