3D data processing method and device, equipment, storage medium and product

By performing characterization, feature transformation, quantization, and encoding processing on 3D digital human data, coded information is generated, solving the problem of low data transmission efficiency in 3D digital human data and achieving more efficient data transmission.

CN121125960APending Publication Date: 2025-12-12CHINA MOBILE COMM LTD RES INST +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510511118.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

When displaying a 3D digital human on a client side, existing technologies require the transmission of a large amount of 3D Gaussian sphere data, resulting in low data transmission efficiency.

Method used

By performing characterization, feature transformation, quantization, and encoding processing on 3D data of 3D digital humans based on 3D Gaussian sputtering, encoded information is generated, reducing the amount of data transmitted.

Benefits of technology

It reduces the amount of data transmitted and improves data transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121125960A_ABST
    Figure CN121125960A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a 3D data processing method and device, equipment, a storage medium and a computer program product. The 3D data processing method comprises the steps that 3D data of a 3D digital human based on 3D Gaussian sputtering 3DGS is acquired; representation generation processing, feature transformation processing, quantization processing and coding processing are sequentially carried out on the 3D data, coding information is obtained, and the coding processing comprises entropy coding processing and / or video coding processing. According to the invention, the data transmission efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video data processing, in particular to a 3D data processing method and device, equipment, storage medium and computer program product. BACKGROUND

[0002] Digital human is a virtual role generated by using a computer and related technologies. The digital human has similar appearance characteristics and behaviors to humans. In the fields of games, films, education and the like, digital humans have a wide range of applications. Among them, 3D digital humans have more extensive application potential in advertising, medical treatment, education and many other fields.

[0003] In the related art, a 3D Gaussian Splatting (3DGS) is used to represent a 3D digital human. When a client needs to display a 3D digital human, a Gaussian sphere corresponding to the 3DGS is transmitted to the client, so that the client can render the 3D Gaussian sphere to obtain a corresponding 3D digital human and display it. Since millions of 3D Gaussian spheres are needed on a 360-degree scene, the storage space exceeds 1 GB. When transmitting 1 GB of data to the client for rendering, the data transmission efficiency is reduced. SUMMARY

[0004] To solve the above technical problems, the embodiments of the present application provide a 3D data processing method and device, equipment, storage medium and computer program product, which can reduce the amount of data transmission and improve the data transmission efficiency.

[0005] The technical solution of the present application is implemented as follows:

[0006] The embodiments of the present application provide a 3D data processing method, which comprises:

[0007] Obtaining 3D data of a 3D digital human based on a 3D Gaussian Splatting (3DGS);

[0008] Performing representation generation processing, feature transformation processing, quantization processing and encoding processing on the 3D data in sequence to obtain encoding information, wherein the encoding processing comprises entropy encoding processing and / or video encoding processing.

[0009] The embodiments of the present application provide a 3D data processing method, which comprises:

[0010] Obtaining encoding information;

[0011] Performing decoding processing, inverse quantization processing and rendering processing on the encoding information to obtain the 3D digital human, wherein the decoding comprises entropy decoding processing and / or video decoding processing.

[0012] The embodiment of the present application provides a 3D data processing device, which comprises:

[0013] The first obtaining unit is configured to obtain 3D data of a 3D digital person based on three-dimensional Gaussian sputtering (3DGS).

[0014] The first processing unit is configured to sequentially perform feature generation processing, feature transformation processing, quantization processing and encoding processing on the 3D data to obtain encoding information, wherein the encoding processing comprises entropy encoding processing and / or video encoding processing.

[0015] The embodiment of the present application provides a client, which comprises:

[0016] The second obtaining unit is configured to obtain the encoding information.

[0017] The second processing unit is configured to perform decoding processing, inverse quantization processing and rendering processing on the encoding information to obtain the 3D digital person, wherein the decoding comprises entropy decoding processing and / or video decoding processing.

[0018] The embodiment of the present application provides a 3D data processing device, which comprises:

[0019] The first memory, the first processor and the first communication bus, the first memory communicates with the first processor through the first communication bus, the first memory stores a 3D data processing program executable by the first processor, when the 3D data processing program is executed, the first processor executes the 3D data processing method applied to the 3D data processing device.

[0020] The embodiment of the present application provides a client, which comprises:

[0021] The second memory, the second processor and the second communication bus, the second memory communicates with the second processor through the second communication bus, the second memory stores a 3D data processing program executable by the second processor, when the 3D data processing program is executed, the second processor executes the 3D data processing method applied to the client.

[0022] The embodiment of the present application provides a storage medium, which stores a computer program, and is applied to a 3D data processing device and a client, and is characterized by the following: when the computer program is executed by the first processor, the 3D data processing method applied to the 3D data processing device is realized; and when the computer program is executed by the second processor, the 3D data processing method applied to the client is realized.

[0023] This application also provides a computer program product, including a computer program that can be executed by a first processor in a 3D data processing device to complete the steps of the aforementioned 3D data processing method applied in a 3D data processing device, and the computer program can be executed by a second processor in a client to complete the steps of the aforementioned 3D data processing method applied in a client.

[0024] This application provides a 3D data processing method, apparatus, device, storage medium, and computer program product. The 3D data processing method includes: acquiring 3D data of a 3D digital human based on 3D Gaussian sputtering 3DGS; sequentially performing characterization generation processing, feature transformation processing, quantization processing, and encoding processing on the 3D data to obtain encoded information. The encoding processing includes entropy encoding processing and / or video encoding processing. Using the above method, the 3D data processing apparatus obtains multiple reference points by performing characterization generation processing, feature transformation processing, quantization processing, and encoding processing on the 3D data of the 3D digital human based on 3D Gaussian sputtering. The data volume of the reference points is less than the data volume of the 3D Gaussian sphere, resulting in less encoded information obtained by encoding the information corresponding to multiple reference points. This reduces the amount of data transmitted, thereby improving data transmission efficiency. Attached Figure Description

[0025] Figure 1 A method flow for 3D data processing provided in this application embodiment Figure One ;

[0026] Figure 2 A method flow for 3D data processing provided in this application embodiment Figure Two ;

[0027] Figure 3 A schematic diagram illustrating an exemplary 3D data processing method provided in an embodiment of this application;

[0028] Figure 4 This is a schematic diagram of the composition structure of a 3D data processing device provided in an embodiment of this application;

[0029] Figure 5 This is a schematic diagram of the composition structure of a 3D data processing device provided in an embodiment of this application;

[0030] Figure 6 A schematic diagram of the composition structure of a client provided in an embodiment of this application. Figure One ;

[0031] Figure 7 A schematic diagram of the composition structure of a client provided in an embodiment of this application. Figure Two . Detailed Implementation

[0032] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.

[0033] This application provides a 3D data processing method, which is applied to a 3D data processing device. Figure 1 A flowchart of a 3D data processing method provided in this application embodiment is shown below. Figure 1 As shown, 3D data processing methods may include:

[0034] S101. Acquire 3D data of 3D digital human based on 3D Gaussian sputtering 3DGS.

[0035] The 3D data processing method provided in this application embodiment is applicable to scenarios where 3D digital human data based on 3D Gaussian sputtering 3DGS is processed.

[0036] In the embodiments of this application, the 3D data processing device can be implemented in various forms. For example, the 3D data processing device described in this application may include a core network, a server, or a cloud, or other devices, and the specific embodiments of this application do not limit this.

[0037] In this embodiment, the 3D data processing device can acquire 3D data of a 3D digital human based on 3D Gaussian sputtering 3DGS, or obtain 3D data of a 3D digital human based on 3D Gaussian sputtering 3DGS from other devices, or obtain 3D data of a 3D digital human based on 3D Gaussian sputtering 3DGS through other means. The specific method by which the 3D data processing device obtains 3D data of a 3D digital human based on 3D Gaussian sputtering 3DGS can be determined according to the actual situation, and this embodiment does not limit it.

[0038] In this embodiment of the application, the 3D data can specifically be sparse point cloud data of digital human signals based on 3DGS.

[0039] S102. The 3D data is sequentially processed by representation generation, feature transformation, quantization and encoding to obtain encoded information. The encoding process includes entropy encoding and / or video encoding.

[0040] In this embodiment of the application, after the 3D data processing device acquires the 3D data of the 3D digital human based on three-dimensional Gaussian sputtering 3DGS, it sequentially performs characterization generation processing, feature transformation processing, quantization processing and encoding processing on the 3D data to obtain encoded information.

[0041] In this embodiment of the application, the 3D data processing device sequentially performs characterization generation processing, feature transformation processing, quantization processing, and encoding processing on 3D data to obtain encoded information. The process includes: performing characterization generation processing on the 3D data to obtain a voxel mesh, multiple reference points, and multiple reference point features obtained after the characterization generation processing; performing feature transformation processing based on the voxel mesh and multiple reference points to obtain hierarchical detail information corresponding to the multiple reference points; performing quantization processing on the multiple reference points and multiple reference point features to obtain quantized reference points and quantized reference point features; and encoding the quantized reference points and quantized reference point features to obtain encoded information.

[0042] It should be noted that encoding the quantized reference points includes encoding the quantized reference points based on hierarchical detail information.

[0043] In the embodiments of this application, multiple reference points and multiple reference point features correspond one-to-one, with each reference point specifically corresponding to one reference point feature.

[0044] In the embodiments of this application, after characterizing and generating 3D data, a voxel grid, multiple reference points, and multiple sets of reference point attributes corresponding to the multiple reference points can be obtained. The multiple sets of reference point attributes include multiple reference point features, multiple scaling factors, and multiple sets of offsets.

[0045] In this embodiment, feature transformation processing based on voxel mesh and multiple reference points can be performed using a hierarchical detail model; alternatively, other methods can be used to perform feature transformation processing based on voxel mesh and multiple reference points. The specific method of feature transformation processing based on voxel mesh and multiple reference points can be determined according to the actual situation, and this embodiment does not limit it.

[0046] In this embodiment of the application, multiple reference point features are quantized to obtain quantized reference point features; multiple reference points are quantized to obtain quantized reference points.

[0047] It should be noted that multiple reference points specifically refer to multiple reference point location information. The location information of multiple reference points is quantized to obtain the quantized reference point location information.

[0048] In the embodiments of this application, the method of encoding the quantized reference point and the quantized reference point features can be determined according to the actual situation, and the embodiments of this application do not limit this.

[0049] In this embodiment, the process of the 3D data processing device acquiring the voxel mesh, multiple reference points, and multiple reference point features obtained after characterization generation processing further includes: the 3D data processing device acquiring multiple sets of offsets corresponding to the multiple reference points obtained after characterization generation processing; correspondingly, the process of the 3D data processing device quantizing the multiple reference points and multiple reference point features to obtain quantized reference points and quantized reference point features further includes: acquiring observation points; quantizing the multiple sets of offsets and observation points to obtain quantized observation points and multiple sets of quantized offsets; correspondingly, the process of the 3D data processing device encoding the quantized reference points and quantized reference point features to obtain encoded information includes: acquiring the task requirements transmitted by the client and / or the network conditions of the transmission network; adjusting the encoder parameters based on the task requirements and / or network conditions to obtain adjusted encoder parameters; and encoding the quantized information based on the adjusted encoder parameters to obtain encoded information.

[0050] It should be noted that the quantized information includes the quantized reference point, the quantized reference point features, the quantized observation point, and multiple sets of quantized offsets.

[0051] In the embodiments of this application, multiple reference points and multiple sets of offsets correspond one-to-one, that is, one reference point corresponds to one set of offsets.

[0052] In this embodiment, observation points can be obtained from user-inputted information, from other devices, or through other means. The specific method of obtaining observation points can be determined according to the actual situation, and this embodiment does not limit it.

[0053] In this embodiment of the application, the observation point here specifically refers to the location information of the observation point.

[0054] In this embodiment of the application, multiple sets of offsets are quantized to obtain multiple sets of quantized offsets; observation points are quantized to obtain quantized observation points.

[0055] It should be noted that if the observation point is the location information of the observation point, then the location information of the observation point can be quantified to obtain the quantified observation point.

[0056] In this embodiment, the task requirements can be set for the client when rendering a 3D digital human. The task requirements include whether the rendering accuracy of the 3D digital human is higher than a first preset accuracy or lower than a second preset accuracy. If the rendering accuracy is higher than the first preset accuracy, the amount of data transmitted during encoding can be increased; if the rendering accuracy is lower than the second preset accuracy, the amount of data transmitted during encoding can be reduced, thereby improving transmission efficiency.

[0057] It should be noted that the first preset precision is greater than or equal to the second preset precision.

[0058] It should be noted that the network requirements can be the network conditions detected by the 3D data processing device when information is to be transmitted to the client. When the network conditions are identified as high-quality, the amount of data transmitted during encoding can be increased; when the network conditions are not identified as high-quality, and the rendering task of the 3D digital human is maximized, the amount of data transmitted during encoding can be reduced, thereby improving transmission efficiency.

[0059] In this embodiment, the number of multiple reference points generated during the characterization generation process can be adjusted according to task requirements and / or network conditions. For example, if the rendering accuracy of the 3D digital human is lower than the second preset accuracy or the network condition is not identified as a high-quality network, the number of multiple reference points generated can be reduced; if the rendering accuracy of the 3D digital human is higher than the first preset accuracy or the network condition is identified as a high-quality network, the number of multiple reference points generated can be increased.

[0060] In this embodiment, the process by which the 3D data processing device encodes quantized information based on adjusted encoder parameters to obtain encoded information includes: entropy encoding of quantized reference points based on adjusted encoder parameters and hierarchical detail information to obtain reference point encoded information; entropy encoding of quantized observation points and multiple sets of quantized offsets based on adjusted encoder parameters to obtain entropy encoded information; video encoding of quantized reference point features based on adjusted encoder parameters to obtain video encoded information; compression processing of the video encoded information to obtain compressed video encoded information; and using the entropy encoded information, reference point encoded information, and compressed video encoded information as encoded information.

[0061] In this embodiment, the adjusted encoder parameters can be the encoding parameters of the entropy encoder obtained by adjusting the encoding parameters of the entropy encoder. The entropy encoder obtained by adjusting the encoding parameters of the entropy encoder can be used to entropy encode the quantized reference points based on hierarchical detail information to obtain reference point encoding information. Alternatively, other methods can be used to entropy encode the quantized reference points based on hierarchical detail information to obtain reference point encoding information. The specific method of entropy encoding the quantized reference points based on hierarchical detail information to obtain reference point encoding information can be determined according to the actual situation, and this embodiment does not limit it.

[0062] In this embodiment, the entropy encoder obtained by adjusting the encoding parameters of the entropy encoder can be used to entropy encode the quantized observation points and multiple sets of quantized offsets to obtain entropy encoding information; other methods can also be used to entropy encode the quantized observation points and multiple sets of quantized offsets to obtain entropy encoding information; the specific method of entropy encoding the quantized observation points and multiple sets of quantized offsets to obtain entropy encoding information can be determined according to the actual situation, and this embodiment does not limit it.

[0063] In this embodiment, the adjusted encoder parameters can be the encoding parameters of the video encoder obtained after adjusting the encoding parameters of the video encoder. The video encoder obtained after adjusting the encoding parameters of the video encoder can be used to perform video encoding on the quantized reference point features to obtain video encoding information. Alternatively, other methods can be used to perform video encoding on the quantized reference point features to obtain video encoding information. The specific method of performing video encoding on the quantized reference point features to obtain video encoding information can be determined according to the actual situation, and this embodiment does not limit it.

[0064] In this embodiment, the video encoder obtained by adjusting the encoding parameters of the video encoder can be used to compress the video encoded information, or other compression methods can be used to compress the video encoded information. The specific method of compressing the video encoded information can be determined according to the actual situation, and this embodiment does not limit it.

[0065] It should be noted that the encoding information includes entropy encoding information, reference point encoding information, and compressed video encoding information.

[0066] In this embodiment of the application, the process by which the 3D data processing device uses entropy coding information, reference point coding information, and compressed video coding information as coding information includes: acquiring a multilayer perceptron and network parameters in the multilayer perceptron; performing entropy coding on the network parameters to obtain first coding information; and using the first coding information, entropy coding information, reference point coding information, and compressed video coding information as coding information.

[0067] It should be noted that the multilayer perceptron includes a first multilayer perceptron, a second multilayer perceptron, a third multilayer perceptron, and a fourth multilayer perceptron; the network parameters in the multilayer perceptron include the network parameters in the first multilayer perceptron, the second multilayer perceptron, the third multilayer perceptron, and the fourth multilayer perceptron.

[0068] In this embodiment, the multilayer perceptron and its network parameters can be obtained from other devices; the multilayer perceptron can also be trained locally and its network parameters can be obtained locally; or the multilayer perceptron and its network parameters can be obtained through other means. The specific method of obtaining the multilayer perceptron and its network parameters can be determined according to the actual situation, and this embodiment does not limit it.

[0069] In this embodiment, the network parameters can be entropy encoded using an entropy encoder to obtain the first encoded information; alternatively, other methods can be used to entropy entropy encode the network parameters to obtain the first encoded information. The specific method of entropy encoding the network parameters to obtain the first encoded information can be determined according to the actual situation, and this embodiment does not limit it.

[0070] In this embodiment of the application, the process of the 3D data processing device performing characterization and generation processing on 3D data includes: converting 3D data into a voxel mesh; using the center point of each voxel in the voxel mesh as a reference point to obtain multiple reference points and multiple sets of reference point attributes corresponding to the multiple reference points.

[0071] It should be noted that multiple sets of reference point attributes include multiple reference point features, multiple scaling factors, and multiple sets of offsets.

[0072] In this embodiment, a structure-from-motion (SFM) can be used to generate a voxel mesh from 3D data, or other methods can be used to convert 3D data into a voxel mesh. The specific method of converting 3D data into a voxel mesh can be determined according to the actual situation, and this embodiment does not limit it.

[0073] In the embodiments of this application, multiple reference points and multiple sets of reference point attributes correspond one-to-one, that is, one reference point corresponds to one set of reference point attributes. A set of reference point attributes includes a reference point feature, a scaling factor, and a set of offsets.

[0074] In this embodiment of the application, the process of obtaining multiple reference points by using the center point of each voxel in the voxel mesh as a reference point includes: using the center point of each voxel in the voxel mesh as a reference point to obtain multiple initial reference points; obtaining the task requirements transmitted by the client; and adjusting the number of multiple initial reference points based on the task requirements to obtain multiple reference points.

[0075] In this embodiment, the number of reference points during the characterization generation process can be adjusted based on the client's task requirements to obtain multiple reference points. For example, if the rendering accuracy of the 3D digital human is lower than a second preset accuracy, the number of generated reference points is reduced (i.e., the number of initial reference points is reduced to obtain multiple reference points); if the rendering accuracy of the 3D digital human is higher than a first preset accuracy, the number of generated reference points is increased (the number of initial reference points is increased to obtain multiple reference points).

[0076] In this embodiment of the application, the process by which the 3D data processing device performs feature transformation processing based on a voxel mesh and multiple reference points to obtain hierarchical detail information corresponding to the multiple reference points includes: determining the tree depth based on the multiple reference points; establishing an octree based on the voxel mesh and the tree depth; and assigning the multiple reference points to the octree based on their positions to obtain hierarchical detail information corresponding to the multiple reference points.

[0077] In this embodiment of the application, the tree depth can be determined based on the depth of field range when multiple reference points are photographed from the observation point.

[0078] In the embodiments of this application, the method of building an octree based on voxel mesh and tree depth can be determined according to the actual situation, and the embodiments of this application do not limit it.

[0079] In this embodiment of the application, the hierarchical detail information can be the hierarchical information of multiple reference points in the octree.

[0080] In this embodiment, the process by which the 3D data processing device quantizes multiple reference points and multiple reference point features to obtain quantized reference points and quantized reference point features includes: acquiring a quantization step size and a preset quantization method; quantizing multiple reference points based on the quantization step size to obtain quantized reference points; and quantizing multiple reference point features using the preset quantization method to obtain quantized reference point features. Correspondingly, the process by which the 3D data processing device quantizes multiple sets of offsets and observation points to obtain quantized observation points and multiple sets of quantized offsets includes: quantizing multiple sets of offsets and observation points based on the quantization step size to obtain multiple sets of quantized offsets and quantized observation points.

[0081] In this embodiment, the quantization step size can be the step size information configured in the 3D data processing device, the step size information transmitted to the 3D data processing device from other devices, or the step size information obtained by the 3D data processing device through other means. The specific way in which the 3D data processing device obtains the quantization step size can be determined according to the actual situation, and this embodiment does not limit it.

[0082] It should be noted that the specific parameter value of the quantization step size can be determined according to the actual situation, and this application embodiment does not limit it.

[0083] In this embodiment of the application, multiple reference points are quantized based on the quantization step size to obtain the quantized reference points in the manner shown in formula (1):

[0084] X V =(X v ) / s (1)

[0085] It should be noted that s is the quantization step size, and X... v For any reference point, X V This refers to the quantized reference point obtained after quantizing any given reference point.

[0086] In this embodiment, the preset quantization method can be information configured in the 3D data processing device, information transmitted to the 3D data processing device from other devices, or information obtained by the 3D data processing device through other means. The specific way in which the 3D data processing device obtains the preset quantization method can be determined according to the actual situation, and this embodiment does not limit it.

[0087] It should be noted that the preset quantization method can be a way of converting information into a preset format, such as converting it into a 10-bit integer format compatible with the input format of the video encoder; the preset quantization method can also be other methods, and the specific preset quantization method can be determined according to the actual situation. This application embodiment does not limit this.

[0088] In this embodiment, multiple reference points can be quantized based on a quantization step size. This allows for the quantization of multiple sets of offsets and observation points, resulting in multiple sets of quantized offsets and observation points. Alternatively, other quantization methods can be used to quantize multiple sets of offsets and observation points based on a quantization step size. The specific method for quantizing multiple sets of offsets and observation points based on a quantization step size can be determined based on actual circumstances, and this embodiment does not limit this approach.

[0089] In this embodiment of the application, the process of a 3D data processing device determining tree depth based on multiple reference points includes: drawing a histogram of the shooting scene distances corresponding to multiple reference points in a voxel grid; determining the confidence interval of the histogram; and determining the tree depth based on the distance range corresponding to the confidence interval.

[0090] In this embodiment, a histogram of the shooting scene distances corresponding to multiple reference points in a voxel grid can be drawn using a histogram drawing device. Alternatively, other methods can be used to draw the histogram of the shooting scene distances corresponding to multiple reference points in a voxel grid. The specific method for drawing the histogram of the shooting scene distances corresponding to multiple reference points in a voxel grid can be determined according to the actual situation, and this embodiment does not limit it.

[0091] In this embodiment, the confidence interval can be configured information, user-input information, or information obtained through other means. The specific method for determining the confidence interval of the histogram can be determined according to the actual situation, and this embodiment does not limit it.

[0092] In the embodiments of this application, after determining the confidence interval, the distance range corresponding to the confidence interval can be determined.

[0093] For example, after determining the camera position, the positions of each point cloud (multiple reference points) relative to the camera are different. After obtaining the positions of each point relative to the camera, a histogram can be plotted. After determining the confidence interval of the histogram, a confidence interval of (0.05, 0.95) is selected from the histogram to obtain the distance range corresponding to the confidence interval. It should be noted that, It is the maximum value within this distance range. It is the minimum value within this distance range.

[0094] In this embodiment of the application, the method for determining the tree depth based on the distance range is as shown in formula (2):

[0095]

[0096] It should be noted that k is the tree depth. It is the maximum value within the distance range. It is the minimum value within the distance range.

[0097] In this embodiment, the 3D data processing device uses the center point of each voxel in the voxel mesh as a reference point to obtain multiple reference points and multiple sets of reference point attributes corresponding to the multiple reference points. The 3D data processing device also determines multiple relative distances and multiple viewing directions between the multiple reference points and the observation point. Based on the multiple reference points, multiple sets of offsets and multiple scaling factors, it determines multiple sets of 3D Gaussian positions corresponding to the multiple reference points. Using a multilayer perceptron, it determines multiple batches of 3D Gaussian attributes corresponding to the multiple reference points based on the features of the multiple reference points, multiple relative distances and multiple viewing directions.

[0098] In this embodiment of the application, multiple relative distances and multiple viewing directions between multiple reference points and observation points can be determined using formulas (3)-(4):

[0099] δ vc =|| xv -x c ||2 (3)

[0100]

[0101] It should be noted that x v Let x be the position of any one of multiple reference points. c δ represents the location of the observation point. vc The relative distance to any reference point. This refers to the viewing direction corresponding to any reference point.

[0102] It should be noted that there is a one-to-one correspondence between multiple reference points and multiple relative distances; specifically, one reference point corresponds to one relative distance. There is also a one-to-one correspondence between multiple reference points and multiple viewing directions; specifically, one reference point corresponds to one viewing direction.

[0103] In this embodiment, multiple reference point features, multiple relative distances, and multiple viewpoint directions can be input into a multilayer perceptron to obtain multiple batches of 3D Gaussian attributes.

[0104] In this embodiment, the 3D data processing device utilizes a multilayer perceptron to determine multiple batches of 3D Gaussian attributes corresponding to multiple reference points based on multiple reference point features, multiple relative distances, and multiple viewing directions. This process includes: inputting multiple reference point features, multiple relative distances, and multiple viewing directions into a first multilayer perceptron to obtain multiple sets of opacities corresponding to multiple sets of 3D Gaussian positions; inputting multiple relative distances, multiple viewing directions, and multiple reference point features into a second multilayer perceptron to obtain multiple sets of colors corresponding to multiple sets of 3D Gaussian positions; inputting multiple relative distances, multiple viewing directions, and multiple reference point features into a third multilayer perceptron to obtain multiple sets of four-element data points corresponding to multiple sets of 3D Gaussian positions; inputting multiple relative distances, multiple viewing directions, and multiple reference point features into a fourth multilayer perceptron to obtain multiple sets of scaling features corresponding to multiple sets of 3D Gaussian positions; and using the multiple sets of opacities, multiple sets of colors, multiple sets of four-element data points, and multiple sets of scaling features as multiple batches of 3D Gaussian attributes.

[0105] In the embodiments of this application, any batch of 3D Gaussian attributes includes a set of opacities, a set of colors, a set of four elements, and a set of scaling features.

[0106] In this embodiment of the application, the multilayer perceptron includes a first multilayer perceptron for determining opacity, a second multilayer perceptron for determining color, a third multilayer perceptron for determining four elements, and a fourth multilayer perceptron for determining scaling features.

[0107] In this embodiment, multiple reference point features, multiple relative distances, and multiple viewing directions are input into the first multilayer perceptron to obtain multiple sets of opacities corresponding to multiple 3D Gaussian positions, as shown in formula (5):

[0108]

[0109] It should be noted that F α Let {α0,...,α} be the network parameters in the first multilayer perceptron. k-1} represents a set of opacities corresponding to any reference point, f v For any reference point feature corresponding to any one of multiple reference points, δ vc The relative distance to any one of multiple reference points. This refers to the viewing direction corresponding to any one of the multiple reference points.

[0110] In this embodiment, multiple relative distances, multiple viewing directions, and multiple reference point features are input into the second multilayer perceptron to obtain multiple sets of colors corresponding to multiple 3D Gaussian positions, as shown in formula (6):

[0111]

[0112] It should be noted that F c Here are the network parameters in the second multilayer perceptron, {c0,...,c k-1} represents a set of colors corresponding to any reference point, f v For any reference point feature corresponding to any one of multiple reference points, δ vc The relative distance to any one of multiple reference points. This refers to the viewing direction corresponding to any one of the multiple reference points.

[0113] In this embodiment, multiple relative distances, multiple viewing directions, and multiple reference point features are input into the third multilayer perceptron to obtain multiple sets of four elements corresponding to multiple 3D Gaussian positions, as shown in formula (7):

[0114]

[0115] It should be noted that F q Let {q0,...,q} be the network parameters in the third multilayer perceptron. k-1} represents a set of four elements corresponding to any reference point, f v For any reference point feature corresponding to any one of multiple reference points, δ vc The relative distance to any one of multiple reference points. This refers to the viewing direction corresponding to any one of the multiple reference points.

[0116] In this embodiment, multiple relative distances, multiple viewing directions, and multiple reference point features are input into the fourth multilayer perceptron to obtain multiple sets of scaling features corresponding to multiple 3D Gaussian positions, as shown in formula (8):

[0117]

[0118] It should be noted that F s Let {s0,...,s} be the network parameters in the fourth multilayer perceptron. k-1} represents a set of scaling features corresponding to any reference point, f v For any reference point feature corresponding to any one of multiple reference points, δ vc The relative distance to any one of multiple reference points. This refers to the viewing direction corresponding to any one of the multiple reference points.

[0119] In this embodiment of the application, the process by which the 3D data processing device determines multiple sets of 3D Gaussian positions corresponding to multiple reference points based on multiple reference points, multiple sets of offsets, and multiple scaling factors includes: acquiring the position of any one reference point among the multiple reference points, as well as any set of offsets and any scaling factor corresponding to the position of any one reference point; determining the product of any scaling factor and each offset in any set of offsets to obtain any set of products; determining the sum of any reference point position and each product in any set of products to obtain any set of 3D Gaussian positions corresponding to the position of any reference point; until multiple sets of 3D Gaussian positions corresponding to multiple reference points are obtained.

[0120] In this embodiment of the application, after obtaining the position of any one of the multiple reference points, as well as any set of offsets and any scaling factor corresponding to the position of any one reference point, the method for determining any set of 3D Gaussian positions corresponding to the position of any one reference point based on the position of any one reference point, any set of offsets and any scaling factor is as shown in formula (9):

[0121]

[0122] It should be noted that x v For the position of any one of multiple reference points, It is a set of learnable offsets corresponding to any given reference point (i.e., a set of offsets corresponding to any given reference point), l v It is a scaling factor corresponding to any given reference point.

[0123] In this embodiment of the application, the 3D data processing device sequentially performs characterization generation processing, feature transformation processing, quantization processing and encoding processing on the 3D data. After obtaining the encoded information, it transmits the encoded information to the client so that the client can decode the encoded information and render it to obtain a 3D digital human.

[0124] In this embodiment, the 3D data processing device can transmit encoded information to the client via a transmission network. The specific method of transmitting encoded information to the client can be determined according to the actual situation, and this embodiment does not limit it.

[0125] Understandably, the 3D data processing device performs characterization, feature transformation, quantization, and encoding processing on the 3D data of the 3D digital human based on three-dimensional Gaussian sputtering to obtain multiple reference points. The amount of data for each reference point is less than the amount of data for the 3D Gaussian sphere, resulting in less encoded information obtained by encoding the information corresponding to multiple reference points, thus reducing the amount of data transmitted and improving data transmission efficiency.

[0126] This application embodiment also provides a 3D data processing method, which is applied to a client-side application. Figure 2 A flowchart of a 3D data processing method is provided as an embodiment of this application, such as... Figure 2 As shown, 3D data processing methods may include:

[0127] S201. Obtain encoding information.

[0128] The 3D data processing method provided in this application embodiment is applicable to scenarios where a 3D digital human is rendered based on encoded information.

[0129] In the embodiments of this application, the client can be implemented in various forms. For example, the client described in this application may include devices such as mobile phones, cameras, tablet computers, laptops, PDAs, personal digital assistants (PDAs), portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as devices such as digital TVs, desktop computers, servers, etc. The specific embodiments of this application are not limited in this regard.

[0130] In this embodiment of the application, the encoded information of multiple reference points transmitted by the 3D data processing device can be received to obtain the encoded information.

[0131] In the embodiments of this application, the encoded information can be information obtained by video encoding, information obtained by entropy encoding, or information obtained by other encoding methods. The specific information can be determined according to the actual situation, and the embodiments of this application do not limit it.

[0132] S202. Decode, dequantize, and render the encoded information to obtain a 3D digital human. Decoding includes entropy decoding and / or video decoding.

[0133] In this embodiment of the application, after the client obtains the encoded information, it can perform decoding, dequantization, and rendering processes on the encoded information to obtain a 3D digital human.

[0134] In this embodiment, the process of the client performing decoding, dequantization, and rendering on the encoded information to obtain a 3D digital human includes: decompressing the compressed video encoded information in the encoded information to obtain video encoded information; performing video decoding on the video encoded information to obtain quantized reference point features; performing entropy decoding on the first encoded information, reference point encoded information, and entropy encoded information in the encoded information to obtain network parameters in the multilayer perceptron, quantized reference points, quantized observation points, and multiple sets of quantized offsets; performing dequantization on the quantized reference point features, quantized reference points, quantized observation points, and multiple sets of quantized offsets to obtain multiple reference points, multiple reference point features, observation points, and multiple sets of offsets; and performing rendering on the multiple reference points, multiple reference point features, observation points, multiple sets of offsets, and network parameters to obtain a 3D digital human.

[0135] In the embodiments of this application, decompression is the reverse of compression.

[0136] In the embodiments of this application, video decoding is the reverse of video encoding.

[0137] In this embodiment, entropy decoding is the reverse of entropy encoding.

[0138] In the embodiments of this application, inverse quantization is the reverse of quantization.

[0139] In this embodiment, the process of obtaining a 3D digital human by rendering based on multiple reference points, multiple reference point features, observation points, multiple sets of offsets, and network parameters includes: obtaining multiple relative distances between multiple reference points and observation points, the tree depth of the octree, and maximum value information; determining multiple rendering layer details corresponding to multiple reference points based on multiple relative distances, tree depth, and maximum value information; and performing rendering based on multiple rendering layer details, multiple reference points, multiple reference point features, observation points, multiple sets of offsets, and network parameters to obtain a 3D digital human.

[0140] It should be noted that the maximum value information is the maximum value within the distance range determined by the confidence interval of the histogram, which is a histogram of the shooting scene distance drawn based on multiple reference points in the voxel grid.

[0141] In the embodiments of this application, multiple relative distances between multiple reference points and observation points can be determined, that is, multiple relative distances between multiple reference points and observation points can be obtained.

[0142] In this embodiment, the tree depth can be determined based on multiple reference points, i.e., the tree depth is obtained. Specifically, a histogram of the shooting scene distances corresponding to multiple reference points in the voxel mesh is drawn; the confidence interval of the histogram is determined; and the tree depth is determined based on the distance range corresponding to the confidence interval.

[0143] The maximum value information can be determined from the distance range corresponding to the confidence interval.

[0144] In this embodiment, the client can determine multiple batches of 3D Gaussian attributes (including multiple sets of opacity, multiple sets of color, multiple sets of four elements and multiple sets of scaling features) corresponding to multiple reference points based on network parameters, multiple reference point features, multiple relative distances and multiple view directions, and thus render a 3D digital human based on the multiple batches of 3D Gaussian attributes.

[0145] It should be noted that the client can also determine multiple relative distances and multiple viewing directions between multiple reference points and observation points.

[0146] In this embodiment of the application, the method for determining multiple rendering level details corresponding to multiple reference points based on multiple relative distances, tree depths, and maximum value information is as shown in formula (10):

[0147]

[0148] It should be noted that, For any one of multiple rendering level details, Ψ is the rounding operation, which rounds the estimated level of detail L. *Convert to an integer in the range [0, K-1]. Then, only the level of detail needs to be no more than [0, K-1]. The reference points are used to determine the level of detail from level 0 to level 2. The final reconstruction or rendering is then performed at each level. Specifically, during the rendering process, a corresponding neural Gaussian function is generated from the aforementioned benchmark points that meet the requirements. Then, a tile rasterization method is used to achieve efficient scene rendering. This process renders progressively from low to high level, improving the level of detail of the scene by rendering benchmark points with high detail levels. The above method significantly reduces the number of primitives required to render the corresponding view, greatly improving rendering efficiency.

[0149] For example, such as Figure 3The process involves: acquiring 3D data of a 3D digital human based on 3D Gaussian sputtering 3DGS (digital human signal based on 3DGS); performing characterization generation processing on the 3D data to obtain a voxel mesh, multiple reference points, and features of these reference points; performing feature transformation processing based on the voxel mesh and multiple reference points to obtain hierarchical detail information corresponding to these reference points; quantizing these reference points and their features to obtain quantized reference points and features; quantizing multiple sets of offsets and observation points (camera positions) to obtain quantized observation points and multiple sets of quantized offsets; acquiring the task requirements transmitted by the client (terminal-side task requirements) and / or the network conditions of the transmission network; adjusting the encoder parameters based on the task requirements and / or network conditions to obtain the adjusted encoder parameters; and then applying the adjusted encoder parameters to the network conditions. The hierarchical detail information is used to entropy encode the quantized reference points (reference point positions) to obtain reference point encoded information. Based on the adjusted encoder parameters, the quantized observation points (camera positions) and multiple sets of quantized offsets are entropy encoded to obtain entropy encoded information. Based on the adjusted encoder parameters, the quantized reference point features are video encoded to obtain video encoded information. The video encoded information is compressed to obtain compressed video encoded information. The multilayer perceptron (multilayer perceptrons include a first, second, third, and fourth multilayer perceptron) and its network parameters (MLP network parameters) are acquired. The network parameters are entropy encoded to obtain first encoded information. The first encoded information, entropy encoded information, reference point encoded information, and compressed video encoded information are used as encoded information (bitstream). The encoded information is transmitted to the client (user) via the network for decoding and rendering to obtain a 3D digital human.The client obtains the encoded information; decompresses the compressed video encoded information to obtain the video encoded information; performs video decoding on the video encoded information to obtain the quantized reference point features; performs entropy decoding on the first encoded information, reference point encoded information, and entropy encoded information to obtain the network parameters (MLP network parameters), quantized reference points, quantized observation points, and multiple sets of quantized offsets in the multilayer perceptron; and performs dequantization on the quantized reference point features, quantized reference points (reference point positions), quantized observation points (camera positions), and multiple sets of quantized offsets (offsets) to obtain the multilayer perceptron network parameters (MLP network parameters), quantized reference points, quantized observation points (camera positions), and multiple sets of quantized offsets (offsets). The system identifies several reference points (reference point locations), multiple reference point features, an observation point (camera position), and multiple sets of offsets. It acquires multiple relative distances between these reference points and the observation point, the depth of the octree, and maximum value information. The maximum value information is the maximum value within a distance range determined by the confidence interval of a histogram, which is a histogram of shooting scene distances drawn from multiple reference points in a voxel grid. Based on the multiple relative distances, tree depth, and maximum value information, it determines multiple rendering layer details corresponding to the multiple reference points. Rendering is then performed based on these rendering layer details, multiple reference points, multiple reference point features, the observation point, multiple sets of offsets, and network parameters to obtain a 3D digital human.

[0150] It is understandable that when a client receives encoded information from multiple reference points transmitted by a 3D data processing device, the amount of data from each reference point is less than that from a 3D Gaussian sphere. This results in a smaller amount of encoded information obtained by encoding the information corresponding to multiple reference points, which reduces the amount of data transmitted and thus improves data transmission efficiency.

[0151] Based on the same inventive concept as the 3D data processing method applied in the 3D data processing device described above, this application provides a 3D data processing device 1, corresponding to a 3D data processing method. Figure 4 A schematic diagram of the composition structure of a 3D data processing device provided in this application embodiment. Figure One The 3D data processing device 1 may include:

[0152] The first acquisition unit 11 is used to acquire 3D data of the 3D digital human based on 3D Gaussian sputtering 3DGS;

[0153] The first processing unit 12 is used to sequentially perform characterization generation processing, feature transformation processing, quantization processing and encoding processing on the 3D data to obtain encoded information. The encoding processing includes entropy encoding processing and / or video encoding processing.

[0154] In some embodiments of this application, the first processing unit 12 is configured to perform characterization generation processing on the 3D data to obtain a voxel mesh, multiple reference points, and multiple reference point features obtained after the characterization generation processing; perform feature transformation processing based on the voxel mesh and the multiple reference points to obtain hierarchical detail information corresponding to the multiple reference points; perform quantization processing on the multiple reference points and the multiple reference point features to obtain quantized reference points and quantized reference point features; and encode the quantized reference points and the quantized reference point features to obtain the encoded information, wherein encoding the quantized reference points includes encoding the quantized reference points based on the hierarchical detail information.

[0155] In some embodiments of this application, the first acquisition unit 11 is used to acquire multiple sets of offsets corresponding to the plurality of reference points obtained after the characterization generation process;

[0156] Accordingly, the first acquisition unit 11 is used to acquire the observation point;

[0157] The first processing unit 12 is used to quantize the multiple sets of offsets and the observation points to obtain quantized observation points and multiple sets of quantized offsets.

[0158] Accordingly, the device also includes an adjustment unit;

[0159] The first acquisition unit 11 is used to acquire the task requirements transmitted by the client and / or the network conditions of the transmission network;

[0160] The adjustment unit is used to adjust the encoder parameters based on the task requirements and / or the network conditions to obtain the adjusted encoder parameters.

[0161] The first processing unit 12 is used to encode the quantized information based on the adjusted encoder parameters to obtain the encoded information. The quantized information includes the quantized reference point, the quantized reference point features, the quantized observation point, and the multiple sets of quantized offsets.

[0162] In some embodiments of this application, the apparatus further includes a compression unit and a determination unit;

[0163] The first processing unit 12 is configured to perform entropy encoding on the quantized reference point based on the adjusted encoder parameters and the hierarchical detail information to obtain reference point encoding information; perform entropy encoding on the quantized observation point and the multiple sets of quantized offsets based on the adjusted encoder parameters to obtain entropy encoding information; and perform video encoding on the quantized reference point features based on the adjusted encoder parameters to obtain video encoding information.

[0164] The compression unit is used to compress the video encoding information to obtain compressed video encoding information;

[0165] The determining unit is used to use the entropy coding information, the reference point coding information, and the compressed video coding information as the coding information.

[0166] In some embodiments of this application, the first acquisition unit 11 is used to acquire a multilayer perceptron and network parameters in the multilayer perceptron; the multilayer perceptron includes a first multilayer perceptron, a second multilayer perceptron, a third multilayer perceptron, and a fourth multilayer perceptron.

[0167] The first processing unit 12 is used to entropy encode the network parameters to obtain first encoded information;

[0168] The determining unit is used to take the first encoding information, the entropy encoding information, the reference point encoding information, and the compressed video encoding information as the encoding information.

[0169] In some embodiments of this application, the apparatus further includes a conversion unit;

[0170] The conversion unit is used to convert the 3D data into the voxel mesh;

[0171] The determining unit is used to take the center point of each voxel in the voxel grid as a reference point to obtain the plurality of reference points and the plurality of reference point attributes corresponding to the plurality of reference points; the plurality of reference point attributes include the plurality of reference point features, the plurality of scaling factors and the plurality of offsets.

[0172] In some embodiments of this application, the determining unit is used to take the center point of each voxel in the voxel grid as a reference point to obtain multiple initial reference points;

[0173] The first acquisition unit 11 is used to acquire the task requirements transmitted by the client;

[0174] The adjustment unit is used to adjust the number of the plurality of initial reference points based on the task requirements, thereby obtaining the plurality of reference points.

[0175] In some embodiments of this application, the apparatus further includes an establishment unit and an allocation unit;

[0176] The determining unit is used to determine the tree depth based on the plurality of reference points;

[0177] The establishment unit is used to establish an octree based on the voxel grid and the tree depth;

[0178] The allocation unit is used to allocate the multiple reference points to the octree based on their positions, thereby obtaining the hierarchical detail information corresponding to the multiple reference points.

[0179] In some embodiments of this application, the first acquisition unit 11 is used to acquire the quantization step size and the preset quantization method;

[0180] The first processing unit 12 is configured to perform quantization processing on the plurality of reference points based on the quantization step size to obtain the quantized reference points; and to perform quantization processing on the features of the plurality of reference points using a preset quantization method to obtain the quantized reference point features.

[0181] Accordingly, the first processing unit 12 is used to quantize the multiple sets of offsets and the observation points based on the quantization step size, so as to obtain the multiple sets of quantized offsets and the quantized observation points.

[0182] In some embodiments of this application, the apparatus further includes a drawing unit;

[0183] The drawing unit is used to draw a histogram of the shooting scene distances corresponding to multiple reference points in the voxel grid;

[0184] The determining unit is used to determine the confidence interval of the histogram; and to determine the tree depth based on the distance range corresponding to the confidence interval.

[0185] In some embodiments of this application, the determining unit is used to determine multiple relative distances and multiple viewing directions between the multiple reference points and the observation point; determine multiple sets of 3D Gaussian positions corresponding to the multiple reference points based on the multiple reference points, the multiple sets of offsets and the multiple scaling factors; and use a multilayer perceptron to determine multiple batches of 3D Gaussian attributes corresponding to the multiple reference points based on the features of the multiple reference points, the multiple relative distances and the multiple viewing directions.

[0186] In some embodiments of this application, the device further includes an input unit;

[0187] The input unit is used to input the plurality of reference point features, the plurality of relative distances, and the plurality of viewing directions into a first multilayer perceptron to obtain a plurality of opacities corresponding to the plurality of 3D Gaussian positions; input the plurality of relative distances, the plurality of viewing directions, and the plurality of reference point features into a second multilayer perceptron to obtain a plurality of colors corresponding to the plurality of 3D Gaussian positions; input the plurality of relative distances, the plurality of viewing directions, and the plurality of reference point features into a third multilayer perceptron to obtain a plurality of four elements corresponding to the plurality of 3D Gaussian positions; and input the plurality of relative distances, the plurality of viewing directions, and the plurality of reference point features into a fourth multilayer perceptron to obtain a plurality of scaling features corresponding to the plurality of 3D Gaussian positions.

[0188] The determining unit is used to take the multiple sets of opacity, the multiple sets of colors, the multiple sets of four elements, and the multiple sets of scaling features as the multiple batches of 3D Gaussian attributes.

[0189] In some embodiments of this application, the first acquisition unit 11 is used to acquire the position of any one of the plurality of reference points, as well as any set of offsets corresponding to the position of the reference point and any scaling factor corresponding to the position of the reference point.

[0190] The determining unit is used to determine the product of any scaling factor and each offset in any set of offsets to obtain any set of products; determine the sum of any reference point position and each product in any set of products to obtain any set of 3D Gaussian positions corresponding to any reference point position; until multiple sets of 3D Gaussian positions corresponding to the multiple reference points are obtained.

[0191] In some embodiments of this application, the apparatus further includes a transmission unit;

[0192] The transmission unit is used to transmit the encoded information to the client so that the client can decode the encoded information and render it to obtain the 3D digital human.

[0193] It should be noted that, in practical applications, the first acquisition unit 11 and the first processing unit 12 can be implemented by the first processor 13 on the 3D data processing device, specifically by a CPU (Central Processing Unit), MPU (Microprocessor Unit), DSP (Digital Signal Processor), or FPGA (Field Programmable Gate Array), etc.; the data storage can be implemented by the first memory 14 on the 3D data processing device.

[0194] This application also provides a 3D data processing device, such as... Figure 5 As shown, the 3D data processing device includes: a first processor 13, a first memory 14, and a first communication bus 15. The first memory 14 communicates with the first processor 13 through the first communication bus 15. The first memory 14 stores programs executable by the first processor 13. When the program is executed, the 3D data processing method described above is executed by the first processor 13.

[0195] In practical applications, the first memory 14 can be volatile memory, such as random-access memory (RAM); or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the first processor 13.

[0196] This application provides a computer-readable storage medium having a computer program thereon, which, when executed by a first processor 13, implements the 3D data processing method as described above.

[0197] For example, embodiments of this application also provide a computer program product, including a computer program that can be executed by a first processor 13 in a 3D data processing device to complete the steps described in the aforementioned 3D data processing method.

[0198] Understandably, the 3D data processing device performs characterization, feature transformation, quantization, and encoding processing on the 3D data of the 3D digital human based on three-dimensional Gaussian sputtering to obtain multiple reference points. The amount of data for each reference point is less than the amount of data for the 3D Gaussian sphere, which reduces the amount of encoded information obtained by encoding the information corresponding to multiple reference points. This reduces the amount of data transmitted and thus improves data transmission efficiency.

[0199] Based on the same inventive concept as the 3D data processing method applied to the client described above, this application provides a client 2 corresponding to a 3D data processing method; Figure 6 A schematic diagram of the composition structure of a client provided in an embodiment of this application. Figure One The client 2 may include:

[0200] The second acquisition unit 21 is used to acquire encoded information;

[0201] The second processing unit 22 is used to perform decoding, dequantization, and rendering on the encoded information to obtain the 3D digital human. The decoding includes entropy decoding and / or video decoding.

[0202] In some embodiments of this application, the second processing unit 22 is configured to decompress the compressed video encoding information in the encoding information to obtain video encoding information; perform video decoding on the video encoding information to obtain quantized reference point features; perform entropy decoding on the first encoding information, reference point encoding information, and entropy encoding information in the encoding information to obtain network parameters in the multilayer perceptron, quantized reference points, quantized observation points, and multiple sets of quantized offsets; perform dequantization on the quantized reference point features, the quantized reference points, the quantized observation points, and the multiple sets of quantized offsets to obtain multiple reference points, multiple reference point features, observation points, and multiple sets of offsets; and perform rendering processing based on the multiple reference points, the multiple reference point features, the observation points, the multiple sets of offsets, and the network parameters to obtain the 3D digital human.

[0203] In some embodiments of this application, the client further includes a second determining unit;

[0204] The second acquisition unit 21 is used to acquire multiple relative distances between the multiple reference points and the observation point, the tree depth of the octree, and the maximum value information; the maximum value information is the maximum value in the distance range determined according to the confidence interval of the histogram, and the histogram is a histogram of shooting scene distances drawn according to multiple reference points in the voxel grid;

[0205] The second determining unit is used to determine multiple rendering layer details corresponding to the multiple reference points based on the multiple relative distances, the tree depth, and the maximum value information;

[0206] The second processing unit 22 is used to perform rendering processing based on the multiple rendering level details, the multiple reference points, the multiple reference point features, the observation point, the multiple sets of offsets and the network parameters to obtain the 3D digital human.

[0207] It should be noted that, in practical applications, the second acquisition unit 21 and the second processing unit 22 mentioned above can be implemented by the second processor 23 on the client, specifically by a CPU (Central Processing Unit), MPU (Microprocessor Unit), DSP (Digital Signal Processor), or FPGA (Field Programmable Gate Array), etc.; the data storage mentioned above can be implemented by the second memory 24 on the client.

[0208] This application also provides a client, such as... Figure 7 As shown, the client includes: a second processor 23, a second memory 24, and a second communication bus 25. The second memory 24 communicates with the second processor 23 through the second communication bus 25. The second memory 24 stores programs executable by the second processor 23. When the program is executed, the 3D data processing method described above is executed by the second processor 23.

[0209] In practical applications, the aforementioned second memory 24 can be volatile memory, such as random-access memory (RAM); or non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the second processor 23.

[0210] This application provides a computer-readable storage medium having a computer program thereon, which, when executed by a second processor 23, implements the 3D data processing method as described above.

[0211] For example, embodiments of this application also provide a computer program product, including a computer program that can be executed by a second processor 23 in a client to complete the steps described in the aforementioned 3D data processing method.

[0212] It is understandable that when a client receives encoded information from multiple reference points transmitted by a 3D data processing device, the amount of data from each reference point is less than that from a 3D Gaussian sphere. This results in a smaller amount of encoded information obtained by encoding the information corresponding to multiple reference points, which reduces the amount of data transmitted and thus improves data transmission efficiency.

[0213] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.

[0214] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure One One or more processes and / or boxes Figure One A device that provides the functions specified in one or more boxes.

[0215] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure One One or more processes and / or boxes Figure One The function specified in one or more boxes.

[0216] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure One One or more processes and / or boxes Figure One The steps of the function specified in one or more boxes.

[0217] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application.

Claims

1. A 3D data processing method, characterized in that, The method includes: Acquire 3D data of 3D digital humans based on 3D Gaussian sputtering 3DGS; The 3D data is sequentially subjected to characterization generation processing, feature transformation processing, quantization processing, and encoding processing to obtain encoded information. The encoding processing includes entropy encoding processing and / or video encoding processing.

2. The method according to claim 1, characterized in that, The 3D data is sequentially subjected to characterization generation processing, feature transformation processing, quantization processing, and encoding processing to obtain encoded information, including: The 3D data is subjected to characterization generation processing to obtain the voxel mesh, multiple reference points, and multiple reference point features obtained after the characterization generation processing. Based on the voxel grid and the plurality of reference points, feature transformation processing is performed to obtain the hierarchical detail information corresponding to the plurality of reference points; The multiple reference points and their features are quantized to obtain quantized reference points and their features. The quantized reference point and its features are encoded to obtain the encoded information. Encoding the quantized reference point includes encoding the quantized reference point based on the hierarchical detail information.

3. The method according to claim 2, characterized in that, The process of obtaining the voxel mesh, multiple reference points, and multiple reference point features obtained after the characterization generation process also includes: Obtain multiple sets of offsets corresponding to the multiple reference points obtained after the representation generation process; Accordingly, the step of quantizing the plurality of reference points and the plurality of reference point features to obtain quantized reference points and quantized reference point features further includes: Obtain observation points; The multiple sets of offsets and the observation points are quantized to obtain quantized observation points and multiple sets of quantized offsets; Accordingly, encoding the quantized reference point and the quantized reference point features to obtain the encoded information includes: Obtain the client's task requirements and / or the network conditions of the transmission network; Adjust the encoder parameters based on the task requirements and / or the network conditions to obtain the adjusted encoder parameters; The quantized information is encoded based on the adjusted encoder parameters to obtain the encoded information, which includes the quantized reference point, the quantized reference point features, the quantized observation point, and the multiple sets of quantized offsets.

4. The method according to claim 3, characterized in that, The process of encoding the quantized information based on the adjusted encoder parameters to obtain the encoded information includes: Based on the adjusted encoder parameters and the hierarchical detail information, the quantized reference point is entropy encoded to obtain reference point encoding information. Based on the adjusted encoder parameters, entropy coding is performed on the quantized observation points and the multiple sets of quantized offsets to obtain entropy coding information. Based on the adjusted encoder parameters, the quantized reference point features are video encoded to obtain video encoded information; The video encoding information is compressed to obtain compressed video encoding information; The entropy coding information, the reference point coding information, and the compressed video coding information are used as the coding information.

5. The method according to claim 4, characterized in that, The step of using the entropy coding information, the reference point coding information, and the compressed video coding information as the coding information includes: Acquire a multilayer perceptron and the network parameters in the multilayer perceptron; the multilayer perceptron includes a first multilayer perceptron, a second multilayer perceptron, a third multilayer perceptron, and a fourth multilayer perceptron; The network parameters are entropy encoded to obtain the first encoded information; The first encoding information, the entropy encoding information, the reference point encoding information, and the compressed video encoding information are used as the encoding information.

6. The method according to claim 2, characterized in that, The characterization and generation process for the 3D data includes: The 3D data is converted into the voxel mesh; Using the center point of each voxel in the voxel grid as a reference point, multiple reference points and multiple sets of reference point attributes corresponding to the multiple reference points are obtained; the multiple sets of reference point attributes include multiple reference point features, multiple scaling factors and multiple sets of offsets.

7. The method according to claim 6, characterized in that, The step of using the center point of each voxel in the voxel grid as a reference point to obtain the plurality of reference points includes: Using the center point of each voxel in the voxel grid as a reference point, multiple initial reference points are obtained; Obtain the task requirements transmitted by the client; The number of the multiple initial reference points is adjusted based on the task requirements to obtain the multiple reference points.

8. The method according to claim 2, characterized in that, The feature transformation process based on the voxel mesh and the plurality of reference points to obtain the hierarchical detail information corresponding to the plurality of reference points includes: The tree depth is determined based on the aforementioned reference points; An octree is constructed based on the voxel grid and the tree depth; Based on the positions of the multiple reference points, the multiple reference points are assigned to the octree to obtain the hierarchical detail information corresponding to the multiple reference points.

9. The method according to claim 3, characterized in that, The step of quantizing the plurality of reference points and their features to obtain quantized reference points and their features includes: Obtain the quantization step size and preset quantization method; The multiple reference points are quantized based on the quantization step size to obtain the quantized reference points. The features of the multiple reference points are quantized using a preset quantization method to obtain the quantized reference point features; Accordingly, the quantization of the multiple sets of offsets and the observation points to obtain quantized observation points and multiple sets of quantized offsets includes: The multiple sets of offsets and the observation points are quantized based on the quantization step size to obtain the multiple sets of quantized offsets and the quantized observation points.

10. The method according to claim 8, characterized in that, Determining the tree depth based on the multiple reference points includes: Draw a histogram of the shooting scene distances corresponding to multiple reference points in the voxel grid; Determine the confidence intervals of the histogram; The tree depth is determined based on the distance range corresponding to the confidence interval.

11. The method according to claim 6, characterized in that, After obtaining the plurality of reference points and the plurality of sets of reference point attributes corresponding to the plurality of reference points by taking the center point of each voxel in the voxel mesh as a reference point, the method further includes: Determine multiple relative distances and multiple viewing directions between the multiple reference points and observation points; Based on the multiple reference points, the multiple sets of offsets, and the multiple scaling factors, determine the multiple sets of 3D Gaussian positions corresponding to the multiple reference points; Using a multilayer perceptron, multiple batches of 3D Gaussian attributes corresponding to the multiple reference points are determined based on the features of the multiple reference points, the multiple relative distances, and the multiple viewpoint directions.

12. The method according to claim 11, characterized in that, The method of using a multilayer perceptron to determine multiple batches of 3D Gaussian attributes corresponding to the multiple reference points based on the features of the multiple reference points, the multiple relative distances, and the multiple viewpoint directions includes: The multiple reference point features, the multiple relative distances, and the multiple viewpoint directions are input into the first multilayer perceptron to obtain multiple sets of opacities corresponding to the multiple sets of 3D Gaussian positions; The multiple relative distances, multiple viewing directions, and multiple reference point features are input into the second multilayer perceptron to obtain multiple sets of colors corresponding to the multiple sets of 3D Gaussian positions; The multiple relative distances, multiple viewing directions, and multiple reference point features are input into the third multilayer perceptron to obtain multiple sets of four elements corresponding to the multiple sets of 3D Gaussian positions; The multiple relative distances, multiple viewing directions, and multiple reference point features are input into the fourth multilayer perceptron to obtain multiple sets of scaling features corresponding to the multiple sets of 3D Gaussian positions; The multiple sets of opacity, multiple sets of color, multiple sets of four elements, and multiple sets of scaling features are used as the multiple batches of 3D Gaussian attributes.

13. The method according to claim 11, characterized in that, The step of determining multiple sets of 3D Gaussian positions corresponding to the multiple reference points based on the multiple reference points, the multiple sets of offsets, and the multiple scaling factors includes: Obtain the position of any one of the plurality of reference points, as well as any set of offsets and any scaling factor corresponding to the position of any one reference point. Determine the product of any scaling factor with each offset in any set of offsets to obtain any set of products; The sum of each product of any reference point position and any set of products is determined to obtain any set of 3D Gaussian positions corresponding to any reference point position; this process continues until multiple sets of 3D Gaussian positions corresponding to the multiple reference points are obtained.

14. The method according to claim 1, characterized in that, After sequentially performing characterization generation processing, feature transformation processing, quantization processing, and encoding processing on the 3D data to obtain encoded information, the method further includes: The encoded information is transmitted to the client so that the client can decode the encoded information and render it to obtain the 3D digital human.

15. A 3D data processing method, characterized in that, The method includes: Obtain encoding information; The encoded information is decoded, dequantized, and rendered to obtain a 3D digital human. The decoding includes entropy decoding and / or video decoding.

16. The method according to claim 15, characterized in that, The process of decoding, dequantizing, and rendering the encoded information to obtain a 3D digital human includes: The compressed video encoding information in the encoding information is decompressed to obtain the video encoding information; The video encoding information is subjected to video decoding processing to obtain quantized reference point features; Entropy decoding is performed on the first encoded information, the reference point encoded information and the entropy encoded information in the encoded information to obtain the network parameters, the quantized reference point, the quantized observation point and multiple sets of quantized offsets in the multilayer perceptron. The quantized reference point features, the quantized reference points, the quantized observation points, and multiple sets of quantized offsets are dequantized to obtain multiple reference points, multiple reference point features, observation points, and multiple sets of offsets. The 3D digital human is obtained by rendering based on the multiple reference points, the features of the multiple reference points, the observation point, the multiple sets of offsets and the network parameters.

17. The method according to claim 16, characterized in that, The process of rendering based on the multiple reference points, the features of the multiple reference points, the observation point, the multiple sets of offsets, and the network parameters to obtain the 3D digital human includes: The system acquires multiple relative distances between the reference points and the observation points, the tree depth of the octree, and the maximum value information. The maximum value information is the maximum value in the distance range determined by the confidence interval of the histogram, and the histogram is a histogram of shooting scene distances drawn based on multiple reference points in the voxel grid. Based on the multiple relative distances, the tree depth, and the maximum value information, multiple rendering layer details corresponding to the multiple reference points are determined; The 3D digital human is obtained by rendering based on the multiple rendering layer details, the multiple reference points, the multiple reference point features, the observation point, the multiple sets of offsets, and the network parameters.

18. A storage medium having a computer program stored thereon, used in a 3D data processing device and a client, characterized in that, When the computer program is executed by a first processor in a 3D data processing device, it implements the method according to any one of claims 1 to 14; when the computer program is executed by a second processor in a client, it implements the method according to any one of claims 15 to 17.

19. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the first processor, it implements the method according to any one of claims 1 to 14; when the computer program is executed by the second processor, it implements the method according to any one of claims 15 to 17.