Digital-human processing method and apparatus, and device, medium and product

By performing semantic analysis and representation generation on 3D data, combined with tensor decomposition and encoding compression techniques, the transmission challenges posed by the large data volume of 3D digital humans were solved, achieving efficient encoding and low-bandwidth transmission.

WO2026001469A1PCT designated stage Publication Date: 2026-01-02CHINA MOBILE COMM LTD RES INST +1

Patent Information

Application Number
PCT/CN2025/096693
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-24
Filing Date
2025-05-22
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

The large volume of 3D data and the resulting surge in bandwidth requirements for transmitting or storing 3D digital humans pose challenges to real-time applications. Achieving efficient encoding and low-bandwidth transmission of 3D digital humans is a pressing technical problem that needs to be solved.

Method used

By performing semantic analysis and representation generation on 3D data based on neural radiation field representation, feature transformation, quantization and encoding are performed. Tensor decomposition algorithm is used to reduce data dimensionality, and Metadata encoder and video encoder are used for compression to generate compact representation and encoding results.

Benefits of technology

While ensuring the quality of the rendered view, it achieves efficient encoding of 3D digital humans, reduces the amount of data required for transmission, and improves transmission efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025096693_02012026_PF_FP_ABST
    Figure CN2025096693_02012026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present disclosure are a digital-human processing method and apparatus, and a device, a medium and a product. The method comprises: performing semantic analysis processing and / or representation generation processing on 3D data which is represented on the basis of a neural radiance field; and performing feature transformation, quantization processing and encoding processing on processed data.
Need to check novelty before this filing date? Find Prior Art

Description

Digital human processing methods, devices, equipment, media and products

[0001] Cross-references to related applications

[0002] This disclosure is based on and claims priority to Chinese Patent Application No. 202410818067.X, filed on June 24, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure relates to the technical field of video and 3D data transmission, and more specifically, to a digital human processing method, apparatus, device, medium, and product. Background Technology

[0004] Digital humans are virtual characters created using computer technology and artificial intelligence, possessing human appearance or behavior. These virtual characters have a wide range of applications, such as corporate promotion, education and training, virtual actors and hosts, intelligent customer service, virtual exhibitions and tour guides, IP character customization, and virtual conferences. Recently, 3D digital humans have become a research hotspot, boasting advantages such as high realism, flexibility, and cost-effectiveness, and are expected to find widespread application in entertainment, education, commerce, and science.

[0005] Compared to 2D data, 3D data is much larger, and the bandwidth required to transmit or store 3D digital human data has increased dramatically, posing a great challenge to real-time application scenarios. Achieving efficient encoding and low-bandwidth transmission of digital humans is crucial, and how to achieve efficient encoding and transmission of 3D digital humans is a technical problem that urgently needs to be solved. Summary of the Invention

[0006] This disclosure provides at least one digital human processing method, apparatus, device, medium, and product.

[0007] In a first aspect, embodiments of this disclosure provide a digital human processing method, executed by an encoding end, comprising:

[0008] Semantic analysis and / or representation generation are performed on 3D data based on neural radiation field representation. The processed data is then subjected to feature transformation, quantization, and encoding.

[0009] In some embodiments, the method further includes:

[0010] Obtain feedback on network conditions to adjust encoding parameters and / or encoding modes.

[0011] In some embodiments, the method further includes:

[0012] Obtain the first requirement determined by the terminal side based on downstream tasks;

[0013] Based on the first requirement, determine the matching encoding parameters and / or adjust the bitrate control strategy.

[0014] In some embodiments, the method further includes:

[0015] The first encoded result after the encoding process is transmitted to the decoding end; wherein, based on the quality feedback on the terminal side, it is determined whether to activate or deactivate the quality enhancement module to enhance the quality of the rendering viewpoint or 3D model.

[0016] In some implementations, when the rendering viewpoint is video, the quality enhancement module includes at least one of the following functions: frame interpolation, video super-resolution, motion blur removal, and noise reduction.

[0017] In some implementations, when the 3D model is a point cloud model or a mesh model, the quality enhancement module includes at least one of the following functions: upsampling, completion, denoising, and frame rate upconversion.

[0018] In some implementations, the semantic analysis and / or representation generation processing of the 3D data based on neural radiation field representation includes:

[0019] Semantic analysis is performed on the 3D data to obtain the two-dimensional background data and digital human data of the 3D digital human;

[0020] The digital human data is subjected to representation generation processing to obtain a compact representation of the 3D digital human; wherein the compact representation includes at least: tensor plane and network model parameters;

[0021] The processed data is determined based on the two-dimensional background data and the compact representation.

[0022] In some embodiments, the process of performing characterization generation on the digital human data to obtain a compact representation of the 3D digital human includes:

[0023] The digital human data is decomposed to obtain the neural radiation field and external parameters of the 3D digital human; wherein, the external parameters include at least: camera parameters and driving parameters of the 3D model of the 3D digital human;

[0024] The neural radiation field is converted into a feature grid;

[0025] The feature mesh features are decomposed using a tensor decomposition algorithm to obtain multiple tensor planes and network model parameters of a multilayer perceptron.

[0026] The compact representation is determined based on the multiple tensor planes and the network model parameters.

[0027] In some embodiments, the method further includes:

[0028] The external parameters and the network model parameters in the compact representation are compressed using a Metadata encoder to obtain a second encoding result, which is then transmitted to the decoding end.

[0029] In some embodiments, the method further includes:

[0030] The two-dimensional background data is compressed and encoded using a first video encoder to obtain a third encoding result;

[0031] The third encoding result is transmitted to the decoding end.

[0032] In some implementations, the encoding process for the processed data includes:

[0033] The quantized feature map obtained after the quantization process is encoded using a second video encoder.

[0034] Secondly, this disclosure provides another digital human processing method, executed by a decoding end, including:

[0035] The encoding end sends a first encoding result and a second encoding result; wherein, the first encoding result is obtained by the encoding end performing feature transformation, quantization and encoding on the processed data, the processed data is obtained by the encoding end performing semantic analysis and / or representation generation on 3D data based on neural radiation field representation, and the second encoding result is obtained by the encoding end encoding and compressing the network model parameters and external parameters of the neural radiation field, the external parameters including at least: camera parameters and driving parameters of the 3D model of the 3D digital human;

[0036] The first encoding result and the second encoding result are decoded respectively to obtain the first decoding result and the second decoding result;

[0037] Based on the first and second decoding results, reconstruction rendering is performed to obtain the rendering viewpoint and / or 3D model of the 3D digital human, and the rendering viewpoint and / or 3D model is displayed on the terminal side.

[0038] In some embodiments, the method further includes:

[0039] Obtain the third encoding result transmitted from the encoding end;

[0040] The third encoding result is decoded using a first video decoder to obtain the third decoding result; wherein, the third encoding result is obtained by encoding and compressing the two-dimensional background data in the 3D data by the encoding end;

[0041] The step of reconstructing and rendering based on the first decoding result and the second decoding result to obtain the rendering viewpoint and / or 3D model of the 3D digital human includes: reconstructing and rendering based on the first decoding result, the second decoding result and the third decoding result to obtain the rendering viewpoint and / or 3D model of the 3D digital human.

[0042] In some embodiments, the step of decoding the first encoding result and the second encoding result to obtain the first decoding result and the second decoding result includes:

[0043] The first encoding result is decoded using a representation decoder to obtain the first decoding result;

[0044] The second encoded result is decoded using a Metadata decoder to obtain the second decoded result.

[0045] In some embodiments, the step of using a representation decoder to decode the first encoding result to obtain the first decoding result includes:

[0046] The first encoding result is decoded using the second video decoder in the representation decoder to obtain a quantized feature map;

[0047] The quantized feature map is dequantized to obtain another feature map;

[0048] Based on the feature map and the network model parameters in the second decoding result, feature reconstruction is performed to obtain the first decoding result.

[0049] Thirdly, embodiments of this disclosure provide a digital human processing device, disposed at an encoding end, comprising:

[0050] The processing unit is used to perform semantic analysis and / or characterization generation on 3D data based on neural radiation field characterization.

[0051] The feature transformation, quantization, and encoding unit is used to perform feature transformation, quantization, and encoding on the processed data.

[0052] Fourthly, embodiments of this disclosure provide a digital human processing device, disposed at a decoding end, comprising:

[0053] An acquisition unit is used to acquire a first encoding result and a second encoding result sent by an encoding end; wherein, the first encoding result is obtained by the encoding end performing feature transformation, quantization processing and encoding processing on the processed data, the processed data is obtained by the encoding end performing semantic analysis processing and / or representation generation processing on 3D data based on neural radiation field representation, and the second encoding result is obtained by the encoding end encoding and compressing the network model parameters and external parameters of the neural radiation field, the external parameters including at least: camera parameters and driving parameters of the 3D model of the 3D digital human;

[0054] A decoding unit is used to decode the first encoding result and the second encoding result respectively to obtain a first decoding result and a second decoding result;

[0055] The reconstruction rendering unit is used to perform reconstruction rendering based on the first decoding result and the second decoding result to obtain the rendering viewpoint and / or 3D model of the 3D digital human, and to display the rendering viewpoint and / or 3D model on the terminal side.

[0056] Fifthly, embodiments of this disclosure also provide an electronic device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps in any of the possible implementations of the first aspect or the second aspect described above are performed.

[0057] In a sixth aspect, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of any of the possible implementations of the first or second aspect described above.

[0058] In a seventh aspect, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of any possible implementation of the first aspect or the second aspect.

[0059] This disclosure provides a digital human processing method, apparatus, device, medium, and product. In this embodiment, firstly, semantic analysis and / or representation generation processing are performed on 3D data based on neural radiation field characterization to obtain processed data; feature transformation is performed on the processed data to obtain a feature map of the 3D digital human; the feature map is sequentially quantized to obtain a quantized feature map; and the quantized feature map is encoded to obtain a first encoding result.

[0060] In the above embodiments, the method of performing feature transformation, quantization and encoding on the processed 3D data based on neural radiation field representation can achieve efficient encoding of 3D digital humans while ensuring the quality of the rendered view, reducing the amount of data required to transmit 3D digital humans, thereby improving the efficient transmission of 3D digital humans.

[0061] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0062] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0063] Figure 1 shows a flowchart of a digital human processing method provided in an embodiment of this disclosure;

[0064] Figure 2 shows a flowchart illustrating the encoding process of an encoder provided in an embodiment of this disclosure;

[0065] Figure 3 shows a flowchart of a frame-by-frame characterization generation method provided in an embodiment of this disclosure;

[0066] Figure 4 shows a flowchart of a keyframe-based representation generation provided by an embodiment of this disclosure;

[0067] Figure 5 shows a flowchart of another digital human processing method provided in an embodiment of this disclosure;

[0068] Figure 6 shows a decoding flowchart of a characterization decoder provided in an embodiment of this disclosure;

[0069] Figure 7 shows a schematic diagram of a digital human processing system provided in an embodiment of the present disclosure;

[0070] Figure 8 shows a schematic diagram of a digital human processing device provided in an embodiment of the present disclosure;

[0071] Figure 9 shows a schematic diagram of another digital human processing device provided in an embodiment of the present disclosure;

[0072] Figure 10 shows a schematic diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0073] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0074] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0075] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0076] Digital humans are virtual characters created using computer technology and artificial intelligence, possessing human appearance or behavior. These virtual characters have a wide range of applications, such as corporate promotion, education and training, virtual actors and hosts, intelligent customer service, virtual exhibitions and tour guides, IP character customization, and virtual conferences. Recently, 3D digital humans have become a research hotspot, boasting advantages such as high realism, flexibility, and cost-effectiveness, and are expected to find widespread application in entertainment, education, commerce, and science.

[0077] Compared to 2D data, 3D data is much larger, and the bandwidth required to transmit or store 3D digital human data has increased dramatically, posing a great challenge to real-time application scenarios. Achieving efficient encoding and low-bandwidth transmission of digital humans is crucial, and how to achieve efficient encoding and transmission of 3D digital humans is a technical problem that urgently needs to be solved.

[0078] Based on the above research, this disclosure provides a digital human processing method, apparatus, device, medium, and product. In this disclosure, firstly, semantic analysis and / or representation generation processing are performed on 3D data based on neural radiation field representation; subsequently, feature transformation, quantization, and encoding processing can be performed on the processed data.

[0079] In the above embodiments, by performing feature transformation, quantization, and encoding on the processed 3D data based on neural radiation field characterization, efficient encoding of 3D digital humans can be achieved while ensuring the quality of the rendered view, reducing the amount of data required for transmitting 3D digital humans, thereby improving the efficient transmission of 3D digital humans.

[0080] To facilitate understanding of this embodiment, a digital human processing method disclosed in this disclosure will first be described in detail. The executing entity of the digital human processing method provided in this disclosure is generally an electronic device with a certain computing power. In some possible implementations, the digital human processing method can be implemented by a processor calling computer-readable instructions stored in memory.

[0081] Referring to Figure 1, which is a flowchart of a digital human processing method provided in an embodiment of this disclosure, the method is applied to the encoding end and includes steps S101 to S102, wherein:

[0082] S101: Perform semantic analysis and / or characterization generation on 3D data based on neural radiation field characterization.

[0083] In this embodiment of the disclosure, the decoding end can acquire an input signal, which may contain 3D digital data; wherein, the 3D data may include digital human data of the 3D digital human and two-dimensional background data corresponding to the 3D digital human, such as 2D background data and / or material data corresponding to the 3D digital human.

[0084] Here, the neural radiation field is used to indicate the color and density of each point on the 3D digital human in each viewing direction.

[0085] Neural Radiance Fields (NeRF) takes the position coordinates of a sampling point in 3D space and the viewing direction as input, and uses a multilayer perceptron (MLP) to predict information about the sampling point, including density, color, and other information. In essence, a NeRF is a function that maps any 3D position coordinates (x, y, z) and viewing direction d to its volume density σ and view-dependent color c (color), supporting differentiable rays for volume rendering.

[0086] In this embodiment of the disclosure, after obtaining 3D data, semantic analysis and / or representation generation processing can be performed on the 3D data to obtain processed data; wherein, by performing semantic analysis processing on the 3D data, 2D background data and / or material data corresponding to the 3D digital human, as well as digital human data based on neural radiation field representation, can be obtained.

[0087] After obtaining digital human data through semantic analysis, representation generation processing can be performed on the digital human data to obtain a compact representation. At this point, a compact representation of the processed data can be obtained based on 2D background data and / or source material, etc. Through representation generation processing, high-dimensional data can be decomposed into combinations of low-dimensional data, thereby reducing the complexity of data computation and improving computational efficiency. That is to say, in this embodiment of the disclosure, the data dimension of the processed data is lower than the dimension of the 3D data based on neural radiation fields.

[0088] S102: Perform feature transformation, quantization, and encoding on the processed data.

[0089] After obtaining the processed data, feature transformation can be performed on the processed data to obtain multiple feature maps of the 3D digital human, wherein the multiple feature maps can form a corresponding feature map sequence.

[0090] Subsequently, the feature maps can be sequentially quantized to obtain quantized feature maps; and the quantized feature maps can then be encoded. Specifically, a second video encoder can be used to encode the quantized feature maps obtained after the quantization process.

[0091] In the above embodiments, by performing feature transformation, quantization, and encoding on the processed 3D data based on neural radiation field characterization, efficient encoding of 3D digital humans can be achieved while ensuring the quality of the rendered view, reducing the amount of data required for transmitting 3D digital humans, thereby improving the efficient transmission of 3D digital humans.

[0092] In an optional implementation, the above steps of performing semantic analysis and / or representation generation on 3D data based on neural radiation field representation include the following steps:

[0093] Step S11: Perform semantic analysis on the 3D data to obtain the two-dimensional background data and digital human data of the 3D digital human;

[0094] Step S12: Perform a representation generation process on the digital human data to obtain a compact representation of the 3D digital human; wherein the compact representation includes at least: tensor plane and network model parameters;

[0095] Step S13: Determine the processed data based on the two-dimensional background data and the compact representation.

[0096] In this embodiment of the disclosure, after acquiring the input signal, the encoding end can perform semantic analysis processing on the input signal to obtain semantic analysis results. This semantic analysis processing can separate the foreground and background of the input signal, thereby separating data containing digital human data and data containing two-dimensional background data. The two-dimensional background data can be 2D background data and / or source material data of the 3D digital human.

[0097] It should be noted that if the input signal does not contain two-dimensional background data, then digital human data can be directly obtained through semantic analysis.

[0098] After obtaining digital human data through semantic analysis, representation generation processing can be performed on the digital human data to obtain a compact representation. For example, the digital human data can be decomposed into a tensor plane and network model parameters, where the network model parameters can be understood as the model parameters of a multilayer perceptron (MLP) corresponding to the neural radiation field. Finally, the two-dimensional background data and the compact representation can be determined as the processed data described above.

[0099] Based on the technical solutions described in steps S11 to S13 above, the method further includes the following steps:

[0100] Step S14: Use the first video encoder to compress and encode the two-dimensional background data to obtain the third encoding result;

[0101] Step S15: Transmit the third encoding result to the decoding end.

[0102] In this embodiment of the disclosure, when the two-dimensional background data is determined based on semantic analysis processing, the two-dimensional background data can be compressed and encoded using a first video encoder to obtain a third encoding result; then, the third encoding result can be transmitted to the decoding end.

[0103] Here, the following first video encoders can be used to compress and encode the two-dimensional background data: H.264, H.265, H.266, AVS3, AV1, VP9, ​​etc. Alternatively, a neural network-based first video encoder can also be used to compress and encode the two-dimensional background data. No specific limitations are made here; the choice is based on what is feasible.

[0104] The above processing method can separate digital human data and two-dimensional background data. Different encoding methods can be used for different types of signals. Therefore, the above separation method provides a technical basis for data compression.

[0105] In an optional implementation, the above steps perform characterization generation processing on the digital human data to obtain a compact representation of the 3D digital human, specifically including the following steps:

[0106] S21, decompose the digital human data to obtain the neural radiation field and external parameters of the 3D digital human; wherein, the external parameters include at least: camera parameters and driving parameters of the 3D model of the 3D digital human.

[0107] In this embodiment of the disclosure, the digital human data includes neural radiation fields and external parameters; wherein, the external parameters may include camera parameters and driving parameters of the 3D model of the 3D digital human. For example, the driving parameters may be understood as the pose information of a specified skeleton of the 3D digital human, or the pose information of the facial feature points of the 3D digital human.

[0108] Here, the digital human data can be decomposed based on the data attributes of the various data contained within it, thereby obtaining the neural radiation field and external parameters. In specific implementation, the digital human data can be decomposed into neural radiation field and external parameters according to the data name of each data type.

[0109] The above processing method enables further decomposition of digital human data, thereby obtaining neural radiation fields and external parameters, providing a basis for the encoding and compression of neural radiation fields.

[0110] Step S22: Convert the neural radiation field into a feature grid;

[0111] Step S23: Decompose the feature mesh using a tensor decomposition algorithm to obtain multiple tensor planes and network model parameters of a multilayer perceptron;

[0112] Step S24: Determine the compact representation based on the plurality of tensor planes and the network model parameters.

[0113] As described above, the neural radiation field is essentially a function that maps any 3D position coordinates (x, y, z) and viewing direction d to its volume density σ and view-related color c (color). Therefore, such a function can be modeled using a conventional feature grid with multi-channel features (per voxel).

[0114] Based on this, after separating the neural radiation field of the 3D digital human from the digital human data, the neural radiation field can be converted into a 3D voxel grid of features, i.e., a feature grid. This 3D voxel grid can then be further decomposed to generate a more compact representation.

[0115] In practice, tensor decomposition algorithms can be used to decompose the 3D feature voxel mesh, resulting in the following compact representation: multiple tensor planes, target vectors, and network model parameters of a multilayer perceptron. By using tensor decomposition algorithms to decompose the 3D feature voxel mesh, the data dimensionality can be reduced, thereby reducing the computational load.

[0116] Here, since the tensor plane accounts for a large proportion of the training model, the tensor plane can be converted into a feature map compatible with the format of the second video encoder, namely the feature map of the 3D digital human described in S102 above.

[0117] In the embodiments of this disclosure, the tensor decomposition algorithm may employ CANDECOMP / PARAFAC (CP) decomposition or vector-matrix (VM) decomposition. In addition, other tensor decomposition algorithms may be employed. For example, the neural radiation field may be converted into multiple planes, and then feature map conversion and video encoding processing may be performed on these multiple planes.

[0118] Below, using the VM algorithm as an example, we will introduce the decomposition principle of 3D feature voxel mesh:

[0119] Here, the 3D feature voxel grid can be viewed as a four-dimensional tensor of size (X,Y,Z,C), where (X,Y,Z) represents the resolution size of the 3D feature voxel grid and C represents the dimension of the feature vector stored in each 3D feature voxel grid.

[0120] For example, for any 3D tensor T∈R I×J×K It can be decomposed into a set of matrices and vectors. The above 3D tensor can be decomposed into:

[0121] For the r-th component of each tensor pattern, the factorization vectors of the three tensor patterns are respectively defined as follows: Let R be the matrix factor. The superscripts 1, 2, and 3 represent tensor modes. These three tensor modes naturally correspond to the three spatial axes X, Y, and Z. Therefore, the 3D scene can be directly represented using X, Y, and Z. Under this premise, considering R1 = R2 = R3 = R for most scenes, the above formula can be further simplified to the following formula:

[0122] In this embodiment of the disclosure, a tensor decomposition algorithm is used to reduce the dimensionality of high-dimensional data, decomposing a three-dimensional tensor into multiple low-rank tensor components, i.e., vector and matrix representations. In actual processing, the four-dimensional tensor can be represented by these low-rank tensor components, thereby reducing the storage space of this data and improving the computation speed.

[0123] Furthermore, the 3D feature voxel mesh is segmented into a geometric mesh and an appearance mesh via feature channels. Then, the volume density σ and view-dependent color c are modeled separately using the tensor decomposition algorithm described above. Specifically, it is divided into a geometric mesh G via feature channels. σ ∈R I×J×K and appearance grid G c ∈R I×J×K×P Where I, J, and K represent the feature resolutions of the X, Y, and Z axes, respectively, and P represents the number of channels for the appearance features. Here, the volume density σ and the view-dependent color c can be modeled separately, as shown in the following formulas:

[0124] Here, because the appearance mesh has one more dimension than the geometric mesh—the feature channel dimension—its rank is lower. Therefore, vector b is used during decomposition. r express.

[0125] Through the operations described above, the 3D feature voxel mesh can be decomposed into multiple tensor planes, vectors, and MLP network model parameters. Here, the tensor planes account for 99.2% of the training model size. Therefore, the tensor planes can be converted into feature maps compatible with video encoding formats, and then compressed using a second video encoder. This process is described in the following embodiments.

[0126] Based on the technical solutions described in steps S21 to S24 above, the method further includes the following steps:

[0127] Step S31: Use the Metadata encoder to compress the external parameters and the network model parameters in the compact representation to obtain a second encoding result, and transmit the second encoding result to the decoding end.

[0128] In this embodiment of the disclosure, after obtaining the compact representation through the decomposition described above, different encoders can be used to encode the compact representation. For example, for the tensor plane in the compact representation, a representation encoder can be used to perform feature transformation, quantization, and encoding on the tensor plane in the compact representation to obtain a first encoding result. In addition, a metadata encoder can be used to encode and compress the network model parameters and extrinsic parameters in the compact representation to obtain a second encoding result; and the second encoding result is transmitted to the decoding end.

[0129] Here, network model parameters can be understood as network parameters such as MLP, while external parameters include camera parameters (pose, intrinsic and extrinsic parameters, etc.), motion parameters, and other data. Among them, motion parameters refer to the application scenario where only keyframes (refer to Figure 4) and motion parameters are transmitted, and digital human-driven operation is achieved at the decoding end.

[0130] Metadata encoders can use Huffman coding, arithmetic coding, etc. Note that there are no restrictions on the encoder here, as long as it is a lossless compression method, it is within the protection scope of this application.

[0131] In the above embodiments, by compressing the parameters of the tensor plane and the network model, the efficient encoding of the 3D digital human can be achieved while ensuring the quality of the rendered view, reducing the amount of data required to transmit the 3D digital human, thereby improving the efficient transmission of the 3D digital human.

[0132] The encoding process at the encoding end is described below with reference to Figure 2. This specific process can be described as follows:

[0133] Step S1: Convert each tensor plane into a feature map to obtain multiple feature maps;

[0134] Step S2: Perform feature quantization on each feature map to obtain multiple quantized feature maps;

[0135] Step S3: Use the second video encoder to encode multiple quantized feature maps to obtain the first encoding result.

[0136] Here, for each trained neural radiation field, N tensor planes can be decomposed, where each of the N tensor planes contains N / 2 volume densities σ and a plane corresponding to N / 2 colors c. Therefore, feature map sequences can be generated from the N tensor planes, and then each feature map sequence can be compressed individually using a second video encoder.

[0137] Here, to generate feature maps, the tensor plane of each frame is first converted into a single feature map. Then, these single feature maps are arranged in frame order to form a feature map sequence. Finally, the feature map sequence is compressed using a second video encoder, as shown in Figure 3.

[0138] After obtaining a compact representation through the representation generation process, the feature map first undergoes a feature transformation, converting the tensor plane composed of multiple channels into a single feature map. Then, each value of the single feature map is quantized, converting it from a 32-bit floating-point format to a 10-bit integer format compatible with the input format of the second video encoder, resulting in the quantized feature map described above. Finally, the quantized feature map is encoded using the second video encoder to obtain a first encoding result, which is then output.

[0139] For example, linear quantization can be performed using the minimum and maximum values ​​of the feature map. The specific quantization formula is described below:

[0140] It should be noted that, in addition to the quantization method described above, other quantization methods can also be used. This disclosure does not specifically limit these methods; any quantization method that can convert the quantized data into a 10-bit integer format compatible with the input format of the second video encoder is within the scope of protection of this application. Subsequently, the quantized feature map is compressed using a second video encoder, such as H.264, H.265, H.266, AVS3, AV1 / 2, VP9, ​​etc., to obtain the first encoding result.

[0141] In this embodiment of the disclosure, as shown in FIG3, the tensor plane in the compact representation obtained by decomposing the neural radiation field of each frame can be converted into the corresponding feature map. In this case, the extrinsic parameters do not include driving parameters. Alternatively, as shown in FIG4, only the tensor plane in the compact representation obtained by decomposing the neural radiation field of the keyframe can be converted into the corresponding feature map. In this case, driving parameters are included in the extrinsic parameters.

[0142] In an optional implementation, based on the embodiment corresponding to FIG1, the method further includes:

[0143] Step S41: Obtain feedback network conditions for adjusting encoding parameters and / or encoding mode.

[0144] In specific implementation, network conditions can be obtained; wherein, the network conditions are the transmission conditions of the network transmission module between the encoding end and the decoding end; then, the encoding parameters and / or encoding mode are adjusted according to the network conditions; wherein, the encoding parameters and / or encoding mode are information for performing the encoding process.

[0145] In this embodiment of the disclosure, before compression and compaction characterization, the network conditions of the network transmission module can be obtained, for example, the network transmission speed of the network transmission module can be determined. Then, the encoding parameters and / or encoding mode can be adjusted according to the network conditions, wherein the encoding parameters and / or encoding mode are the parameters required to perform the encoding process in the above steps.

[0146] For example, encoding parameters that match the network conditions can be determined; these encoding parameters can be parameters such as encoding parameters and bitrate. Here, the mapping relationship between network conditions and encoding parameters can be pre-defined. For instance, the value range of the network conditions and the corresponding parameter values ​​for those ranges can be set. Alternatively, the condition value points of the network conditions and the corresponding parameter values ​​for those points can be set. Another example is the construction of a mapping function that indicates the mapping relationship between network conditions and encoding parameters. Here, the network conditions can be substituted into this mapping function to obtain the encoding parameters that match the network conditions.

[0147] At this point, the encoding parameters corresponding to the network conditions, i.e., the encoding parameters and / or encoding mode, can be determined according to the mapping relationship described above.

[0148] In this embodiment of the disclosure, after determining the encoding parameters and / or encoding mode that match the network conditions, the quantized feature map can be encoded according to the encoding parameters and / or encoding mode to obtain the first encoding result.

[0149] Through the above processing method, the encoding parameters can be dynamically adjusted according to the network conditions of the network transmission module, so as to achieve efficient transmission of 3D digital humans.

[0150] In an optional implementation, based on the embodiment corresponding to FIG1, the method further includes:

[0151] Step S51: Obtain the first requirement determined by the terminal side based on the downstream task;

[0152] Step S52: Determine matching encoding parameters and / or adjust the bitrate control strategy based on the first requirement.

[0153] In this embodiment of the disclosure, the terminal side's requirements for the rendering viewpoint and / or 3D model of the 3D digital human can also be obtained, namely the first requirements, such as the playback resolution, frame rate, accuracy of 3D data, playback network status (e.g., playback via mobile data or playback via access to a wireless network), and playback environment.

[0154] At this point, the matching encoding parameters and / or bitrate control strategies can be determined based on the terminal's primary requirements for the 3D digital human.

[0155] Here, the first requirement contains at least one sub-requirement of a requirement dimension. For each sub-requirement of a dimension, corresponding requirement value points and requirement value ranges can be set. The mapping relationship between requirement value points and encoding parameters and bitrate control strategies, as well as the mapping relationship between requirement value ranges and encoding parameters and bitrate control strategies, can also be set.

[0156] In practical implementation, the required value range for each sub-requirement can be set, along with the corresponding encoding parameter values ​​and rate control strategies. Alternatively, a required value point for each sub-requirement can be set, along with the corresponding encoding parameter values ​​and rate control strategies. Another example is the construction of a mapping function that indicates the mapping relationship between sub-requirements, encoding parameters, and rate control strategies. Here, the required value of each sub-requirement can be substituted into this mapping function to obtain the matching encoding parameters and rate control strategies.

[0157] At this point, the coding parameters and / or rate control strategy matching the network conditions can be determined according to the mapping relationship described above. Then, the coding process can be executed according to the matching coding parameters and / or rate control strategy to obtain the first coding result.

[0158] Through the above processing method, the encoding parameters and / or bit rate control strategy can be dynamically adjusted according to the real-time transmission status of the network transmission module, so as to achieve efficient transmission of 3D digital humans.

[0159] In an optional implementation, based on the embodiment described in FIG1, the method further includes the following steps:

[0160] The first encoded result after encoding processing is transmitted to the decoding end; wherein, based on the quality feedback on the terminal side, it is determined whether to activate or deactivate the quality enhancement module to enhance the quality of the rendering viewpoint or 3D model.

[0161] Here, the terminal side is used to provide feedback to the quality enhancement module on the quality of the rendered viewpoint and / or 3D model obtained after decoding the first encoding result, so that the quality enhancement module can perform quality enhancement processing on the rendered viewpoint and / or 3D model after being activated.

[0162] In this embodiment of the disclosure, after the encoding end obtains a first encoding result, it can transmit the first encoding result to the decoding end. Then, the decoding end can perform decoding processing on the first encoding result to obtain a first decoding result. Finally, based on the first decoding result, the rendering viewpoint and / or 3D model of the 3D digital human can be reconstructed and rendered. Finally, the rendering viewpoint and / or 3D model can be displayed on the terminal side.

[0163] Here, a quality enhancement module can be set at the edge node. This module is used to enhance the quality of the rendered viewpoint and / or 3D model. The terminal can provide feedback on the quality of the rendered viewpoint and / or 3D model to the quality enhancement module. After being activated, the quality enhancement module can perform quality enhancement on the rendered viewpoint and / or 3D model and then send the enhanced version back to the terminal for display.

[0164] Here, the quality feedback can be determined based on the playback effect of the rendered viewpoint and / or 3D model on the terminal at historical moments. The quality feedback can be automatically generated on the terminal or entered by the user on the terminal; this feedback can be numerical or textual description, and there are no specific limitations on the type of information. For example, the quality feedback could be: average, good, bad, very good, etc.

[0165] After obtaining the quality feedback, it can be determined whether to perform quality enhancement processing on the rendering viewpoint and / or 3D model based on the feedback. For example, if the quality feedback is "poor," then it is determined that quality enhancement processing is needed for the rendering viewpoint and / or 3D model. In this case, the quality enhancement processing of the rendering viewpoint and / or 3D model can be performed according to the specified enhancement method.

[0166] Here, you can select an enhancement method that matches the quality feedback as the designated enhancement method; in addition, you can also select an enhancement method that is preset by the user as the designated enhancement method.

[0167] In this embodiment of the disclosure, after obtaining the quality feedback from the terminal side, it is also possible to detect whether the quality enhancement function is enabled. If the quality enhancement function is enabled, the step of determining whether to perform quality enhancement processing on the rendering viewpoint and / or 3D model based on the quality feedback is executed.

[0168] When the rendering viewpoint is video, the quality enhancement module includes at least one of the following functions: frame interpolation, video super-resolution, motion blur removal, and noise reduction.

[0169] When the 3D model is a point cloud model or a mesh model, the quality enhancement module includes at least one of the following functions: upsampling, completion, denoising, and frame rate upconversion.

[0170] Referring to Figure 5, which is a flowchart of a digital human processing method provided in an embodiment of this disclosure, the method includes steps S501 to S503, wherein:

[0171] S501: Obtain the first encoding result and the second encoding result sent by the encoding end; wherein, the first encoding result is obtained by the encoding end performing feature transformation, quantization processing and encoding processing on the processed data, the processed data is obtained by the encoding end performing semantic analysis processing and / or representation generation processing on 3D data based on neural radiation field representation, and the second encoding result is obtained by the encoding end encoding and compressing the network model parameters and external parameters of the neural radiation field, the external parameters including at least: camera parameters and driving parameters of the three-dimensional model of the 3D digital human.

[0172] In this embodiment, the decoding end can acquire an input signal, which may contain 3D digital data. This 3D data may include 3D digital human data and corresponding 2D background data, such as 2D background data and / or source material data. Next, the encoding end can perform feature transformation on the processed data to obtain a feature map of the 3D digital human. Finally, the feature map is sequentially quantized to obtain a quantized feature map, and the quantized feature map is encoded to obtain a first encoding result.

[0173] S502: Decode the first encoding result and the second encoding result respectively to obtain the first decoding result and the second decoding result.

[0174] S503: Based on the first decoding result and the second decoding result, perform reconstruction rendering to obtain the rendering viewpoint and / or 3D model of the 3D digital human, and display the rendering viewpoint and / or 3D model through the terminal side.

[0175] After obtaining the first encoding result and the second encoding result, the encoding end can transmit the first encoding result and the second encoding result to the decoding end; then, the decoding end can perform decoding processing on the first encoding result and the second encoding result respectively to obtain the first decoding result and the second decoding result, and render the rendering viewpoint and / or 3D model of the 3D digital human based on the first decoding result and the second decoding result. Finally, the rendering viewpoint and / or 3D model of the 3D digital human is transmitted to the terminal side to display the rendering viewpoint and / or 3D model of the 3D digital human on the terminal side.

[0176] In the above embodiments, by performing feature transformation, quantization, and encoding on the processed 3D data based on neural radiation field characterization, efficient encoding of 3D digital humans can be achieved while ensuring the quality of the rendered view, reducing the amount of data required for transmitting 3D digital humans, thereby improving the efficient transmission of 3D digital humans.

[0177] In an optional implementation, the method further includes the following steps: obtaining a third encoding result transmitted by the encoding end; using a first video decoder to decode the third encoding result to obtain a third decoding result; wherein the third encoding result is obtained by the encoding end encoding and compressing the two-dimensional background data in the 3D data.

[0178] In this embodiment of the disclosure, the encoding end can use a first video encoder to compress and encode the two-dimensional background data to obtain a third encoding result; and transmit the third encoding result to the decoding end. After obtaining the third encoding result, the decoding end can use a first video decoder to decode the third encoding result to obtain a third decoding result.

[0179] Here, a first video decoder is used to decode the third encoded result, obtaining a third decoded result. The first video decoder is of the same type as the first video encoder, and the decoding process performed by the first video decoder is the reverse process of video encoding. Specifically, decoding the third encoded result yields decoded 2D background data and / or source material data. Finally, the first decoded result, the third decoded result, and the second decoded result can be used for reconstruction rendering to obtain the rendering viewpoint and / or 3D model.

[0180] At this point, reconstruction rendering can be performed based on the first decoding result, the second decoding result, and the third decoding result to obtain the rendering viewpoint and / or 3D model of the 3D digital human.

[0181] In an optional implementation, the first encoding result and the second encoding result are decoded respectively to obtain a first decoding result and a second decoding result, specifically including:

[0182] Step S61: Use a representation decoder to decode the first encoding result to obtain the first decoding result;

[0183] Step S62: Use the Metadata decoder to decode the second encoding result to obtain the second decoding result.

[0184] In this embodiment, after obtaining the first encoding result and the second encoding result, the decoding end can use a representation decoder to decode the first encoding result to obtain the first decoding result, and use a metadata decoder to decode the second encoding result to obtain the second decoding result. Here, the metadata decoder follows the same technical approach as the metadata encoder; the metadata decoder is the reverse process of metadata encoding. The first decoding result output by the metadata decoder includes network model parameters such as MLP, camera parameters (pose, intrinsic and extrinsic parameters, etc.), driving parameters, and other data.

[0185] In an optional implementation, the above steps involve using a representation decoder to decode the first encoding result to obtain the first decoding result, specifically including:

[0186] Step S621: Use the second video decoder in the representation decoder to decode the first encoding result to obtain the quantized feature map;

[0187] Step S622: Perform inverse quantization on the quantized feature map to obtain a new feature map;

[0188] Step S623: Based on the feature map and the network model parameters in the second decoding result, feature reconstruction is performed to obtain the first decoding result.

[0189] In this embodiment of the disclosure, as shown in Figure 6, in specific implementation, the first encoding result is decoded by the second video decoder to obtain a quantized feature map; wherein, the second video decoder is of the same type as the video encoder in the representation encoder, and is the reverse process of encoding. Then, the quantized feature map undergoes dequantization processing to obtain a dequantized feature map; wherein, dequantization processing can be understood as the reverse process of quantization. Afterwards, this feature map can be input to the representation reconstruction module for representation reconstruction processing, and decoded together with the Metadata decoder to obtain network parameters such as MLP, and recover the tensor-based neural radiation field, that is, the feature reconstruction result.

[0190] After obtaining the decoded neural radiation field, a 3D model of a 3D digital human can be rendered from any viewpoint or reconstructed based on the needs of the terminal side. Here, the selected rendering method can be the NeRF differentiable volumetric rendering method.

[0191] In the above embodiments, by compressing the parameters of the tensor plane and the network model, the efficient encoding of the 3D digital human can be achieved while ensuring the quality of the rendered view, reducing the amount of data required to transmit the 3D digital human, thereby improving the efficient transmission of the 3D digital human.

[0192] The digital human processing system will be described in detail below with reference to Figure 7. As shown in Figure 7, the digital human processing system includes an encoding end, a decoding end, and a terminal side.

[0193] At the encoding end, the input signal may include a 3D digital human, corresponding 2D background data, and / or source material data, where the background or source material is not mandatory. Therefore, the input first passes through a semantic analysis module to decouple the digital human from the background or source material. Then, the digital human data enters the representation generation module to generate a compact representation of the digital human. The basic compact representation (tensor plane) is then compressed by the representation encoder; while the network model parameters and external parameters (camera parameters, driving parameters, etc.) are encoded as metadata by the metadata encoder. The background or source material is compressed by the first video encoder. Afterward, the encoding results from the first video encoder, the representation encoder, and the metadata encoder are encapsulated and distributed to the decoding end via a network transmission module.

[0194] Here, the network conditions of the network transmission module are fed back to the encoding end via signal A, and the encoding end adjusts the encoding parameters and / or encoding method based on the network conditions.

[0195] At the decoding end, the received bitstream undergoes preliminary parsing. For example, 2D background data or source material data streams are decoded by the first video decoder, the representation stream sent by the representation encoder is decoded by the representation decoder, and the metadata data stream is decoded by the metadata decoder. Then, the decoding results from the first video decoder, the representation decoder, and the metadata decoder are input to the reconstruction rendering module to synthesize the final output, thereby presenting the rendering viewpoint and / or 3D model of the 3D digital human on the terminal.

[0196] Here, based on the quality feedback C from the terminal side, it is also possible to choose whether to enable the quality enhancement module (i.e., the quality enhancement module in the above embodiment) to enhance the quality of the rendered viewpoint or 3D model. In addition to the above process, the terminal side can also feed back its first requirement to the encoding end through signal B based on downstream tasks, thereby better selecting encoding parameters and adjusting bitrate control strategies, etc.

[0197] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0198] Based on the same inventive concept, this disclosure also provides a digital human processing device corresponding to the digital human processing method. Since the principle of the device in this disclosure for solving the problem is similar to the digital human processing method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0199] Referring to Figure 8, which is a schematic diagram of a digital human processing device according to an embodiment of this disclosure, the device includes: a processing unit 81, a feature transformation, quantization, and encoding unit 82; wherein...

[0200] The processing unit 81 is used to perform semantic analysis and / or characterization generation on 3D data based on neural radiation field characterization.

[0201] The feature transformation unit 82 is used to perform feature transformation, quantization, and encoding on the processed data.

[0202] In one possible implementation, the device is further configured to: acquire feedback network conditions for adjusting encoding parameters and / or encoding modes.

[0203] In one possible implementation, the device is further configured to: acquire a first requirement determined by the terminal side based on the downstream task; and determine matching encoding parameters and / or adjust the bitrate control strategy based on the first requirement.

[0204] In one possible implementation, the device is further configured to: transmit the first encoding result after the encoding process to the decoding end; wherein, based on the quality feedback from the terminal side, it is determined whether to activate or deactivate the quality enhancement module to enhance the quality of the rendering viewpoint or 3D model.

[0205] In one possible implementation, when the rendering viewpoint is video, the quality enhancement module includes at least one of the following functions: frame interpolation, video super-resolution, motion blur removal, and noise reduction.

[0206] In one possible implementation, when the 3D model is a point cloud model or a mesh model, the quality enhancement module includes at least one of the following functions: upsampling, completion, denoising, and frame rate upconversion.

[0207] In one possible implementation, the processing unit 81 is further configured to: perform semantic analysis processing on the 3D data to obtain two-dimensional background data and digital human data of the 3D digital human; perform representation generation processing on the digital human data to obtain a compact representation of the 3D digital human; wherein the compact representation includes at least: tensor plane and network model parameters; and determine the processed data based on the two-dimensional background data and the compact representation.

[0208] In one possible implementation, the processing unit 81 is further configured to: decompose the digital human data to obtain the neural radiation field and external parameters of the 3D digital human; wherein the external parameters include at least: camera parameters and driving parameters of the 3D model of the 3D digital human; convert the neural radiation field into a feature mesh; decompose the mesh features of the feature mesh using a tensor decomposition algorithm to obtain multiple tensor planes and network model parameters of a multilayer perceptron; and determine the compact representation based on the multiple tensor planes and the network model parameters.

[0209] In one possible implementation, the device is further configured to: compress the external parameters and the network model parameters in the compact representation using a Metadata encoder to obtain a second encoding result, and transmit the second encoding result to the decoding end.

[0210] In one possible implementation, the device is further configured to: compress and encode the two-dimensional background data using a first video encoder to obtain a third encoding result; and transmit the third encoding result to a decoding end.

[0211] In one possible implementation, the quantization and encoding unit is further configured to: encode the quantized feature map obtained after the quantization process using a second video encoder.

[0212] Referring to Figure 9, which is a schematic diagram of a digital human processing device according to an embodiment of this disclosure, the device includes: an acquisition unit 91, a decoding unit 92, and a reconstruction rendering unit 93; wherein...

[0213] The acquisition unit 91 is used to acquire a first encoding result and a second encoding result sent by the encoding end; wherein, the first encoding result is obtained by the encoding end performing feature transformation, quantization processing and encoding processing on the processed data, the processed data is obtained by the encoding end performing semantic analysis processing and / or representation generation processing on 3D data based on neural radiation field representation, and the second encoding result is obtained by the encoding end encoding and compressing the network model parameters and external parameters of the neural radiation field, the external parameters including at least: camera parameters and driving parameters of the 3D model of the 3D digital human;

[0214] Decoding unit 92 is used to decode the first encoding result and the second encoding result respectively to obtain the first decoding result and the second decoding result;

[0215] The reconstruction rendering unit 93 is used to perform reconstruction rendering based on the first decoding result and the second decoding result to obtain the rendering viewpoint and / or 3D model of the 3D digital human, and to display the rendering viewpoint and / or 3D model on the terminal side.

[0216] In one possible implementation, the device is further configured to: acquire a third encoding result transmitted from the encoding end; decode the third encoding result using a first video decoder to obtain a third decoding result; wherein the third encoding result is obtained by the encoding end encoding and compressing the two-dimensional background data in the 3D data; the reconstruction rendering unit 93 is configured to: perform reconstruction rendering based on the first decoding result, the second decoding result and the third decoding result to obtain the rendering viewpoint and / or 3D model of the 3D digital human.

[0217] In one possible implementation, the decoding unit is further configured to: use a representation decoder to decode the first encoding result to obtain the first decoding result; and use a metadata decoder to decode the second encoding result to obtain the second decoding result.

[0218] In one possible implementation, the decoding unit is further configured to: decode the first encoding result using the second video decoder in the representation decoder to obtain a quantized feature map; perform inverse quantization on the quantized feature map to obtain a feature map; and perform feature reconstruction based on the feature map and the network model parameters in the second decoding result to obtain the first decoding result.

[0219] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0220] Corresponding to the digital human processing method in Figure 1, this disclosure also provides an electronic device 1000, as shown in Figure 10, which is a schematic diagram of the structure of the electronic device 1000 provided in this disclosure, including:

[0221] The system includes a processor 101, a memory 102, and a bus 103. The memory 102 stores execution instructions and includes main memory 1021 and external memory 1022. The main memory 1021, also called internal memory, temporarily stores computational data in the processor 101, as well as data exchanged with external memory such as a hard disk. The processor 101 exchanges data with the external memory 1022 through the main memory 1021. When the electronic device 1000 is running, the processor 101 communicates with the memory 102 through the bus 103, causing the processor 101 to execute the following instructions:

[0222] Semantic analysis and / or representation generation are performed on 3D data based on neural radiation field representation. The processed data is then subjected to feature transformation, quantization, and encoding.

[0223] Alternatively, follow these steps:

[0224] The encoding end sends a first encoding result and a second encoding result; wherein, the first encoding result is obtained by the encoding end performing feature transformation, quantization and encoding on the processed data, the processed data is obtained by the encoding end performing semantic analysis and / or representation generation on 3D data based on neural radiation field representation, and the second encoding result is obtained by the encoding end encoding and compressing the network model parameters and external parameters of the neural radiation field, the external parameters including at least: camera parameters and driving parameters of the 3D model of the 3D digital human;

[0225] The first encoding result and the second encoding result are decoded respectively to obtain the first decoding result and the second decoding result;

[0226] Based on the first and second decoding results, reconstruction rendering is performed to obtain the rendering viewpoint and / or 3D model of the 3D digital human, and the rendering viewpoint and / or 3D model is displayed on the terminal side.

[0227] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the digital human processing method described in the above-described method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.

[0228] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the digital human processing method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0229] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0230] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0231] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0232] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0233] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0234] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A digital human processing method, executed by an encoding end, comprising: Semantic analysis and / or representation generation are performed on 3D data based on neural radiation field representation. The processed data is then subjected to feature transformation, quantization, and encoding.

2. The method according to claim 1, wherein, The method further includes: Obtain feedback on network conditions to adjust encoding parameters and / or encoding modes.

3. The method according to claim 1 or 2, wherein, The method further includes: Obtain the first requirement determined by the terminal side based on downstream tasks; Based on the first requirement, determine the matching encoding parameters and / or adjust the bitrate control strategy.

4. The method according to any one of claims 1 to 3, wherein, The method further includes: The first encoded result after the encoding process is transmitted to the decoding end; wherein, based on the quality feedback on the terminal side, it is determined whether to activate or deactivate the quality enhancement module to enhance the quality of the rendering viewpoint or 3D model.

5. The method according to claim 4, wherein, When the rendering viewpoint is video, the quality enhancement module includes at least one of the following functions: frame interpolation, video super-resolution, motion blur removal, and noise reduction.

6. The method according to claim 4, wherein, When the 3D model is a point cloud model or a mesh model, the quality enhancement module includes at least one of the following functions: upsampling, completion, denoising, and frame rate upconversion.

7. The method according to any one of claims 1 to 6, wherein, The semantic analysis and / or representation generation processing of 3D data based on neural radiation field representation includes: Semantic analysis is performed on the 3D data to obtain the two-dimensional background data and digital human data of the 3D digital human; The digital human data is subjected to representation generation processing to obtain a compact representation of the 3D digital human; wherein the compact representation includes at least: tensor plane and network model parameters; The processed data is determined based on the two-dimensional background data and the compact representation.

8. The method according to claim 7, wherein, The process of generating a representation from the digital human data to obtain a compact representation of the 3D digital human includes: The digital human data is decomposed to obtain the neural radiation field and external parameters of the 3D digital human; wherein, the external parameters include at least: camera parameters and driving parameters of the 3D model of the 3D digital human; The neural radiation field is converted into a feature grid; The feature mesh features are decomposed using a tensor decomposition algorithm to obtain multiple tensor planes and network model parameters of a multilayer perceptron. The compact representation is determined based on the multiple tensor planes and the network model parameters.

9. The method according to claim 8, wherein, The method further includes: The external parameters and the network model parameters in the compact representation are compressed using a Metadata encoder to obtain a second encoding result, which is then transmitted to the decoding end.

10. The method according to any one of claims 7 to 9, wherein, The method further includes: The two-dimensional background data is compressed and encoded using a first video encoder to obtain a third encoding result; The third encoding result is transmitted to the decoding end.

11. The method according to any one of claims 1 to 10, wherein, The encoding process for the processed data includes: The quantized feature map obtained after the quantization process is encoded using a second video encoder.

12. A digital human processing method, executed by a decoding end, comprising: The encoding end sends a first encoding result and a second encoding result; wherein, the first encoding result is obtained by the encoding end performing feature transformation, quantization and encoding on the processed data, the processed data is obtained by the encoding end performing semantic analysis and / or representation generation on 3D data based on neural radiation field representation, and the second encoding result is obtained by the encoding end encoding and compressing the network model parameters and external parameters of the neural radiation field, the external parameters including at least: camera parameters and driving parameters of the 3D model of the 3D digital human; The first encoding result and the second encoding result are decoded respectively to obtain the first decoding result and the second decoding result; Based on the first and second decoding results, reconstruction rendering is performed to obtain the rendering viewpoint and / or 3D model of the 3D digital human, and the rendering viewpoint and / or 3D model is displayed on the terminal side.

13. The method according to claim 12, wherein, The method further includes: Obtain the third encoding result transmitted from the encoding end; The third encoding result is decoded using a first video decoder to obtain a third decoding result; wherein, the third encoding result is obtained by encoding and compressing the two-dimensional background data in the 3D data by the encoding end; The step of reconstructing and rendering based on the first decoding result and the second decoding result to obtain the rendering viewpoint and / or 3D model of the 3D digital human includes: reconstructing and rendering based on the first decoding result, the second decoding result and the third decoding result to obtain the rendering viewpoint and / or 3D model of the 3D digital human.

14. The method according to claim 12, wherein, The step of decoding the first encoding result and the second encoding result respectively to obtain the first decoding result and the second decoding result includes: The first encoding result is decoded using a representation decoder to obtain the first decoding result; The second encoded result is decoded using a Metadata decoder to obtain the second decoded result.

15. The method according to claim 14, wherein, The step of using a representation decoder to decode the first encoding result to obtain the first decoding result includes: The first encoding result is decoded using the second video decoder in the representation decoder to obtain a quantized feature map; The quantized feature map is dequantized to obtain another feature map; Based on the feature map and the network model parameters in the second decoding result, feature reconstruction is performed to obtain the first decoding result.

16. A digital human processing device, disposed at an encoding end, comprising: The processing unit is used to perform semantic analysis and / or characterization generation on 3D data based on neural radiation field characterization. The feature transformation, quantization, and encoding unit is used to perform feature transformation, quantization, and encoding on the processed data.

17. A digital human processing device, disposed at a decoding end, comprising: An acquisition unit is used to acquire a first encoding result and a second encoding result sent by an encoding end; wherein, the first encoding result is obtained by the encoding end performing feature transformation, quantization processing and encoding processing on the processed data, the processed data is obtained by the encoding end performing semantic analysis processing and / or representation generation processing on 3D data based on neural radiation field representation, and the second encoding result is obtained by the encoding end encoding and compressing the network model parameters and external parameters of the neural radiation field, the external parameters including at least: camera parameters and driving parameters of the 3D model of the 3D digital human; A decoding unit is used to decode the first encoding result and the second encoding result respectively to obtain a first decoding result and a second decoding result; The reconstruction rendering unit is used to perform reconstruction rendering based on the first decoding result and the second decoding result to obtain the rendering viewpoint and / or 3D model of the 3D digital human, and to display the rendering viewpoint and / or 3D model on the terminal side.

18. An electronic device comprising: The device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is in operation, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, they perform the steps of the digital human processing method as described in any one of claims 1 to 11, or perform the steps of the digital human processing method as described in any one of claims 12 to 15.

19. A computer-readable storage medium having a computer program stored thereon, wherein, The computer program, when run by a processor, performs the steps of the digital human processing method as described in any one of claims 1 to 11, or performs the steps of the digital human processing method as described in any one of claims 12 to 15.

20. A computer program product, said computer program product being stored in a storage medium, wherein, The program product is executed by at least one processor to implement the steps of the digital human processing method as claimed in any one of claims 1 to 11, or to implement the steps of the digital human processing method as claimed in any one of claims 12 to 15.

Citation Information

Patent Citations

  • Construction method and device of deformable neural radiation field network

    CN115909015A

  • Nerve radiation field digital human generation method, system and device

    CN116778045A

  • Virtual image synthesis method, synthesis device, equipment and medium

    CN117115331A

  • Hash coding-based digital human head reconstruction method, system, equipment and medium

    CN117893650A

  • Digital human processing method, device, equipment, medium and product

    CN118798211A

Cited By

  • Food safety interactive science popularization method and system based on AI digital human

    CN121918702A