New posture and new view angle human body image rendering method based on semantic decoupling
By segmenting the human body into semantic regions and using cross-attention to propagate textures, the method addresses semantic distortion and texture confusion in dynamic 3D rendering, achieving improved rendering quality and efficiency.
Patent Information
- Application Number
- CN202510400729.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-15
AI Technical Summary
The existing monocular image-driven dynamic human body rendering method has semantic distortion and texture confusion problems in invisible areas, making it difficult to achieve high-quality rendering of new perspectives and new posture human body image.
By dividing the human body image into several semantic regions, using the cross attention mechanism and mask autoencoder, the semantics and textures of the invisible regions are predicted based on the semantics and texture information of the visible regions, and the image is rendered using ray sampling technology.
It effectively alleviates the problems of semantic distortion and texture confusion, and achieves higher quality new perspectives and new postures of human body images, with small model size and fast inference speed.
Smart Images

Figure CN120318399A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of dynamic 3D human body rendering, and in particular to a method for rendering human body images with new poses and new viewpoints based on semantic decoupling. Background Art
[0002] Dynamic 3D human body rendering technology can render human body images with new viewpoints or new poses from 2D images, and has important application values in technologies such as augmented reality and virtual reality. Traditional methods usually rely on videos or dense multi-view images, and need to optimize a model for each specific human body separately, which is difficult to be applied to real-time scenarios. Recent methods can be based on sparse viewpoints (3-4 viewpoints evenly distributed around the human body), and directly render human body images with new viewpoints or new poses through single-step reasoning, improving the flexibility of human body rendering. However, these methods perform poorly in the case of only a single image, because they mainly rely on directly querying and aggregating information from multiple viewpoints. The latest methods have started to turn to single-image-driven dynamic 3D human body rendering, and these innovative methods have realized the image synthesis of new viewpoints and new poses, and have the potential for real-time applications.
[0003] Although the single-eye image-driven dynamic human body image rendering method has made certain progress, the semantic distortion in the invisible area is still a major obstacle to achieving high-quality human body rendering. Existing methods often fail to fully consider the human body semantic information, resulting in incorrect texture information being obtained in the invisible area, thus causing semantic distortion and texture confusion. To improve the human body rendering effect, this method proposes to divide the human body into several semantic regions. In the same semantic region, the visible area of the human body is used as the texture prior for the invisible area, and texture completion is performed based on the similarity of textures between the two. Introducing semantic information helps to improve semantic rationality and enhance texture realism.
[0004] This method trains a model based on a multi-view human body image dataset. The model first infers the semantic labels of the invisible areas of the human body, completes the semantic assignment of the invisible areas, and ensures semantic rationality. Then, within the same semantic region, the texture of the visible area is accurately propagated to the invisible area to restore the texture of the invisible area, effectively alleviating the problem of texture confusion. This strategy finally demonstrates better image rendering quality than existing methods. Summary of the Invention
[0005] The purpose of the present invention is to overcome the deficiencies of the prior art, and propose a method for rendering human body images with new poses and new viewpoints based on semantic decoupling. This method can accurately infer the semantic structure of the invisible area, and then restore the texture in each semantic region respectively, solve the problems of semantic distortion texture and texture confusion, and finally demonstrate excellent human body image rendering quality under new viewpoints and new poses, with better image rendering quality, lighter model volume and faster inference speed.
[0006] To achieve the above object, the technical solution provided by the present invention is: a method for rendering a human body image with a new pose and a new perspective based on semantic decoupling, comprising the following steps:
[0007] Step1: Obtain a human body image as the original image and the original camera parameters for shooting the original image; use a human body semantic parsing algorithm based on the original image to obtain the original semantic segmentation map of the human body, and use a human body shape estimation algorithm based on the original camera parameters and the original image to obtain the original mesh model of the human body; preset a target camera parameter and a target mesh model of the human body, which are respectively used to provide new perspective and new pose information, wherein the target mesh model needs to have the same topological structure as the original mesh model;
[0008] Step2: Based on the original camera parameters and the original mesh model, use rasterization technology to divide the vertices of the original mesh model into visible vertices and invisible vertices;
[0009] Step3: Divide the original image into several semantic regions according to the original semantic segmentation map, and each semantic region corresponds to a semantic label; project the visible vertices onto the original semantic segmentation map based on the original camera parameters to obtain the semantic label of the visible vertex. Based on the assumption that there is an association between the semantic labels of different vertices, design a semantic propagation module based on the cross-attention mechanism to correct the semantic label of the visible vertex to eliminate the error introduced by the original semantic segmentation map, and at the same time predict the semantic label of the invisible vertex;
[0010] Step4: Based on the original semantic segmentation map, use an image encoder to extract the feature map and feature vector of each semantic region of the original image;
[0011] Step5: Project the visible vertices onto the feature map corresponding to the semantic region with the same semantic label as the vertex based on the original camera parameters to obtain the feature vector of the visible vertex. Based on the assumption that vertices with the same semantic label are similar in style, design a feature propagation module based on the masked autoencoder. For vertices with the same semantic label, update the feature vector of the visible vertex based on the feature vector of the visible vertex and the feature vector corresponding to the semantic region with the same semantic label, and predict the feature vector of the invisible vertex;
[0012] Step6: Bind the feature vectors of all vertices to the target mesh model, and based on the target camera parameters, use ray sampling technology to define a set of sampling points in three-dimensional space. For any sampling point, calculate the feature vector of the sampling point based on the feature vector of the target mesh model vertex and the distance between the vertex and the sampling point, then predict the color and volume density of the sampling point based on the feature vector of the sampling point, and finally use volume rendering technology to render the human body image based on the color and volume density of the sampling point to obtain the target image.
[0013] Furthermore, the specific operation steps of Step 3 are as follows:
[0014] Step 3-1: Given the set of vertex coordinates {v obs,i}, the set of vertex numbers {i}, and the subsets of the above two sets that only contain visible vertices and the original camera parameters C obs , where v obs,i represents the coordinates of the i-th vertex. When the i-th vertex is a visible vertex, i is denoted as i vis , represents the coordinates of the i vis -th vertex; project the visible vertex coordinates onto the original semantic segmentation map M obs to obtain the semantic label corresponding to this vertex
[0015]
[0016] In the formula, π(·) represents the perspective projection operation, and Φ(·) represents the nearest neighbor interpolation operation;
[0017] Step 3-2: Given the set of vertex coordinates {v obs,i}, the set of semantic labels {s i}, the set of vertex numbers {i}, and the subsets of the above three sets that only contain visible vertices where s i represents the semantic label of the i-th vertex, represents the semantic label of the i vis -th vertex. The semantic labels of visible vertices are obtained from Step 3-1, and the semantic labels of invisible vertices are initialized as the "unknown" class. Based on the semantic propagation module composed of the cross-attention mechanism and the fully connected layer, predict the probabilities that all vertices belong to J kinds of semantic labels where j represents the semantic label number:
[0018]
[0019] In the formula, Emb cls (·) represents the learnable semantic encoding, Emb xyz (·) represents the sine position encoding, [·] represents the concatenation operation, γ(·) represents the learnable position encoding, represent different fully connected layers respectively. Q, K, and V represent the query, key, and value in the attention mechanism respectively. Attn(·) represents the cross-attention mechanism, and softmax(·) represents the softmax activation function;
[0020] Step3-3: Given the probabilities that vertex i of the original mesh model belongs to J semantic labels Determine the final semantic label of this vertex based on the highest value among the probabilities
[0021]
[0022] where p i,j represents the probability that the i-th vertex belongs to the j-th semantic label.
[0023] Furthermore, the training process of the semantic propagation module is as follows:
[0024] First, construct pseudo-ground truth semantic labels: For the same human body, obtain L sets of original images and original camera parameter data taken around the human body. For each set of data, obtain the semantic labels of the visible vertices of the mesh model by Step3-1; after each vertex obtains semantic labels from different sets of data, select the semantic label that appears the most times as the pseudo-ground truth semantic label of this vertex. Second, train the semantic propagation module: Given the set of pseudo-ground truth semantic labels {s i ′} of the vertices and the probabilities that the vertices belong to the pseudo-ground truth semantic labels represents the probability that the semantic label of vertex i is s i ′, use the cross-entropy loss to train the model:
[0025]
[0026] Furthermore, the specific operation steps of Step 4 are as follows:
[0027] Step4-1: Given the original semantic segmentation map M obs , obtain the binary mask M obs,j corresponding to the semantic region of each semantic label j:
[0028]
[0029] where (u, v) represents the pixel coordinates, u is the horizontal coordinate, and v is the vertical coordinate, that is, the (u, v)-th pixel of M obs ;
[0030] Step4-2: Given the original human body image I obs , extract the image segmented by the binary mask M obs,j , and use the image encoder to obtain the feature map F j corresponding to the semantic region of each semantic label j and the feature vector f cls,j :
[0031] F j = εimg (I obs ⊙M obs,j ),f cls,j =ε cls (I obs ⊙M obs,j )
[0032] In the formula, ⊙ represents the Hadamard product, and ε img (·) represents the encoder composed of the Conv_1, BN_1, and ReLU_1 layers of ResNet-18, and ε cls (·) represents ResNet-18.
[0033] Furthermore, the specific operation steps of Step5 are as follows:
[0034] Step5-1: Given the set of vertex coordinates {v obs,i} of the original mesh model, the set of final semantic labels and the subsets of the above two sets that only contain visible vertices and the original camera parameters C obs , project the visible vertex coordinates onto the feature map of the semantic region corresponding to the final semantic label of the vertex to obtain the feature vector f corresponding to the vertex: i vis :
[0035]
[0036] In the formula, Ψ(·) represents the bilinear interpolation operation;
[0037] Step5-2: Perform the following operations for each semantic label: Given the set of vertex coordinates belonging to the semantic label j, the set of vertex feature vectors {f i j}, the set of vertex numbers {i j}, where when the i-th vertex belongs to the semantic label j, i is denoted as i j , represents the coordinate of the i j -th vertex, and f i j represents the feature vector of the i j -th vertex. The feature vectors of visible vertices are obtained from Step5-1, and the feature vectors of invisible vertices are initialized as the shared learnable feature vector f init , and based on the feature propagation module of the masked autoencoder, predict the updated feature vectors
[0038]
[0039] In the formula, represents the feature vector after the vertices are encoded, and respectively represent the feature vectors of the encoded visible vertices and invisible vertices; ε feat (·) represents the encoder component of the masked autoencoder, which consists of 3 Transformer blocks. Each block has an input dimension of 192, contains 3 feature heads, and the internal MLP has an intermediate dimension of 384; represents the decoder component of the masked autoencoder, which consists of 3 Transformer blocks. Each block has an input dimension of 96, contains 3 feature heads, and the internal MLP intermediate layer has a dimension of 192; represents the feature vector of the updated i j th vertex;
[0040] Step5-3: Combine the sets of feature vectors of the vertices with J semantic labels to obtain the set of feature vectors of the vertices
[0041] Furthermore, the specific operation steps of Step6 are as follows:
[0042] Step6-1: Based on the target camera parameters, use the ray sampling technique to define a set of sampling points in the three-dimensional space. The specific operations are as follows: First, based on the target camera parameters C tgt define a cluster of rays in the three-dimensional space, where each ray has the focus of the camera corresponding to the target camera parameters as the vertex and points towards a pixel on the imaging plane of the camera. Calculate the axis-aligned bounding box based on the target mesh model, intercept the line segment where the ray passes through the axis-aligned bounding box, sample a preset number of sampling points at equal intervals on the line segment, and finally retain the sampling points whose distance from the target mesh model is less than d;
[0043] Step6-2: Bind the feature vectors of the vertices of the original mesh model to the vertices {v tgt,i} of the target mesh model. For each sampling point x, retrieve K nearest target mesh model vertices {v tgt,k}, where k represents the serial number of the retrieved vertex. According to the distance between the sampling point and the retrieved vertex, perform a weighted sum on the vertex feature vectors where represents the feature vector of the kth retrieved vertex, to obtain the feature f(x) of the sampling point:
[0044]
[0045] Where MLP(·) represents the multi-layer perceptron used to fuse feature vectors and sinusoidal position encoding, w k For vertex v tgt,k The assigned weight,∈,represents a very small constant to prevent division by 0;
[0046] Step 6-3: Given the sampling point feature f(x), predict the color c(x) and volume density σ(x) information of the sampling point:
[0047] c(x),σ(x)=MLP NeRF (f(x))
[0048] In the formula, MLP NeRF (·) represents a multilayer perceptron used to predict information;
[0049] Step 6-4: Given the color and volume density information of each sampling point, use volume rendering technology to render the color value and opacity of each pixel of the target image, and obtain the final predicted target image based on the color value. Get the final predicted object mask based on opacity
[0050] Furthermore, the training process of the image encoder and feature propagation module is as follows:
[0051] First, construct a training data set: for the same human body, obtain L groups of original images and original camera parameter data taken around the human body. The cameras corresponding to the original camera parameters are arranged in order around the human body. Given the original camera parameter set Among them C l Represents the original camera parameters corresponding to the lth group of data; the randomly sampled view angle m∈{1,2,...,L} provides the original camera parameters C obs =C m ; As the training cycle increases, the perspective n with a larger difference from the perspective m is gradually selected to provide the target camera parameter C tgt =C n , n selection strategy is formally expressed as follows:
[0052]
[0053] Where, e is the current training cycle, q is the maximum difference between the viewing angle corresponding to the original camera parameters and the viewing angle corresponding to the target camera parameters, {e q} is an arithmetic progression about the period number, e q+1 -e q =1; e1, e q 、e q+1 They represent the minimum number of training cycles corresponding to the maximum difference in viewing angles of 1, q, and q+1, respectively; is a discrete uniform distribution within the numerical range (m - q, m + q), p is the value of the discrete uniform distribution, and mod represents the arithmetic modulo operation;
[0054] Secondly, train the model: Given the predicted target image and the target mask as well as the real target image I tgt , use the human semantic parsing algorithm based on the target image to obtain the target semantic segmentation map of the human body, and perform value quantization on the target semantic segmentation Figure 2 to obtain the target mask M tgt , use the L2 distance, structural dissimilarity D-SSIM, and learned perceptual image patch similarity LPIPS to construct an image loss function Use the L1 distance to construct a mask loss function to obtain the final loss function Train the model:
[0055]
[0056] In the formula, ‖·‖2 and ‖·‖ respectively represent the L2 distance and the L1 distance, and λ1, λ2, and λ3 are the loss function weight coefficients.
[0057] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0058] 1. The method of the present invention can, based on a given single image, obtain human body images at any new perspective and new pose through inference only without optimization.
[0059] 2. The method of the present invention effectively alleviates the semantic distortion and texture confusion phenomena existing in the prior methods, and can truly restore the semantic structure and texture details in the human body image.
[0060] 3. In the actual inference process, the model size of the method of the present invention is significantly smaller than that of the current method, and the inference speed is significantly better than that of the current method.
[0061] Generally speaking, the method of the present invention can decouple different semantic regions of the human body, separately restore the texture of the invisible regions for different semantic regions, effectively alleviate the semantic distortion and texture confusion problems, and can achieve better rendering results of human body images at new perspectives and new poses with a smaller model size and a faster inference speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is a schematic logical flow diagram of the method of the present invention.
[0063] Figure 2 is a schematic diagram of human body image rendering proposed by the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0064] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0065] As Figure 1 and Figure 2 shown, this embodiment discloses a new pose and new perspective human body image rendering method based on semantic decoupling. The specific implementation steps include:
[0066] Step1: Obtain a human body image as the original image and the original camera parameters for shooting the original image; use a human body semantic parsing algorithm based on the original image to obtain the original semantic segmentation map of the human body, and use a human body shape estimation algorithm based on the original camera parameters and the original image to obtain the original mesh model of the human body; preset a target camera parameter and a target mesh model of the human body, which are respectively used to provide new perspective and new pose information, where the target mesh model needs to have the same topological structure as the original mesh model.
[0067] Step2: Based on the original camera parameters and the original mesh model, use rasterization technology to divide the vertices of the original mesh model into visible vertices and invisible vertices.
[0068] Step3: Divide the original image into several semantic regions according to the original semantic segmentation map, and each semantic region corresponds to a semantic label; project the visible vertices onto the original semantic segmentation map based on the original camera parameters to obtain the semantic label of the visible vertex. Based on the assumption that there is an association between the semantic labels of different vertices, design a semantic propagation module based on the cross-attention mechanism to correct the semantic label of the visible vertex to eliminate the error introduced by the original semantic segmentation map, and at the same time predict the semantic label of the invisible vertex; the specific operation steps are as follows:
[0069] Step3-1: Given the vertex coordinate set {v obs,i} of the original mesh model, the vertex sequence number set {i}, and the subsets of the above two sets that only contain visible vertices and the original camera parameter C obs , where, v obs,i represents the coordinate of the i-th vertex. When the i-th vertex is a visible vertex, i is represented as i vis , represents the coordinate of the i vis -th vertex; project the visible vertex coordinates onto the original semantic segmentation map M obs to obtain the semantic label corresponding to the vertex
[0070]
[0071] In the formula, π(·) represents the perspective projection operation, and Φ(·) represents the nearest neighbor interpolation operation;
[0072] Step3-2: Given the set of vertex coordinates {v obs,i}, the set of semantic labels {s i}, the set of vertex numbers {i}, and subsets of the above three sets that only contain visible vertices {i vis}, where s i represents the semantic label of the i-th vertex, represents the semantic label of the i vis -th vertex. The semantic labels of visible vertices are obtained in Step3-1, and the semantic labels of invisible vertices are initialized to the "unknown" class. Based on the semantic propagation module composed of a cross-attention mechanism and a fully connected layer, predict the probabilities that all vertices belong to J types of semantic labels where j represents the semantic label number:
[0073]
[0074] In the formula, Emb cls (·) represents the learnable semantic encoding, Emb xyz (·) represents the sine position encoding, [·] represents the concatenation operation, γ(·) represents the learnable position encoding, represent different fully connected layers respectively, Q, K, and V represent the query, key, and value in the attention mechanism respectively, Attn(·) represents the cross-attention mechanism, and softmax(·) represents the softmax activation function;
[0075] Step3-3: Given the probabilities that vertex i of the original mesh model belongs to J types of semantic labels Determine the final semantic label of this vertex based on the highest value in the probabilities
[0076]
[0077] In the formula, p i,j represents the probability that the i-th vertex belongs to the j-th semantic label;
[0078] The training process of the semantic propagation module is as follows:
[0079] First, construct pseudo-ground-truth semantic labels: For the same human body, obtain L groups of original images and original camera parameter data taken around the human body. For each group of data, obtain the semantic labels of the visible vertices of the mesh model by Step3-1; after each vertex obtains semantic labels from different groups of data, select the semantic label that appears the most times as the pseudo-ground-truth semantic label of this vertex. Second, train the semantic propagation module: Given the set of pseudo-ground-truth semantic labels {s i'} and the probability that the vertex belongs to the pseudo-ground truth semantic label Indicates that the semantic label of vertex i is s i 's probability, using cross-entropy loss Train the model:
[0080]
[0081] Step4: Based on the original semantic segmentation map, use the image encoder to extract the feature map and feature vector of each semantic region of the original image. The specific operation steps are as follows:
[0082] Step4-1: Given the original semantic segmentation map M obs , obtain the binary mask M of the semantic region corresponding to each semantic label j obs,j :
[0083]
[0084] In the formula, (u, v) represents the pixel coordinates, u is the horizontal coordinate, and v is the vertical coordinate, that is, the (u, v)th pixel of M obs ;
[0085] Step4-2: Given the original human body image I obs , extract the image segmented by the binary mask M obs,j , and use the image encoder to obtain the feature map F j and the feature vector f cls,j of the semantic region corresponding to each semantic label j:
[0086] F j = ε img (I obs ⊙ M obs,j ), f cls,j = ε cls (I obs ⊙ M obs,j )
[0087] In the formula, ⊙ represents the Hadamard product, and ε img (·) represents the encoder composed of the Conv_1, BN_1, and ReLU_1 layers of ResNet-18, and ε cls (·) represents ResNet-18.
[0088] Step5: Project the visible vertices onto the feature map corresponding to the semantic region with the same semantic label based on the original camera parameters to obtain the feature vectors of the visible vertices; Based on the assumption that vertices with the same semantic label are similar in style, design a feature propagation module based on a masked autoencoder. For vertices with the same semantic label, update the feature vectors of the visible vertices and predict the feature vectors of the invisible vertices based on the feature vectors of the visible vertices among them and the feature vectors corresponding to the semantic regions with the same semantic label. The specific operation steps are as follows:
[0089] Step5-1: Given the set of vertex coordinates {v obs,i} of the original mesh model, the final semantic label set and the subsets of the above two sets that only contain visible vertices and the original camera parameters C obs , project the visible vertex coordinates onto the feature map of the semantic region corresponding to the final semantic label of this vertex to obtain the corresponding feature vector f of this vertex: i vis :
[0090]
[0091] where Ψ(·) represents the bilinear interpolation operation;
[0092] Step5-2: Perform the following operations for each semantic label respectively: Given the set of vertex coordinates of vertices belonging to semantic label j, the set of vertex feature vectors {f i j}, and the set of vertex numbers {i j}, where when the i-th vertex belongs to semantic label j, i can be additionally represented as i j , represents the coordinates of the i j -th vertex, and f i j represents the feature vector of the i j -th vertex. The feature vectors of the visible vertices are obtained from Step5-1, and the feature vectors of the invisible vertices are initialized as the shared learnable feature vector f init . Based on the feature propagation module of the masked autoencoder, predict the updated feature vectors
[0093]
[0094] of all vertices. Denote the feature vector after encoding of the vertices, and respectively denote the feature vectors of the visible vertices and the invisible vertices after encoding; ε feat (·) represents the encoder component of the masked autoencoder, which consists of 3 Transformer blocks. Each block has an input dimension of 192 dimensions, contains 3 feature heads, and the intermediate dimension of the internal MLP is 384 dimensions; represents the decoder component of the masked autoencoder, which consists of 3 Transformer blocks. Each block has an input dimension of 96 dimensions, contains 3 feature heads, and the intermediate layer dimension of the internal MLP is 192 dimensions; Denote the updated feature vector of the i j -th vertex;
[0095] Step5-3: Merge the sets of feature vectors of the vertices with J semantic labels to obtain the set of feature vectors of the vertices
[0096] Step6: Bind the feature vectors of all vertices to the target mesh model. Based on the target camera parameters, use the ray sampling technique to define a set of sampling points in the three-dimensional space. For any sampling point, calculate the feature vector of the sampling point based on the feature vectors of the vertices of the target mesh model and the distance between the vertex and the sampling point. Then, predict the color and volume density of the sampling point based on the feature vector of the sampling point. Finally, render the human body image using the volume rendering technique based on the color and volume density of the sampling point to obtain the target image; the specific operation steps are as follows:
[0097] Step6-1: Based on the target camera parameters, use the ray sampling technique to define a set of sampling points in the three-dimensional space. The specific operation is as follows: First, define a cluster of rays in the three-dimensional space based on the target camera parameters C tgt where each ray has the focus of the camera corresponding to the target camera parameters as the vertex and points towards a pixel on the imaging plane of the camera. Calculate the axis-aligned bounding box based on the target mesh model, intercept the line segment where the ray passes through the axis-aligned bounding box, sample a preset number of sampling points at equal intervals on the line segment, and finally retain the sampling points whose distance from the target mesh model is less than d;
[0098] Step6-2: Bind the feature vectors of the vertices of the original mesh model to the vertices {v tgt,i} of the target mesh model. For each sampling point x, retrieve K nearest vertices {v tgt,k} of the target mesh model, where k represents the serial number of the retrieved vertex. According to the distance between the sampling point and the retrieved vertex, perform weighted summation on the vertex feature vectors where represents the feature vector of the k-th retrieved vertex, to obtain the feature f(x) of the sampling point:
[0099]
[0100] Wherein, MLP(·) represents a multi-layer perceptron for fusing feature vectors and sinusoidal position encoding, and w k is the weight assigned to vertex v tgt,k , ∈ represents a very small constant to prevent division by zero;
[0101] Step6-3: Given the feature f(x) of the sampling point, predict the color c(x) and volume density σ(x) information of the sampling point:
[0102] c(x), σ(x) = MLP NeRF (f(x))
[0103] Wherein, MLP NeRF (·) represents a multi-layer perceptron for predicting information;
[0104] Step6-4: Given the color and volume density information of each sampling point, use volume rendering technology to render the color value and opacity of each pixel of the target image, and obtain the finally predicted target image based on the color value Obtain the finally predicted target mask based on the opacity
[0105] The training process of the image encoder and the feature propagation module is as follows:
[0106] First, construct a training data set: For the same human body, obtain L groups of original images and original camera parameter data taken around the human body. The cameras corresponding to the original camera parameters are arranged in an orderly manner around the human body, and a set of original camera parameters is given where C l represents the original camera parameter corresponding to the l-th group of data; randomly sample a viewing angle m ∈ {1, 2,..., L} to provide the original camera parameter C obs = C m ; As the training cycle increases, gradually select a viewing angle n with a larger viewing angle difference from the viewing angle m to provide the target camera parameter C tgt = C n , and the selection strategy of n is formally expressed as follows:
[0107]
[0108] Wherein, e is the current training cycle, q is the maximum difference between the viewing angle corresponding to the original camera parameter and the viewing angle corresponding to the target camera parameter, {e q} is an arithmetic sequence about the number of cycles, e q+1 - e q = 1; e1, e q , e q+1respectively represent the minimum number of training cycles corresponding to the maximum difference in viewing angles of 1, q, and q + 1; is a discrete uniform distribution within the numerical range (m - q, m + q), p is the value taken by the discrete uniform distribution, and mod represents the arithmetic modulo operation.
[0109] Secondly, train the model: Given the predicted target image and the target mask as well as the real target image I tgt , use the human semantic parsing algorithm based on the target image to obtain the target semantic segmentation map of the human body, and binarize the target semantic segmentation Figure 2 value to obtain the target mask M tgt , use the L2 distance, the structural dissimilarity D - SSIM, and the learned perceptual image patch similarity LPIPS to construct the image loss function Use the L1 distance to construct the mask loss function to obtain the final loss function Train the model:
[0110]
[0111] In the formula, ‖·‖2 and ‖·‖ respectively represent the L2 distance and the L1 distance, and λ1, λ2, λ3 are the loss function weight coefficients, taking 0.05, 0.05, and 0.1 respectively.
[0112] The THUman dataset contains 100 different instances. Each instance contains human body images of 20 poses captured from 24 different viewpoints. Each viewpoint provides the corresponding camera intrinsic and extrinsic parameters, and each human action provides the corresponding human body mesh model. In the experiment, 90 instances are selected as the training set and 10 instances are selected as the test set. For new viewpoint generation, in the test set, for each instance, one of the 5th, 13th, and 21st viewpoints of the 1st, 3rd, 5th, 7th, and 9th poses is selected as the input, and the other odd viewpoints of the same pose are used as the output; for new pose generation, in the test set, for each instance, one of the 4th, 10th, and 16th viewpoints of the 1st pose is selected as the input, and the even viewpoints of the 3rd, 5th, 7th, and 9th poses are used as the output.
[0113] The ZJU-MoCap dataset contains 9 different instances. Each instance contains human body images of approximately 500 poses captured from 20 different perspectives. Each perspective provides corresponding camera intrinsic and extrinsic parameters, and each human body motion provides a corresponding human body mesh model. In the experiment, 6 instances were selected as the training set and 3 instances as the test set. For new perspective generation, in the test set, for each instance, one of the 5th, 11th, and 17th perspectives of the 1st, 21st, 41st, …, 481st poses was selected as the input, and the other odd-numbered perspectives of the same pose were used as the output; for new pose generation, in the test set, for each instance, one of the 5th, 11th, and 17th perspectives of the 1st pose was selected as the input, and the odd-numbered perspectives of the 21st, 41st, …, 481st poses were used as the output.
[0114] In this experiment, the peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and learned perceptual image patch similarity (LPIPS) were used to measure the quality of image generation. To verify the effectiveness of the method, this experiment was compared with the NHP, MPS-NeRF, SHERF, and TransHuman methods, where the NHP and MPS-NeRF were modified according to the practice methods applicable to monocular image-driven provided by SHERF. The results are shown in the following table:
[0115] Table 1 - Analysis Table of Experimental Results of Image Rendering Quality
[0116]
[0117]
[0118] The experimental results can fully prove the effectiveness of this method. By predicting the semantic information of the invisible region based on the semantic information of the visible region of the human body, the human body is divided into different semantic regions; then for each semantic region, the texture information of the invisible region is predicted based on the texture information of the visible region. It effectively solves the problems of semantic distortion and texture confusion existing in the current methods and obtains more ideal human body image rendering results under new perspectives and new poses.
[0119] To verify the advantages of the designed model of this method in terms of model size and inference speed, this experiment was compared with the NHP, MPS-NeRF, SHERF, and TransHuman methods. The results are shown in the following table:
[0120] Table 2 - Analysis Table of Experimental Results of Model Size and Inference Speed
[0121] NHP MPS-NeRF TransHuman SHERF This method Model size (MB) 62.25 158.12 83.91 238.9 55.79 Inference speed (FPS) 0.151 0.778 0.708 1.504 4.007
[0122] The experimental results can fully prove that this method is more lightweight and has a faster inference speed.
[0123] Experimental conclusion: Aiming at the problems of semantic distortion and texture confusion in invisible regions caused by the lack of attention to human semantics in existing methods, the present invention proposes a new pose and new perspective human image rendering method based on semantic decoupling. Experimental evaluations on the public datasets THUman and ZJU-MoCap show that the use of this method effectively alleviates the above problems, achieving more ideal image rendering results, and outperforming existing methods in three classic image quality evaluation metrics: PSNR, SSIM, and LPIPS. The method of the present invention has good application value and application prospects in the real-time applications of virtual reality and mixed reality, and is worthy of promotion.
[0124] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention shall be equivalent replacement methods and are all included in the protection scope of the present invention.
Claims
1. A new pose and new perspective human body image rendering method based on semantic decoupling, characterized in that It includes the following steps: Step1: Obtain a human body image as the original image and the original camera parameters for shooting the original image; Based on the original image, use the human body semantic parsing algorithm to obtain the original semantic segmentation map of the human body. Based on the original camera parameters and the original image, use the human body shape estimation algorithm to obtain the original mesh model of the human body; Preset a target camera parameter and a target mesh model of the human body, which are used to provide new perspective and new pose information respectively, where the target mesh model needs to have the same topological structure as the original mesh model; Step2: Based on the original camera parameters and the original mesh model, use the rasterization technology to divide the vertices of the original mesh model into visible vertices and invisible vertices; Step3: Divide the original image into several semantic regions according to the original semantic segmentation map, and each semantic region corresponds to a semantic label; Project the visible vertices onto the original semantic segmentation map based on the original camera parameters to obtain the semantic labels of the visible vertices. Based on the assumption that there is an association between the semantic labels of different vertices, design a semantic propagation module based on the cross-attention mechanism to correct the semantic labels of the visible vertices to eliminate the errors introduced by the original semantic segmentation map, and at the same time predict the semantic labels of the invisible vertices; Step4: Based on the original semantic segmentation map, use the image encoder to extract the feature maps and feature vectors of each semantic region of the original image; Step5: Project the visible vertices onto the feature map corresponding to the semantic region with the same semantic label as the vertex based on the original camera parameters to obtain the feature vector of the visible vertex; Based on the assumption that the styles of different vertices with the same semantic label are similar, design a feature propagation module based on the masked autoencoder. For vertices with the same semantic label, based on the feature vectors of the visible vertices among them and the feature vectors corresponding to the semantic regions with the same semantic label, update the feature vectors of the visible vertices and predict the feature vectors of the invisible vertices; Step6: Bind the feature vectors of all vertices to the target mesh model. Based on the target camera parameters, use the ray sampling technology to define a set of sampling points in the three-dimensional space. For any sampling point, calculate the feature vector of the sampling point based on the feature vectors of the vertices of the target mesh model and the distance between the vertex and the sampling point. Then, based on the feature vector of the sampling point, predict the color and volume density of the sampling point. Finally, based on the color and volume density of the sampling point, use the volume rendering technology to render the human body image to obtain the target image.
2. The method for rendering a human body image with a new pose and a new perspective based on semantic decoupling according to claim 1, wherein The specific operation steps of Step3 are as follows: Step3-1: Given the set of vertex coordinates {v obs,i}, the set of vertex numbers {i}, and subsets of the above two sets that only contain visible vertices {i vis} and the original camera parameters C obs , where v obs,i represents the coordinates of the i-th vertex. When the i-th vertex is a visible vertex, i is denoted as i vis , represents the coordinates of the i vis -th vertex; project the visible vertex coordinates onto the original semantic segmentation map M obs to obtain the semantic label corresponding to the vertex In the formula, π(·) represents the perspective projection operation, and Φ(·) represents the nearest neighbor interpolation operation; Step3-2: Given the set of vertex coordinates {v obs,i}, the set of semantic labels {s i}, the set of vertex numbers {i}, and subsets of the above three sets that only contain visible vertices {i vis}, where s i represents the semantic label of the i-th vertex, represents the semantic label of the i vis -th vertex. The semantic labels of visible vertices are obtained from Step3-1, and the semantic labels of invisible vertices are initialized to the "unknown" class. Based on the semantic propagation module composed of a cross-attention mechanism and a fully connected layer, predict the probabilities that all vertices belong to J types of semantic labels where j represents the semantic label number: where, Emb cls (·) represents learnable semantic encoding, Emb xyz (·) represents sinusoidal position encoding, [·] represents the concatenation operation, γ(·) represents learnable position encoding, respectively represent different fully connected layers, Q, K, and V respectively represent the query, key, and value in the attention mechanism, Attn(·) represents the cross-attention mechanism, and softmax(·) represents the softmax activation function; Step3-3: Given the probabilities that vertex i of the original mesh model belongs to J semantic labels respectively Determine the final semantic label of this vertex based on the highest value among the probabilities where p i,j represents the probability that the i-th vertex belongs to the j-th semantic label.
3. The method for rendering a human body image with a new pose and a new perspective based on semantic decoupling according to claim 2, wherein The training process of the semantic propagation module is as follows: First, construct pseudo - true semantic labels: For the same human body, obtain L sets of original images and original camera parameter data taken around the human body. For each set of data, obtain the semantic labels of the visible vertices of the mesh model by Step3 - 1; after each vertex obtains semantic labels from different sets of data, select the semantic label that appears most frequently as the pseudo - true semantic label of this vertex. Secondly, train the semantic propagation module: Given the set of pseudo - true semantic labels {s i ′} of the vertices and the probability that the vertex belongs to the pseudo - true semantic label, i representing the probability that the semantic label of vertex i is s ′, use the cross - entropy loss to train the model:
4. The method for rendering a human body image with a new pose and a new perspective based on semantic decoupling according to claim 3, wherein The specific operation steps of Step 4 are as follows: Step4-1: Given the original semantic segmentation map M obs , obtain the binary mask M of the semantic region corresponding to each semantic label j obs,j : where (u, v) represents the pixel coordinates, u is the horizontal coordinate, and v is the vertical coordinate, that is, the (u, v)-th pixel of M obs ; Step4-2: Given the original human body image I obs , extract the image after segmentation by the binary mask M obs,j . Use the image encoder to obtain the feature map F j of the semantic region corresponding to each semantic label j and the feature vector f cls,j : F j = ε img (I obs ⊙ M obs,j ), f cls,j = ε cls (I obs ⊙ M obs,j ) In the formula, ⊙ represents the Hadamard product, and ε img (·) represents the encoder composed of the Conv_1, BN_1, and ReLU_1 layers of ResNet-18, and ε cls (·) represents ResNet-18.
5. The method for rendering a human body image with a new pose and a new perspective based on semantic decoupling according to claim 4, characterized in that The specific operation steps of Step5 are as follows: Step5-1: Given the set of vertex coordinates {v obs,i} of the original mesh model, the set of final semantic labels , and the subsets of the above two sets that only contain visible vertices , and the original camera parameter C obs , project the visible vertex coordinates onto the feature map of the semantic region corresponding to the final semantic label of this vertex to obtain the feature vector f corresponding to this vertex i vis : In the formula, Ψ(·) represents the bilinear interpolation operation; Step5-2: Perform the following operations for each semantic label: Given the set of vertex coordinates belonging to semantic label j The set of vertex feature vectors {f i j}, the set of vertex serial numbers {i j}, where when the i-th vertex belongs to semantic label j, i is denoted as i j , denotes the coordinates of the i j -th vertex, and f i j denotes the feature vector of the i j -th vertex. It can be seen that the feature vectors of visible vertices are obtained from Step5-1, and the feature vectors of invisible vertices are initialized as the shared learnable feature vector f init . Based on the feature propagation module of the masked autoencoder, predict the updated feature vectors of all vertices In the formula, represents the feature vector after the vertices are encoded, and respectively represent the feature vectors of the visible vertices and the invisible vertices after encoding; ε feat (·) represents the encoder component of the masked autoencoder, which consists of 3 Transformer blocks. Each block has an input dimension of 192, contains 3 feature heads, and the internal MLP has an intermediate dimension of 384; represents the decoder component of the masked autoencoder, which consists of 3 Transformer blocks. Each block has an input dimension of 96, contains 3 feature heads, and the internal MLP has an intermediate layer dimension of 192; represents the j feature vector of the updated i-th vertex; Step5-3: Combine the feature vector sets of the vertices of the J semantic tags to obtain the feature vector set of the vertices 6. The method for rendering a new pose and new perspective human body image based on semantic decoupling according to claim 5, wherein The specific operation steps of Step6 are as follows: Step6-1: Based on the target camera parameters, use the ray sampling technique to define a set of sampling points in the three-dimensional space. The specific operations are as follows: First, based on the target camera parameters C tgt Define a cluster of rays in the three-dimensional space, where each ray has the focus of the camera corresponding to the target camera parameters as the vertex and points towards a pixel on the imaging plane of the camera. Calculate the axis-aligned bounding box based on the target mesh model, intercept the line segments where the rays pass through the axis-aligned bounding box, sample a preset number of sampling points at equal intervals on the line segments, and finally retain the sampling points whose distance from the target mesh model is less than d; Step6-2: Bind the feature vectors of the vertices of the original mesh model to the vertices {v tgt,i} of the target mesh model. For each sampling point x, retrieve K nearest vertices {v tgt,k} of the target mesh model, where k represents the serial number of the retrieved vertices. According to the distance between the sampling point and the retrieved vertices, perform weighted summation on the vertex feature vectors to obtain the feature f(x) of the sampling point: represents the feature vector of the k-th retrieved vertex Where MLP(·) represents the multi-layer perceptron used to fuse feature vectors and sinusoidal position encoding, w k For vertex v tgt,k The assigned weight,∈,represents a very small constant to prevent division by 0; Step6-3: Given the feature f(x) of the sampling point, predict the color c(x) and volume density σ(x) information of the sampling point: c(x), σ(x) = MLP NeRF (f(x)) where MLP NeRF (·) represents a multi-layer perceptron for predicting information; Step6-4: Given the color and volume density information of each sampling point, use volume rendering technology to render the color value and opacity of each pixel of the target image, and obtain the finally predicted target image based on the color value Obtain the finally predicted target mask based on the opacity 7. The method for rendering a human body image with a new pose and a new perspective based on semantic decoupling according to claim 6, characterized in that, The training processes of the image encoder and the feature propagation module are as follows: First, construct a training data set: For the same human body, obtain L groups of original images and original camera parameter data taken around the human body. The cameras corresponding to the original camera parameters are arranged in order around the human body, and a set of original camera parameters is given. where C l represents the original camera parameter corresponding to the l-th group of data; randomly sample a view m ∈ {1, 2,..., L} to provide the original camera parameter C obs = C m ; as the training cycle increases, gradually select a view n with a larger view angle difference from view m to provide the target camera parameter C tgt = C n , and the view selection strategy is formally expressed as follows: Where \(e\) is the current training cycle, \(q\) is the maximum difference between the viewing angles corresponding to the original camera parameters and the viewing angles corresponding to the target camera parameters, \(\{e q \}\) is an arithmetic progression with respect to the number of cycles, \(e q+1 - e q = 1\); \(e_1\), \(e q \), \(e q+1 \) respectively represent the minimum training cycle numbers corresponding to the maximum viewing angle differences of 1, \(q\), and \(q + 1\); is a discrete uniform distribution within the numerical range \((m - q, m + q)\), \(p\) is the value of the discrete uniform distribution, and mod represents the arithmetic modulo operation; Secondly, train the model: Given the predicted target image and the target mask as well as the real target image I tgt , use the human semantic parsing algorithm based on the target image to obtain the target semantic segmentation map of the human body, and binarize the target semantic segmentation map to obtain the target mask M tgt , use the L2 distance, the structural dissimilarity D-SSIM, and the learned perceptual image patch similarity LPIPS to construct an image loss function Use the L1 distance to construct a mask loss function to obtain the final loss function Train the model: In the formula, ‖·‖2 and ‖·‖ respectively represent the L2 distance and the L1 distance, and λ1, λ2, and λ3 are the weight coefficients of the loss function.