A Point Cloud Video Upsampling Method Based on Feature Fine-Tuning
The feature tuning-based point cloud video upsampling method addresses computational challenges in real-time point cloud reconstruction by optimizing feature extraction and reconstruction stages, resulting in efficient and accurate dense point cloud generation.
Patent Information
- Application Number
- CN202211503922.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-11-29
AI Technical Summary
The existing point cloud video upsampling method is computationally expensive and difficult to achieve real-time processing, resulting in huge consumption of point cloud video transmission bandwidth.
The point cloud video upsampling method based on feature fine-tuning is adopted. Through feature extraction, feature fine-tuning and point cloud reconstruction stages, the similarity between point cloud video frames is used to reduce the calculation amount and design a lightweight model for frame-by-frame upsampling.
Efficient point cloud video upsampling is achieved, reducing computational volume, reducing bandwidth consumption, generating accurate dense point clouds, and improving processing efficiency.
Smart Images

Figure CN115761185B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision, and particularly relates to a method for upsampling point cloud videos based on feature fine-tuning. Background Art
[0002] Due to the huge volume of point cloud files, the transmission of volumetric videos requires a large amount of bandwidth consumption. However, current point cloud compression-based methods are difficult to meet the requirements of real-time transmission. Therefore, people have proposed using the method of upsampling point cloud videos. When transmitting at the server side, only low-resolution point clouds need to be transmitted, and then after receiving at the client side, the original dense point clouds are reconstructed through an upsampling model. This technology can be used simultaneously with compression-based methods to obtain video files with a smaller volume, and can avoid the huge bandwidth consumption in the transmission of point-based volumetric videos.
[0003] However, the computational complexity of the learning-based static point cloud upsampling model is too large, and it is difficult to meet the requirements of real-time processing. It is necessary to find a method that is more efficient than the method of upsampling frame by frame using the static point cloud upsampling method, and can reconstruct accurate dense point clouds within a short time. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides a method for upsampling point cloud videos based on feature fine-tuning, so as to reduce the computational complexity of the most computationally intensive feature extraction part, make the model more efficient and lightweight, and achieve the purpose of upsampling point cloud videos.
[0005] To achieve the above purpose, the technical solution of the present invention is as follows:
[0006] A method for upsampling point cloud videos based on feature fine-tuning includes a model training process and a model inference process. The model training process includes using the sparsely sampled sparse point cloud video in the training dataset as the input of the model, and obtaining the output point cloud video after upsampling; during the model training process, comparing the output point cloud video with the original dense point cloud video in the training dataset that has not been sparsely sampled, and using the reconstruction loss and the repulsion loss as loss functions to repeatedly train the model iteratively; the model inference process includes inputting any sparse point cloud video into the trained model, and obtaining the corresponding dense point cloud video after upsampling;
[0007] The process of upsampling is as follows:
[0008] S1. Feature extraction stage: The sparse point cloud video is segmented into several segments at a fixed interval. For each obtained segment, the point cloud features of the first frame of each segment are extracted using a feature extractor; then the point cloud features of the first frame are expanded to obtain the expanded point cloud features;
[0009] S2. Feature fine-tuning stage: For the current frame point cloud in each segment except the first frame, input the current frame point cloud, its previous frame point cloud, and the expanded point cloud features of the previous frame into the feature fine-tuning unit. After feature fine-tuning, obtain the point cloud features of the current frame, and then expand the point cloud features of the current frame to obtain the expanded point cloud features;
[0010] S3. Point cloud reconstruction stage: Perform point cloud reconstruction on the expanded point cloud features of each frame to obtain rough point cloud coordinates, and perform noise reduction and correction on the rough point cloud coordinates to obtain the output point cloud video after upsampling.
[0011] In the above solution, the feature extraction stage specifically includes the following process:
[0012] S11. Divide the sparse point cloud video into several segments at a fixed interval of T frames. Each segment is S = {s i} ∈ R N ×d , where s i represents the i-th frame point cloud, i represents the index of each frame in the segment, each frame point cloud has N points, and each point is a d-dimensional vector. Here, only the x, y, and z coordinates of the point cloud are considered, so d = 3;
[0013] S12. For the first frame s 1 ∈ R N×d in each segment, use the feature extractor Feat(x) of the MPU to extract the feature f 1 of the first frame s 1 ∈ R N×C , and C is the dimension of each feature vector.
[0014] In the above solution, the feature fine-tuning stage specifically includes the following process:
[0015] S21. For the i-th frame point cloud s i , first use a multi-layer perceptron to expand the features of the previous frame point cloud s i-1 , increase the dimension of the feature vector from the original N×C dimension to the rN×C dimension, where r represents the upsampling rate, to obtain the expanded point cloud features Input the current frame point cloud s i ∈ R N×d , the previous frame point cloud s i-1 ∈ R N×d , and the expanded point cloud features into the feature fine-tuning unit at the same time;
[0016] S22. Use as the feature template of the current frame point cloud s i , and use and the current frame point cloud si Perform splicing and merging, and input it into a multi-layer perceptron with a single 128-dimensional hidden layer to obtain the intermediate feature result to be searched.
[0017] S23. For the current frame of point cloud s i For each point in it, search for K nearest neighbor points in the previous frame of point cloud s i-1 Use the K-nearest neighbor search algorithm to search for K nearest neighbor points, and gather these points together to obtain the K-nearest neighbor point set. At the same time, retrieve according to the index of each neighbor point in to obtain the corresponding point cloud features.
[0018] S24. Input the obtained in S23 and the current frame of point cloud s i into the local position encoding unit. For each point Use the K-nearest neighbor search algorithm to find K nearest neighbor points. After local position encoding operation, obtain the local position encoding.
[0019]
[0020] Among them, "+" represents the merging and splicing operation of coordinate dimensions. represents the residual between K neighbor points and their center point, and MLP represents a multi-layer perceptron with a single 128-dimensional hidden layer;
[0021] Then, aggregate all the local position encodings to obtain the encoding of the i-th frame of point cloud that contains both absolute distance and relative distance information. U represents the union operation of the set, which represents the operation of vector merging or splicing here;
[0022] S25. Merge the obtained in S24 and the obtained in S23 through a multi-layer perceptron α(·) with a single 128-dimensional hidden layer, and then input the result into another multi-layer perceptron β(·) with a single 128-dimensional hidden layer to obtain the final feature of the current i-th frame of point cloud.
[0023]
[0024] Among them, "+" represents the vector merging and splicing operation.
[0025] In the above solution, the point cloud reconstruction stage specifically includes the following process:
[0026] S31. Reconstruct the expanded features through a multi-layer perceptron with three hidden layers of 128 dimensions, reconstruct the rN×C-dimensional features into rN×3-dimensional coordinates to obtain rough point cloud coordinates.
[0027] S32. For each point in , use the K-nearest neighbor search algorithm to find K nearest neighbor points among all points in . Aggregate the K nearest neighbor points obtained for each point into At the same time, find the point cloud features of the corresponding point cloud through the index among all point cloud features obtained in S25
[0028] S33. Input the point cloud features obtained in S32 into the self-attention unit for self-attention mechanism operation to obtain self-attention features
[0029]
[0030] Here, Q, K, and V are different encoding results obtained by inputting the features into three different multi-layer perceptrons with one hidden layer of 128 dimensions respectively. d K is the dimension of the feature vector K, and Softmax() represents the normalized exponential function;
[0031] At the same time, input into another multi-layer perceptron to obtain the attention matrix:
[0032]
[0033] where, M is the attention matrix, and g(·) is a multi-layer perceptron with one hidden layer of 128 dimensions;
[0034] S34. Pass the aggregated point cloud coordinates obtained in S32 through a multi-layer perceptron γ(·) with one hidden layer of 128 dimensions to obtain intermediate features
[0035]
[0036] where, γ(·) is a multi-layer perceptron network with one hidden layer of 128 dimensions;
[0037] Then, multiply the intermediate features by the attention matrix M obtained in S33 to obtain the features selected according to the features, add them to the self-attention features obtained in S33 , and at the same time use residual connection with the point cloud features obtained in S32 Finally, the final corrected point cloud coordinates h are reconstructed through a multi-layer perceptron δ(·) with two hidden layers of 128 dimensions. i :
[0038]
[0039] In the above scheme, the reconstruction loss in the model training process is as follows:
[0040] The obtained rough point cloud result and the point cloud coordinates h after noise reduction and correction i are simultaneously used to calculate the chamfer distance and the Hausdorff distance with the original dense point cloud G, and these two methods are used to measure the predicted rough point cloud result and the predicted point cloud coordinates h i and the similarity between the original dense point cloud coordinates G, and the reconstruction loss function for the two stages is obtained as:
[0041]
[0042] Among them, represents the rough point cloud without noise reduction and correction, h i represents the dense point cloud finally output by the model, G represents the correct original dense point cloud of the i-th frame, and λ1 and λ2 are parameters used to control the importance of the two stages;
[0043] L CD (·) represents the chamfer distance:
[0044]
[0045] L HD (·) represents the Hausdorff distance:
[0046]
[0047] Among them, S1 and S2 represent two point clouds that need to be calculated, selected from G, h i ; x ∈ S1 means that x belongs to the points in the point cloud S1, and y ∈ S2 means that y belongs to the points in the point cloud S2.
[0048] In the above scheme, the repulsion loss in the model training process is as follows:
[0049]
[0050] Among them, L rep is the repulsion loss function, r represents the upsampling rate, N represents the number of points in each frame of the point cloud, K(i) represents the K nearest neighbors of x i of x iDenote the i-th point in the final dense point cloud composed of rN points, x j Denote the point x i The j-th nearest neighbor point of Decreases as t increases, Decreases as t increases, where h is a hyperparameter.
[0051] In the above solution, the L2 regularization method is used in the model training process. The model is trained by minimizing the combined loss function using an end-to-end method. The overall loss function of the model is:
[0052]
[0053] Where, L rec Is the reconstruction loss function, L rep Is the repulsive loss function, λ rec 、λ rep And λ reg Are the proportionality coefficients of L rec 、L rep And w respectively, where w represents the weights of the network.
[0054] In the above solution, the specific method of the model inference process is as follows:
[0055] (1) Fix the parameters of the upsampling model D that has completed training;
[0056] (2) Take the sparse point cloud video PointCloud sparse To be upsampled as the input data and input it into the upsampling model D to generate the corresponding dense point cloud video PointCloud dense :
[0057] PointCloud dense = D(x).
[0058] Through the above technical solution, a point cloud video upsampling method based on feature fine-tuning provided by the present invention has the following beneficial effects:
[0059] (1) By designing a point cloud video upsampling method, the present invention solves the problem of huge bandwidth consumption required in point-based volumetric video transmission.
[0060] (2) The present invention utilizes the characteristic of high similarity between adjacent frames of the point cloud video and uses the feature fine-tuning method to reduce the computational amount of the most computationally intensive feature extraction part, making the model more efficient and lightweight.
[0061] (3) The present invention designs a point cloud video upsampling method based on feature fine-tuning by utilizing inter-frame redundancy, achieving a more efficient video upsampling method than processing each frame using a static point cloud upsampling model.
[0062] (4) The present invention uses a two-stage method. First, it generates a rough point cloud that pays more attention to object contours, and then through a correction and denoising unit, it generates a more accurate and dense point cloud that pays more attention to local features. In this way, the model efficiently realizes the upsampling effect of the point cloud video and finally obtains a more accurate predicted point cloud. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for description in the embodiments or the prior art.
[0064] Figure 1 Schematic diagram of a point cloud video upsampling method based on feature fine-tuning disclosed in an embodiment of the present invention;
[0065] Figure 2 Schematic diagram of the feature extractor of the MPU;
[0066] Figure 3 Schematic diagram of a multi-layer perceptron;
[0067] Figure 4 Schematic diagram of the calculation of a neuron;
[0068] Figure 5 Schematic diagram of a feature fine-tuning unit;
[0069] Figure 6 Schematic diagram of a noise reduction and correction unit. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention.
[0071] The present invention provides a point cloud video upsampling method based on feature fine-tuning, including a model training process and a model inference process. The model training process includes using the sparsely sampled sparse point cloud video in the training dataset as the input of the model, and obtaining the output point cloud video after upsampling; during the model training process, comparing the output point cloud video with the original dense point cloud video in the training dataset that has not been sparsely sampled, and using the reconstruction loss and the repulsion loss as loss functions to iteratively train the model; the model inference process includes inputting any sparse point cloud video into the trained model, and obtaining the corresponding dense point cloud video after upsampling.
[0072] Among them, as Figure 1As shown below, the upsampling process is as follows:
[0073] S1. Feature extraction stage: The sparse point cloud video is segmented into several segments at a fixed interval. For each obtained segment, the point cloud features of the first frame of each segment are extracted using a feature extractor; then the point cloud features of the first frame are expanded to obtain the expanded point cloud features.
[0074] The feature extraction stage specifically includes the following process:
[0075] S11. The sparse point cloud video is segmented into several segments at a fixed interval of T frames, and each segment is S = {s i} ∈ R N ×d , where s i represents the point cloud of the i-th frame, i represents the index of each frame in the segment, each frame of the point cloud has N points, and each point is a d-dimensional vector. Here, only the x, y, and z coordinates of the point cloud are considered, so d = 3;
[0076] S12. For the first frame s 1 ∈ R N×d of each segment, the feature extractor Feat(x) of the MPU as shown in Figure 2 is used to extract the feature f 1 of the first frame s 1 ∈ R N×C , and C is the dimension of each feature vector.
[0077] S2. Feature fine-tuning stage: For the current frame point cloud of each segment except the first frame, the current frame point cloud, its previous frame point cloud, and the expanded point cloud features of the previous frame are input into the feature fine-tuning unit together. After feature fine-tuning, the point cloud features of the current frame are obtained, and then the point cloud features of the current frame are expanded to obtain the expanded point cloud features.
[0078] The feature fine-tuning stage specifically includes the following process:
[0079] S21. For the point cloud s i of the i-th frame, first use a multi-layer perceptron to expand the features of the previous frame point cloud s i-1 , and increase the dimension of the feature vector from the original N×C dimension to the rN×C dimension, where r represents the upsampling rate, to obtain the expanded point cloud features
[0080] The multi-layer perceptron used here is a classic fully connected neural network model, which includes an input layer, a hidden layer, and an output layer, Figure 3An example of a multi-layer perceptron is illustrated, where the input layer has four neurons, the hidden layer consists of two fully connected layers in the middle, and each hidden layer has six neurons, and the output layer consists of two neurons.
[0081] Among them, the calculation of neurons can be exemplified by the first neuron in the first hidden layer, as Figure 4 shown, and can be expressed as:
[0082]
[0083] Among them, represents the value of the i-th neuron in the j-th layer, w i,j represents the weight of the j-th neuron in the i-th layer, b i represents the bias of the i-th neuron, and σ represents the activation function. Here, the Sigmoid function is used, that is:
[0084]
[0085] In each multi-layer perceptron, the weights are learned by optimizing the loss function during the training process, and the number of neurons in the input and output layers of each multi-layer perceptron varies with the dimensions of the input and output vectors, while the middle hidden layer uses a structure of 128 neurons and the number of layers is 1 - 3 layers.
[0086] Then, the point cloud s i ∈R N×d of the current frame, the point cloud s i-1 ∈R N×d of the previous frame, and the point cloud features after dilation are simultaneously input into the feature fine-tuning unit, as Figure 5 shown.
[0087] S22. Take as the feature template of the point cloud s i of the current frame, splice and merge with the point cloud s i of the current frame, and input it into a multi-layer perceptron with a hidden layer of 128 dimensions to obtain the intermediate feature result to be searched
[0088] S23. For each point in the point cloud s i of the current frame, search for K nearest neighbor points in the point cloud s i-1 of the previous frame using the K-nearest neighbor search algorithm, gather these points together to obtain the K-nearest neighbor point set At the same time, retrieve according to the index of each neighbor point in to obtain the corresponding point cloud features
[0089] S24. Take the result obtained in S23 and the current frame point cloud s i and input them into the local position encoding unit. For each point use the K-nearest neighbor search algorithm to find K nearest neighbor points and obtain the local position encoding through local position encoding operation
[0090]
[0091] where, "+" represents the merging and splicing operation of coordinate dimensions represents the residuals between the K nearest neighbor points and their center point, and MLP represents a multi-layer perceptron with a single hidden layer of 128 dimensions;
[0092] Then, aggregate all the local position encodings to obtain the encoding of the i-th frame point cloud that contains both absolute distance and relative distance information U represents the union operation of sets, and here it represents the operation of vector merging or splicing;
[0093] S25. Merge the result obtained in S24 with the result obtained in S23 through a multi-layer perceptron α(·) with a single hidden layer of 128 dimensions, and then input the result into another multi-layer perceptron β(·) with a single hidden layer of 128 dimensions to obtain the final feature of the current i-th frame point cloud
[0094]
[0095] where, "+" represents the vector merging and splicing operation.
[0096] s3. Point cloud reconstruction stage: Reconstruct the point cloud features after dilation for each frame to obtain rough point cloud coordinates, and input the rough point cloud coordinates into Figure 6 the denoising and correction unit shown in the figure for denoising and correction to obtain the upsampled output point cloud video.
[0097] The point cloud reconstruction stage specifically includes the following process:
[0098] S31. Reconstruct the dilated features through a multi-layer perceptron with three hidden layers of 128 dimensions, and reconstruct the rN×C-dimensional features into rN×3-dimensional coordinates to obtain rough point cloud coordinates
[0099] S32. For each point in obtained in S31, use the K-nearest neighbor search algorithm in Find the K nearest neighbor points among all points, and aggregate the K nearest neighbor points obtained for each point into At the same time, all the point cloud features obtained through the index in S25 Find the point cloud features corresponding to the point cloud in
[0100] S33. Input the point cloud features obtained in S32 into the self-attention unit for self-attention mechanism operation to obtain self-attention features
[0101]
[0102] Here, Q, K, and V are different encoding results obtained by inputting the features into three different hidden layers, respectively, for a multi-layer perceptron with 128 dimensions in one layer. d K is the dimension of the feature vector K, and Softmax() represents the normalized exponential function;
[0103] At the same time, input into another multi-layer perceptron to obtain the attention matrix:
[0104]
[0105] where M is the attention matrix, and g(·) is a multi-layer perceptron with 128 dimensions in one hidden layer;
[0106] S34. Pass the aggregated point cloud coordinates obtained in S32 through a multi-layer perceptron γ(·) with 128 dimensions in one hidden layer to obtain intermediate features
[0107]
[0108] where γ(·) is a multi-layer perceptron network with 128 dimensions in one hidden layer;
[0109] Then, multiply the intermediate features by the attention matrix M obtained in S33 to obtain the features selected according to the features, and add them to the self-attention features obtained in S33 At the same time, use the residual connection of the point cloud features obtained in S32 Finally, reconstruct through a multi-layer perceptron δ(·) with 128 dimensions in two hidden layers to obtain the final corrected point cloud coordinates h i :
[0110]
[0111] Model training process:
[0112] The training process of the model takes the point cloud in the training data after sparse sampling as the input data. After inputting the data into the model established by S1 - S3, the output point cloud is obtained. This output point cloud is compared with the original dense point cloud in the training data that has not undergone sparse sampling, and after calculating the error, the network parameters are updated using the error backpropagation algorithm.
[0113] To better learn the distribution characteristics of the point cloud, the model uses a two - stage reconstruction loss as the loss function.
[0114] The reconstruction loss in the model training process is as follows:
[0115] The obtained rough point cloud result is calculated for the chamfer distance and Hausdorff distance simultaneously with the point cloud coordinates h after noise reduction and correction i and the original dense point cloud G. These two methods are used to measure the similarity between the predicted rough point cloud result and the predicted point cloud coordinates h i and the original dense point cloud coordinates G, and the two - stage reconstruction loss function is obtained as:
[0116]
[0117] Among them, represents the rough point cloud without noise reduction and correction, h i represents the dense point cloud finally output by the model, G represents the correct original dense point cloud of the i - th frame, and λ1 and λ2 are parameters used to control the importance of the two stages;
[0118] L CD (·) represents the chamfer distance:
[0119]
[0120] L HD (·) represents the Hausdorff distance:
[0121]
[0122] Among them, S1, S2 represent two point clouds that need to be calculated, selected from G, h i ; x ∈ S1 means x belongs to the points in point cloud S1, and y ∈ S2 means y belongs to the points in point cloud S2.
[0123] In the early stage of the training phase, the model learns the approximate contour features of the object and does not need to integrate too many local detail features of the object. Therefore, it pays more attention to the prediction effect of the rough point cloud, and smaller λ1 and λ2 are used in this stage. As the training time increases, the model can well understand the local detail features, and the final reconstruction effect of the model at this stage becomes more important. Therefore, we use larger λ1 and λ2.
[0124] Only the reconstruction loss may cluster the r points generated by a point. The model adds the repulsion loss to make these points more evenly and dispersedly distributed. The repulsion loss in the model training process is as follows:
[0125]
[0126] where L rep is the repulsion loss function, r represents the upsampling rate, N represents the number of points in each frame of the point cloud, K(i) represents the K nearest neighbors of x i , x i represents the i-th point in the dense point cloud composed of rN points finally obtained, x j represents the j-th nearest neighbor of the point x i , and decreases as t increases, decreases as t increases, and h is a hyperparameter.
[0127] In addition to the reconstruction loss and the repulsion loss, the L2 regularization method is also used in training to avoid overfitting. The model uses an end-to-end method to minimize the mixed loss function for training. The overall loss function of the model is:
[0128]
[0129] where L rec is the reconstruction loss function, L rep is the repulsion loss function, λ rec , λ rep and λ reg are the proportionality coefficients of L rec , L rep and w respectively, and w represents the weights of the network.
[0130] Model inference process:
[0131] The specific method of the model inference process is as follows:
[0132] (1) Fix the parameters of the upsampling model D that has completed training;
[0133] (2) The sparse point cloud video PointCloud to be upsampled sparseAs input data, it is input into the upsampling model D to generate the corresponding dense point cloud video PointCloud dense :
[0134] PointCloud dense = D(x).
[0135] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A point cloud video upsampling method based on feature fine-tuning, characterized in that, It includes a model training process and a model inference process. The model training process includes using the sparsely sampled sparse point cloud video in the training dataset as the input of the model, and obtaining the output point cloud video after upsampling; During the model training process, the output point cloud video is compared with the original dense point cloud video in the training dataset that has not been sparsely sampled, and the reconstruction loss and the repulsion loss are used as loss functions to iteratively train the model repeatedly. The model inference process includes inputting any sparse point cloud video into the trained model, and obtaining the dense point cloud video corresponding to the input sparse point cloud video after upsampling; The process of the upsampling is as follows: S1. Feature extraction stage: The sparse point cloud video is segmented into several segments at a fixed interval. For each obtained segment, a feature extractor is used to extract the point cloud features of the first frame of each segment; then the point cloud features of the first frame are expanded to obtain the expanded point cloud features; S2. Feature fine-tuning stage: For the current frame point cloud except the first frame in each segment, the current frame point cloud, its previous frame point cloud, and the expanded point cloud features of the previous frame are input into the feature fine-tuning unit together. After feature fine-tuning, the point cloud features of the current frame are obtained, and then the point cloud features of the current frame are expanded to obtain the expanded point cloud features; S3. Point cloud reconstruction stage: The expanded point cloud features of each frame are used for point cloud reconstruction to obtain rough point cloud coordinates, and the rough point cloud coordinates are denoised and corrected to obtain the output point cloud video after upsampling.
2. The method for upsampling point cloud videos based on feature fine-tuning according to claim 1, wherein The feature extraction stage specifically includes the following process: S11. Divide the sparse point cloud video into several segments at a fixed interval of T frames. Each segment is S = {s i} ∈ R N×d , where s i represents the i-th frame of the point cloud, i represents the index of each frame in the segment, each frame of the point cloud has N points, and each point is a d-dimensional vector. Here, only the x, y, and z coordinates of the point cloud are considered, so d = 3; S12. For the first frame s in each segment 1 ∈R N×d , use the feature extractor Feat(x) of the MPU to extract the feature f 1 of the first frame s 1 ∈R N×C , where C is the dimension of each feature vector.
3. The method for upsampling point cloud videos based on feature fine-tuning according to claim 2, wherein The feature fine-tuning stage specifically includes the following process: S21. For the point cloud s of the i-th frame i , first use a multi-layer perceptron to expand the features of the previous frame of point cloud s i-1 . Expand the dimension of the feature vector from the original N×C dimension to rN×C dimension, where r represents the upsampling rate, to obtain the expanded point cloud features . Input the point cloud s of the current frame i ∈R N×d , the point cloud s of the previous frame i-1 ∈R N×d , and the expanded point cloud features into the feature fine-tuning unit simultaneously; S22. Take as the feature template of the current frame point cloud s i . Concatenate with the current frame point cloud s i , and input the result into a multi-layer perceptron with a single 128-dimensional hidden layer to obtain the intermediate feature result to be searched S23. For each point in the current frame of point cloud s i search for K nearest neighbor points in the previous frame of point cloud s i-1 using the K-nearest neighbor search algorithm, and gather these points together to obtain the K-nearest neighbor point set At the same time, retrieve according to the index of each neighbor point in to obtain the corresponding point cloud features S24. Take the result obtained in S23 and the current frame point cloud s i as inputs to the local position encoding unit. For each point use the K-nearest neighbor search algorithm to find K nearest neighbor points and obtain the local position encoding through local position encoding operations Among them, "+" represents the merging and splicing operation of coordinate dimensions, represents the residuals of K nearest neighbor points and their center points, and MLP represents a multi-layer perceptron with a hidden layer of 128 dimensions; Then, all local position encodings are aggregated together to obtain the encoding of the point cloud of the i-th frame that contains both absolute distance and relative distance information U represents the union operation of sets, which here represents the operation of vector merging or concatenation; S25. Combine the obtained in S24 with the obtained in S23 through a multi-layer perceptron α(·) with one hidden layer of 128 dimensions, and then input the result into another multi-layer perceptron β(·) with one hidden layer of 128 dimensions to obtain the feature of the current i-th frame of point cloud finally. Among them, "+" represents the merging and splicing operation of vectors.
4. A method for upsampling point cloud videos based on feature fine-tuning according to claim 3, characterized in that, The point cloud reconstruction stage specifically includes the following process: S31. Reconstruct the expanded features through a multi-layer perceptron with three hidden layers of 128 dimensions, reconstruct the features of rN×C dimensions into coordinates of rN×3 dimensions, and obtain rough point cloud coordinates S32. For each point in , use the K-nearest neighbor search algorithm to find K nearest neighbor points among all points in . Aggregate the K nearest neighbor points obtained for each point into . At the same time, find the point cloud features of the corresponding point cloud among all the point cloud features obtained through indexing in S25 S33. Input the point cloud features obtained in S32 into the self-attention unit for self-attention mechanism operation to obtain self-attention features Here, Q, K, and V are different encoded results obtained by respectively inputting features into three different hidden layers of a multi-layer perceptron with 128 dimensions in one layer, and d K is the dimension of the feature vector K, and Softmax() represents the normalized exponential function; Meanwhile, it is input into another multi-layer perceptron to obtain an attention matrix: Among them, M is the attention matrix, and g(·) is a multi-layer perceptron with a single hidden layer of 128 dimensions; S34. The aggregated point cloud coordinates obtained in S32 pass through a multi-layer perceptron γ(·) with a single hidden layer of 128 dimensions to obtain intermediate features Among them, γ(·) is a multi-layer perceptron network with a single hidden layer of 128 dimensions; Then, multiply the intermediate feature by the attention matrix M obtained in S33 to get the feature selected according to the feature, and add it to the self-attention feature obtained in S33. At the same time, use the residual connection to connect the point cloud feature obtained in S32. Finally, reconstruct through a multi-layer perceptron δ(·) with two hidden layers of 128 dimensions to obtain the finally corrected point cloud coordinates h i :
5. A method for upsampling point cloud videos based on feature fine-tuning according to claim 1, characterized in that The reconstruction loss of the model training process is as follows: The obtained rough point cloud results and the point cloud coordinates h after noise reduction and correction i At the same time, calculate the chamfer distance and Hausdorff distance with the original dense point cloud G, and use these two methods to measure the predicted rough point cloud results and the predicted point cloud coordinates h i The similarity between the original dense point cloud coordinates G, and the reconstruction loss function of the two stages is obtained as follows: Among them, represents the rough point cloud without noise reduction correction, h i represents the dense point cloud finally output by the model, G represents the correct original dense point cloud of the i-th frame, and λ1 and λ2 are parameters used to control the importance of the two stages; L CD (·) represents the chamfer distance: L HD (·) represents the Hausdorff distance: Among them, S1 and S2 represent two point clouds that need to be calculated, selected from G, h i ; x ∈ S1 means that x belongs to the points in the point cloud S1, and y ∈ S2 means that y belongs to the points in the point cloud S2.
6. A method for upsampling point cloud videos based on feature fine-tuning according to claim 1, characterized in that The repulsion loss of the model training process is as follows: Among them, L rep is the rejection loss function, r represents the upsampling rate, N represents the number of points in each frame of point cloud, and K(i) represents the K i nearest neighbors of x i represents the i-th point in the final dense point cloud composed of rN points, x j represents the j-th nearest neighbor of point x i , decreases as t increases, decreases as t increases, and h is a hyperparameter.
7. A method for upsampling point cloud videos based on feature fine-tuning according to claim 1, characterized in that, The L2 regularization method is used in the model training process. The model uses an end-to-end method to minimize the combined loss function for training. The overall loss function of the model is: Among them, L rec is the reconstruction loss function, and L rep is the rejection loss function. λ rec , λ rep and λ reg are the proportionality coefficients of L rec , L rep and w respectively, where w represents the weights of the network.
8. A method for upsampling point cloud videos based on feature fine-tuning according to claim 1, characterized in that The specific method of the model inference process is as follows: (1) Fix the parameters of the upsampling model D that has completed training; (2) The sparse point cloud video PointCloud that needs to be upsampled sparse is used as input data and input into the upsampling model D to generate the corresponding dense point cloud video PointCloud dense : PointCloud dense = D(x).
Citation Information
Patent Citations
Point cloud up-sampling method based on deep learning
CN111724478A
Dynamic point cloud geometric compression method based on scene flow network and time entropy model
CN114025146A