A pose estimation optimization method based on surround viewpoints

Through the pose estimation optimization method under the surround viewpoint, the U-shaped codec network and 3D convolutional network are used to solve the problem of pose estimation error accumulation in multiple views, and accurate camera pose estimation is achieved, improving the quality of three-dimensional reconstruction.

CN114170316BActive Publication Date: 2025-05-06HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111528516.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-13
Publication Date
2025-05-06
Estimated Expiration
2041-12-13

AI Technical Summary

Technical Problem

In the prior art, in multi-view situations, especially in surrounding viewpoints, rough pose estimation leads to error accumulation, resulting in large deviations in camera poses, affecting the quality of three-dimensional reconstruction.

Method used

A pose estimation optimization method based on circumferential viewpoint is proposed. Through a U-shaped codec network and an improved 3D convolutional network, combining coarse camera pose and feature point information, iteratively updates the pose information to achieve accurate camera pose estimation.

Benefits of technology

Through accurate camera pose estimation, the quality of the three-dimensional reconstruction is improved, the offset and distortion in the reconstruction results are reduced, and the fineness of the model and the accuracy of the texture geometry are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114170316B_ABST
    Figure CN114170316B_ABST
Patent Text Reader

Abstract

The present invention discloses a pose estimation optimization method based on surround viewpoints. The present invention takes sparse viewpoint images and corresponding rough camera poses as input, uses a codec neural network to extract features of different scales of the image sequence, performs transposed convolution on the features of different scales and multiplies them two by two, maps and fuses them into 3D feature voxels according to the rough camera pose, extracts the spatial coordinates, density and color information of the feature points through a 3D convolutional network as input to an MLP network, regresses to obtain the pose offset error, updates the control viewpoint pose and regresses to the main viewpoint coordinate system. The present invention preprocesses the input multi-viewpoint images, uses a neural network, and integrates the two tasks of feature matching and pose optimization based on multiple views to obtain an accurate camera pose. The pose estimation optimization method for surround views proposed by the present invention improves the accuracy of pose estimation and greatly improves the quality of three-dimensional reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of posture estimation, and in particular to a posture estimation optimization method based on surround viewpoints. Background Art

[0002] Realizing the reproduction of the real world has always been a hot topic in computer vision research. An image is a projection of a real object in the three-dimensional world on a two-dimensional plane. Existing methods can reconstruct three-dimensional models from images, but the overall algorithm is time-consuming and susceptible to inaccurate depth estimation, which leads to large deviations in camera pose estimation and poor model reconstruction quality. In recent years, with the popularity of depth cameras, three-dimensional reconstruction based on RGB-D cameras has also continued to develop, but there are still shortcomings such as low camera depth resolution and inaccurate pose estimation, which leads to problems such as insufficient precision of the reconstructed model and distortion of texture geometry.

[0003] Fine-tuning the rough pose can achieve higher quality 3D reconstruction. The traditional pose estimation algorithm first solves the pose transformation matrix between adjacent frames through the feature point matching algorithm, and obtains the camera coordinates by continuous extension. However, in the case of multiple views, especially in the surround viewpoint, since only a rough pose estimation is performed, as the number of viewpoints increases, the error accumulates continuously, causing a large offset in the subsequent pose, resulting in unsatisfactory reconstruction results. Summary of the invention

[0004] The purpose of the present invention is to propose a pose estimation optimization method based on surround viewpoints to address the deficiencies of the prior art. The present invention inputs a sequence of surround viewpoint images and a rough camera pose, and outputs an accurate camera pose, thereby achieving pose estimation and optimization of surround viewpoints.

[0005] In order to achieve the above object, the technical solution of the present invention comprises the following steps:

[0006] The input of the present invention is a sequence of pictures of surrounding viewpoints arranged in order and a rough camera pose, and the output is a precise camera pose, which realizes the pose estimation and optimization of the surrounding viewpoints. The specific implementation steps are as follows:

[0007] Step 1. Input is a sequence of images of surrounding viewpoints and a rough camera pose. The image sequence includes the main viewpoint image I0 and other viewpoint images I n, where n is 1 to N, representing the nth other viewpoint picture; assuming that the size of all viewpoint pictures is S*S, and other viewpoint pictures are sorted according to the shooting angle transformation; starting from the main viewpoint picture, the real postures of all viewpoint pictures can be connected to form a closed loop; the picture sequence is input into a U-type codec network, using a 3*3 convolution kernel and upsampling by transposed convolution. The U-type codec network can encode information of features at different levels, thereby realizing estimation optimization from different scales; in addition, after extracting the feature points of the viewpoint pictures, feature points with a matching number greater than or equal to 2 times will be filtered out;

[0008] Step 2. In the U-type codec network of step 1, feature extraction is performed on viewpoint images at different scales to obtain feature maps, which are then projected to corresponding positions through rough camera poses and fused into a 3D feature voxel;

[0009] Step 3. After the 3D feature voxels obtained in step 2, feature extraction is performed through the improved 3D convolutional network; and after each feature extraction, the number of channels is continuously reduced by half to obtain feature voxels of size S*S*32; finally, the spatial coordinates (x, y, z), density (ρ), and image color (R, G, B) of the feature points are regressed, and a 1*7 matrix is ​​constructed as the input of the four-layer MLP network. The pose offset is solved by the global mean square error, and finally the new transformed pose P′ of the viewpoint image I1 is regressed. T0-1 With the depth information D′1, by updating the camera pose P′1;

[0010] Step 4. The updated camera pose P′1, the viewpoint image I1, the third viewpoint image I2 and the rough camera pose P2 corresponding to the viewpoint image I2 obtained in step 3 are input into the U-type codec neural network to extract features and filter out feature points with a matching number greater than or equal to 2; after rough camera pose estimation projection mapping, the 3D feature voxels that meet the requirements are constructed;

[0011] Step 5. For the 3D feature voxels obtained in step 4, use the improved 3D convolutional network to extract features; finally regress the feature point spatial coordinates (x, y, z), density (ρ), and image color (R, G, B), and construct a 1*7 matrix as the input of the four-layer MLP network; solve the pose offset error through the global mean square error, and finally regress to obtain the new transformed pose P′ of the viewpoint image I2 T1-2 With the depth information D′2, update the camera pose P′2;

[0012] Step 6. Select viewpoint image I2 as the main viewpoint and the fourth viewpoint I3 as the reference viewpoint, and update the camera pose P′3 of viewpoint image I3 according to steps 3, 4, and 5; according to this rule, use the updated new viewpoint image as the main viewpoint, iteratively add the next adjacent viewpoint image and the rough pose to obtain the corrected precise camera transfer matrix; when the nth viewpoint image obtains the updated camera pose, it indicates that the camera poses of all viewpoints have been optimized.

[0013] Furthermore, the rough camera pose is obtained by using colmap to perform pose estimation.

[0014] Further, the step 2 is specifically implemented as follows:

[0015] The feature map scale of the 13th layer in the U-type codec network in step 1 is set to one-fourth of the original image, which is an estimate of the maximum scale; the feature map scale of the 16th layer is set to one-half of the original image, which is an estimate of the intermediate scale; the feature map scale of the 19th layer is set to be equal to the original image, as an estimate of the finest scale; the feature map scale F of the (S / 4) scale is 0-S / 4 Perform transposed convolution and add two by two to obtain the feature map F of (S / 2) scale 0-S / 2 , the feature map F of (S / 2) scale 0-S / 2 Perform transposed convolution and add two by two to get the feature map F of scale S 0-S ; Then, the 3D feature voxel is constructed by projection mapping based on the rough camera pose according to formula (1), with a size of S*S*128. In formula (1), f is the focal length of the camera, u and v are arbitrary coordinate points in the image coordinate system, and x w ,y w ,z w Represents the three-dimensional coordinate point in the world coordinate system, z c Represents the z-axis value of the camera coordinates, that is, the distance from the target to the camera;

[0016]

[0017] The overall structure of the improved 3D convolutional network includes three convolutional layers, using a small convolution kernel of 3*3*3 and a step size of 1*1*1; two pooling layers are 2*2*2 with a step size of 2*2*2; finally, the spatial coordinates (x, y, z) are obtained from the extracted three-dimensional feature points, and according to the color information in the corresponding RGB image, the pixel color information is obtained according to the weight distribution method of 70% for the main view and 30% for the control view, and is represented by (R, G, B). Finally, the image density (ρ) is added to form a 1*7 matrix as the input of the four-layer MLP network;

[0018] The pose offset error is solved by global mean square error, and the new transformed pose P′ is finally regressed. T0-1and depth information D′1; the camera pose can represent the rotation and translation of the viewpoint, as shown in Formula 2. The camera poses of the two viewpoint images I0 and I1 are represented by P0 and P1 respectively. The camera poses are all 4*4 matrices. The two viewpoint images are transformed by the transformation matrix P T To achieve conversion, P T It is composed of a 3*3 rotation matrix R and a 3*1 translation matrix T. For convenience, it is supplemented with 0 and 1 to form a 4*4 matrix to achieve cascade multiplication. The camera pose P′1 of the viewpoint image I1 is updated according to formula (2). Since a large viewpoint transfer angle will reduce the number of matching feature points, which has a great impact on the final result, two adjacent frames are selected for correction. At the same time, after completing the correction step, it returns to the main view I0 and unifies the camera pose comparison object.

[0019]

[0020] The beneficial effects of the present invention are as follows:

[0021] The present invention proposes a new neural network model, which connects the matching feature points with the camera pose correction by simulating light, thereby achieving accurate camera pose estimation optimization and improving the quality of 3D reconstruction.

[0022] In the scene of surrounding viewpoints, the present invention extracts the corresponding feature points in each viewpoint through the feature extraction network from the input picture sequence, combines the rough camera pose information, builds a rough 3D model through the 3D convolution network, simulates the light penetrating the three-dimensional feature points from different viewpoints, obtains the volume density and RGB value, combines the spatial coordinates as the network input, and iteratively updates the pose information. The back-projection error of the feature points is solved through the codec network to achieve the optimization effect of the pose, and finally obtains the accurate camera pose. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 Diagram of the overall neural network model architecture.

[0024] Figure 2 Diagram of the encoder-decoder neural network structure.

[0025] Figure 3 Feature extraction network structure diagram.

[0026] Figure 4 Four-layer MLP network structure diagram.

[0027] Figure 5 Detailed diagram of the neural network as a whole.

[0028] Figure 6 For comparison effect diagram. DETAILED DESCRIPTION

[0029] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0030] Step 1. Input is a sequence of surrounding viewpoint images and a rough camera pose. The image sequence includes the main viewpoint image I0 (numbered 0) and other viewpoint images I n (Nviews, viewpoint 1, 2, ..., n, ...N), assuming that the size of all viewpoint images is S*S, and other viewpoint images are sorted according to the shooting angle transformation; starting from the main viewpoint, the real poses of all viewpoints can be connected to form a closed loop. The rough camera pose is obtained by using colmap for pose estimation. The picture sequence is input into a U-type codec network, using a 3*3 convolution kernel and transposed convolution to achieve upsampling. The network can encode information of features at different levels, thereby achieving estimation optimization from different scales. In addition, after extracting the feature points of the picture, the feature points that are matched more than or equal to 2 times will be filtered out. The feature points that are matched more than or equal to 2 times have large errors, which have a bad impact on the results, so they need to be filtered out.

[0031] Step 2. In step 1, the neural network extracts features from images at different scales to obtain feature maps, which are then projected to corresponding positions through rough camera poses and fused into a 3D feature voxel. The specific implementation is as follows:

[0032] The feature map scale of the 13th layer in the U-type codec network in step 1 is set to one-fourth of the original image, which is the maximum scale estimate; the feature map scale of the 16th layer is set to one-half of the original image, which is the intermediate scale estimate; the feature map scale of the 19th layer is set to be equal to the original image, which is the finest scale estimate. 0-S / 4 Perform transposed convolution and add two by two to obtain the feature map F of (S / 2) scale 0-S / 2 , the feature map F of (S / 2) scale 0-S / 2 Perform transposed convolution and add two by two to get the feature map F of scale S 0-S Then, the 3D feature voxel is constructed by projection mapping based on the rough camera pose according to formula (1), with a size of S*S*128. In formula (1), f is the focal length of the camera, u and v are arbitrary coordinate points in the image coordinate system, and x is w ,y w ,z w Represents the three-dimensional coordinate point in the world coordinate system, z c Indicates the z-axis value of the camera coordinates, that is, the distance from the target to the camera.

[0033]

[0034] Step 3. After the 3D feature voxels obtained in step 2, feature extraction is performed through the improved 3D convolutional network. The overall structure of the improved 3D convolutional network includes three convolutional layers, using a small convolution kernel of 3*3*3 and a step size of 1*1*1; two pooling layers are 2*2*2 with a step size of 2*2*2. Finally, the spatial coordinates (x, y, z) are obtained from the extracted three-dimensional feature points, and the pixel color information is obtained according to the color information in the corresponding RGB image, and the weight distribution method of 70% for the main view and 30% for the control view is used to represent (R, G, B). Finally, the image density (ρ) is added to form a 1*7 matrix as the input of the four-layer MLP network.

[0035] The pose offset error is solved by global mean square error, and the new transformed pose P′ is finally regressed. T0-1 The camera pose can represent the rotation and translation of the viewpoint, as shown in Formula 2. The camera poses of the two viewpoint images I0 and I1 are represented by P0 and P1 respectively. The camera poses are all 4*4 matrices. The two viewpoint images are transformed by the transformation matrix P T To achieve conversion, P T It is composed of a 3*3 rotation matrix R and a 3*1 translation matrix T. For convenience, it is supplemented with 0 and 1 to achieve cascade multiplication. The camera pose P′1 of the viewpoint image I1 is updated according to formula (2). Since a large viewpoint transfer angle will reduce the number of matching feature points and have a greater impact on the final result, two adjacent frames are selected for correction. At the same time, after completing the correction step, it returns to the main view I0 and unifies the camera pose comparison object.

[0036]

[0037] Furthermore, the three-layer improved 3D convolutional network is used, and after each feature extraction, the number of channels is continuously reduced by half to obtain a feature voxel of size S*S*32. Finally, the spatial coordinates (x, y, z), density (ρ), and image color (R, G, B) of the feature points are regressed, and a 1*7 matrix is ​​constructed as the input of the four-layer MLP network. The pose offset is solved by the global mean square error, and finally the new transformed pose P′ of I1 is regressed. T0-1 And depth information D′1.

[0038] Step 4. The updated camera pose P′1, viewpoint image I1, the third viewpoint image I2 and the rough camera pose P2 corresponding to the viewpoint image I2 obtained in step 3 are input into the U-type codec neural network to extract features and filter out feature points with a matching number of times greater than or equal to 2. After the rough camera pose estimation projection mapping, the conforming 3D feature voxels are constructed.

[0039] Step 5. For the 3D feature voxels obtained in step 4, use the improved 3D convolutional network to extract features. Finally, the spatial coordinates (x, y, z), density (ρ), and image color (R, G, B) of the feature points are regressed, and a 1*7 matrix is ​​constructed as the input of the four-layer MLP network. The pose offset error is solved by the global mean square error, and the new transformed pose P′ of the viewpoint image I2 is finally regressed. T1-2 With the depth information D′2, the camera pose P′2 is updated by formula (2).

[0040] Step 6. Then select viewpoint image I2 as the main viewpoint and the fourth viewpoint I3 as the reference viewpoint, and update the camera pose P′3 of viewpoint image I3 according to steps 3 and 4. According to this rule, take the updated new viewpoint image as the main viewpoint, iteratively add the next adjacent viewpoint image and the rough pose, and obtain the corrected precise camera transfer matrix. When the nth viewpoint obtains the updated camera pose, it means that the camera poses of all viewpoints have been optimized.

[0041] Example:

[0042] like Figure 1 As shown in the figure, the pose estimation optimization method based on the surround viewpoint has the following general framework:

[0043] Input image sequence, the main view image is I0, and other view images are I n (n=0,1...N), first passes through an encoding and decoding neural network, and extracts and fuses features of different levels at different scales, then maps the rough camera pose to the 3D space to construct 3D feature voxels, and converts the voxels to 3D cost voxels, finally extracts the spatial coordinates, density, and color information of the feature points and inputs them into the MLP network, solves the pose offset error through the mean square error, and finally regresses to obtain the camera rotation pose.

[0044] Figure 2 It is a U-type encoding and decoding neural network structure. The input image sequence is I0 for the main view image and I for other view images. n (n=0,1...N), in order to encode and decode feature information at different levels. The 19th layer is equal to the original image size, as the most detailed scale estimate, the 16th layer feature map scale is half of the original image, as the intermediate scale estimate; the 13th layer feature map scale is one quarter of the original image, as the largest scale estimate.

[0045] The neural network of the codec is defined as follows: the size of the convolution kernel is 3*3, the pooling layer is maxpooling, and the upsampling method is transposed convolution. In the case of the same scale, the number of channels remains unchanged. In the (S / 4) scale, (S / 2) scale, and S scale, the number of channels is 1024, 512, and 256 respectively. The multi-view image of size S*S is input into the codec neural network. First, it passes through three layers of convolution to obtain the feature map S1 of S*S*256; the feature map S1 is maxpooled and doubled, and after three layers of convolution, the feature map S2 of (S / 2)*(S / 2)*512 is obtained; the feature map S2 is maxpooled and doubled, and after three layers of convolution, the feature map S3 of (S / 4)*(S / 4)*1024 is obtained; the feature map S3 is transposed convolution to obtain S tran3 , S tran3 Divide them into 512 pairs and add them two by two to obtain the feature map S4 of (S / 2)*(S / 2)*512; perform transpose convolution on the feature map S4 and divide it into 256 pairs and add them two by two to obtain the feature map S5 of S*S*256.

[0046] Figure 3 It is a network structure diagram that is converted into 3D feature voxels after feature extraction and optimized. First, the output feature map S5 of the encoding and decoding neural network is mapped to different positions according to the corresponding rough posture and converted into 3D voxels. The 3D voxels are extracted through three layers of 3D convolutional layers. After these steps, the number of 3D voxel channels is reduced to one-fourth of the original number. Then, according to the variance measurement method, the reduced 3D voxels are converted into cost voxels. The cost voxels are regressed and multiplied pixel by pixel with the corresponding scale feature map S3 and upsampled. The 3D voxels of size (S / 2)*(S / 2) are updated. Then, the 3D feature voxels of scale S*S are obtained by the same method, and the three-dimensional coordinates, density and color information of the feature points are obtained as the subsequent network input.

[0047] Figure 4 This is a diagram of the four-layer MLP network structure. Consider a feature point information as a patch, and each patch size is 1*7. Two viewpoint images are input each time for matching. Assuming there are M feature points in total, the final MLP network input size is a tensor of 1*7M*2. The hidden layer is two fully connected layers, and the pose offset Loss is finally regressed to update the viewpoint camera pose, and the accurate pose is obtained as the network output.

[0048] Figure 5 This is a detailed neural network structure diagram of the present invention. The general structure and each module have been described above.

[0049] like Figure 6As shown in the figure, figure a is a mesh model reconstructed based on a rough pose, figure b is a mesh model reconstructed based on an optimized pose, and figure c is an RGB image of the real object. It can be seen that the ghosting effect in figure b has been significantly improved and is more consistent with the real object.

Claims

1. A pose estimation optimization method based on surround viewpoints, characterized in that The input is a sequence of images of surrounding viewpoints and a rough camera pose arranged in order, and the output is a precise camera pose, which realizes the pose estimation and optimization of the surrounding viewpoints. The specific implementation steps are as follows: Step 1. Input is a sequence of images of surrounding viewpoints and a rough camera pose. The image sequence includes the main viewpoint image I0 and other viewpoint images I n , where n is 1 to N, representing the nth other viewpoint picture; assuming that the size of all viewpoint pictures is S*S, and other viewpoint pictures are sorted according to the shooting angle transformation; starting from the main viewpoint picture, the real postures of all viewpoint pictures can be connected to form a closed loop; the picture sequence is input into a U-type codec network, using a 3*3 convolution kernel and upsampling by transposed convolution. The U-type codec network can encode information of features at different levels, thereby realizing estimation optimization from different scales; in addition, after extracting the feature points of the viewpoint pictures, feature points with a matching number greater than or equal to 2 times will be filtered out; Step 2. In the U-type codec network of step 1, feature extraction is performed on viewpoint images at different scales to obtain feature maps, which are then projected to corresponding positions through rough camera poses and fused into a 3D feature voxel; Step 3. After the 3D feature voxels obtained in step 2, feature extraction is performed through the improved 3D convolutional network; and after each feature extraction, the number of channels is continuously reduced by half to obtain feature voxels of size S*S*32; finally, the spatial coordinates (x, y, z), density ρ, and image color R, G, B of the feature points are regressed, and a 1*7 matrix is ​​constructed as the input of the four-layer MLP network. The pose offset is solved by the global mean square error, and finally the new transformed pose P′ of the viewpoint image I1 is regressed. T0-1 With the depth information D′1, update the camera pose P′1; Step 4. The updated camera pose P′1, the viewpoint image I1, the third viewpoint image I2 and the rough camera pose P2 corresponding to the viewpoint image I2 obtained in step 3 are input into the U-type codec neural network to extract features and filter out feature points with a matching number greater than or equal to 2; After rough camera pose estimation and projection mapping, the corresponding 3D feature voxels are constructed; Step 5. For the 3D feature voxels obtained in step 4, use the improved 3D convolutional network to extract features; Finally, the spatial coordinates (x, y, z), density ρ, and image color R, G, B of the feature points are obtained by regression, and a 1*7 matrix is ​​constructed as the input of a four-layer MLP network; the pose offset error is solved by the global mean square error, and the new transformed pose P′ of the viewpoint image I2 is finally obtained by regression. T1-2 With the depth information D′2, update the camera pose P′2; Step 6. Select viewpoint image I2 as the main viewpoint and the fourth viewpoint I3 as the reference viewpoint, and update the camera pose P′3 of viewpoint image I3 according to steps 3, 4, and 5; according to this rule, use the updated new viewpoint image as the main viewpoint, iteratively add the next adjacent viewpoint image and the rough pose to obtain the corrected precise camera transfer matrix; when the nth viewpoint image obtains the updated camera pose, it indicates that the camera poses of all viewpoints have been optimized.

2. The method for optimizing posture estimation based on surround viewpoints according to claim 1, characterized in that The rough camera pose is obtained by using colmap to perform pose estimation.

3. The method for optimizing posture estimation based on surround viewpoints according to claim 1, characterized in that Step 2 is implemented as follows: The feature map scale of the 13th layer in the U-type codec network in step 1 is set to one-fourth of the original image, which is an estimate of the maximum scale; the feature map scale of the 16th layer is set to one-half of the original image, which is an estimate of the intermediate scale; the feature map scale of the 19th layer is set to be equal to the original image, as an estimate of the finest scale; the feature map scale F of S / 4 is set to be equal to the original image, as an estimate of the finest scale; 0-s / 4 Perform transposed convolution and add two by two to get the feature map F of scale S / 2 0-s / 2 , the feature map F of scale S / 2 0-s / 2 Perform transposed convolution and add two by two to get the feature map F of scale S 0-S ; Then, the 3D feature voxel is constructed by projection mapping based on the rough camera pose according to formula (1), with a size of S*S*128. In formula (1), f is the focal length of the camera, u and v are any coordinate points in the image coordinate system, and x w ,y w , z w Represents the three-dimensional coordinate point in the world coordinate system, z c Represents the z-axis value of the camera coordinates, that is, the distance from the target to the camera; 4. The method for optimizing posture estimation based on surround viewpoints according to claim 3 is characterized in that The overall structure of the improved 3D convolutional network includes three convolutional layers, using a small convolution kernel of 3*3*3 and a step size of 1*1*1; two pooling layers are 2*2*2 with a step size of 2*2*2; finally, the spatial coordinates (x, y, z) are obtained from the extracted three-dimensional feature points, and according to the color information in the corresponding RGB image, the pixel color information is obtained according to the weight distribution method of 70% for the main view and 30% for the control view, and R, G, B is used to represent it. Finally, the image density ρ is added to combine into a 1*7 matrix as the input of the four-layer MLP network; The pose offset error is solved by global mean square error, and the new transformed pose P′ is finally regressed. T0-1 and depth information D′1; the camera pose can represent the rotation and translation of the viewpoint, as shown in Formula 2. The camera poses of the two viewpoint images I0 and I1 are represented by P0 and P1 respectively. The camera poses are all 4*4 matrices. The two viewpoint images are transformed by the transformation matrix P T To achieve conversion, P T It is composed of a 3*3 rotation matrix R and a 3*1 translation matrix T. For convenience, 0 and 1 are used to complete it into a 4*4 matrix to achieve cascade multiplication; Update the camera pose P′1 of the viewpoint image I1 according to formula (2); since a large viewpoint shift angle will reduce the number of matching feature points and have a greater impact on the final result, two adjacent frames are selected for correction. After completing the correction step, return to the main view I0 to unify the camera pose comparison object;

Citation Information

Patent Citations

  • Flower 3D reconstruction method based on ORB and U-net

    CN109215117A

  • Multi-view view depth estimation method based on detail information preservation

    CN112907641A