A 3D reconstruction method, system and medium for a monocular RGB camera

The two-dimensional features are extracted by feature pyramids and three-dimensional reconstruction is carried out using two-dimensional encoders and voxel cycle modules. The problems of scale inconsistency and non-smooth connections in the existing technology are solved, and efficient and accurate three-dimensional reconstruction effect is achieved.

CN114399601BActive Publication Date: 2025-05-30SUN YAT SEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210060366.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-19
Publication Date
2025-05-30
Estimated Expiration
2042-01-19

AI Technical Summary

Technical Problem

The existing three-dimensional reconstruction methods have problems such as inconsistent scales and non-smooth connections in depth estimation and local detail processing, resulting in insufficient reconstruction speed and accuracy.

Method used

Three-scale two-dimensional features are extracted through feature pyramids to reduce interference; 3D volume is obtained using two-dimensional encoder and backprojection, retaining high-dimensional semantic information; introducing voxel loop module for feature reconstruction and aggregation to improve reconstruction smoothness; using 3D encoding decoder to extract and regress to obtain TSDF reconstruction results.

Benefits of technology

The three-dimensional reconstruction speed and accuracy are improved, and the smoother reconstruction effect is achieved. It can directly perform three-dimensional reconstruction from the video stream and obtain label information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114399601B_ABST
    Figure CN114399601B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional reconstruction method, system and medium for a monocular RGB camera. The method includes: selecting key frames from an image sequence in a video captured by the monocular RGB camera; extracting features from the key frames through a feature pyramid, and performing a residual operation to obtain two-dimensional features; importing the two-dimensional features into a two-dimensional encoder to extract two-dimensional structural features, and performing back-projection to obtain a 3D volume; introducing a voxel cycle module to reconstruct the 3D volume to obtain a reconstructed 3D volume; performing feature aggregation on the reconstructed 3D volumes of all key frames to obtain a fusion result; inputting the fusion result into a 3D encoding and decoding network to extract and regress to obtain a TSDF reconstruction result. The present invention extracts features of different scales, retains high-dimensional semantic information, and reduces the amount of computation; by introducing a voxel cycle module, the three-dimensional feature information is reconstructed and feature aggregation is performed to obtain information of different dimensions, making the reconstruction effect smoother, and improving the speed and accuracy of three-dimensional reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of three-dimensional reconstruction, and particularly relates to a three-dimensional reconstruction method, system and medium for a monocular RGB camera. Background Art

[0002] Existing three-dimensional reconstruction methods generally include several steps: 1. Select frames with suitable disparity from a continuous video frame sequence as key frames; 2. Use a method such as a sliding window to perform depth estimation on the key frames to obtain depth information for each frame; 3. Convert it into point cloud data; 4. Finally, fuse all of them together through the TSDF algorithm to achieve three-dimensional reconstruction. And the most important of these is the depth estimation step in 2. However, the depth data predicted by the depth estimation module will not be completely consistent. Although the pose on each camera is very accurate, since the standard images of each are estimated separately, it will lead to the problem of inconsistent scales; at the same time, since the depth information of multiple frames is calculated first and then fused together, when dealing with local details, there will be a problem that the connection between the fusions of two frames is not smooth, and the reconstruction of local details cannot be well achieved; therefore, a three-dimensional reconstruction method that can improve both the reconstruction speed and the reconstruction accuracy is needed. Summary of the Invention

[0003] The main purpose of the present invention is to overcome the disadvantages and deficiencies of the prior art, and provide a three-dimensional reconstruction method, system and medium for a monocular RGB camera. This method extracts three-scale two-dimensional features through a feature pyramid, reducing interference; obtains a 3D volume through a two-dimensional encoder and back-projection, retaining high-dimensional semantic information; introduces a voxel recurrent module for feature reconstruction and feature aggregation to make the reconstruction effect smoother; uses a 3D encoder-decoder to extract and regress to obtain the TSDF reconstruction result, improving the three-dimensional reconstruction speed and accuracy.

[0004] To achieve the above object, the present invention adopts the following technical solutions:

[0005] On the one hand, the present invention provides a three-dimensional reconstruction method for a monocular RGB camera, including the following steps:

[0006] Select key frames from the image sequence in the video captured by the monocular RGB camera;

[0007] Extract features from the key frames through a feature pyramid, and obtain two-dimensional features using a residual operation;

[0008] Import the two-dimensional features into a two-dimensional encoder to extract two-dimensional structural features, and perform back-projection to obtain a 3D volume;

[0009] Introduce a voxel loop module to reconstruct the 3D volume and obtain the reconstructed 3D volume;

[0010] Aggregate the reconstructed 3D volumes of all key frames to obtain a fusion result;

[0011] Input the fusion result into the 3D encoder-decoder to extract and regress to obtain the TSDF reconstruction result.

[0012] As a preferred technical solution, the two-dimensional structural features include channels, width, and height; the three-dimensional features include channels, width, height, and depth.

[0013] As a preferred technical solution, the key frames are feature-extracted through a feature pyramid and two-dimensional features are obtained through residual operations, specifically:

[0014] The feature pyramid includes a first feature extractor, a second feature extractor, a third feature extractor, a first upsampling module, a second upsampling module, and a third upsampling module, which are connected in sequence;

[0015] For the input key frames, the feature extractor sequentially extracts feature maps in a bottom-up manner;

[0016] The upsampling module upsamples the feature maps obtained from the previous layer in a top-down manner;

[0017] During upsampling, the first upsampling module performs a residual operation on the feature map extracted by the third feature extractor; the second upsampling module performs a residual operation on the feature map output by the first upsampling module and the feature map extracted by the second feature extractor; the third upsampling module performs a residual operation on the feature map output by the second upsampling module and the feature map extracted by the first feature extractor; the feature maps output by each upsampling module are obtained to obtain two-dimensional features.

[0018] As a preferred technical solution, the 3D volume is obtained by back-projection, specifically:

[0019] Let It be the image sequence of the video. Each two-dimensional feature output by each upsampling module in the feature pyramid is passed through a two-dimensional encoder to obtain the corresponding two-dimensional structural feature F(t)=F(It). Given the camera internal parameter P t and the external parameter K t In the case of, the two-dimensional structural feature is back-projected to obtain a three-dimensional voxel volume, that is, the 3D volume:

[0020]

[0021]

[0022] Where V t is the voxel volume, F t is the two-dimensional structural feature, P t is the internal parameter matrix of the camera, K t is the external parameter matrix of the camera, Π is the perspective matrix, : is the slicing operator, is the coordinate of the pixel on the image, and (i, j, k) is the coordinate of the voxel in the world space.

[0023] As a preferred technical solution, the voxel loop module combines the obtained 3D volume and the historical 3D volume, and predicts a new 3D volume through an MLP;

[0024] The voxel loop module includes a first voxel loop module and a second voxel loop module; the first voxel loop module is located after the back-projection operation with a high resolution of the output 3D volume; the second voxel loop module is located after the back-projection operation with a medium resolution of the output 3D volume;

[0025] The first voxel loop module converts the 3D volume with a high resolution into the same dimension as the 3D volume with a medium resolution through upsampling;

[0026] The second voxel loop module converts the 3D volume with a medium resolution and the converted 3D volume with a high resolution into the same dimension as the 3D volume with a low resolution through upsampling;

[0027] Then, the 3D volume with a low resolution and the converted 3D volume with a medium resolution are superimposed to obtain the reconstructed 3D volume.

[0028] As a preferred technical solution, the obtaining of the fusion result is specifically:

[0029] Aggregate the features of the reconstructed 3D volumes of all key frames, and then perform weighted averaging of the 3D volumes over the entire image sequence for fusion. The formula is:

[0030]

[0031]

[0032] Where, is the voxel volume corresponding to the fusion of t key frames, is the binary mask corresponding to the fusion of t key frames. The binary mask W t (i, j, k) ∈ (0, 1) indicates whether the voxel is inside or outside the camera's frustum.

[0033] As a preferred technical solution, the fusion result is input into a 3D encoding and decoding device, and the TSDF reconstruction result is extracted and regressed, specifically as follows:

[0034] The 3D encoding and decoding device includes an encoder and a decoder; the encoder includes a first encoding layer, a second encoding layer, and a third encoding layer; each encoding layer sequentially includes two-layer residual blocks and a downsampling layer; the decoder includes a first decoding layer, a second decoding layer, and a third decoding layer; each decoding layer sequentially includes an upsampling layer and two-layer residual blocks; a convolutional block is connected after the third decoding layer for outputting the reconstruction result; the decoder and the encoder are connected by three-layer residual blocks for skip connection;

[0035] For the input fusion result, the encoder sequentially performs residual operations and downsampling in a bottom-up manner;

[0036] The decoder performs upsampling and residual operations on the residual result obtained from the previous layer in a top-down manner;

[0037] During upsampling, the residual of the downsampling result of the third encoding layer and the residual result of the previous layer are input into the first decoding layer for upsampling and residual operations; the residual of the downsampling result of the second encoding layer and the residual result of the previous layer are input into the second decoding layer for upsampling and residual operations; the residual of the downsampling result of the first encoding layer and the residual result of the previous layer are input into the third decoding layer for upsampling and residual operations. Finally, the residual result passes through the convolutional block after the third decoding layer, and the TSDF reconstruction result is obtained through regression using the tanh activation function.

[0038] As a preferred technical solution, the size of the residual block is 3×3×3; the downsampling layer is implemented using a convolutional block with a size of 3×3×3 and a stride of 2; the upsampling layer uses trilinear interpolation and a convolutional block with a size of 1×1×1;

[0039] A batch normalization and a ReLU activation function are connected after the convolutional blocks in the encoder and the decoder;

[0040] The size of the convolutional block after the third decoder is 1×1×1.

[0041] On the other hand, the present invention provides a three-dimensional reconstruction system for a monocular RGB camera, which is applied to the above-mentioned three-dimensional reconstruction method for a monocular RGB camera, and includes a key frame selection module, a two-dimensional feature extraction module, a three-dimensional feature extraction module, a voxel loop module, a feature fusion module, and a result obtaining module;

[0042] The key frame selection module is used to select key frames from the image sequence in the video captured by the monocular RGB camera;

[0043] The two-dimensional feature extraction module is used to extract features from key frames through a feature pyramid and obtain two-dimensional features using residual operations;

[0044] The three-dimensional feature extraction module is used to import two-dimensional features into a two-dimensional encoder to extract two-dimensional structural features and perform back-projection to obtain a 3D volume;

[0045] The voxel cycle module is used to reconstruct the 3D volume to obtain a reconstructed 3D volume;

[0046] The feature fusion module is used to perform feature aggregation on the reconstructed 3D volumes of all key frames to obtain a fusion result;

[0047] The result obtaining module is used to input the fusion result into a 3D encoding and decoding network to extract and regress to obtain a TSDF reconstruction result.

[0048] The present invention also provides a computer-readable storage medium storing a program, which when executed by a processor, implements the above-mentioned three-dimensional reconstruction method for a monocular RGB camera.

[0049] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0050] 1. The present invention extracts features of different scales through a feature pyramid, has high resolution, not only reduces the amount of calculation but also retains high-dimensional semantic information, and improves the reconstruction effect through a process from coarse to fine;

[0051] 2. The present invention introduces a voxel cycle module, weights the information of the current feature and the information of the historical three-dimensional feature and then performs prediction through an MLP, which can jointly process the information of the previous key frame and the current key frame, making the reconstruction effect smoother;

[0052] 3. The present invention does not require depth data, can directly perform three-dimensional reconstruction through a video stream, and can perform semantic segmentation without additional calculation; while obtaining the TSDF reconstruction result, label information can also be obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0054] Figure 1It is a flow framework diagram of a 3D reconstruction method for a monocular RGB camera according to an embodiment of the present invention;

[0055] Figure 2 It is a schematic structural diagram of a feature pyramid according to an embodiment of the present invention;

[0056] Figure 3 It is a specific example diagram of feature extraction of a feature pyramid according to an embodiment of the present invention;

[0057] Figure 4 It is a schematic structural diagram of a voxel loop module according to an embodiment of the present invention;

[0058] Figure 5 It is a schematic diagram of steps of a 3D encoder-decoder according to an embodiment of the present invention;

[0059] Figure 6 It is a structural diagram of a 3D reconstruction system for a monocular RGB camera according to an embodiment of the present invention;

[0060] Figure 7 It is a result diagram of a computer-readable storage medium according to an embodiment of the present invention. Specific embodiments

[0061] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.

[0062] Mentioning "embodiment" in the present application means that the specific features, structures or characteristics described in combination with the embodiment may be included in at least one embodiment of the present application. The phrase appears at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art understand explicitly and implicitly that the embodiments described in the present application can be combined with other embodiments.

[0063] Embodiment

[0064] As Figure 1 shown, in an embodiment of the present application, a 3D reconstruction method for a monocular RGB camera is provided, including the following steps:

[0065] S1. Select key frames from the image sequence in the video captured by the monocular RGB camera;

[0066] S2. Extract features from the key frames through a feature pyramid, and obtain two-dimensional features using residual operations;

[0067] S3. Import the two-dimensional features into the two-dimensional encoder to extract two-dimensional structural features, and perform back-projection to obtain a 3D volume;

[0068] S4. Introduce a voxel cycle module to reconstruct the 3D volume to obtain a reconstructed 3D volume;

[0069] S5. Aggregate the reconstructed 3D volumes of all key frames to obtain a fusion result;

[0070] S6. Input the fusion result into the 3D encoding and decoding network to extract and regress to obtain the TSDF reconstruction result.

[0071] In this embodiment, one video frame is selected as a key frame every nine frames, thereby reducing the computational burden.

[0072] More specifically, as Figure 2 shown, feature extraction is performed through a feature pyramid to extract coarse-grained features at three levels: coarse, medium, and fine. In order to reduce the impact caused by reducing interference during deconvolution, residual operations are used to stack the previous convolution results together, and then two-dimensional features are output for each layer. The specific steps of step S2 are as follows:

[0073] The feature pyramid includes a first feature extractor, a second feature extractor, a third feature extractor, a first upsampling module, a second upsampling module, and a third upsampling module;

[0074] For the input key frame, the feature extractor sequentially extracts feature maps in a bottom-up manner;

[0075] The upsampling module upsamples the feature map obtained from the previous layer in a top-down manner;

[0076] During upsampling, the first upsampling module performs a residual operation on the feature map extracted by the third feature extractor; the second upsampling module performs a residual operation on the feature map output by the first upsampling module and the feature map extracted by the second feature extractor; the third upsampling module performs a residual operation on the feature map output by the second upsampling module and the feature map extracted by the first feature extractor; the feature maps output by each upsampling module are obtained to obtain two-dimensional features.

[0077] The feature pyramid is mainly composed of the superposition of convolution operations. For example, the feature extractor and others are all convolution operations.

[0078] As Figure 3 shown, in this embodiment, a key frame image with 3 channels, a height of 480, and a width of 640 is input into the feature pyramid for feature extraction;

[0079] Features with 24 channels, 120 in height, and 160 in width are obtained through the first feature extractor Figure 1 ;

[0080] Features with 40 channels, 60 in height, and 80 in width are obtained through the second feature extractor Figure 2 ;

[0081] Features with 80 channels, 30 in height, and 40 in width are obtained through the third feature extractor Figure 3 ;

[0082] The first upsampling module upsamples the features Figure 3 and outputs features with 40 channels, 60 in height, and 80 in width Figure 4 ;

[0083] The second upsampling module upsamples the features Figure 4 and performs a residual operation with the features Figure 2 and outputs features with 24 channels, 120 in height, and 160 in width Figure 5 ;

[0084] The third upsampling module upsamples the features Figure 5 and performs a residual operation with the features Figure 1 and outputs features with 3 channels, 480 in height, and 640 in width Figure 6 ;

[0085] Finally, the feature maps with three different resolutions output by the feature pyramid are used as the two-dimensional features of the key frame images

[0086] More specifically, in order to obtain the structural features of each scene graph, the two-dimensional features are imported into a two-dimensional encoder to extract their two-dimensional structural features. The structural features are the features of a two-dimensional image, including information such as channels, width, and height. Then, the two-dimensional structural features are back-projected to obtain a 3D volume, which includes information on channels, width, height, and depth; Step S3 is specifically as follows

[0087] Let It be the image sequence of the video. The two-dimensional features output by each upsampling module in the feature pyramid are each passed through a two-dimensional encoder to obtain the corresponding two-dimensional structural features F(t) = F(It). Given the camera intrinsic parameters P t and extrinsic parameters K t the two-dimensional structural features are back-projected to obtain a three-dimensional voxel volume, i.e., 3D volume

[0088]

[0089]

[0090] where V t is the voxel volume, F t is the two-dimensional structural feature, P t is the internal parameter matrix of the camera, K t is the external parameter matrix of the camera, Π is the perspective matrix, : is the slicing operator, is the coordinate of the pixel on the image, and (i, j, k) is the coordinate of the voxel in the world space.

[0091] The 3D volume is obtained by back-projection according to the two-dimensional structural feature of the two-dimensional feature. There is an additional depth data based on the two-dimensional feature, which ensures semantic richness and makes the reconstruction effect smoother.

[0092] More specifically, in order to improve the consistency of the reconstruction effect of each frame, a voxel loop module is introduced. As Figure 4 shown, the voxel loop module in step S4 combines the obtained 3D volume with the historical 3D volume and predicts a new 3D volume through an MLP;

[0093] There are two voxel loop modules, namely the first voxel loop module and the second voxel loop module; the first voxel loop module is located after the back-projection operation with a high-resolution output 3D volume; the second voxel loop module is located after the back-projection operation with a medium-resolution output 3D volume;

[0094] The first voxel loop module upsamples the 3D volume with high resolution to convert it into the same dimension as the 3D volume with medium resolution;

[0095] The second voxel loop module upsamples the 3D volume with medium resolution and the converted 3D volume with high resolution to convert them into the same dimension as the 3D volume with low resolution;

[0096] Then, the 3D volume with low resolution is superimposed with the converted 3D volume with medium resolution to obtain the reconstructed 3D volume.

[0097] By upsampling the 3D volumes at the high and medium resolution levels to convert them into the same dimension as the 3D volume with low resolution, and then superimposing them with the 3D volume with low resolution to obtain the reconstructed 3D volume, through three levels of coarse-grainedness, different-dimensional information can be better obtained, and a better reconstruction effect can be achieved.

[0098] More specifically, in step S5, the specific method for obtaining the fusion result is:

[0099] Aggregate the features of the reconstructed 3D volumes of all key frames, and then fuse the 3D volumes by weighted averaging over the entire image sequence. The formula is as follows:

[0100]

[0101]

[0102] where, is the voxel volume corresponding to the fusion of t key frames, is the binary mask corresponding to the fusion of t key frames. Use the binary mask W t (i, j, k) ∈ (0, 1) indicates whether the voxel is inside or outside the camera's frustum.

[0103] More specifically, as Figure 5 shown, in step S6, the 3D encoder-decoder includes an encoder and a decoder; the encoder includes a first encoding layer, a second encoding layer, and a third encoding layer; each encoding layer sequentially contains two residual blocks (1) and a downsampling layer (2); the decoder includes a first decoding layer, a second decoding layer, and a third decoding layer; each decoding layer sequentially contains an upsampling layer (3) and two residual blocks (1); a MLP convolutional block (4) is connected after the third decoding layer for outputting the reconstruction result; at the same time, the decoder and the encoder are connected by three residual blocks (1) for skip connection;

[0104] For the input fusion result, the encoder sequentially performs residual operations and downsampling in a bottom-up manner;

[0105] The decoder upsamples and performs residual operations on the residual result obtained from the previous layer in a top-down manner;

[0106] During upsampling, the residual of the downsampling result of the third encoding layer and the residual result of the previous layer are input into the first decoding layer for upsampling and residual operations; the residual of the downsampling result of the second encoding layer and the residual result of the previous layer are input into the second decoding layer for upsampling and residual operations; the residual of the downsampling result of the first encoding layer and the residual result of the previous layer are input into the third decoding layer for upsampling and residual operations. Finally, the residual result passes through the MLP convolutional block after the third decoding layer, and the tsdf reconstruction result is obtained by regression using the tanh activation function.

[0107] In this embodiment, each layer of the encoder and decoder uses a two-layer 3×3×3 residual block; the downsampling layer is implemented using a convolutional block with a size of 3×3×3 and a stride of 2; the upsampling layer uses trilinear interpolation and a convolutional block with a size of 1×1×1 to change the feature dimension; each time downsampling is performed, the feature dimension is doubled, and each time upsampling is performed, the feature dimension is halved; after all convolutional blocks in the 3D encoder-decoder, a batch normalization and a ReLU activation function are connected; the size of the convolutional block after the third decoding layer is 1×1×1.

[0108] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously.

[0109] Based on the same idea as the three-dimensional reconstruction method of a monocular RGB camera in the above embodiment, the present invention also provides a three-dimensional reconstruction system of a monocular RGB camera, which can be used to execute the three-dimensional reconstruction method of a monocular RGB camera described above. For the sake of convenience of description, in the structural schematic diagram of an embodiment of the three-dimensional reconstruction system of a monocular RGB camera, only the parts related to the embodiment of the present invention are shown. Those skilled in the art can understand that the illustrated structure does not constitute a limitation on the device, and may include more or fewer components than those illustrated, or combine certain components, or arrange different components.

[0110] As Figure 6 shown, in another embodiment of the present application, a three-dimensional reconstruction system of a monocular RGB camera is provided, and the system at least includes the following modules:

[0111] A key frame selection module, configured to select key frames from the image sequence in the video captured by the monocular RGB camera;

[0112] A two-dimensional feature extraction module, configured to perform feature extraction on the key frames through a feature pyramid and obtain two-dimensional features using a residual operation;

[0113] A three-dimensional feature extraction module, configured to import the two-dimensional features into a two-dimensional encoder to extract two-dimensional structural features and perform backprojection to obtain a 3D volume;

[0114] A voxel loop module, configured to reconstruct the 3D volume to obtain a reconstructed 3D volume;

[0115] A feature fusion module, configured to perform feature aggregation on the reconstructed 3D volumes of all key frames to obtain a fusion result;

[0116] A result acquisition module is used to input the fusion result into a 3D encoding and decoding device, extract and regress to obtain the TSDF reconstruction result.

[0117] It should be noted that the 3D reconstruction system of a monocular RGB camera according to the present invention corresponds one-to-one with the 3D reconstruction method of a monocular RGB camera according to the present invention. The technical features and beneficial effects described in the embodiments of the above 3D reconstruction method of a monocular RGB camera are all applicable to the embodiments of the 3D reconstruction system of a monocular RGB camera. For specific content, reference can be made to the description in the method embodiments of the present invention, which will not be elaborated here. This is hereby declared.

[0118] In addition, in the implementation manner of the 3D reconstruction system of a monocular RGB camera in the above embodiments, the logical division of each program module is only for illustration. In actual applications, according to needs, for example, considering the configuration requirements of corresponding hardware or the convenience of software implementation, the above functions can be allocated to different program modules to complete, that is, the internal structure of the 3D reconstruction system of a monocular RGB camera is divided into different program modules to complete all or part of the functions described above.

[0119] As Figure 7 shown, in one embodiment, a computer-readable storage medium is provided, which stores a program in a memory. When the program is executed by a processor, the 3D reconstruction method of a monocular RGB camera described above can be implemented. Specifically:

[0120] Select key frames from the image sequence in the video captured by the monocular RGB camera;

[0121] Extract features from the key frames through a feature pyramid, and use residual manipulation to obtain two-dimensional features;

[0122] Import the two-dimensional features into a two-dimensional encoder to extract two-dimensional structural features, and perform back-projection to obtain a 3D volume;

[0123] Introduce a voxel loop module to reconstruct the 3D volume to obtain a reconstructed 3D volume;

[0124] Aggregate the features of the reconstructed 3D volumes of all key frames to obtain a fusion result;

[0125] Input the fusion result into a 3D encoding and decoding device, extract and regress to obtain the TSDF reconstruction result.

[0126] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0127] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0128] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications made without departing from the spirit and principle of the present invention should be equivalent replacement methods and are all included in the protection scope of the present invention.

Claims

1. A three-dimensional reconstruction method for a monocular RGB camera, characterized in that, it includes the following steps: Select key frames from the image sequence in the video captured by the monocular RGB camera; Extract features from the key frames through a feature pyramid, and use residual operations to obtain two-dimensional features; Import the two-dimensional features into a two-dimensional encoder to extract two-dimensional structural features, and perform back-projection to obtain a 3D volume; Introduce a voxel recurrent module to reconstruct the 3D volume to obtain a reconstructed 3D volume; The voxel recurrent module combines the obtained 3D volume with the historical 3D volume, and predicts a new 3D volume through an MLP; The voxel recurrent module includes a first voxel recurrent module and a second voxel recurrent module; the first voxel recurrent module is located after the back-projection operation with a high resolution of the output 3D volume; the second voxel recurrent module is located after the back-projection operation with a medium resolution of the output 3D volume; The first voxel recurrent module upsamples the 3D volume with a high resolution to convert it into the same dimension as the 3D volume with a medium resolution; The second voxel recurrent module upsamples the 3D volume with a medium resolution and the converted 3D volume with a high resolution to convert them into the same dimension as the 3D volume with a low resolution; Then, the 3D volume with a low resolution is superimposed with the converted 3D volume with a medium resolution to obtain a reconstructed 3D volume; Aggregate the features of the reconstructed 3D volumes of all key frames to obtain a fusion result. Specifically: Aggregate the features of the reconstructed 3D volumes of all key frames, and then perform weighted averaging of the 3D volumes over the entire image sequence for fusion. The formula is: Among them, is the voxel volume corresponding to the fusion of t key frames, is the binary mask corresponding to the fusion of t key frames. The binary mask W t (i, j, k) ∈ (0, 1) indicates whether the voxel is inside or outside the camera's frustum; Input the fusion result into a 3D encoder-decoder to regress and obtain the TSDF reconstruction result; the 3D encoder-decoder includes an encoder and a decoder; the encoder includes a first encoding layer, a second encoding layer, and a third encoding layer; each encoding layer sequentially includes two residual blocks and a downsampling layer; the decoder includes a first decoding layer, a second decoding layer, and a third decoding layer; each decoding layer sequentially includes an upsampling layer and two residual blocks; a convolutional block is connected after the third decoding layer for outputting the reconstruction result; the decoder and the encoder are connected by three residual blocks for skip connection.

2. A three-dimensional reconstruction method for a monocular RGB camera according to claim 1, characterized in that, the two-dimensional structural features include channels, width, and height; the 3D volume includes channels, width, height, and depth.

3. A three-dimensional reconstruction method for a monocular RGB camera according to claim 1, characterized in that, the extraction of features from the key frames through a feature pyramid and the performance of residual operations to obtain two-dimensional features are specifically: The feature pyramid includes a first feature extractor, a second feature extractor, a third feature extractor, a first upsampling module, a second upsampling module, and a third upsampling module, which are connected in sequence; For the input key frames, the feature extractor sequentially extracts feature maps in a bottom-up manner; The upsampling module upsamples the feature maps obtained from the previous layer in a top-down manner; During upsampling, the first upsampling module performs a residual operation on the feature maps extracted by the third feature extractor; the second upsampling module performs a residual operation on the feature maps output by the first upsampling module and the feature maps extracted by the second feature extractor; the third upsampling module performs a residual operation on the feature maps output by the second upsampling module and the feature maps extracted by the first feature extractor; the feature maps output by each upsampling module are obtained to obtain two-dimensional features.

4. A three-dimensional reconstruction method for a monocular RGB camera according to claim 3, wherein, The specific process of performing back-projection to obtain a 3D volume is as follows: Let \(I_t\) be the image sequence of the video. For each two-dimensional feature output by each upsampling module in the feature pyramid, a two-dimensional encoder is used to obtain the corresponding two-dimensional structural feature \(F(t)=F(I_t)\). Given the camera intrinsic parameter \(P\) t and the extrinsic parameter \(K\) t , the two-dimensional structural feature is back-projected to obtain a three-dimensional voxel volume, i.e., 3D volume: where V t is the voxel volume, F t is the two-dimensional structural feature, P t is the intrinsic matrix of the camera, K t is the extrinsic matrix of the camera, Π is the perspective matrix, : is the slicing operator, is the coordinate of the pixel on the image, and (i, j, k) is the coordinate of the voxel in the world space.

5. A three-dimensional reconstruction method for a monocular RGB camera according to claim 4, wherein, The specific process of inputting the fusion result into a 3D encoder-decoder and regressing to obtain a TSDF reconstruction result is as follows: For the input fusion result, the encoder sequentially performs residual operations and downsampling in a bottom-up manner; The decoder upsamples and performs residual operations on the residual results obtained from the previous layer in a top-down manner; During upsampling, the residual of the downsampling result of the third encoding layer and the residual result of the previous layer are input into the first decoding layer for upsampling and residual operations; the residual of the downsampling result of the second encoding layer and the residual result of the previous layer are input into the second decoding layer for upsampling and residual operations; the residual of the downsampling result of the first encoding layer and the residual result of the previous layer are input into the third decoding layer for upsampling and residual operations. Finally, the residual result passes through the convolutional block after the third decoding layer, and a TSDF reconstruction result is obtained by using the tanh activation function for regression.

6. A three-dimensional reconstruction method for a monocular RGB camera according to claim 5, wherein, The size of the residual block is 3×3×3; the downsampling layer is implemented using a convolutional block with a size of 3×3×3 and a stride of 2; the upsampling layer uses trilinear interpolation and a convolutional block with a size of 1×1×1; Batch normalization and relu activation functions are connected after the convolutional blocks in both the encoder and the decoder; The size of the convolutional block after the third decoding layer is 1×1×1.

7. A three-dimensional reconstruction system for a monocular RGB camera, wherein, It is applied to a three-dimensional reconstruction method for a monocular RGB camera according to any one of claims 1-6, and includes a key frame selection module, a two-dimensional feature extraction module, a three-dimensional feature extraction module, a voxel loop module, a feature fusion module, and a result obtaining module; The key frame selection module is used to select key frames from the image sequence in the video captured by the monocular RGB camera; The two-dimensional feature extraction module is used to extract features from the key frames through a feature pyramid and obtain two-dimensional features using residual operations; The three-dimensional feature extraction module is used to import the two-dimensional features into a two-dimensional encoder to extract two-dimensional structural features and perform back-projection to obtain a 3D volume; The voxel recurrent module is used to reconstruct the 3D volume to obtain the reconstructed 3D volume; The voxel recurrent module combines the obtained 3D volume with the historical 3D volume and predicts a new 3D volume through an MLP; The voxel recurrent module includes a first voxel recurrent module and a second voxel recurrent module; the first voxel recurrent module is located after the back-projection operation with a high-resolution output 3D volume; the second voxel recurrent module is located after the back-projection operation with a medium-resolution output 3D volume; The first voxel recurrent module upsamples the 3D volume with high resolution to convert it into the same dimension as the 3D volume with medium resolution; The second voxel recurrent module upsamples the 3D volume with medium resolution and the converted 3D volume with high resolution to convert them into the same dimension as the 3D volume with low resolution; Then, the 3D volume with low resolution is superimposed on the converted 3D volume with medium resolution to obtain the reconstructed 3D volume; The feature fusion module is used to aggregate the features of the reconstructed 3D volumes of all key frames to obtain a fusion result. Specifically: Aggregate the features of the reconstructed 3D volumes of all key frames, and then perform weighted averaging of the 3D volumes over the entire image sequence for fusion. The formula is: Among them, is the voxel volume corresponding to the fusion of t key frames, is the binary mask corresponding to the fusion of t key frames. Using the binary mask W t (i, j, k) ∈ (0, 1) indicates whether the voxel is inside or outside the camera's frustum; The result acquisition module is used to input the fusion result into a 3D encoder-decoder to regress and obtain the TSDF reconstruction result; the 3D encoder-decoder includes an encoder and a decoder; the encoder includes a first encoding layer, a second encoding layer, and a third encoding layer; each encoding layer sequentially includes two residual blocks and a downsampling layer; the decoder includes a first decoding layer, a second decoding layer, and a third decoding layer; each decoding layer sequentially includes an upsampling layer and two residual blocks; a convolutional block is connected after the third decoding layer for outputting the reconstruction result; the decoder and the encoder are connected by three residual blocks for skip connection.

8. A computer-readable storage medium stores a program, Characterized in that, When the program is executed by a processor, it implements the three-dimensional reconstruction method of a monocular RGB camera according to any one of claims 1-6.

Citation Information

Patent Citations

  • Method of monocular depth estimation based on joint self-attention mechanism

    AU2020103715A4

  • Three-dimensional reconstruction method and related device and equipment

    CN112750201A