Deep learning three-dimensional sparse reconstruction method and system suitable for planetary rover image
By integrating the deep neural network DPFeat with an attention mechanism and incremental motion recovery structure technology, the problem of 3D reconstruction failure in planetary rover images with weak texture and large illumination variations was solved, achieving high-precision sparse 3D reconstruction and providing a reliable data foundation.
Patent Information
- Application Number
- CN202310587367.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-23
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2043-05-23
AI Technical Summary
Existing technologies struggle to effectively extract reliable key points from planetary rover images, leading to 3D reconstruction failures, especially in cases of weak texture and significant lighting variations.
A deep neural network DPFeat with an attention mechanism was designed for the extraction and description of image key points. Combined with a robust gross error removal method, the image was registered using incremental motion reconstruction technique. Finally, the object space 3D point coordinates and camera position and attitude were calculated using bundle adjustment technique.
This method enables high-precision and robust sparse 3D reconstruction on planetary rover images with weak texture and large illumination variations, solving the problem of 3D reconstruction failure caused by weak texture and varying illumination in traditional methods and providing a reliable data foundation.
Smart Images

Figure CN116664855B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of planetary three-dimensional reconstruction, and particularly relates to a deep learning three-dimensional sparse reconstruction method and system suitable for planetary rover images. BACKGROUND
[0002] Using images collected by a planetary rover to construct an accurate three-dimensional terrain model of a planet is one of the most direct and effective methods for studying the terrain of the planet, and is also an important part of positioning and navigation of the planetary rover. The three-dimensional reconstruction result will provide basic data for control point research, high-precision mapping, and subsequent detection tasks. At present, three-dimensional reconstruction technology based on photogrammetry has been widely used in planetary terrain mapping. However, due to the characteristics of weak texture and large illumination variation of planetary rover images, it is difficult to extract reliable key points, and three-dimensional reconstruction is likely to fail, resulting in the inability to perform subsequent terrain analysis.
[0003] Previous researchers use a motion recovery structure algorithm to process planetary images for estimating camera poses or generating photogrammetric products such as orthophotos and digital elevation models. One of the necessary steps of the motion recovery structure method is the matching of homonymous points, which aims to identify and correspond to points with the same or similar structure from two or more images using a local feature method with invariance. The accuracy of homonymous point matching directly affects the accuracy of three-dimensional reconstruction. However, due to the texture of the planetary surface image being neither rich nor clear, lacking obvious recognizable features, and the brightness varying greatly, it poses a great challenge to the robustness of image matching.
[0004] Traditional manually extracted local feature methods are usually divided into two stages, first detecting key points, and then calculating local descriptors for each key point, among which the most typical and commonly used is the SIFT algorithm. These methods set various standards for screening and describing key points through prior knowledge, but due to excessive reliance on human experience, they are difficult to adapt to different scenes, and once they encounter scenes that cannot be handled, it becomes very complex to adjust parameters. When the imaging conditions change dramatically, or the texture is too uniform, the accuracy and reliability of these local feature methods will be severely affected. SUMMARY
[0005] The purpose of the present application is to achieve high-precision and robust sparse three-dimensional reconstruction of planetary terrain. A new deep neural network that integrates an attention mechanism is designed for key point extraction and description of images. In combination with a robust gross error rejection method, accurate and reliable homonymous points are obtained. Incremental motion recovery structure technology is used to register the images, and finally, the beam adjustment technology is used to solve the three-dimensional point coordinates in the object space and the position and attitude of the camera.
[0006] Based on the above technical problems, the present application adopts the following technical solutions:
[0007] A deep learning three-dimensional sparse reconstruction method suitable for planetary exploration vehicle images, comprising the following steps:
[0008] Step one, import the pre-processed planetary exploration vehicle images and other auxiliary data;
[0009] Step two, use a deep convolutional neural network DPFeat that integrates attention mechanism to extract and describe the key points of each imported image, obtaining a key point set containing the two-dimensional image coordinates of the key points on each image and the corresponding descriptors;
[0010] Step three, match the key points according to the two-dimensional image coordinates and descriptor information of the key points, extract the homonymous points in the images with overlap, and eliminate the incorrect homonymous points while matching more correct homonymous points;
[0011] Step four, use the coordinates of the key points and the corresponding relationship of the homonymous points obtained by matching to calculate the accurate three-dimensional point coordinates in the object space and the position and attitude of the camera, and then restore the three-dimensional structure of the terrain.
[0012] Further, in step one, the data of the exploration vehicle needs to be converted from the original format at the time of publication to a commonly used image format, and the camera position, attitude, and internal parameter information obtained from other data sources are also imported as initial values for reconstruction.
[0013] Further, in step two, the deep convolutional neural network DPFeat used to extract and describe the key points replaces the up-sampling layer in L2-Net with a dilated convolution, and after the image is processed by the main network, two branches are generated, one for generating the image coordinates of the key points and the other for generating the corresponding descriptors of the key points. Each branch integrates attention mechanism to improve the weight of the key feature map and reduce the weight of the irrelevant feature map, achieving efficient use of high-dimensional feature maps.
[0014] Further, the two branches of the DPFeat network are supervised by loss functions during network training, including similarity loss, extreme value loss, and reliability loss.
[0015] Further, the similarity loss formula is as follows:
[0016]
[0017] Wherein, I and I' represent two images respectively, U represents the corresponding relationship of the same point between the two images, P represents a set containing tiles overlapping with each other, p represents a tile in the set, i and j represent the horizontal and vertical coordinates of a pixel in the tile p respectively, S is the probability map of the key points of the image I generated by the network, S' U The probability map of the key points of the image I' after transformation according to U, the cosim function calculates the similarity of the input quantity by cosine, the greater the value is, the more similar the input quantity is, the higher the probability of the key point is, and the lower the loss value is;
[0018] The extreme loss formula is as follows:
[0019]
[0020] Wherein, N represents the neighborhood size, S ij represents the value of the point with horizontal and vertical coordinates i and j on the probability map of the key points, the max and mean functions are used to calculate the maximum value and the average value respectively, the greater the difference of the point is, the smaller the extreme loss is;
[0021] The reliability loss formula is as follows:
[0022]
[0023] Wherein, B represents the number of tiles in a batch, R represents the reliability map, R ij represents the value of the point with horizontal and vertical coordinates i and j on the reliability map, p ij represents the value of the point with horizontal and vertical coordinates i and j on the tile p, is a differentiable function for supervised learning descriptor.
[0024] Further, in step three, the improved gross error elimination method is used to eliminate the wrong homonymic points,
[0025] Wherein, the improved gross error elimination method is as follows:
[0026] Firstly, the image is normalized, then the image is divided into multiple blocks, the reliability of the matching result in each block is calculated, the normal value and the gross error in each block are classified respectively, then the gross error is merged, and finally the gross error is eliminated.
[0027] Further, in step four, based on the incremental motion recovery structure technology, the bundle adjustment technology is used to calculate the accurate object three-dimensional point coordinates and the position and attitude of the camera, and then the three-dimensional structure of the terrain is recovered.
[0028] Further, the incremental motion structure-from-motion technology firstly selects two adjacent images for initialization, then uses the key point matching relationship of the triangular points in the registered images to register other images to the current model, the newly registered images not only cover the existing object points, but also increase the coverage range through triangulation, and the calculation result is optimized by using the bundle adjustment, and the outlier measurement values are removed; after the iterative calculation, the accurate object point coordinates and the position and attitude information of the camera are output.
[0029] The application also provides a deep learning three-dimensional sparse reconstruction system suitable for planetary rover image, comprising:
[0030] An image data import module is used to import the preprocessed planetary rover image and other auxiliary data;
[0031] A key point extraction and description module is used to extract and describe the key points of each imported image by using a deep convolutional neural network DPFeat fused with an attention mechanism, so as to obtain a key point set containing the two-dimensional image coordinates of the key points on each image and corresponding descriptors;
[0032] A key point matching module is used to match the key points according to the positions and descriptor information of the key points, so as to extract the homonymous points in the images with mutual overlaps and remove the incorrect homonymous points while matching more correct homonymous points;
[0033] A solving module is used to solve the accurate object three-dimensional point coordinates and the position and attitude of the camera by using the coordinates of the key points and the corresponding relationship of the homonymous points obtained by matching, and then recover the three-dimensional structure of the terrain.
[0034] Further, the system is used to solve the problems of weak texture and large illumination change of the planetary rover image data, and realizes the sparse three-dimensional reconstruction of the planetary surface and the solving of the position and attitude of the camera.
[0035] Compared with the prior art, the application has the following beneficial effects:
[0036] The application designs a new deep neural network fused with an attention mechanism for image key point extraction and description, uses a robust gross error elimination method to obtain accurate and reliable homonymous points, and then registers the images by using the incremental motion structure-from-motion technology, solves the object three-dimensional point coordinates and the position and attitude of the camera by using the bundle adjustment technology, realizes the high-precision and robust sparse three-dimensional reconstruction of the planetary terrain, and solves the problems that the traditional feature extraction and matching method fails to process the planetary rover image due to weak texture, changing illumination and other factors, and then leads to three-dimensional reconstruction failure. BRIEF DESCRIPTION OF DRAWINGS
[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed to be used in the embodiments of the present application will be briefly introduced as follows:
[0038] Figure 1 A method effect diagram suitable for deep learning three-dimensional sparse reconstruction of planetary exploration vehicle image in the embodiments of the present application;
[0039] Figure 2 A schematic diagram of the method suitable for deep learning three-dimensional sparse reconstruction of planetary exploration vehicle image in the embodiments of the present application;
[0040] Figure 3 A network structure schematic diagram of the deep neural network DPFeat in the embodiments of the present application,
[0041] Figure 4 An effect comparison diagram of the embodiments of the present application and the general method in the image homonym point extraction task;
[0042] Figure 5 An effect comparison diagram of the embodiments of the present application and the general method in the image homonym point matching task;
[0043] Figure 6 An effect comparison diagram of the embodiments of the present application and the general method in the planetary surface sparse three-dimensional reconstruction task. DETAILED DESCRIPTION
[0044] In order to better understand the technical solutions of the present application, the present application will be further described in detail in combination with the drawings and embodiments.
[0045] The present application proposes to apply the deep neural network and the incremental motion recovery structure technology in the three-dimensional sparse reconstruction of planetary exploration vehicle image. Specifically, a new type of deep neural network which fuses attention mechanism is designed for image key point extraction and description, and a robust gross error elimination method is combined, so that accurate and reliable homonym points can be obtained even on the weak texture and strong light change exploration vehicle image, and then the incremental motion recovery structure technology is used to register the images one by one, perform triangulation and bundle adjustment, and finally recover the camera pose and attitude and the three-dimensional coordinates of the object point, as shown in Figure 1 The advantage of this reconstruction method is that it is fully automatic, and only the images with mutual overlap in a certain area need to be input, without other information, the position and attitude of the camera can be directly recovered, the camera internal parameter can be solved, and the accurate and reliable three-dimensional coordinates of the ground point can be obtained, thereby providing a data basis for subsequent other processing and analysis.
[0046] Therefore, the key improvement of the present application is that: in view of the problem that the traditional feature extraction and matching method fails to process the images of planetary exploration vehicles due to weak texture, changing light and other factors, a deep neural network for image key point extraction and description is designed, which uses the strong generalization ability of the deep convolutional network to extract sufficient and reliable key points from the images of the exploration vehicles; then a rough error elimination method based on prior information is used to accurately match the image pairs; finally, the incremental motion recovery structure technology based on the principle of photogrammetry is used to obtain the camera pose and parameters, and the three-dimensional terrain structure through iterative calculation optimization.
[0047] Embodiment 1
[0048] Referring to Figure 2 , the embodiment provides a deep learning three-dimensional sparse reconstruction method suitable for planetary exploration vehicle images, comprising the following steps:
[0049] Step one: import the pre-processed planetary exploration vehicle images and other auxiliary data: the format of the planetary exploration vehicle images published is not the commonly used image format, so the image needs to be converted into the commonly used image format by means of the related processing method; in addition, other auxiliary data such as camera angle, height, focal length and the like information can be contained, which can be used as initial values in subsequent calculation.
[0050] Step two: for each imported exploration vehicle image, use the deep convolutional neural network DPFeat to extract and describe the key points, to obtain a key point set containing the two-dimensional image coordinates of the key points on each image and the corresponding descriptors.
[0051] As Figure 3 shown, DPFeat is a deep convolutional neural network, the backbone network of which is based on the L2-Net network, and the upsampling layer in the network is replaced by a dilated convolution, which is conducive to obtaining more sparse prediction values; after the backbone network, there are two branch networks, one for generating image coordinates of key points and the other for generating descriptors corresponding to the key points, each branch network fuses an attention mechanism on the basis of the convolutional network, which improves the weight of the key feature map and reduces the weight of the irrelevant feature map, realizes the efficient use of high-dimensional feature maps, and finally achieves the purpose of extracting sufficient and reliable key points on the images of planetary exploration vehicles with weak texture and strong light variation.
[0052] The attention mechanism is fused in the branch network because the ordinary convolution itself is difficult to model the channel relationship between feature maps, and the attention mechanism has good universality and efficiency. For a feature map γ with a dimension of WxHxC, where W is the width, H is the height, and C is the number of channels, a tensor z with a dimension of (1x1xC) is first obtained through a global average pooling layer, so that the network can collect global information; then, based on the global information, weights of different channels are generated to perform weighted processing on the channels.
[0053] The two branches of the DPFeat network are supervised by a loss function during network training, including similarity loss, extreme value loss, and reliability loss.
[0054] The similarity loss formula is as follows:
[0055]
[0056] where I and I' represent two images, U represents the corresponding relationship between the same named points of the two images, P represents a patch, S is the probability map of the key points of the image I generated by the network, S' U is the probability map of the key points of the image I' after transformation according to U, the cosim function calculates the similarity of the input quantity by cosine, the greater the value is, the more similar it is, and the probability of the key point is also higher, and the loss value is lower;
[0057] The extreme value loss formula is as follows:
[0058]
[0059] where N represents the neighborhood size, the max and mean functions are used to find the maximum and average value respectively, the greater the difference of this point is, the smaller the extreme value loss is;
[0060] The reliability loss formula is as follows:
[0061]
[0062] where B represents the number of patches in a batch, R represents the reliability map, is a differentiable function for supervising the learning of the descriptor.
[0063] Using the DPFeat network, a large number of key points with reasonable distribution can be extracted from the sparse texture detection vehicle image, such as Figure 4As shown, the key points extracted by the general method are concentrated in the areas with rich image texture, and are unevenly distributed, which increases the difficulty and instability of matching and solving; the improved algorithm proposed in the application can extract uniform key points on the whole image, and can obtain a considerable number of key points even on a flat ground lacking texture, so that the three-dimensional reconstruction process is more stable.
[0064] Step three: matching the key points according to the positions of the key points and the descriptor information, so as to extract the homonymous points in the images with mutual overlap, and using the improved gross error elimination method, eliminating the wrong homonymous points and matching more correct homonymous points.
[0065] In step three, the image is first normalized, then the image is divided into multiple blocks, the reliability of the matching results in each block is calculated, the normal values and the gross errors in each block are classified respectively, and then the gross errors are merged, so that the gross errors are finally eliminated.
[0066] Figure 5 The results of key point extraction and matching using the general method and the method proposed in the application are shown in the figures. The feature points obtained by the general algorithm are not only unevenly distributed, but also have many matching errors; using the improved algorithm proposed in the application, more key points can be extracted, which are more reasonably distributed, and the accuracy of key point matching is also significantly improved, which can greatly help the three-dimensional reconstruction.
[0067] Step four: on the basis of the key points and the matching results obtained in step three, the incremental motion recovery structure technology first selects two adjacent images for initialization, then uses the matching relationship between the key points of the triangular points in the registered images to register other images to the current model, the newly registered images not only cover the existing object points, but also increase the coverage range through triangulation, on this basis, the bundle adjustment optimization is used to calculate the results, and the outliers are eliminated; after iterative calculation, the accurate object point coordinates and the position and attitude information of the camera are output.
[0068] Figure 6 The results of sparse three-dimensional reconstruction using the general method and the method proposed in the application using the images of the planetary exploration vehicle are shown in the figures. The results obtained by the general algorithm have many gross errors, which leads to the wrong estimation of the position and attitude of some cameras, and because the number of matching key points is small, the three-dimensional reconstruction result is obviously incomplete in many areas, and the overall result is sparse; when the improved algorithm proposed in the application is used for three-dimensional reconstruction, the position and attitude of all cameras can be correctly estimated, and because the key point extraction and description algorithm based on deep learning is used, more dense homonymous point results are obtained, and the final three-dimensional reconstruction result is also denser.
[0069] Through the above process, the application realizes key improvements:
[0070] The deep learning three-dimensional sparse reconstruction method for planetary rover images described in steps two to four solves the problem that traditional feature extraction and matching methods fail when processing planetary rover images due to weak texture, changing illumination, and other factors, thereby causing three-dimensional reconstruction to fail, and realizes high-precision and robust sparse three-dimensional reconstruction of planetary terrain, which can provide reliable data basis for planetary exploration missions.
[0071] In specific implementation, the method proposed by the technical solution of the application can be automatically run by a person skilled in the art using computer software technology, and the system device of the method, such as a computer readable storage medium storing the corresponding computer program of the technical solution of the application and a computer device including the corresponding computer program, should also be within the protection scope of the application.
[0072] Embodiment 2
[0073] The embodiment provides a deep learning three-dimensional sparse reconstruction system suitable for planetary rover images, comprising:
[0074] An image data import module is used to import preprocessed planetary rover images and other auxiliary data;
[0075] A key point extraction and description module adopts a deep convolutional neural network DPFeat that fuses an attention mechanism to extract and describe key points of each imported image to obtain a key point set containing two-dimensional image coordinates of key points on each image and corresponding descriptors;
[0076] A key point matching module matches key points according to the positions of the key points and the descriptor information, thereby extracting homonymous points in images that overlap each other and eliminating incorrect homonymous points while matching more correct homonymous points;
[0077] A solving module uses the coordinates of the key points and the corresponding relationship of the homonymous points obtained by matching to solve accurate object three-dimensional point coordinates and the position and pose of the camera, and further restores the three-dimensional structure of the terrain.
[0078] The system in the above embodiment is used to solve the problem of weak texture and large illumination variation of planetary rover image data, and realizes sparse three-dimensional reconstruction of the planetary surface and solving of the position and pose of the camera.
[0079] In some possible embodiments, a deep learning three-dimensional sparse reconstruction system suitable for planetary rover images is provided, comprising a processor and a memory, the memory being configured to store program instructions, and the processor being configured to invoke the program instructions stored in the memory to execute a deep learning three-dimensional sparse reconstruction method suitable for planetary rover images as described above.
[0080] In some possible embodiments, a deep learning three-dimensional sparse reconstruction system suitable for planetary rover images is provided, comprising a readable storage medium, and a computer program is stored on the readable storage medium, and the computer program is configured to implement a deep learning three-dimensional sparse reconstruction method suitable for planetary rover images as described above when executed.
[0081] The above description is merely preferred specific embodiments of the present application. The scope of the present application is not limited to the above description, and any simple changes or equivalent replacements within the scope of the present application disclosed herein are included in the scope of the present application.
Claims
1. A deep learning three-dimensional sparse reconstruction method suitable for planetary rover imagery, characterized in that, Comprising the following steps: Step one, import the pre-processed planetary rover image and other auxiliary data; Step two, adopt the deep convolutional neural network DPFeat which fuses attention mechanism to extract and describe the key points of each imported image, and obtain a key point set which contains the two-dimensional image coordinates of the key points on each image and the corresponding descriptors; in the step two, the deep convolutional neural network DPFeat used for extracting and describing the key points replaces the up-sampling layer in L2-Net with a dilated convolution, and after the image is processed by the backbone network, two branches are generated, one is used to generate the image coordinates of the key points, and the other is used to generate the descriptors corresponding to the key points, each branch fuses the attention mechanism to improve the weight of the key feature maps and reduce the weight of the irrelevant feature maps, and realizes the efficient use of high-dimensional feature maps; the two branches of the DPFeat network are supervised by loss functions during network training, including similarity loss, extreme loss and reliability loss; The similarity loss formula is as follows: wherein, and represent two images, represents the correspondence of the same point between the two images, represents a tile, the image generated by the network the probability map of the key point, represents the image the probability map of the key point according to the transformed probability map, The function calculates the similarity of the input quantity by cosine, the greater the similarity, the higher the probability of the key point, and the lower the loss value. The extreme loss formula is as follows: wherein represents a neighborhood size, represents the value of a point on the keypoint probability map with horizontal and vertical coordinates i and j, respectively, and max and mean are functions for finding the maximum and mean value, respectively, the greater the difference between the points, the smaller the extremum loss. The reliability loss formula is as follows: wherein represents the number of tiles in a batch, represents a reliability map, represents the value of a point on the reliability map with i and j as horizontal and vertical coordinates, respectively, represents the value of a point on tile p with i and j as horizontal and vertical coordinates, respectively, is a differentiable function for supervised learning descriptors; Step three, match the key points according to the two-dimensional image coordinates of the key points and the descriptor information, so as to extract the homonymic points in the images with overlap and eliminate the wrong homonymic points while matching more correct homonymic points; Step four, use the coordinates of the key points and the corresponding relationship of the homonymic points obtained by matching to calculate the accurate object side three-dimensional point coordinates and the position and attitude of the camera, and then restore the three-dimensional structure of the terrain. 2.The deep learning three-dimensional sparse reconstruction method for images of a planetary rover according to claim 1, wherein: In step one, the rover data needs to be converted from the original format when published into a commonly used image format, and the camera position, attitude, internal parameter information obtained from other data sources are also imported as initial values for reconstruction. 3.The deep learning three-dimensional sparse reconstruction method for images of a planetary rover according to claim 1, wherein: In step three, the improved gross error elimination method is used to eliminate the wrong homonymic points, wherein the improved gross error elimination method is specifically as follows: Firstly, the image is normalized, then the image is divided into multiple blocks, the reliability of the matching results in each block is calculated, the normal value and the gross error in each block are classified respectively, and then the gross error is merged to finally realize the elimination of the gross error. 4.The deep learning three-dimensional sparse reconstruction method for image of planetary rover according to claim 1, wherein: In the step four, based on the incremental motion recovery structure technology, the bundle adjustment technology is used to calculate the accurate object side three-dimensional point coordinates and the position and attitude of the camera, and then the three-dimensional structure of the terrain is restored.
5. The deep learning three-dimensional sparse reconstruction method suitable for planetary rover images according to claim 4, characterized in that: The incremental motion recovery structure technology firstly selects two adjacent images for initialization, then uses the key point matching relationship of the triangular points in the registered images to register other images to the current model, the newly registered images not only cover the existing object side points, but also increase the coverage range through triangulation, and on this basis, the bundle adjustment optimization calculation result is used to eliminate the outlying measurement values; after the iterative calculation, the accurate object side point coordinates and the position and attitude information of the camera are output. 6.A deep learning three-dimensional sparse reconstruction system suitable for planetary rover imagery, characterized in that, Comprising: An image data import module is configured to import preprocessed images of a planetary rover and other auxiliary data; A key point extraction and description module is configured to extract and describe key points of each imported image by using a deep convolutional neural network DPFeat with an attention mechanism, to obtain a key point set including two-dimensional image coordinates of the key points of each image and corresponding descriptors; A key point matching module is configured to match the key points according to the positions and descriptor information of the key points, to extract homonymous points in images with overlaps, and to eliminate incorrect homonymous points and match more correct homonymous points; A solving module is configured to solve accurate object coordinates and positions and poses of the camera by using the coordinates of the key points and the corresponding relationship of the homonymous points obtained by matching, to recover a three-dimensional structure of a terrain. The deep learning three-dimensional sparse reconstruction system suitable for images of a planetary rover is configured to perform the steps of the deep learning three-dimensional sparse reconstruction method suitable for images of a planetary rover in any one of claims 1-5.
7. The deep learning three-dimensional sparse reconstruction system for image of planetary rover according to claim 6, wherein: The deep learning three-dimensional sparse reconstruction system suitable for images of a planetary rover is configured to perform the steps of the deep learning three-dimensional sparse reconstruction method suitable for images of a planetary rover in any one of claims 1-5. The deep learning three-dimensional sparse reconstruction system suitable for images of a planetary rover is configured to perform the steps of the deep learning three-dimensional sparse reconstruction method suitable for images of a planetary rover in any one of claims 1-5. The deep learning three-dimensional sparse reconstruction system suitable for images of a planetary rover is configured to perform the steps of the deep learning three-dimensional sparse reconstruction method suitable for images of a planetary rover in any one of claims 1-5. The deep learning three-dimensional sparse reconstruction system suitable for images of a planetary rover is configured to perform the steps of the deep learning three-dimensional sparse reconstruction method suitable for images of a planetary rover in any one of claims 1-5.