Stable multi-modal data fusion spacecraft three-dimensional reconstruction method based on deep learning

By using multimodal data fusion and deep learning technologies, and leveraging data from lidar and cameras, the problems of insufficient accuracy and computational resource consumption in satellite 3D reconstruction have been solved, achieving efficient and accurate 3D reconstruction results, especially with outstanding performance in the space environment.

CN120852699APending Publication Date: 2025-10-28INNOVATION ACAD FOR MICROSATELLITES OF CAS +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510659174.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing satellite 3D reconstruction schemes suffer from insufficient accuracy, susceptibility to interference, and high computational resource consumption, failing to meet the demand for efficient and accurate 3D reconstruction, especially in complex space environments.

Method used

A deep learning-based multimodal data fusion method is adopted, which uses LiDAR and camera to acquire multiple sparse depth images and color images. Feature extraction and fusion are performed through CSPN++ network. The information from the sparse depth images and color images is combined to perform depth completion, generating a dense depth image. Finally, a 3D model is generated through point cloud registration and fusion.

Benefits of technology

It improves the accuracy and efficiency of 3D reconstruction, can generate high-quality 3D models in complex environments, reduces computing resource consumption, and is suitable for resource-constrained satellite systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852699A_ABST
    Figure CN120852699A_ABST
Patent Text Reader

Abstract

The invention discloses a stable multi-modal data fusion spacecraft three-dimensional reconstruction method based on deep learning. The method comprises the following steps: acquiring a plurality of sparse depth images and color images of a spacecraft under different visual angles by using a laser radar and a camera; preprocessing the sparse depth image and the color image under the same visual angle; in combination with the preprocessed color image, complementing the preprocessed sparse depth image by adopting a CSPN + + network to obtain a dense depth image, and further converting the dense depth image into a point cloud image; repeating the process to obtain all point cloud images under different visual angles; and carrying out coarse registration and fine registration on the point cloud images under different visual angles by adopting an SIFT feature matching algorithm and an iterative closest point (ICP) algorithm in sequence so as to obtain a complete three-dimensional point cloud of the spacecraft, thereby carrying out three-dimensional model reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of three-dimensional reconstruction, specifically relating to a stable multimodal data fusion method for three-dimensional reconstruction of spacecraft based on deep learning. Background Technology

[0002] Current satellite 3D reconstruction solutions on the market are mainly divided into active and passive methods. Although methods combining LiDAR and camera images with deep learning have made breakthroughs in certain application scenarios, these solutions still have obvious shortcomings, including the inability to provide high-precision, dense 3D detail reconstruction, poor adaptability to special environments (such as low light, complex backgrounds, etc.), and high computational resource requirements, which cannot meet the needs of efficient and accurate 3D reconstruction for satellites.

[0003] Specifically, active methods such as LiDAR and structured light sensors present challenges. While LiDAR can quickly acquire relatively accurate depth information, its point clouds are sparse and lack detail, failing to provide sufficiently detailed 3D structures, especially in complex satellite maintenance tasks where it cannot meet the high-precision reconstruction requirements. Structured light sensors have extremely limited working distances, making them unsuitable for practical applications of 3D reconstruction of space satellites. Passive methods, such as sparse 3D point clouds generated from camera images and SfM technology, face severe challenges in the space environment. Due to special lighting conditions, low contrast, and complex backgrounds, image quality is often severely affected, leading to noise, uneven illumination, and insufficient satellite surface features (such as low-texture or repetitive texture areas). This significantly impacts the accuracy of 3D reconstruction, resulting in large errors in the generated point clouds and hindering the accurate reconstruction of satellite surface details. Although deep learning-driven image 3D reconstruction methods (such as the MVSNet series) have made some progress in recent years, these methods typically require substantial computational resources, especially when processing high-resolution images and high-resolution point clouds, posing a significant challenge to the resources of satellite systems. In a limited satellite computing environment, deep learning models may result in excessively long processing times and an inability to output reconstruction results in a timely manner, thereby affecting the on-orbit service efficiency of spacecraft. Summary of the Invention

[0004] This invention proposes a stable multimodal data fusion method for spacecraft 3D reconstruction based on deep learning, which solves the technical problems of insufficient accuracy, susceptibility to interference, and high computational resource consumption in existing satellite 3D reconstruction schemes.

[0005] This invention can be achieved through the following technical solutions:

[0006] A stable multimodal data fusion method for spacecraft 3D reconstruction based on deep learning includes the following steps:

[0007] Step 1: Use lidar and cameras to obtain multiple sparse depth and color images of the spacecraft from different perspectives;

[0008] Step 2: Preprocess the sparse depth image and color image from the same viewpoint;

[0009] Step 3: Combine the preprocessed color image with the CSPN++ network to complete the preprocessed sparse depth image, obtain the dense depth image, and then convert it into a point cloud image.

[0010] Step 4: Repeat steps 2 and 3 to obtain all point cloud images from different perspectives;

[0011] Step 5: The SIFT feature matching algorithm and the Iterative Closest Point (ICP) algorithm are used to perform coarse and fine registration on point cloud images from different perspectives to obtain a complete 3D point cloud of the spacecraft, which is then used for 3D model reconstruction.

[0012] Furthermore, in step three, the CSPN++ network is used to extract features from the preprocessed sparse depth image and color image respectively, and then the features are fused using the following equation to obtain the dense depth image F(x,y).

[0013] F(x,y)=α·F I (x,y)+(1-α)·F D (x,y)

[0014] Among them, F D (x,y) represents the depth features extracted from the preprocessed sparse depth image, α represents the fusion weight, and F I (x,y) represents the color features extracted from the preprocessed color image.

[0015] Firstly, the spatial propagation mechanism in the CSPN++ network is used to perform depth completion on the aforementioned fused feature F(x,y). The known depth values ​​are extended to unknown regions through the propagation process between adjacent pixels. During propagation, the network uses spatial and color information to infer the depth values. The propagation equation can be expressed as:

[0016]

[0017] Where D(x′,y′) is the depth value at the neighboring pixel (x′,y′); W(x,y,x′,y′) is the propagation weight between pixels (x,y), which is usually determined by image features (such as texture and color) and spatial distance; f(I(x,y)) is the color information extracted from the color image I(x,y), used to guide the completion of depth information; and N(x,y) is the neighborhood set of pixel (x,y).

[0018] To enhance the network's understanding of the spatial structure between pixels, CSPN++ employs positional encoding. This is achieved by converting the coordinates (x, y) of each pixel into a high-dimensional representation p(x, y). Positional encoding can be constructed as follows:

[0019] p(x,y)=[sin(W1x+b1),cos(W1x+b1),sin(W2y+b2),cos(W2y+b2)]

[0020] Where W1 and W2 are the learned weights, b1 and b2 are the bias terms, and (x, y) are the coordinates of the pixel.

[0021] Furthermore, color features F extracted from the preprocessed color image are used using a convolutional neural network. I (x,y), depth features F extracted from the preprocessed sparse depth image. D (x,y) represents the depth information corresponding to each coordinate point in the sparse depth image.

[0022] Furthermore, record the camera's intrinsic parameter matrix. Among them, f x and f y Indicates the camera's focal length, (c x ,c y () indicates the optical center of the camera, which is the center point of the image.

[0023] The following formula is used to iterate through each pixel in the dense depth image F(x,y), and the depth value D(x,y) of each pixel (x,y) is converted into three-dimensional spatial coordinates (X,Y,Z) to obtain the point cloud representation of the dense depth image F(x,y), which is the point cloud image.

[0024]

[0025] Z = D(x,y)

[0026] Firstly, noise may be introduced during the depth completion process. Therefore, the bilateral filtering method in the data preprocessing stage is used again for noise reduction to remove noise while preserving the edge information of the image as much as possible.

[0027] Furthermore, in step five, voxel mesh filtering is applied to the 3D point clouds acquired from each viewpoint to reduce redundancy and error accumulation; then, point cloud registration and fusion are performed to obtain a complete 3D point cloud; finally, Delaunay triangulation is used for surface reconstruction to generate a 3D mesh model.

[0028] The beneficial technical effects of this invention are as follows:

[0029] The 3D reconstruction method of this invention optimizes the data flow and processing process from a structural perspective by fusing multimodal data, using deep learning-guided completion, optimizing point cloud registration and fusion, and using lightweight network design. This reduces errors and processing time in traditional methods and improves efficiency and accuracy.

[0030] (1) Multimodal data fusion: Fully utilize the advantages of lidar and camera methods. LiDAR data provides accurate depth information to effectively guide camera images, enhancing the main satellite region in the image and filtering out background interference such as the Earth; camera images are used to supplement the details of the sparse point cloud of lidar. By taking the best of both worlds and complementing each other, the combination of the two can generate high-quality dense point clouds.

[0031] (2) Introducing a series of deep learning methods for depth completion, resulting in robust reconstruction: Compared to traditional geometric completion methods, the CSPN++ network can learn from training data and better understand the contextual information in the image. Through deep learning, it intelligently completes sparse point clouds, filling in missing depth information, thereby generating more realistic and accurate 3D reconstructions.

[0032] (3) Strong anti-interference ability and adaptability to complex environment: The lighting in the space environment is complex, and traditional camera images are often severely affected. This invention can effectively remove these interferences by introducing the depth information of the lidar, ensuring the accuracy and quality of 3D reconstruction. Especially in low texture areas and high dynamic scenes, this solution can give full play to its anti-interference ability.

[0033] (4) Low computational resource consumption and improved real-time performance: By adopting the CSPN++ network with excellent resource utilization efficiency and combining optimized processing of sparse point cloud and image data, efficient computing can be achieved in resource-constrained environments. Attached Figure Description

[0034] Figure 1 This is a schematic diagram of the overall process of the present invention. Detailed Implementation

[0035] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings and preferred embodiments.

[0036] like Figure 1As shown, this invention provides a stable multimodal data fusion method for spacecraft 3D reconstruction based on deep learning. First, synchronous acquisition by LiDAR and cameras provides high-precision depth information and rich texture details for subsequent processing. Then, using the sparse depth map obtained from the LiDAR, a dense depth map is generated through nearest-neighbor interpolation. Next, the foreground region is extracted based on a set depth threshold to generate a foreground mask. Finally, the mask is used to filter out background interference from the camera image. After background removal, the image enters a filtering and enhancement stage, further improving image quality through techniques such as denoising, normalization, and image smoothing. Next, the core step of depth completion is performed, guiding the generation of a dense depth image through a CSPN++ network. After denoising and filtering, the dense depth image can be converted into a high-quality point cloud image, which is then further registered and fused with point cloud data. Surface reconstruction is then performed on the point cloud model to obtain the final 3D mesh model. In this way, by fusing LiDAR and camera data and combining deep learning technology, the technical problems of insufficient accuracy, susceptibility to interference, and high computational resource consumption in existing satellite 3D reconstruction schemes are solved, thereby improving the accuracy and efficiency of 3D reconstruction. This invention can effectively utilize the high-precision depth information provided by lidar and the rich details of camera images to optimize the performance of traditional methods in complex environments, especially in the complex space environment with severely uneven lighting, ensuring the generation of fine and accurate 3D network models.

[0037] Specifically as follows:

[0038] Step 1: Use lidar and cameras to obtain multiple sparse depth and color images of the spacecraft from different perspectives to achieve multimodal data acquisition.

[0039] The multimodal data acquisition section utilizes LiDAR and a camera to acquire depth information and image data, respectively. The LiDAR generates precise sparse point cloud data, i.e., a sparse depth image, by emitting laser pulses and measuring their reflection time, providing high-precision spatial depth information. The camera, on the other hand, captures the color and texture information of the scene, obtaining a color image to provide rich visual data. To ensure accurate fusion of the two data sets, synchronous acquisition is required. This invention employs software timestamp alignment technology to ensure precise temporal alignment between the LiDAR and camera data.

[0040] Step 2: Preprocess the sparse depth image and color image from the same viewpoint.

[0041] ①Use sparse depth images obtained by lidar as a reference to filter out space background interference in the images.

[0042] Since the depth map obtained by LiDAR contains only limited effective depth information, it needs to be interpolated to generate a denser depth map, thereby effectively filtering out background interference. In the sparse depth map, regions where no depth value can be obtained are marked as invalid values ​​(such as NaN or 0), and then a coarse dense depth map is generated using the nearest neighbor interpolation method.

[0043] Nearest neighbor interpolation involves finding the nearest known point (a point with a depth value) for each unknown point (i.e., a point without a depth value) in a sparse depth image, and then assigning the depth value of the known point to the unknown point.

[0044] Suppose we have a sparse depth image D, where the coordinates of a known point are (x, y, y). i ,y i ), with a depth value of z i For any unknown point (x, y), its depth value z is calculated using the following formula:

[0045] z(x,y)=z i* (x i* ,y i* )

[0046] Among them, (x i* ,y i* () represents the known point closest to the unknown point (x, y) that satisfies:

[0047]

[0048] After obtaining a rough dense depth image, a depth threshold is selected to separate the foreground and background. If depth_threshold is used as the threshold, a boolean array can be created to set the foreground part with a depth value less than depth_threshold to True, and the part with a depth value greater than the threshold and invalid values ​​to False, thereby generating a foreground mask and filtering out background interference.

[0049] ②Pixel intensity normalization and smoothing

[0050] Pixel normalization converts the pixel values ​​of a color image to a standard range, usually [0,1] or [-1,1]. This facilitates subsequent processing, alleviates overexposure and underexposure caused by severe uneven lighting in space, and avoids affecting the stability of calculations due to pixel value ranges that are too large or too small.

[0051] Normalization is typically performed using the following formula:

[0052] Among them, I original It is the pixel value in the original pixel, I min and I maxThese are the minimum and maximum values ​​of all pixels in the image, I. normalized It is the normalized pixel value.

[0053] Image smoothing is commonly used to reduce noise in color images. This invention uses a bilateral filtering method that combines spatial distance (distance between pixel locations) and pixel value similarity (difference between pixel values) into a weighted average. This effectively removes noise while preserving edge details, thereby enhancing the edges of space images and providing cleaner, clearer input for subsequent networks.

[0054] The output of a bilateral filter can be expressed as:

[0055]

[0056] Among them, I filtered (x,y) is the filtered pixel value, Ω is the neighborhood region of the current pixel (x,y), and f s (||(x′,y′)-(x,y)||) is the spatial weight function, usually a Gaussian function, which measures the spatial distance between pixels. r (||I(x′,y′)-I(x,y)||) is the color / brightness similarity weight function, usually a Gaussian function, which measures the difference in pixel values. p It is a normalization factor used to ensure that the sum of the weights is 1.

[0057] The biggest advantage of bilateral filtering is its ability to effectively remove noise while preserving the image's edge information. Since edges typically correspond to sharp changes in pixel values, bilateral filtering reduces the smoothness at edges, making it ideal for preserving image edges, especially effective in removing high-frequency noise.

[0058] Step 3: Combine the preprocessed color image with the CSPN++ network to complete the preprocessed sparse depth image, obtain the dense depth image, and then convert it into a point cloud image.

[0059] 1. Sparse depth image completion

[0060] To obtain denser and more accurate depth information for satellite 3D reconstruction, this invention employs the CSPN++ network. The main objective of CSPN++ is to recover complete depth information from sparse depth maps and corresponding color images. It primarily consists of an input module, a feature fusion module, a propagation module, and an output module.

[0061] Sparse depth map D(x) input to CSPN++ i ,y i ) represents the sparse points sampled in the image, where (x) i ,y iD(x) represents the pixel coordinates in the image. i ,y i (x) represents the depth value of that point. Depth maps contain many missing regions in space, meaning some pixels (x) are missing. j ,y j The depth value D(x) j ,y j The color image I(x,y) is unknown, while the input color image I(x,y) provides color information for each pixel. This information will help the network utilize the texture features in the image when completing the depth map.

[0062] First, the input sparse depth map D(x) i ,y i The color image I(x,y) and its corresponding color image F will pass through a feature extraction module. For color images, a convolutional neural network is typically used to extract color features F. I (x,y), i.e., F I (x,y)=

[0063] For sparse depth maps, CNN(I(x,y)) directly processes the depth information, which is the depth feature.

[0064] After passing through the feature fusion module, sparse depth information and image features are fused to obtain a comprehensive feature representation F(x,y), which is used for depth completion. This fusion process can be implemented through a simple weighted summation or a more complex attention mechanism: F(x,y) = α·F I (x,y)+(1-α)·F D (x,y), where F D (x,y) are the depth features extracted from the sparse depth map, and α is the fusion weight. After obtaining the fusion features, CSPN++ guides the depth information to propagate along the neighborhood in the image space through convolution operations combined with the spatial propagation mechanism. Through multiple iterations, the unknown area is gradually updated according to the depth information of the neighboring pixels, thereby achieving depth completion.

[0065] Finally, the network outputs a completed depth map D(x,y), which is a complete depth map derived from the sparse depth map and image information to obtain a dense depth image.

[0066] The core of CSPN++ is the spatial propagation mechanism, which utilizes the known sparse depth information D(x) in the depth map. i ,y iThe network uses spatial and color information to gradually fill in missing depth information. The spatial propagation process is as follows: Assuming D(x,y) represents the depth value of the depth image at pixel (x,y), the goal is to extend the known depth value to the unknown region through the propagation process of neighboring pixels. During propagation, the network uses spatial and color information to infer the depth value. The propagation equation can be expressed as:

[0067]

[0068] Where D(x′,y′) is the depth value at the neighboring pixel (x′,y′); W(x,y,x′,y′) is the propagation weight between pixels (x,y), which is usually determined by image features (such as texture and color) and spatial distance; f(I(x,y)) is the color information extracted from the color image I(x,y), used to guide the completion of depth information; and N(x,y) is the neighborhood set of pixel (x,y).

[0069] Through the above propagation equations, the depth value D(x,y) of the current pixel is gradually inferred using the depth information of neighboring pixels and the color information of the color image, until the missing parts of the entire depth map are filled. Compared with the CSPN network, CSPN++ further improves effectiveness and efficiency by learning adaptive convolutional kernel size and propagation iterations, thereby dynamically allocating the context and computational resources required for each pixel according to the request.

[0070] To enable the network to understand the spatial relationships between pixels, CSPN++ employs positional encoding. Positional encoding converts the coordinates (x, y) of each pixel into a high-dimensional representation p(x, y), allowing the network to convey spatial information using these codes. Specifically, positional encoding can be constructed as follows:

[0071] p(x,y)=[sin(W1x+b1),cos(W1x+b1),sin(W2y+b2),cos(W2y+b2)]

[0072] Where W1 and W2 are the learned weights, b1 and b2 are bias terms, and (x, y) are the coordinates of the pixel. Positional encoding maps each pixel in the image to a high-dimensional space, helping the network understand spatial dependencies.

[0073] By first performing coordinate encoding to provide spatial context information for each pixel, and then using a spatial propagation mechanism, CSPN++ can gradually complete the missing depth information using sparse depth and color image information, ultimately generating a complete depth map. This network enables multimodal information fusion, effectively processes sparse point clouds from LiDAR, and fully utilizes image information captured by the camera, thereby improving completion accuracy.

[0074] 2. Depth Map Denoising and Point Cloud Generation

[0075] Noise may be introduced during depth completion. Here, we use the bilateral filtering method from the data preprocessing stage to remove noise while preserving the edge information of the image as much as possible.

[0076] To obtain a 3D representation of a satellite in space, the depth information needs to be further converted into point cloud information. This requires a depth map D(x,y) and the camera's intrinsic parameter matrix K. The camera's intrinsic parameter matrix K contains the camera's focal length and optical center position information, and is typically as follows:

[0077]

[0078] Among them, f x and f y It is the camera's focal length (usually measured in pixels), (c x ,c y ) is the optical center of the camera (i.e., the center point of the image).

[0079] Using the camera's intrinsic parameter matrix and depth map information, the following formula can be used to iterate through each pixel in the depth map, converting the depth value D(x,y) of each pixel (x,y) in the depth map into three-dimensional spatial coordinates (X,Y,Z):

[0080]

[0081] Z = D(x,y)

[0082] Ultimately, all the three-dimensional points (X, Y, Z) constitute a point cloud.

[0083] Step 4: Repeat steps 2 and 3 to obtain all point cloud images from different perspectives.

[0084] Step 5: The SIFT feature matching algorithm and the Iterative Closest Point (ICP) algorithm are used to perform coarse and fine registration on point cloud images from different perspectives to obtain a complete 3D point cloud of the spacecraft, which is then used for 3D model reconstruction.

[0085] (1) Point cloud registration and fusion

[0086] After obtaining point cloud representations of satellites from multiple perspectives, it is necessary to align the point cloud data from different perspectives so that they share the same coordinate system. Point cloud registration is divided into two stages: coarse registration and fine registration.

[0087] The goal of the coarse registration stage is to find a rough alignment result. This scheme adopts a feature-based registration method. First, key features such as normal vectors and corner points are extracted from the point cloud. Then, the SIFT feature matching algorithm is used to calculate the preliminary transformation T0 between the point clouds (including the rotation matrix R0 and the translation vector t0).

[0088]

[0089] Where R0 is the rotation matrix, t0 is the translation vector, and T0 is the initial transformation matrix.

[0090] The goal of the fine registration stage is to further reduce registration errors through iterative optimization. This scheme uses the Iterative Closest Point (ICP) algorithm to optimize the registration results by minimizing the distance error between point clouds. The principle is as follows: Assume two point clouds P A ={p i} and P B ={q j}, where p i and q j These are points in point cloud A and point cloud B, respectively. The core of the ICP algorithm is to solve for the transformation matrix by minimizing the objective function.

[0091]

[0092] Where, p i q is a point in point cloud A. i is the matching point in point cloud B, N is the number of point pairs selected in the registration, R and t are the rotation matrix and translation vector respectively, and ||·|| represents the Euclidean distance.

[0093] By iteratively optimizing the objective function, the ICP algorithm will eventually obtain the optimized transformation matrix T. * This allows the two point clouds to be aligned while minimizing the error.

[0094] The effectiveness of point cloud registration can be evaluated by calculating the distance error between each pair of matched points. The registration error is defined as the distance error between each point p in the registered point cloud A. i to its matching point q i Distance:

[0095]

[0096] This error can be used to assess the registration accuracy; the smaller the error, the higher the registration accuracy.

[0097] After completing the point cloud registration, all satellite point clouds need to be merged into a complete 3D satellite point cloud. During the fusion process, redundancy will occur between different satellite point clouds due to the overlap of multiple viewpoints. To solve the redundancy problem, this invention adopts a voxel mesh filtering method. First, the space is divided into multiple voxels, and the points within each voxel are averaged or a representative point is selected to reduce redundancy. The specific operation is as follows:

[0098]

[0099] Where, p avg It is the average position of all points within a voxel, where N is the number of points within the voxel, and p i It refers to each point in a voxel.

[0100] After redundancy removal, a weighted fusion method is used for point cloud fusion. Different weights are assigned to each point based on its quality (e.g., reliability of depth values, viewing angle, sensor accuracy), and the specific operation is as follows:

[0101]

[0102] Where, p fused These are the coordinates of the merged satellite point cloud. wi p is the weight of the i-th point. i It is the position of the i-th point.

[0103] (2) Model Surface Reconstruction

[0104] Although the point set in point cloud data has been aligned and redundancy removed, surface structure information is still lacking. This solution employs the Delaunay triangulation method, connecting point cloud data into a triangular mesh to achieve surface reconstruction.

[0105] In three-dimensional space, Delaunay triangulation involves not only triangles on a two-dimensional plane but also the construction of tetrahedrons. A tetrahedron is a tetrahedral face composed of four points, where each face is a triangle composed of three points. Therefore, the mesh generated by Delaunay triangulation in three-dimensional space consists of multiple tetrahedrons, each of which satisfies the condition that its circumsphere does not contain any other points.

[0106] For a Delaunay triangulation of a 3D point set, the generated tetrahedrons satisfy the following condition: for any tetrahedron T = {p i ,p j ,p k ,p l Its circumscribed ball S ijkl It does not contain any other point p m ∈P:

[0107]

[0108] Among them, S ijkl The circumsphere of tetrahedron T is the sphere that encloses the tetrahedron. The property of the circumsphere ensures the reasonable spatial relationship between the point sets that constitute the Delaunay tetrahedron. Only under the condition of the circumsphere can the point sets meet the requirements of Delaunay triangulation.

[0109] This invention uses the Bowyer-Watson algorithm to progressively insert points, checking at each insertion whether the point is compatible with the circumsphere of the existing tetrahedron. If the newly inserted point violates the Delaunay condition, the tetrahedron is reconstructed to ensure that all generated tetrahedrons satisfy the Delaunay condition.

[0110] In Delaunay triangulation of point clouds, each tetrahedron is formed by connecting adjacent points, and these tetrahedrons describe the point cloud surface as a continuous mesh. The vertices of each tetrahedron form a small surface segment, and multiple tetrahedrons are connected by sharing edges to form the surface of the entire 3D object.

[0111] In some cases, Delaunay triangulation may generate abnormally small or unreasonable tetrahedrons, possibly due to noise or outliers. These unsuitable tetrahedrons can be removed using mesh smoothing or simplification methods. Furthermore, post-processing operations such as point cloud mesh accuracy adjustment and vertex smoothing can yield high-quality satellite surface models, providing accurate 3D data support for subsequent analysis and applications.

[0112] This invention provides a stable multimodal data fusion spacecraft 3D reconstruction system based on deep learning, including a multimodal data acquisition module, a depth map completion module, a depth map denoising and point cloud generation module, and a model surface reconstruction module.

[0113] The multimodal data acquisition module provides the raw data to the data preprocessing module. The data preprocessing and enhancement module processes and filters the raw data, and then sends the processed and enhanced color image and sparse depth image together to the depth map completion module. The CSPN++ network performs depth map completion and outputs continuous dense depth images. This process is repeated multiple times, and the resulting continuous dense depth images are sent to the depth map denoising and point cloud generation module to obtain multiple point cloud images. These point cloud images are then sent to the point cloud registration and fusion module for point cloud registration, redundancy removal, and fusion to obtain a 3D point cloud representation of the satellite. Finally, the 3D point cloud of the satellite is sent to the model surface reconstruction module to convert the satellite point cloud model into a 3D mesh model of the satellite surface, enriching the surface structure information of the model, and finally outputting the result.

[0114] This invention features high reconstruction model accuracy and strong anti-interference capability, making it suitable for tasks such as satellite repair, space debris removal, and detection of non-cooperative target satellites in space, providing them with high-precision three-dimensional data support.

[0115] (1) Satellite Repair in Space: Accurate 3D models are crucial for satellite repair. This solution utilizes a combination of lidar and camera data acquisition to obtain high-precision depth information and image data of the satellite. Sparse depth maps guide background filtering and smoothing denoising, reducing computation for subsequent operations and making the solution more suitable for the complex space environment. Depth map completion and denoising ensure high-precision and continuous point cloud data, helping repair personnel better understand the satellite's shape, damage, and the location of different components. This allows for more precise 3D data reference when designing repair plans, providing data support for subsequent robotic arm grasping, laser ablation, and other operations.

[0116] (2) Space Debris Removal: This solution provides high-precision 3D model data of space debris during removal, facilitating accurate identification, location, and assessment of the size, shape, and hazard of the debris. High-precision depth information of the space debris is acquired via LiDAR and combined with camera images to ensure the integrity and accuracy of the depth information. Through depth map completion and denoising, the generated 3D point cloud clearly reflects the shape, size, and location of the debris, providing high-precision data support for the removal operation. Especially in complex debris environments, the solution effectively removes noise, improves depth map accuracy, and ensures the smooth progress of the mission.

[0117] (3) Detection of non-cooperative target spacecraft: In missions involving the detection of non-cooperative target spacecraft, communication or direct interaction with the target spacecraft is typically impossible. Therefore, high-precision 3D spatial information is required to accurately describe the satellite's appearance and status. This solution utilizes a combination of lidar and cameras to acquire high-precision depth information and image data of the spacecraft. Depth completion is performed using a CSPN++ network, ensuring that even in sparse point cloud conditions, a high-quality 3D model can be generated, aiding in the analysis of the spacecraft's shape, function, and potential threats.

[0118] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A stable multimodal data fusion method for spacecraft 3D reconstruction based on deep learning, characterized in that... Includes the following steps: Step 1: Use lidar and cameras to obtain multiple sparse depth and color images of the spacecraft from different perspectives; Step 2: Preprocess the sparse depth image and color image from the same viewpoint; Step 3: Combine the preprocessed color image with the CSPN++ network to complete the preprocessed sparse depth image, obtain the dense depth image, and then convert it into a point cloud image. Step 4: Repeat steps 2 and 3 to obtain all point cloud images from different perspectives; Step 5: The SIFT feature matching algorithm and the Iterative Closest Point (ICP) algorithm are used to perform coarse and fine registration on point cloud images from different perspectives to obtain a complete 3D point cloud of the spacecraft, which is then used for 3D model reconstruction.

2. The method for stable multimodal data fusion spacecraft 3D reconstruction based on deep learning according to claim 1, characterized in that: In step three, the CSPN++ network is used to extract features from the preprocessed sparse depth image and color image respectively. Then, the features are fused using the following equation to obtain the dense depth image F(x,y). F(x,y)=α·F I (x,y)+(1-α)·F D (x,y) Among them, F D (x,y) represents the depth features extracted from the preprocessed sparse depth image, α represents the fusion weight, and F I (x,y) represents the color features extracted from the preprocessed color image.

3. The method for stable multimodal data fusion spacecraft 3D reconstruction based on deep learning according to claim 2, characterized in that: Color features F extracted from the preprocessed color image using a convolutional neural network. I (x,y), depth features F extracted from the preprocessed sparse depth image. D (x,y) represents the depth information corresponding to each coordinate point in the sparse depth image.

4. The method for stable multimodal data fusion spacecraft 3D reconstruction based on deep learning according to claim 2, characterized in that: Describe the camera's intrinsic parameter matrix Among them, f x and f y Indicates the camera's focal length, (c x ,c y () indicates the optical center of the camera, which is the center point of the image. The following formula is used to iterate through each pixel in the dense depth image F(x,y), and the depth value D(x,y) of each pixel (x,y) is converted into three-dimensional spatial coordinates (X,Y,Z) to obtain the point cloud representation of the dense depth image F(x,y), which is the point cloud image. Z = D(x,y).

5. The method for stable multimodal data fusion spacecraft 3D reconstruction based on deep learning according to claim 1, characterized in that: During preprocessing, the sparse depth image is first interpolated, then a depth threshold is selected to retain pixels with depth values ​​less than the depth threshold and filter out the space background. At the same time, the color image is normalized and smoothed to remove noise and enhance image quality.

6. The method for stable multimodal data fusion spacecraft 3D reconstruction based on deep learning according to claim 1, characterized in that: In step five, a voxel mesh filtering method is used to denoise the complete 3D point cloud, and then Delaunay triangulation is used to reconstruct the surface and generate a 3D mesh model.