An underground space multi-perspective cross-modal 3D point cloud data fusion detection method

Through the multi-view cross-modal 3D point cloud data fusion method, the problem of data acquisition and fusion in complex underground environments is solved, high-precision three-dimensional modeling and detection is realized, and the accuracy and reliability of underground space detection are improved.

CN120107324BActive Publication Date: 2025-07-11CHINA UNIV OF MINING & TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510599986.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-11
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The prior art is difficult to obtain high-precision, multi-angle, and multi-modal three-dimensional data in complex underground environments and effectively integrate the characteristics of different modalities, resulting in insufficient accuracy and reliability of underground space detection.

Method used

Multi-view cross-modal 3D point cloud data fusion method is adopted, and data is collected from different perspectives and orientations through multiple sensors, and improved Transformer module and cross-period deep fusion network are registered. Features are extracted in combination with sparse 3D U-Net modules, and data fusion and supplementation are achieved through self-supervised learning and feature projection models to achieve efficient matching and alignment of multimodal data.

Benefits of technology

It improves the accuracy and reliability of underground space detection, can realize high-precision three-dimensional modeling and data processing in complex environments, and enhances the accuracy of feature matching and data integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107324B_ABST
    Figure CN120107324B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-view cross-modal 3D point cloud data fusion detection method for underground spaces. First, 3D point cloud data of underground spaces are collected from different perspectives and orientations. Then, the collected 3D point cloud data are subjected to data registration. Then, local feature points are extracted, dynamic objects are removed, and feature processing and map construction are carried out. Then, the extracted features are input into an elastic pooling layer to process features of different granularities. Then, the extracted features are input into a self-supervised learning module, and self-supervised learning is performed through unlabeled point cloud data to iteratively generate and fill in missing point cloud data. Finally, instance feature aggregation is used to improve the alignment process before cross-modal matching. The present invention can effectively fuse 3D data obtained in complex underground environments, thereby improving the accuracy and reliability of underground space detection, and can provide a theoretical basis and data support for achieving high-precision detection and 3D modeling in complex underground environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for underground space detection, specifically an underground space detection technology based on multi-view cross-modal 3D point cloud data fusion, belonging to the fields of three-dimensional data processing and underground space target detection. Background Art

[0002] With the development of modern technology, the demand for reconnaissance and detection of underground enclosed spaces is increasing day by day. Traditional detection technologies usually can only obtain underground space data from a single perspective, resulting in information loss or data deviation, and are often restricted by many factors in complex underground environments, such as insufficient light, complex structures, inability to receive GPS signals, etc., making it difficult to effectively perform tasks. In addition, the data forms obtained by different detection technologies (such as lidar, radar imaging, etc.) are quite different, making it difficult to effectively fuse them, resulting in low efficiency of data processing and analysis. Moreover, underground enclosed spaces usually have characteristics of being enclosed, narrow, and having complex three-dimensional structures. The application effect of a single detection technology in complex underground environments is limited, and it is difficult to accurately restore the true structure of the underground space, which poses many challenges to traditional reconnaissance methods. Especially in dark or low-light environments, the use of visual sensors is restricted, and traditional image-based detection methods cannot obtain effective information.

[0003] In addition, existing three-dimensional target detection methods often face problems of feature differences and modality inconsistencies when processing cross-modal data (such as LiDAR and Radar). Different sensors have their own physical characteristics, resulting in large differences in the spatial distribution, resolution, and noise level of the collected data. Directly fusing and matching these data will lead to a decrease in the accuracy of the detection results. In addition, existing research usually relies on single-modal data or simple multi-modal stitching, and cannot make full use of the complementary information of different modal data, making it difficult to maintain high-precision target detection effects in complex environments.

[0004] Therefore, how to obtain high-precision, multi-angle, and multi-modal three-dimensional data in complex underground environments, and through effective fusion, enhance the feature matching between different modalities to improve the accuracy and reliability of underground space detection has become an urgent technical problem to be solved in the industry. Summary of the Invention

[0005] Aiming at the problems existing in the above-mentioned prior art, the present invention provides a multi-view cross-modal 3D point cloud data fusion detection method for underground space, which can effectively fuse the three-dimensional data obtained in complex underground environments, and then improve the accuracy and reliability of underground space detection, and can provide a theoretical basis and data support for realizing high-precision detection and three-dimensional modeling in complex underground environments.

[0006] To achieve the above object, the multi - perspective cross - modal 3D point cloud data fusion detection method for underground space specifically includes the following steps:

[0007] Step1, collect 3D point cloud data of the underground space from different perspectives and orientations through lidar sensors and visual sensors at multiple different positions;

[0008] Step2, use the improved Transformer module to perform data registration on the 3D point cloud data collected in Step1, including designing a cross - period temporal depth fusion network CP - FusionNet and introducing a hierarchical dense positive supervision method to optimize the registration process;

[0009] Step3, extract local feature points, eliminate dynamic objects, and perform feature processing and map construction;

[0010] Step4, extract 3D features through the sparse 3D U - Net module, and input the features extracted in U - Net into the elastic pooling layer to process features of different granularities;

[0011] Step5, input the extracted features into the self - supervised learning module. The model generates labels from the data itself, performs self - supervised learning through unlabeled point cloud data, and iteratively generates and fills in the missing point cloud data;

[0012] Step6, improve the alignment process before cross - modal matching through instance feature aggregation to improve the accuracy of registration and modeling.

[0013] Furthermore, the specific process of Step2 is as follows:

[0014] Step2 - 1, improve the Transformer module: decompose the multi - head attention module in the self - attention mechanism of Transformer into Spatial Heads and Temporal Heads;

[0015] The attention calculation of Spatial Heads is as follows: ;

[0016] In the formula: are the query feature vector, key feature vector, and value feature vector in the spatial dimension respectively; is the result feature vector in the spatial dimension, and its dimension dim is , where n is the number of tokens in the spatial dimension and d is the spatial dimension of each token;

[0017] The attention calculation of Temporal Heads is as follows: ;

[0018] In the formula: are the query feature vector, key feature vector, and value feature vector in the time dimension, respectively; is the result feature vector in the time dimension, and its dimension dim is , where N is the number of tokens in the time dimension and D is the time dimension of each token;

[0019] Merge the Spatial Heads and Temporal Heads parts and linearly map to obtain the final result: ;

[0020] In the formula: Y is the result feature vector; is the weight matrix of the fully connected layer;

[0021] Step2-2, introduce a hierarchical dense positive supervision method to optimize the registration process:

[0022] For the deep fusion matching features, obtain the 3D point cloud information and pose information through two 3D point cloud decoders respectively. The 3D point cloud decoder uses a multi-scale method, and the formula is as follows: ;

[0023] In the formula: Conv is the convolution operation; s is the set of scale ratios;

[0024] Obtain the 3D point cloud information. H is the width of the current frame, W is the height of the current frame. For each pixel point, obtain its three-dimensional coordinates;

[0025] Considering the rotation angle offset, first divide the interval [-180°, 180°] into a set of k sub-intervals with a length of l ; then solve the feature vector obtained by convolution through the softmax function to obtain the expectation of falling into each interval, and finally calculate the rotation angle offset , and the specific formula is as follows: ; In the formula: is the rotation angle offset.

[0026] Furthermore, the specific process of Step3 is as follows:

[0027] Step3-1, select high-curvature points as local feature points by calculating the covariance matrix and curvature of the neighborhood:

[0028] The curvature is estimated through the covariance matrix of the points in the neighborhood. For each point p in the point cloud, calculate the covariance matrix of the points around it , and the specific formula is as follows: ;

[0029] In the formula: N is the number of points in the neighborhood; and are neighboring points; T is the transpose symbol;

[0030] Step3-2, the dynamic object removal formula is as follows: ;

[0031] In the formula: for each point p, if the speed is greater than the set speed threshold , this point is determined to be a dynamic object.

[0032] Furthermore, the specific process of Step4 is as follows:

[0033] Step4-1, assume that the input feature map after the elastic pooling operation is , and the elastic pooling dynamically adjusts the pooling area according to the spatial size and feature resolution, which is expressed as follows: ;

[0034] In the formula: represents the elastic pooling operation;

[0035] Step4-2, in the feature extraction stage and the elastic pooling process, add a context enhancement module based on the attention mechanism to generate attention weights by calculating the correlation between features. The specific formula is as follows: ;

[0036] In the formula: Z is the feature representation after self-attention weighting, containing global context information; D is the time dimension of the feature; Q, K, and V are the input features, representing the query matrix, key matrix, and value matrix respectively.

[0037] Furthermore, the specific process of Step5 is as follows:

[0038] Step5-1, use different perspectives of the same point cloud or enhanced samples as "positive sample pairs", and use data from different point clouds as "negative sample pairs". Train the model to make the feature distances of positive sample pairs closer and the feature distances of negative sample pairs farther. The expression of the contrastive learning loss function is as follows: ;

[0039] In the formula: and respectively represent the feature vectors of positive sample pairs; represents the feature vector of negative sample pairs; represents the similarity calculation;

[0040] Step5-2, based on the self-supervised task, fine-tune the model through a supervised segmentation task, calculate the segmentation loss function, and optimize the network weights through backpropagation. Use the cross-entropy loss function to calculate the segmentation error, as follows: ;

[0041] In the formula: is the true label of the th pixel; is the probability that the th pixel predicted by the model belongs to a certain category;

[0042] Step5-3, after data processing and feature extraction, fill in the missing point cloud data through an iterative generation process, and supplement the missing spatial details of the 3D point cloud. Among them, the model generates new samples by learning the denoising process. The denoising process uses a neural network to predict the mean of the denoising. This network outputs the denoised data by inputting the current noisy image and the time step. The training objective is to minimize the prediction error, and the following loss function is used: ;

[0043] In the formula: is the denoising result predicted by the network; is the true original data.

[0044] Furthermore, the specific process of Step6 is as follows:

[0045] Step6-1, selectively embed the features of the two modalities into the shared latent space through the feature projection model. The projected features of each instance are: ;

[0046] In the formula: the projection matrix ; represents the value of the th feature; b is the bias term;

[0047] Step6-2, aggregate the projected features, and use weighted average: ;

[0048] In the formula: is the weight of the th term, is the value of the th feature after transformation.

[0049] Furthermore, in Step1, after collecting the 3D point cloud data of the underground space from different perspectives and orientations, preprocess the collected original point cloud data, remove the noise points, and apply the filtering algorithm to remove the outliers and unnecessary data.

[0050] Compared with the prior art, the multi-view cross-modal 3D point cloud data fusion detection method for underground space first collects 3D point cloud data of the underground space from multiple perspectives and multiple orientations by arranging multiple sensors including lidar and vision sensors at different positions; registers and aligns data from different perspectives or different time periods; through an improved Transformer module, combines a cross-temporal time series depth fusion network and a hierarchical dense positive supervision method; after data registration, completes feature extraction and map construction by extracting meaningful local features and removing the influence of dynamic objects on map construction; extracts three-dimensional features through a sparse 3D U-Net module; on the basis of feature extraction, introduces a self-supervised learning module, uses unlabeled point cloud data to enhance feature representation, fills in missing point cloud data through an iterative generation process, and supplements missing spatial details; finally, in order to further optimize the matching and alignment process of multi-modal data (such as point cloud and image), designs a feature projection model to map features of different modalities into a shared latent space, ultimately breaks through the problems of three-dimensional semantic segmentation and three-dimensional reconstruction under information missing conditions and multi-agent data stitching, can effectively fuse the three-dimensional data obtained in a complex underground environment, and further improve the accuracy and reliability of underground space detection, and can provide a theoretical basis and data support for realizing high-precision detection and three-dimensional modeling in a complex underground environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a flowchart of the present invention;

[0052] Figure 2 is a flowchart of the multi-view radar feature fusion network adopted by the present invention;

[0053] Figure 3 is a flowchart of the depth fusion network CP-FusionNet of the present invention;

[0054] Figure 4 is the decoder structure in the multi-task joint decoding method based on depth fusion features of the present invention;

[0055] Figure 5 is a flowchart of instance feature aggregation in cross-modal selective matching of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] The present invention will be further described below with reference to the accompanying drawings.

[0057] As Figure 1As shown in the figure, the multi - perspective cross - modal 3D point cloud data fusion detection method for underground space aims to overcome the limitation of single - perspective and single - time - series reconstruction. By using different sensors from multiple locations, orientations, and perspectives, multi - task joint encoding is performed. Combining multiple sensors and long - time - series data analysis and matching, efficient fusion and feature extraction of multi - modal data are achieved. At the same time, to solve the problems of information loss and large semantic segmentation errors, a query selection and recognition strategy is selected, and the matching of previous long - time - series data is used as input and fed into the self - supervised learning module. Combined with the diffusion model to guide points for data reconstruction, and the results are fed back into the joint encoding. Instance feature aggregation is performed on the cross - modal data that has gone through the first two steps. As a necessary step before alignment, through feature projection combined with perceptual information, selective matching, and a common detection head, features of different modalities are mapped into a shared latent space to construct an accurate underground three - dimensional environment perception system. The specific steps are as follows:

[0058] Step1, through lidar sensors and visual sensors at multiple different locations, three - dimensional point cloud data of the underground space are collected from different perspectives and orientations. The original point cloud data collected from different sensors are pre - processed to remove noise points, and filtering algorithms (such as statistical filtering, voxel grid filtering) are applied to remove outliers and unnecessary data.

[0059] Multiple lidar sensors can be arranged at different locations and orientations to achieve high - precision three - dimensional scanning of the underground space. These sensors can accurately measure distance, speed, and object position, generating high - precision point cloud data. The visual sensor can use a high - resolution depth camera, which can capture image data in cooperation with the lidar, providing rich visual information such as texture and color for registration with the point cloud data. The flow chart of the multi - perspective radar feature fusion network is as Figure 2 shown. The data features obtained by the Airborne radar sensor are defined, standardized, and noise - processed through the processing of the key matrix and value matrix. Then, the data features obtained by the ground laser sensor are queried and extracted through the query matrix of the query matrix. For convenient stitching, the two are input into the cross - attention mechanism to calculate the overlap coverage. After processing, a stitching operation is performed to obtain the fused data features.

[0060] The acquisition process can be carried out through sensor arrangement and mobile platform devices. This device can combine static arrangement and dynamic movement. Static sensors are used to continuously monitor specific areas, and dynamic sensors (such as sensors installed on robots or drones) are used to detect a larger range of underground spaces. The mobile platform can include ground robots, unmanned vehicles or aircraft, and these mobile platforms can flexibly adjust the perspective of the sensors to collect data from different directions. There is also a multi-core coprocessor that includes one or more processing cores. The processor is connected to the memory through a bus. The memory is used to store program instructions. When the processor executes the program instructions in the memory, it realizes the training and inference of the deep learning model, especially in computationally intensive tasks such as feature extraction, point cloud registration, and 3D model optimization, and is used to process the previous data acquisition, sensor control, data preprocessing, and simple computational tasks, and cooperate with the GPU and TPU to achieve efficient data flow and processing.

[0061] Considering the large volume of 3D point cloud data, a distributed storage system can be used to store and manage these data. To support efficient data reading, writing, and processing, the storage system has high throughput and low latency. In addition, for the processing of large-scale data sets, through cloud storage, centralized processing and sharing of data at different locations can be achieved, supporting remote operation and collaborative work.

[0062] Step2, use the improved Transformer module to perform data registration on the 3D point cloud data collected in Step1, including designing a cross-temporal temporal depth fusion network CP-FusionNet and introducing a hierarchical dense positive supervision method to optimize the registration process. The flowchart of the cross-temporal temporal depth fusion network CP-FusionNet is as Figure 3 shown. The input original data is first converted into an embedding vector to represent the features of the input data. First, it undergoes layer normalization to accelerate convergence, and then is input into the multi-head attention mechanism for multi-task joint encoding. A dedicated cross-modal attention layer (i.e., an improved self-attention layer) is added to the Transformer. Then, it undergoes layer normalization again before being input into the multi-layer perceptron, which is used to fuse data from different sensors (such as lidar and visual sensors). The decoder among them is a part of the Transformer model and is responsible for generating the output sequence and features. Specifically as follows:

[0063] Step2-1, improve the Transformer module: decompose the multi-head attention module in the self-attention mechanism of the Transformer into two parts: Spatial Heads (spatial part) and Temporal Heads (temporal part).

[0064] The attention calculation of Spatial Heads is as follows: ;

[0065] In the formula: are the query feature vector, key feature vector, and value feature vector in the spatial dimension respectively; is the result feature vector in the spatial dimension, and its dimension dim is , where n is the number of tokens in the spatial dimension and d is the spatial dimension of each token.

[0066] The attention calculation of Temporal Heads is as follows: ;

[0067] In the formula: are the query feature vector, key feature vector, and value feature vector in the temporal dimension respectively; is the result feature vector in the temporal dimension, and its dimension dim is , where N is the number of tokens in the temporal dimension and D is the temporal dimension of each token.

[0068] Compared with the original module, the improved Transformer module combines the spatial part and the temporal part in parallel without adding additional time cost, and then linearly maps to obtain the final result: ;

[0069] In the formula: Y is the result feature vector; is the weight matrix of the fully connected layer.

[0070] Step2-2, introducing a hierarchical dense positive supervision method to optimize the registration process

[0071] For the deep fusion matching features, the present invention respectively obtains 3D point cloud information and pose information through two 3D point cloud decoders. The structure of the 3D point cloud decoder is as Figure 4 shown. First, the decoder performs three-layer different convolutions on the feature map passed from the encoder for different variables i, j, and k in their respective dimensions to form three-dimensional feature representations with different numbers of channels, that is , , . After that, conv represents a merging convolutional layer that fuses the previous multiple feature maps and uses a 1x1 convolutional kernel to integrate the features of different channels to generate a new feature map . Specifically, the 3D point cloud decoder uses a multi-scale method, and the formula is as follows: ;

[0072] In the formula: Conv is the convolution operation; s is the set of scale ratios.

[0073] Obtain 3D point cloud information. Let H be the width of the current frame and W be the height of the current frame. For each pixel point, obtain its three-dimensional coordinates;

[0074] To obtain more accurate pose information, the present invention considers the rotation angle offset. First, divide the interval [-180°, 180°] into a set of k sub-intervals with a length of l ; then solve the eigenvector obtained by convolution through the softmax function to obtain the expectation of falling into each interval, and finally calculate the rotation angle offset , and the specific formula is as follows: ; In the formula: is the rotation angle offset.

[0075] Step3, Extract meaningful local feature points, and eliminate dynamic objects (such as vehicles or pedestrians), and perform feature processing and map construction, specifically as follows:

[0076] Step3-1, Select high-curvature points as local feature points by calculating the covariance matrix and curvature of the neighborhood. These feature points represent important and stable geometric features in the environment. In three-dimensional space, the curvature can be estimated by the covariance matrix of points in the neighborhood. For each point p in the point cloud, first calculate the covariance matrix of the points around it , and the specific formula is as follows: ;

[0077] In the formula: N is the number of points in the neighborhood; and are neighboring points; T is the transpose symbol.

[0078] Step3-2, In a dynamic environment, dynamic objects (such as pedestrians, vehicles, etc.) will interfere with the modeling of the static environment. Therefore, when constructing a map, it is necessary to eliminate these dynamic objects. The formula for eliminating dynamic objects is as follows: ;

[0079] In the formula: For each point p, if the speed is greater than the set speed threshold , determine that this point is a dynamic object.

[0080] After combining the above two operations, an accurate and clean three-dimensional map can be effectively constructed for the underground environment or other application scenarios.

[0081] Step4, Extract 3D features through the sparse 3D U-Net module, and input the features extracted from the U-Net into the elastic pooling layer to flexibly process features of different granularities. The specific steps are as follows:

[0082] Step4-1, Assume that the input feature map The output after the elastic pooling operation is , Elastic pooling dynamically adjusts the pooling region according to the spatial size and feature resolution, as shown below: ;

[0083] Where: represents the elastic pooling operation.

[0084] Step4-2, during the feature extraction stage and the elastic pooling process, add a context enhancement module based on the attention mechanism to generate attention weights by calculating the correlation between features. The specific formula is as follows: ;

[0085] Where: Z is the feature representation after self-attention weighting, containing global context information; D is the time dimension of the feature; Q, K, and V are the input features, representing the query matrix, key matrix, and value matrix respectively. By calculating the relationship between the query, key, and value matrices, the weighted output is generated, as shown in Figure 2 shown.

[0086] Step5, input the extracted features into the self-supervised learning module. The model generates labels from the data itself, enabling effective feature learning without manual annotation. Through self-supervised learning on unlabeled point cloud data, iteratively generate filled missing point cloud data, thereby enhancing feature representation and supplementing spatial details. To achieve the discrimination ability of enhanced feature representation, the model can be achieved by minimizing the loss function. The specific steps are as follows:

[0087] Step5-1, use different perspectives or enhanced samples of the same point cloud as "positive sample pairs", and use data from different point clouds as "negative sample pairs". Train the model to make the feature distances of positive sample pairs closer and the feature distances of negative sample pairs farther. The expression of the contrastive learning loss function is as follows: ;

[0088] Where: and respectively represent the feature vectors of positive sample pairs; represents the feature vector of the negative sample pair; represents the similarity calculation.

[0089] Step5-2, based on the self-supervised task, fine-tune the model through a supervised segmentation task, calculate the segmentation loss function, and optimize the network weights through backpropagation, which can help the model learn richer spatial structure information from unlabeled data. Use the cross-entropy loss function to calculate the segmentation error, as follows: ;

[0090] Where: Given the input x and the corresponding label , is the The true label of a pixel; is the probability that the pixel predicted by the model belongs to a certain category.

[0091] Step5-3, after data processing and feature extraction, fill in the missing point cloud data through an iterative generation process and supplement the missing spatial details of the 3D point cloud. Among them, the model generates new samples by learning the denoising process. The denoising process uses a neural network to predict the mean of the denoising. This network outputs the denoised data by inputting the current noisy image and the time step. The training objective is to minimize the prediction error, and the following loss function is used: ;

[0092] In the formula: is the denoising result predicted by the network; is the true original data. The loss function is the error between the predicted denoising result and the true original data between.

[0093] Step6, in order to improve the accuracy of registration and modeling, improve the alignment process before cross-modal matching through instance feature aggregation to improve the accuracy of registration and modeling. The flow chart of instance feature aggregation is as Figure 5 shown. First, input the feature vectors of the images, combine these feature vectors through splicing and weighted fusion, and the fused features are sent to the shared feature extraction layer for further convolution. After that, the features are assigned to two task branches: the classification task branch (using Softmax output) and the regression task branch (using linear activation output). The corresponding losses generated by each task branch (the classification loss uses cross-entropy and the regression loss uses MSE) are weighted and summed to form the total loss function, which is used to guide the training process of the entire model. The specific steps are as follows:

[0094] Step6-1, for features from different modalities or perspectives, directly aggregating them may result in information loss due to different feature spaces. In this case, selectively embed the features of the two modalities into the shared latent space through the feature projection model. The projected feature of each instance is: ;

[0095] In the formula: the projection matrix ; represents the value of the th feature; b is the bias term;

[0096] Step6-2, aggregate the projected features, using weighted average: ;

[0097] In the formula: is the weight of the th term, is the value of the th feature after transformation.

[0098] The multi-view cross-modal 3D point cloud data fusion detection method for underground space is based on the time-series point cloud registration technology of multi-source collaboration, the three-dimensional semantic segmentation technology of deep convolution and self-supervised learning, the multi-scale heterogeneous data joint calibration and alignment method, and the multi-agent collaborative data stitching and three-dimensional reconstruction technology. It can effectively fuse the three-dimensional data obtained in complex underground environments, and then improve the accuracy and reliability of underground space detection, providing a theoretical basis and data support for achieving high-precision detection and three-dimensional modeling in complex underground environments.

Claims

1. An underground space multi-perspective cross-modal 3D point cloud data fusion detection method, characterized in that Specifically, it includes the following steps: Step1, Collect 3D point cloud data of the underground space from different perspectives and orientations through lidar sensors and vision sensors at multiple different positions; Step2, Use the improved Transformer module to perform data registration on the 3D point cloud data collected in Step1, including designing a cross-temporal temporal depth fusion network CP-FusionNet and introducing a hierarchical dense positive supervision method to optimize the registration process; The specific process is as follows: Step2-1, Improve the Transformer module: Decompose the multi-head attention module in the self-attention mechanism of Transformer into two parts: Spatial Heads and Temporal Heads; The attention calculation of Spatial Heads is as follows: ; In the formula: are the query feature vector, key feature vector, and value feature vector in the spatial dimension, respectively; is the result feature vector in the spatial dimension, and its dimension dim is , where n is the number of tokens in the spatial dimension and d is the spatial dimension of each token; The attention calculation of Temporal Heads is as follows: ; In the formula: are the query feature vector, key feature vector, and value feature vector in the time dimension, respectively; is the result feature vector in the time dimension, and its dimension dim is , where N is the number of tokens in the time dimension and D is the time dimension of each token; Merge the two parts of Spatial Heads and Temporal Heads, and linearly map to get the final result: ; Where: Y is the result feature vector; is the weight matrix of the fully connected layer; Step2-2, Introduce a hierarchical dense positive supervision method to optimize the registration process: For the deep fusion matching features, obtain 3D point cloud information and pose information through two 3D point cloud decoders respectively. The 3D point cloud decoder uses a multi-scale method, and the formula is as follows: ; In the formula: Conv is the convolution operation; s is the set of scale ratios; Obtain the 3D point cloud information. H is the width of the current frame, and W is the height of the current frame. For each pixel point, obtain its three-dimensional coordinates; Considering the rotation angle offset, first divide the interval [-180°, 180°] into a set of k sub-intervals with a length of l ; then solve the expected value of the convolution-obtained feature vector falling into each interval through the softmax function, and finally calculate the rotation angle offset , and the specific formula is as follows: ; where: is the rotation angle offset; Step3, Extract local feature points, eliminate dynamic objects, and perform feature processing and map construction; Step4, Extract 3D features through the sparse 3D U-Net module, and input the features extracted in U-Net into the elastic pooling layer to process features of different granularities; Step5, Input the extracted features into the self-supervised learning module. The model generates labels through the data itself, performs self-supervised learning through unlabeled point cloud data, and iteratively generates point cloud data to fill in the missing parts; Step6, Improve the alignment process before cross-modal matching through instance feature aggregation to improve the accuracy of registration and modeling.

2. The multi-view cross-modal 3D point cloud data fusion detection method for underground space according to claim 1, wherein The specific process of Step3 is as follows: Step 3-1: Select high-curvature points as local feature points by calculating the covariance matrix and curvature of the neighborhood. The curvature is estimated through the covariance matrix of the points within the neighborhood. For each point p in the point cloud, calculate the covariance matrix of the points around it. , and the specific formula is as follows: ; where: N is the number of points in the domain; and are the nearest neighbor points; T is the transpose symbol; The formula for eliminating dynamic objects in Step3-2 is as follows: ; Where: for each point p, if the speed is greater than the set speed threshold it is determined that the point is a moving object.

3. The multi-view cross-modal 3D point cloud data fusion detection method for underground space according to claim 1, wherein The specific process of Step4 is as follows: Step4-1, assume the input feature map The output after the elastic pooling operation is , and the elastic pooling dynamically adjusts the pooling region according to the spatial size and feature resolution, as shown below: ; In the formula: represents an elastic pooling operation; Step4-2, In the feature extraction stage and the elastic pooling process, add a context enhancement module based on the attention mechanism to generate attention weights by calculating the correlation between features. The specific formula is as follows: ; In the formula: Z is the feature representation after self-attention weighting, which contains global context information; D is the time dimension of the feature; Q, K, and V are the input features, representing the query matrix, key matrix, and value matrix respectively.

4. The multi - perspective cross - modal 3D point cloud data fusion detection method for underground space according to claim 1, wherein, The specific process of Step5 is as follows: Step5-1, Use different perspectives or enhanced samples of the same point cloud as "positive sample pairs", and use data from different point clouds as "negative sample pairs". Train the model to make the feature distances of positive sample pairs closer and the feature distances of negative sample pairs farther. The expression of the contrast learning loss function is as follows: ; Wherein: and respectively represent the feature vectors of positive sample pairs; represents the feature vector of a negative sample pair; represents similarity calculation; Step5-2. Based on the self-supervised task, fine-tune the model through a supervised segmentation task, calculate the segmentation loss function, and optimize the network weights through backpropagation. Use the cross-entropy loss function to calculate the segmentation error, as follows: ; Wherein: is the true label of the th pixel; is the probability that the th pixel predicted by the model belongs to a certain category; Step5-3. After data processing and feature extraction, fill in the missing point cloud data through an iterative generation process and supplement the missing spatial details of the 3D point cloud. Among them, the model generates new samples by learning the denoising process. The denoising process uses a neural network to predict the mean of the denoising. This network outputs the denoised data by inputting the current noisy image and the time step. The training objective is to minimize the prediction error. Use the following loss function: ; Wherein: is the denoising result predicted by the network; is the true original data.

5. The multi-view cross-modal 3D point cloud data fusion detection method for underground space according to claim 1, wherein The specific process of Step6 is as follows: Step6-1. Selectively embed the features of the two modalities into the shared latent space through the feature projection model. The projected feature of each instance is: ; Where: projection matrix ; represents the value of the th feature; b is the bias term; Step6-2. Aggregate the projected features and use weighted average: ; Wherein: is the weight of the th term, is the value of the th feature after transformation.

6. The multi-view cross-modal 3D point cloud data fusion detection method for underground space according to claim 1, wherein, In Step1, after collecting the 3D point cloud data of the underground space from different perspectives and orientations, preprocess the collected original point cloud data, remove the noise points, and apply the filtering algorithm to remove the outliers and unnecessary data.

Citation Information

Patent Citations

  • Automatic assembly method based on cross-source point cloud and multi-modal information

    CN117523206A

  • Multi-modal feature fusion reference-point-free cloud quality evaluation method

    CN119478599A