Underground space multi-view cross-modal 3D point cloud data fusion detection method
Through the multi-view cross-modal 3D point cloud data fusion method, combined with the improved Transformer module and the cross-time time-sequence deep fusion network, the problems of data acquisition and feature matching in complex underground environments are solved, and high-precision three-dimensional data fusion and detection effects are achieved.
Patent Information
- Application Number
- CN202510599986.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-05-12
AI Technical Summary
The prior art is difficult to obtain high-precision, multi-angle, and multi-modal three-dimensional data in complex underground environments, and the characteristics differences and modal inconsistencies of data in different modalities lead to a decrease in the accuracy of the detection results.
Multi-view cross-modal 3D point cloud data fusion method is adopted, and data is collected from different perspectives through multiple lidars and vision sensors, data registration is carried out using the improved Transformer module and cross-period timing deep fusion network, local feature points are extracted and dynamic objects are eliminated, feature extraction and data filling are performed through sparse 3D U-Net module and self-supervised learning module, and finally the matching of multimodal data is optimized through the feature projection model.
It realizes the fusion of high-precision three-dimensional data in complex underground environments, improves the accuracy and reliability of underground space detection, and provides theoretical basis and data support for high-precision detection and three-dimensional modeling in complex underground environments.
Smart Images

Figure CN120107324A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an underground space detection method, specifically to an underground space detection technology based on multi-view cross-modal 3D point cloud data fusion, and belongs to the field of three-dimensional data processing and underground space target detection. Background Art
[0002] With the development of modern science and technology, the demand for reconnaissance and detection of underground confined spaces is increasing. Traditional detection technology can usually only obtain underground space data from a single perspective, resulting in information loss or data deviation. It is often subject to many limitations in complex underground environments, such as insufficient light, complex structure, and inability to receive GPS signals, making it difficult to effectively perform tasks. In addition, the data obtained by different detection technologies (such as lidar, radar imaging, etc.) vary greatly in form and are difficult to effectively integrate, resulting in low efficiency in data processing and analysis. In addition, underground confined spaces are usually closed, narrow, and have complex three-dimensional structures. The application effect of a single detection technology in a complex underground environment is limited, and it is difficult to accurately restore the true structure of the underground space, which makes traditional reconnaissance methods face many challenges. Especially in lightless and low-light environments, the use of visual sensors is limited, and traditional image-based detection methods cannot obtain effective information.
[0003] In addition, existing 3D target detection methods often face the problems of feature differences and modality inconsistency when processing cross-modal data (such as LiDAR and Radar). Different sensors have their own physical properties, which leads to large differences in the spatial distribution, resolution and noise level of the collected data. Directly fusing and matching these data will lead to a decrease in the accuracy of the detection results. In addition, existing research usually relies on single-modal data or simple multi-modal splicing, which cannot fully utilize the complementary information of different modal data and is difficult to maintain high-precision target detection effects in complex environments.
[0004] Therefore, how to obtain high-precision, multi-angle, and multi-modal three-dimensional data in complex underground environments, and enhance the feature matching between different modes through effective fusion to improve the accuracy and reliability of underground space detection has become a technical problem that needs to be urgently solved in the industry. Summary of the invention
[0005] In response to the problems existing in the above-mentioned prior art, the present invention provides a multi-perspective and cross-modal 3D point cloud data fusion detection method for underground space, which can effectively fuse the three-dimensional data acquired in a complex underground environment, thereby improving the accuracy and reliability of underground space detection, and can provide a theoretical basis and data support for realizing high-precision detection and three-dimensional modeling in complex underground environments.
[0006] To achieve the above purpose, the local spatial multi-view cross-modal 3D point cloud data fusion detection method specifically includes the following steps: Step 1: Collect 3D point cloud data of underground space from different perspectives and directions through multiple lidar sensors and visual sensors at different locations; Step 2: Use the improved Transformer module to perform data registration on the 3D point cloud data collected in Step 1, including designing a cross-time series deep fusion network CP-FusionNet and introducing a hierarchical dense positive supervision method to optimize the registration process; Step 3, extract local feature points, remove dynamic objects, perform feature processing and map construction; Step 4: Extract 3D features through the sparse 3D U-Net module, input the features extracted from the U-Net into the elastic pooling layer to process features of different granularities; Step 5: Input the extracted features into the self-supervised learning module. The model generates labels through the data itself, performs self-supervised learning through unlabeled point cloud data, and iteratively generates point cloud data to fill in the missing points. Step 6: Improve the alignment process before cross-modal matching by aggregating instance features to improve the accuracy of registration and modeling.
[0007] Furthermore, the specific process of Step 2 is as follows: Step 2-1, improve the Transformer module: decompose the multi-head attention module in the self-attention mechanism in the Transformer into two parts: Spatial Heads and Temporal Heads; The attention calculation of Spatial Heads is as follows: ; Where: They are the query feature vector, key feature vector and value feature vector in the spatial dimension respectively; is the resulting feature vector in the spatial dimension, and its dimension dim is , where n is the number of tokens in the spatial dimension and d is the spatial dimension of each token; The attention calculation of Temporal Heads is as follows: ; Where: They are the query feature vector, key feature vector and value feature vector in the time dimension respectively; is the result feature vector in the time dimension, and its dimension dim is , where N is the number of tokens in the time dimension and D is the time dimension of each token; Combine the Spatial Heads and Temporal Heads and linearly map them to get the final result: ; Where: Y is the result feature vector; is the weight matrix of the fully connected layer; Step 2-2, introduce the hierarchical dense positive supervision method to optimize the registration process: For deep fusion matching features, two 3D point cloud decoders are used to obtain 3D point cloud information and pose information respectively. The 3D point cloud decoder uses a multi-scale approach, and the formula is as follows: ; Where: Conv is the convolution operation; s is the scale ratio set; Get 3D point cloud information, H is the current frame width, W is the current frame height, and for each pixel point, get its 3D coordinates; Considering the rotation angle offset, first divide the interval [-180°, 180°] into k sub-interval sets of length l ; Then the convolution feature vector is solved by the softmax function to be the expectation falling in each interval, and finally the rotation angle offset is calculated , the specific formula is as follows: ; Where: is the rotation angle offset.
[0008] Furthermore, the specific process of Step 3 is as follows: Step 3-1, select high curvature points as local feature points by calculating the covariance matrix and curvature of the neighborhood: The curvature is estimated by the covariance matrix of the points in the neighborhood. For each point p in the point cloud, the covariance matrix of the points around it is calculated , the specific formula is as follows: ; Where: N is the number of points in the field; and are adjacent points; T is the transposition symbol; Step 3-2, the dynamic object culling formula is as follows: ; Where: For each point p, if the speed is greater than the set speed threshold , the point is determined to be a dynamic object.
[0009] Furthermore, the specific process of Step 4 is as follows: Step 4-1, assuming the input feature map The output after the elastic pooling operation is ,Elastic pooling dynamically adjusts the pooling area according to the spatial size and feature resolution, which is expressed as follows: ; Where: Represents an elastic pooling operation; Step 4-2, in the feature extraction stage and elastic pooling process, add a context enhancement module based on the attention mechanism to generate attention weights by calculating the correlation between features. The specific formula is as follows: ; Where: Z is the feature representation weighted by self-attention, which contains global context information; D is the time dimension of the feature; Q, K and V are input features, representing the query matrix, key matrix and value matrix respectively.
[0010] Furthermore, the specific process of Step 5 is as follows: Step 5-1, take different perspectives or enhanced samples of the same point cloud as "positive sample pairs", and take data from different point clouds as "negative sample pairs". Train the model to make the feature distance of the positive sample pairs closer and the feature distance of the negative sample pairs farther. The expression of the comparative learning loss function is as follows: ; Where: and Respectively represent the feature vectors of positive sample pairs; The feature vector representing the negative sample pair; Represents similarity calculation; Step 5-2, based on the self-supervised task, fine-tune the model through the supervised segmentation task, calculate the segmentation loss function, optimize the network weights through back propagation, and use the cross entropy loss function to calculate the segmentation error, as follows: ; Where: It is The true label of each pixel; The model predicts The probability that a pixel belongs to a certain category; Step 5-3, after data processing and feature extraction, the missing point cloud data is filled in through an iterative generation process, and the missing spatial details of the 3D point cloud are supplemented. The model generates new samples by learning the denoising process. The denoising process uses a neural network to predict the denoised mean. This network outputs the denoised data by inputting the current noisy image and time step. The training goal is to minimize the prediction error, using the following loss function: ; Where: is the denoised result predicted by the network; It is the real original data.
[0011] Furthermore, the specific process of Step 6 is as follows: Step 6-1, selectively embed the features of the two modalities into the shared latent space through the feature projection model. The projection feature of each instance is: ; Where: projection matrix ; Indicates The value of the feature; b is the bias term; Step 6-2, aggregate the projected features and use weighted average: ; Where: It is The weight of the item, is the transformed The value of a feature.
[0012] Furthermore, in Step 1, after collecting the three-dimensional point cloud data of the underground space from different perspectives and orientations, the collected raw point cloud data is preprocessed to remove noise points, and a filtering algorithm is applied to remove outliers and unnecessary data.
[0013] Compared with the existing technology, the local underground space multi-view cross-modal 3D point cloud data fusion detection method first collects 3D point cloud data of underground space from multiple viewpoints and multiple directions by arranging multiple sensors including lidar and visual sensors at different locations; registers and aligns data from different viewpoints or different time periods; through the improved Transformer module, combined with the cross-time period time series deep fusion network and the hierarchical dense positive supervision method; after data registration, feature extraction and map construction are completed by extracting meaningful local features and removing the influence of dynamic objects on map construction; through sparse 3D The U-Net module is used to extract three-dimensional features. Based on the feature extraction, a self-supervised learning module is introduced to enhance the feature representation using unlabeled point cloud data. The missing point cloud data is filled in through an iterative generation process, and the missing spatial details are supplemented. Finally, in order to further optimize the matching and alignment process of multimodal data (such as point clouds and images), a feature projection model is designed to map the features of different modalities into a shared latent space, which ultimately breaks through the problems of three-dimensional semantic segmentation and three-dimensional reconstruction and multi-agent data splicing under information missing conditions, and can achieve effective fusion of three-dimensional data obtained in complex underground environments, thereby improving the accuracy and reliability of underground space detection, and providing a theoretical basis and data support for high-precision detection and three-dimensional modeling in complex underground environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 is a flow chart of the present invention; Figure 2 It is a flow chart of the present invention using a multi-view radar feature fusion network; Figure 3 It is a flow chart of the deep fusion network CP-FusionNet of the present invention; Figure 4 It is the decoder structure in the multi-task joint decoding method based on deep fusion features of the present invention; Figure 5 It is a flow chart of instance feature aggregation in cross-modal selective matching of the present invention. DETAILED DESCRIPTION
[0015] The present invention will be further described below in conjunction with the accompanying drawings.
[0016] like Figure 1 As shown in the figure, the local underground space multi-view cross-modal 3D point cloud data fusion detection method overcomes the problem of single-view single-time series reconstruction limitation. It uses different sensors in multiple positions, multiple orientations and multiple viewpoints to perform multi-task joint encoding, and combines multiple sensors and long-time series data analysis and matching to achieve efficient fusion and feature extraction of multi-modal data. At the same time, in order to solve the problems of missing information and large semantic classification errors, the query selection and recognition strategy is selected, and the previous long-time series data matching is used as input, input into the self-supervised learning module, and combined with the diffusion model guide point to reconstruct the data, and the results are fed back to the joint encoding. The instance feature aggregation of the cross-modal data that has undergone the first two steps is a necessary step before alignment. After the feature projection and selective matching combined with the perception information and the common detection head, the features of different modalities are mapped into a shared latent space to build an accurate underground three-dimensional environment perception system. Specifically, it includes the following steps: Step 1: Collect 3D point cloud data of underground space from different perspectives and orientations through multiple lidar sensors and visual sensors at different locations, pre-process the raw point cloud data collected from different sensors, remove noise points, and apply filtering algorithms (such as statistical filtering and voxel grid filtering) to remove outliers and unnecessary data.
[0017] Multiple LiDAR sensors can be arranged in different positions and orientations to achieve high-precision 3D scanning of underground space. These sensors can accurately measure distance, speed and position of objects, and generate high-precision point cloud data. Visual sensors can use high-resolution depth cameras to capture image data in conjunction with LiDAR, providing rich visual information such as texture and color for registration with point cloud data. The flowchart of the multi-view radar feature fusion network is shown below. Figure 2As shown in the figure, the data features obtained by the Airborne radar sensor are defined and standardized through the processing of the key matrix and the value matrix, and the noise is processed. Then, the data features obtained by the ground laser sensor are queried and extracted through the query matrix. In order to facilitate splicing, the two are input into the cross attention mechanism to calculate the overlapping coverage. After processing, the splicing operation is performed to obtain the fused data features.
[0018] The collection process can be achieved through the arrangement of sensors and mobile platform devices, which can be combined with static arrangement and dynamic movement. Static sensors are used to continuously monitor specific areas, and dynamic sensors (such as sensors installed on robots or drones) are used to detect a larger range of underground space. Mobile platforms may include ground robots, unmanned vehicles or aircraft. These mobile platforms can flexibly adjust the sensor's viewing angle to collect data from different directions. There are also multi-core coprocessors that include one or more processing cores. The processor is connected to the memory through a bus. The memory is used to store program instructions. When the processor executes the program instructions in the memory, it realizes the training and reasoning of deep learning models, especially in computationally intensive tasks such as feature extraction, point cloud registration and three-dimensional model optimization. It is also used to process early data collection, sensor control, data preprocessing and simple computing tasks, and cooperates with GPU and TPU to achieve efficient data flow and processing.
[0019] High-performance storage media Considering the large volume of 3D point cloud data, a distributed storage system can be used to store and manage this data. In order to support efficient data reading, writing and processing, the storage system has high throughput and low latency. In addition, for the processing of large-scale data sets, cloud storage can realize centralized processing and sharing of data in different locations, supporting remote operation and collaborative work.
[0020] Step 2, use the improved Transformer module to perform data registration on the 3D point cloud data collected in Step 1, including designing a cross-time series deep fusion network CP-FusionNet, and introducing a layered dense positive supervision method to optimize the registration process. The flowchart of the cross-time series deep fusion network CP-FusionNet is shown in the figure. Figure 3 As shown in the figure, the raw input data is first converted into an embedding vector to represent the features of the input data, firstly it is normalized by layers to accelerate convergence, then it is input into the multi-head attention mechanism for multi-task joint encoding, and a special cross-modal attention layer (i.e. improved self-attention layer) is added to the Transformer, and then it is normalized again by layers before inputting into the multi-layer perceptron, which is used to fuse data from different sensors (such as lidar and visual sensors). The decoder is part of the Transformer model and is responsible for generating output sequences and features. The details are as follows: Step 2-1, improve the Transformer module: decompose the multi-head attention module in the self-attention mechanism in the Transformer into two parts: Spatial Heads (spatial part) and Temporal Heads (temporal part).
[0021] The attention calculation of Spatial Heads is as follows: ; Where: They are the query feature vector, key feature vector and value feature vector in the spatial dimension respectively; is the resulting feature vector in the spatial dimension, and its dimension dim is , where n is the number of tokens in the spatial dimension and d is the spatial dimension of each token.
[0022] Temporal Heads attention is calculated as follows: ; Where: They are the query feature vector, key feature vector and value feature vector in the time dimension respectively; is the result feature vector in the time dimension, and its dimension dim is , where N is the number of tokens in the time dimension and D is the time dimension of each token.
[0023] Compared with the original module, the improved Transformer module combines the spatial part and the temporal part in parallel, so there is no additional time cost to add the two parts, and then linearly map to get the final result: ; Where: Y is the result feature vector; is the weight matrix of the fully connected layer.
[0024] Step 2-2: Introduce a hierarchical dense positive supervision method to optimize the registration process For deep fusion matching features, the present invention obtains 3D point cloud information and pose information respectively through two 3D point cloud decoders. The structure of the 3D point cloud decoder is as follows: Figure 4 As shown in the figure, the decoder first performs three different layers of convolution on the feature map passed from the encoder for different dimensions of variables i, j, and k to form a three-dimensional feature representation with different numbers of channels, that is, , , The following conv represents a merged convolution layer, which fuses the previous multiple feature maps and uses a 1x1 convolution kernel to integrate the features of different channels to generate a new feature map. Specifically, the 3D point cloud decoder uses a multi-scale approach with the following formula: ; Where: Conv is the convolution operation; s is the scale ratio set.
[0025] Get 3D point cloud information, H is the current frame width, W is the current frame height, and for each pixel point, get its 3D coordinates; In order to obtain more accurate posture information, the present invention considers the rotation angle offset and first divides the interval [-180°, 180°] into a set of k sub-intervals of length l. ; Then the convolution feature vector is solved by the softmax function to be the expectation falling in each interval, and finally the rotation angle offset is calculated , the specific formula is as follows: ; Where: is the rotation angle offset.
[0026] Step 3: Extract meaningful local feature points, remove dynamic objects (such as vehicles or pedestrians), perform feature processing and map construction, as follows: Step 3-1, select high curvature points as local feature points by calculating the covariance matrix and curvature of the neighborhood. These feature points represent important and stable geometric features in the environment. In three-dimensional space, curvature can be estimated by the covariance matrix of the points in the neighborhood. For each point p in the point cloud, first calculate the covariance matrix of the points around it. , the specific formula is as follows: ; Where: N is the number of points in the field; and are neighboring points; T is the transposition symbol.
[0027] Step 3-2, in a dynamic environment, dynamic objects (such as pedestrians, vehicles, etc.) will interfere with the static environment modeling. Therefore, when building a map, these dynamic objects need to be removed. The dynamic object removal formula is as follows: ; Where: For each point p, if the speed is greater than the set speed threshold , the point is determined to be a dynamic object.
[0028] The above two steps can be combined to effectively build accurate and clean 3D maps for underground environments or other application scenarios.
[0029] Step 4: Extract 3D features through the sparse 3D U-Net module, input the features extracted from the U-Net into the elastic pooling layer, and flexibly process features of different granularities. The specific steps are as follows: Step 4-1, assuming the input feature map The output after the elastic pooling operation is ,Elastic pooling dynamically adjusts the pooling area according to the spatial size and feature resolution, which is expressed as follows: ; Where: Represents an elastic pooling operation.
[0030] Step 4-2, in the feature extraction stage and elastic pooling process, add a context enhancement module based on the attention mechanism to generate attention weights by calculating the correlation between features. The specific formula is as follows: ; Where: Z is the feature representation weighted by self-attention, which contains global context information; D is the time dimension of the feature; Q, K and V are input features, representing the query matrix, key matrix and value matrix respectively. The weighted output is generated by calculating the relationship between the query, key and value matrices, such as Figure 2 shown.
[0031] Step 5, input the extracted features into the self-supervised learning module. The model generates labels through the data itself, so as to perform effective feature learning without manual labeling. Self-supervised learning is performed through unlabeled point cloud data, and point cloud data that fills in the missing points is iteratively generated to enhance feature representation and supplement spatial details. In order to achieve the distinguishing ability of enhanced feature representation, the model can be achieved by minimizing the loss function. The specific steps are as follows: Step 5-1, take different perspectives or enhanced samples of the same point cloud as "positive sample pairs", and take data from different point clouds as "negative sample pairs". Train the model to make the feature distance of the positive sample pairs closer and the feature distance of the negative sample pairs farther. The expression of the comparative learning loss function is as follows: ; Where: and Respectively represent the feature vectors of positive sample pairs; The feature vector representing the negative sample pair; Represents similarity calculation.
[0032] Step 5-2, based on the self-supervised task, fine-tune the model through the supervised segmentation task, calculate the segmentation loss function, and optimize the network weights through back propagation, which can help the model learn richer spatial structure information from unlabeled data, and use the cross entropy loss function to calculate the segmentation error, as follows: ; Where: given input x and corresponding label , It is The true label of each pixel; The model predicts The probability that a pixel belongs to a certain category.
[0033] Step 5-3, after data processing and feature extraction, the missing point cloud data is filled in through an iterative generation process, and the missing spatial details of the 3D point cloud are supplemented. The model generates new samples by learning the denoising process. The denoising process uses a neural network to predict the denoised mean. This network outputs the denoised data by inputting the current noisy image and time step. The training goal is to minimize the prediction error, using the following loss function: ; Where: is the denoised result predicted by the network; is the real original data. The loss function is the difference between the predicted denoising result and the real original data. The error between .
[0034] Step 6: In order to improve the accuracy of registration and modeling, instance feature aggregation is used to improve the alignment process before cross-modal matching to improve the accuracy of registration and modeling. The flowchart of instance feature aggregation is as follows: Figure 5 As shown in the figure, the feature vector of the image is first input, and these feature vectors are combined through concatenation and weighted fusion. The fused features are sent to the shared feature extraction layer for further convolution, and then the features are assigned to two task branches: the classification task branch (using Softmax output) and the regression task branch (using linear activation output). The corresponding losses generated by each task branch (classification loss uses cross entropy, and regression loss uses MSE) are weighted and summed to form a total loss function, which is used to guide the training process of the entire model. The specific steps are as follows: Step 6-1, for features from different modalities or perspectives, directly aggregating them may cause information loss due to different feature spaces. In this case, the features of the two modalities are selectively embedded into a shared latent space through a feature projection model. The projected features of each instance are: ; Where: projection matrix ; Indicates The value of the feature; b is the bias term; Step 6-2, aggregate the projected features and use weighted average: ; Where: It is The weight of the item, is the transformed The value of a feature.
[0035] The local underground space multi-view cross-modal 3D point cloud data fusion detection method is based on multi-source collaborative temporal point cloud registration technology, deep convolution and self-supervised learning 3D semantic segmentation technology, multi-scale heterogeneous data joint calibration and alignment method, and multi-agent collaborative data stitching and 3D reconstruction technology. It can effectively fuse the 3D data obtained in complex underground environments, thereby improving the accuracy and reliability of underground space detection, and can provide a theoretical basis and data support for achieving high-precision detection and 3D modeling in complex underground environments.
Claims
1. A multi-view cross-modal 3D point cloud data fusion detection method for underground space, characterized in that: The specific steps include: Step 1: Collect 3D point cloud data of underground space from different perspectives and directions through multiple lidar sensors and visual sensors at different locations; Step 2: Use the improved Transformer module to perform data registration on the 3D point cloud data collected in Step 1, including designing a cross-time series deep fusion network CP-FusionNet and introducing a hierarchical dense positive supervision method to optimize the registration process; Step 3, extract local feature points, remove dynamic objects, perform feature processing and map construction; Step 4: Extract 3D features through the sparse 3D U-Net module, input the features extracted from the U-Net into the elastic pooling layer to process features of different granularities; Step 5: Input the extracted features into the self-supervised learning module. The model generates labels through the data itself, performs self-supervised learning through unlabeled point cloud data, and iteratively generates point cloud data to fill in the missing points. Step 6: Improve the alignment process before cross-modal matching by aggregating instance features to improve the accuracy of registration and modeling.
2. The underground space multi-view cross-modal 3D point cloud data fusion detection method according to claim 1 is characterized in that: Step 2 The specific process is as follows: Step 2-1, improve the Transformer module: decompose the multi-head attention module in the self-attention mechanism of the Transformer into two parts: Spatial Heads and Temporal Heads; The attention calculation of Spatial Heads is as follows: ; Where: They are the query feature vector, key feature vector and value feature vector in the spatial dimension respectively; is the resulting feature vector in the spatial dimension, and its dimension dim is , where n is the number of tokens in the spatial dimension and d is the spatial dimension of each token; Temporal Heads attention is calculated as follows: ; Where: They are the query feature vector, key feature vector and value feature vector in the time dimension respectively; is the result feature vector in the time dimension, and its dimension dim is , where N is the number of tokens in the time dimension and D is the time dimension of each token; Combine the Spatial Heads and Temporal Heads and linearly map them to get the final result: ; Where: Y is the result feature vector; is the weight matrix of the fully connected layer; Step 2-2, introduce the hierarchical dense positive supervision method to optimize the registration process: For deep fusion matching features, two 3D point cloud decoders are used to obtain 3D point cloud information and pose information respectively. The 3D point cloud decoder uses a multi-scale approach, and the formula is as follows: ; Where: Conv is the convolution operation; s is the scale ratio set; Get 3D point cloud information, H is the current frame width, W is the current frame height, and for each pixel point, get its 3D coordinates; Considering the rotation angle offset, first divide the interval [-180°, 180°] into k sub-interval sets of length l ; Then the convolution feature vector is solved by the softmax function to be the expectation falling in each interval, and finally the rotation angle offset is calculated , the specific formula is as follows: ; Where: is the rotation angle offset.
3. The underground space multi-view cross-modal 3D point cloud data fusion detection method according to claim 1 is characterized in that: Step 3 The specific process is as follows: Step 3-1, select high curvature points as local feature points by calculating the covariance matrix and curvature of the neighborhood: the curvature is estimated by the covariance matrix of the points in the neighborhood. For each point p in the point cloud, calculate the covariance matrix of the points around it. , the specific formula is as follows: ; Where: N is the number of points in the field; and are neighboring points; T is the transposition symbol; Step 3-2, the dynamic object culling formula is as follows: ; Where: For each point p, if the speed is greater than the set speed threshold , the point is determined to be a dynamic object.
4. The underground space multi-view cross-modal 3D point cloud data fusion detection method according to claim 1, characterized in that: Step 4 The specific process is as follows: Step 4-1, assuming the input feature map The output after the elastic pooling operation is ,Elastic pooling dynamically adjusts the pooling area according to the spatial size and feature resolution, which is expressed as follows: ; Where: Represents an elastic pooling operation; Step 4-2, in the feature extraction stage and elastic pooling process, add a context enhancement module based on the attention mechanism to generate attention weights by calculating the correlation between features. The specific formula is as follows: ; Where: Z is the feature representation weighted by self-attention, which contains global context information; D is the time dimension of the feature; Q, K and V are input features, representing the query matrix, key matrix and value matrix respectively.
5. The underground space multi-view cross-modal 3D point cloud data fusion detection method according to claim 1, characterized in that: The specific process of Step 5 is as follows: Step 5-1, take different perspectives or enhanced samples of the same point cloud as "positive sample pairs", take data from different point clouds as "negative sample pairs", train the model to make the feature distance of the positive sample pairs closer and the feature distance of the negative sample pairs farther, and the expression of the comparative learning loss function is as follows: ; Where: and Respectively represent the feature vectors of positive sample pairs; The feature vector representing the negative sample pair; Represents similarity calculation; Step 5-2, based on the self-supervised task, fine-tune the model through the supervised segmentation task, calculate the segmentation loss function, optimize the network weights through back propagation, and use the cross entropy loss function to calculate the segmentation error, as follows: ; Where: It is The true label of each pixel; The model predicts The probability that a pixel belongs to a certain category; Step 5-3, after data processing and feature extraction, the missing point cloud data is filled in through an iterative generation process, and the missing spatial details of the 3D point cloud are supplemented. The model generates new samples by learning the denoising process. The denoising process uses a neural network to predict the denoised mean. This network outputs the denoised data by inputting the current noisy image and time step. The training goal is to minimize the prediction error, using the following loss function: ; Where: is the denoised result predicted by the network; It is the real original data.
6. The underground space multi-view cross-modal 3D point cloud data fusion detection method according to claim 1, characterized in that: Step 6 The specific process is as follows: Step 6-1, selectively embed the features of the two modalities into the shared latent space through the feature projection model. The projection feature of each instance is: ; Where: projection matrix ; Indicates The value of a feature; b is the bias term; Step 6-2, aggregate the projected features and use weighted average: ; Where: It is The weight of the item, is the transformed The value of a feature.
7. The underground space multi-view cross-modal 3D point cloud data fusion detection method according to claim 1, characterized in that: In Step 1, after collecting the 3D point cloud data of the underground space from different perspectives and orientations, the collected raw point cloud data is preprocessed to remove noise points, and a filtering algorithm is applied to remove outliers and unnecessary data.
Citation Information
Patent Citations
Automatic assembly method based on cross-source point cloud and multi-modal information
CN117523206A
Self-supervised three-dimensional point cloud completion method for underwater target object
CN118470515A
Multi-modal feature fusion reference-point-free cloud quality evaluation method
CN119478599A