A Multimodal Remote Sensing Data Scene Segmentation Method Based on Unbalanced Knowledge Driving
By extracting features from the original data of LiDAR point clouds and images, using unbalanced knowledge-driven methods, combining information from global and category angles, the problem of information loss in multimodal remote sensing data scene segmentation is solved, and the segmentation accuracy is improved.
Patent Information
- Application Number
- CN202211270240.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-18
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-10-18
AI Technical Summary
The prior art fails to fully utilize the information of LiDAR point cloud and image in multimodal remote sensing data scene segmentation, resulting in insufficient information loss and segmentation accuracy, especially in the case of information imbalance.
By directly extracting features from the original data of LiDAR point clouds and images, using an unbalanced knowledge-driven method, combining information from global and category angles, multi-layer perceptrons and gate modules are designed to reduce information loss and improve segmentation accuracy.
Effectively utilize the differences between multimodal data, improve the accuracy of image segmentation, reduce information loss during modal fusion, and achieve more efficient semantic segmentation.
Smart Images

Figure CN115661817B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for multi-modal data scene segmentation, and more particularly to a method for multi-modal remote sensing data scene segmentation based on unbalanced knowledge-driven. Background Art
[0002] With the rapid development of sensors such as optical cameras, radars, and 3D scanners, the big data era has arrived. Multi-modal data for earth observation has become a research frontier in the field of remote sensing, especially in semantic segmentation tasks. Compared with single-modal remote sensing images, multi-modal data can integrate the advantages of single-source data, obtain more diverse feature information, and break through the performance bottleneck of single-modal semantic segmentation.
[0003] Currently, many studies focus on jointly using three-dimensional airborne LiDAR point clouds and two-dimensional satellite image data. To eliminate the structural differences between the two modalities, researchers usually map the high-dimensional LiDAR point clouds to the two-dimensional image space, and then jointly segment the obtained two-dimensional feature map from the point clouds and the image. Although this approach can achieve better results than single-modal methods, it still cannot fully utilize the performance of each modality, because the preprocessing operation of projecting the point clouds from three dimensions to two dimensions will inevitably lose some information, especially geometric structure information. Currently, few studies focus on the problem of using the original LiDAR point clouds to assist the image in two-dimensional segmentation.
[0004] The point density of commonly used airborne LiDAR point clouds is not less than 8 points per square meter, and the resolution of commonly used aerial remote sensing images is about 0.6 meters. It can be seen that generally, the data volume and coordinate accuracy of airborne LiDAR within a unit range far exceed those of aerial images. That is, LiDAR with the same coverage range accommodates a much larger data volume than images. When designing multi-modal interaction algorithms, this inherent information imbalance cannot be ignored. However, so far, almost no studies have focused on the information volume difference between modalities when multi-modal joint action. Although encouraging results have been achieved in the semantic segmentation of single-modal data, the limitations of a single data source still make it difficult to break through the inherent bottleneck of performance. Therefore, multi-modal learning is becoming a hot topic in the field of remote sensing. Current methods usually project high-dimensional features into a low-dimensional space before feature extraction to solve the non-negligible semantic gap between different modalities, which will inevitably cause information loss. Summary of the Invention
[0005] The present invention discloses a method for multi-modal remote sensing data scene segmentation based on unbalanced knowledge-driven. This method directly extracts features from heterogeneous LiDAR point clouds and raw image data, mines cross-modal unbalanced information, and refines the segmentation of the weaker modality (image) using the strong modality (LiDAR point cloud) from both the global and category perspectives. Our method cleverly transforms the differences between multi-modal data to improve the results of single-modal segmentation, fully utilizes the information of both modalities, and reduces information loss during the modality fusion process.
[0006] The technical solution adopted by the present invention is: a method for multi-modal remote sensing data scene segmentation based on unbalanced knowledge-driven, including the following steps:
[0007] Step 1: Crop the LiDAR point clouds and images in the dataset into small pieces of the same size;
[0008] Step 2: Use the KD-tree algorithm to find several nearest points around each point in the point cloud file, and save the found points as a readable file;
[0009] Step 3: Input the point cloud small pieces obtained in Step 1 and the nearest point file obtained in Step 2 into the point cloud encoder to obtain point cloud features at different scales k takes values from 1 to K;
[0010] Step 4: Input the image small pieces obtained in Step 1 into the image encoder to obtain image feature f i k , k takes values from 1 to K;
[0011] Step 5: Input the high-dimensional point cloud feature map obtained in Step 3 into a multi-layer perceptron to obtain a feature map Then perform nearest neighbor interpolation on and concatenate the obtained feature map with and input it into a multi-layer perceptron to obtain a feature map Then repeat the above operation, and finally concatenate the obtained feature map with and input it into a multi-layer perceptron to obtain a feature map Then input into two stacked multi-layer perceptrons to obtain the point cloud endpoint-level segmentation result
[0012] Step 6: Input the high-dimensional image feature map f obtained in Step 4 i K and the after dimensional transformation into the globally knowledge-guided gate module at the same time, then perform bilinear interpolation on the obtained feature map, and concatenate the obtained feature map with f i K-1 and input it into a convolutional block to obtain Input and the dimension-transformed into the globally knowledge-guided gate module simultaneously, and repeat the above operations until obtaining
[0013] Step 7: Input the segmentation result obtained in Step 5 after dimension transformation and the one obtained in Step 6 into the class knowledge-guided gate module simultaneously. The class knowledge-guided gate module processes the feature map from the image side through a convolutional layer, a batch normalization layer, a ReLU activation layer, and a Dropout layer, then performs dimension transformation and transposition, and multiplies it pixel-wise with the coarsely segmented result from the point cloud side after dimension transformation Then, after passing through a Softmax function layer, the resulting output is transposed and multiplied pixel-wise with the dimension-transformed again for pixel-wise multiplication. After the resulting output is dimension-transformed, it is concatenated with to finally obtain the output of the class knowledge-guided gate module
[0014] Step 8: Input the output obtained in Step 7 through a convolutional layer to obtain the final image segmentation result
[0015] Step 9: Calculate the loss function based on the point cloud segmentation obtained in Step 5 and the image segmentation result obtained in Step 8 ;
[0016] Step 10: Use the gradient descent algorithm to backpropagate the loss obtained in Step 9 and update the parameters of the entire network;
[0017] Step 11: Iterate repeatedly from Step 3 to Step 10 until the training ends.
[0018] Furthermore, in Step 3, the point cloud encoder sequentially includes a fully connected layer and four local feature aggregation units, and finally the output of the fully connected layer and the output of each local feature aggregation unit can be obtained The local feature aggregation unit consists of two local spatial encoders, an attention pooling layer, and a random sampling module.
[0019] Furthermore, in Step 3, the local feature aggregation module encodes the coordinate values and features of several nearest neighbor points around each point into a one-dimensional vector and then concatenates them into a two-dimensional feature map. The encoding formula used is:
[0020]
[0021] where r i n represents the coordinate encoding value of the i-th point and its n-th nearest neighbor point, and MLP represents a multi-layer perceptron. represents concatenation, p i and p i n are the x-y-z coordinates of the i-th point and its n-th nearest neighbor point respectively.
[0022] Furthermore, in step 4, the image encoder is stacked by five convolutional blocks. Each convolutional block consists of two units composed of a convolutional layer with a 3×3 convolutional kernel, a batch normalization layer, and a ReLU activation layer. A max pooling layer with a kernel of 2 is passed between every two convolutional blocks.
[0023] Furthermore, the specific implementation method of step 5 is as follows;
[0024] K is taken as 5. The high-dimensional feature map of the point cloud obtained in step 3 is input into the multi-layer perceptron to obtain a feature map Then, is subjected to nearest neighbor interpolation. The obtained feature map is concatenated with and input into the multi-layer perceptron to obtain a feature map Then, is subjected to nearest neighbor interpolation. The obtained feature map is concatenated with and input into the multi-layer perceptron to obtain a feature map Then, is subjected to nearest neighbor interpolation. The obtained feature map is concatenated with and input into the multi-layer perceptron to obtain a feature map Then, is subjected to nearest neighbor interpolation. The obtained feature map is concatenated with and input into the multi-layer perceptron to obtain a feature map Then, is input into two stacked multi-layer perceptrons to obtain the segmentation result at the point cloud endpoint level
[0025] Furthermore, the specific implementation method of step 6 is as follows;
[0026] K is taken as 5. The high-dimensional feature map f of the image obtained in step 4 i 5 and the after dimensional transformation are simultaneously input into the globally knowledge-guided gate module. Then, the obtained feature map is subjected to bilinear interpolation. The obtained feature map is concatenated with f i 4 and input into the convolutional block to obtain Then, and the Input the global knowledge-guided gate module simultaneously, then perform bilinear interpolation on the obtained feature map, and concatenate the obtained feature map with f i 3 After concatenation, input it into the convolutional block to obtain Take and the dimension-transformed Input the global knowledge-guided gate module simultaneously, then perform bilinear interpolation on the obtained feature map, and concatenate the obtained feature map with f i 2 After concatenation, input it into the convolutional block to obtain Take and the dimension-transformed Input the global knowledge-guided gate module simultaneously, then perform bilinear interpolation on the obtained feature map, and concatenate the obtained feature map with f i 1 After concatenation, input it into the convolutional block to obtain
[0027] Furthermore, the global knowledge-guided gate module in step 6 passes the feature map from the point cloud end through the max-pooling layer and average-pooling layer with an output dimension of 1 respectively. After concatenating the obtained feature maps, pass them through the convolutional layer with a convolutional kernel of 7×7, batch normalization layer, ReLU activation layer, and Sigmoid function layer in sequence, and then perform pixel-level multiplication with the feature map from the image end After that, concatenate with The obtained result is the output of the global knowledge-guided gate module;
[0028] The convolutional block consists of two units stacked by convolutional layers with a convolutional kernel of 3×3, batch normalization layers, and ReLU activation layers.
[0029] Furthermore, the kernel size of the convolutional block in the category knowledge-guided gate module in step 7 is 1×1.
[0030] Furthermore, the kernel size of the convolutional layer in step 8 is 1×1.
[0031] Furthermore, the loss function in step 9 includes the point cloud segmentation loss L pc , the image segmentation loss L i and the similarity loss L pi-SC of the point cloud and image segmentation results:
[0032]
[0033]
[0034]
[0035] Where N is the total number of points in the point cloud, M is the number of finally segmented categories, W and H are the width and height of the image, and are the ground truth of the point cloud and image segmentation results respectively, and KL represents the KL divergence;
[0036] The total loss function is:
[0037] L total = L pc + L i + L pi-SC (5).
[0038] Compared with the prior art, the beneficial effects of the present invention are as follows: The present invention proposes a multi-modal remote sensing data scene segmentation method based on unbalanced knowledge-driven. This method utilizes the difference in information volume between heterogeneous data, and refines the segmentation of the weaker modality (image) from the global and category perspectives by using the strong modality (LiDAR point cloud), ultimately improving the accuracy of image segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 : The framework of the multi-modal remote sensing data scene segmentation method designed by the present invention;
[0040] Figure 2 : The structure of the global knowledge-guided gate module used in the present invention;
[0041] Figure 3 : The structure of the category knowledge-guided gate module used in the present invention;
[0042] Figure 4 : Some visualization results of the method of the present invention; DETAILED DESCRIPTION OF THE INVENTION
[0043] The present invention will be further described in detail below with reference to the drawings and embodiments.
[0044] Step 1: Crop the images in the dataset into small pieces of 512×512. Crop the LiDAR point cloud into small pieces with the same size as the image coverage area.
[0045] Step 2: Use the KD-tree algorithm to find several nearest points around each point in the point cloud file, and save the found points as a readable file, such as in the npy format, etc.
[0046] Step 3: Input the point cloud small pieces obtained in Step 1 and the nearest point file obtained in Step 2 into the point cloud encoder. The point cloud encoder sequentially includes a fully connected layer and four local feature aggregation units, and finally the output of the fully connected layer and the output of each local feature aggregation unit can be obtained The local feature aggregation unit consists of two local spatial encoders, an attention pooling layer, and a random sampling module. The local feature aggregation module encodes the coordinates and features of several nearest neighbor points around each point into a one-dimensional vector and then concatenates them into a two-dimensional feature map. The encoding formula is:
[0047]
[0048] where r i n represents the coordinate encoding value of the i-th point and its n-th nearest neighbor point. MLP represents a multi-layer perceptron. represents concatenation. p i and p i n are the x-y-z coordinates of the i-th point and its n-th nearest neighbor point, respectively.
[0049] Step 4: Input the image patch obtained in Step 1 into the image encoder. The image encoder is stacked by five convolutional blocks. Each convolutional block consists of two units composed of a convolutional layer with a 3×3 convolutional kernel, a batch normalization layer, and a ReLU activation layer. There is a max pooling layer with a kernel of 2 between every two convolutional blocks. Finally, the output f i 1 , f i 2 , f i 3 , f i 4 , f i 5 of each convolutional block is obtained.
[0050] Step 5: Input the high-dimensional feature map of the point cloud obtained in Step 3 into a multi-layer perceptron to obtain the feature map Then perform nearest neighbor interpolation on . Concatenate the obtained feature map with and input it into a multi-layer perceptron to obtain the feature map Then perform nearest neighbor interpolation on . Concatenate the obtained feature map with and input it into a multi-layer perceptron to obtain the feature map Then perform nearest neighbor interpolation on . Concatenate the obtained feature map with and input it into a multi-layer perceptron to obtain the feature map Then perform nearest neighbor interpolation on . Concatenate the obtained feature map with and input it into a multi-layer perceptron to obtain the feature map Then input into two stacked multi-layer perceptrons to obtain the point cloud endpoint-level segmentation result
[0051] Step 6: Input the high-dimensional feature map f of the image obtained in Step 4 i 5 and the one after dimensionality transformation (DS) into the globally knowledge-guided gate module (GKG) simultaneously, and then perform bilinear interpolation on the obtained feature map. Concatenate the obtained feature map with f i 4 and input the concatenated result into a convolutional block to obtain Input and the one after dimensionality transformation into the globally knowledge-guided gate module simultaneously, and then perform bilinear interpolation on the obtained feature map. Concatenate the obtained feature map with f i 3 and input the concatenated result into a convolutional block to obtain Input and the one after dimensionality transformation into the globally knowledge-guided gate module simultaneously, and then perform bilinear interpolation on the obtained feature map. Concatenate the obtained feature map with f i 2 and input the concatenated result into a convolutional block to obtain Input and the one after dimensionality transformation into the globally knowledge-guided gate module simultaneously, and then perform bilinear interpolation on the obtained feature map. Concatenate the obtained feature map with f i 1 and input the concatenated result into a convolutional block to obtain The globally knowledge-guided gate module respectively passes the feature map from the point cloud end through a max pooling layer (Maxpooling) with an output dimension of 1 and an average pooling layer (Mean pooling), concatenate the obtained feature maps, and then successively pass through a convolutional layer with a convolution kernel of 7×7 (7×7Conv), a batch normalization layer (BatchNorm), a ReLU activation layer, and a Sigmoid function layer, and then perform pixel-level multiplication with the feature map from the image end and then concatenate with The obtained result is the output of the globally knowledge-guided gate module. The convolutional block consists of two units stacked by a convolutional layer with a convolution kernel of 3×3, a batch normalization layer, and a ReLU activation layer.
[0052] Step 7: Input the point-level segmentation result obtained in Step 5 after dimensionality transformation and the one obtained in Step 6 into the class knowledge-guided gate module (CKG) simultaneously. The class knowledge-guided gate module passes the feature map from the image end After passing through a convolutional layer with a 1×1 convolutional kernel (1×1Conv), a batch normalization layer (BatchNorm), a ReLU activation layer, and a Dropout layer, dimensionality transformation and transposition are performed, and then pixel-wise multiplication is carried out with the coarsely segmented result from the point cloud that has undergone dimensionality transformation. Perform pixel-wise multiplication, and then pass through a Softmax function layer. The resulting output is transposed and then undergoes pixel-wise multiplication again with the that has undergone dimensionality transformation. After the resulting output undergoes dimensionality transformation, it is concatenated, and finally the output of the class knowledge-guided gate module is obtained.
[0053] Step 8: Pass the output obtained in Step 7 through a convolutional layer with a 1×1 convolutional kernel to obtain the final image segmentation result.
[0054] Step 9: Calculate the point cloud segmentation loss L and the image segmentation loss L respectively based on the point cloud segmentation pc obtained in Step 5 and the image segmentation result i obtained in Step 8, as well as the similarity loss L pi -SC between the point cloud and the image segmentation result:
[0055]
[0056]
[0057]
[0058] where N is the total number of points in the point cloud, M is the number of classes in the final segmentation, and W and H are the width and height of the image. and are the ground truths of the point cloud and the image segmentation results respectively. KL represents the KL divergence.
[0059] The total loss function is:
[0060] L total = L pc + L i + L pi-SC (10)
[0061] Step 10: Use the gradient descent algorithm to backpropagate the loss L total obtained in Step 9 and update the parameters of the entire network.
[0062] Step 11: Iterate repeatedly from Step 3 to Step 10 until the training ends.
[0063] Step 12: Testing phase. At the same time, input the remote sensing image and the LiDAR point cloud into the above-mentioned trained overall network to obtain the final 2D segmentation result.
[0064] It should be understood that the above description of the preferred embodiment is relatively detailed, and it should not be considered as a limitation to the protection scope of the present invention patent. Under the inspiration of the present invention, those of ordinary skill in the art can also make substitutions or deformations without departing from the scope protected by the claims of the present invention, and all fall within the protection scope of the present invention. The scope of protection requested by the present invention shall be subject to the appended claims.
Claims
1. A method for scene segmentation of multi-modal remote sensing data based on unbalanced knowledge-driven, characterized in that It includes the following steps: Step 1: Crop the LiDAR point cloud and the image in the dataset into small pieces of the same size; Step 2: Use the KD-tree algorithm to find several nearest points around each point in the point cloud file, and save the found points as a readable file; Step 3: Input the point cloud patches obtained in Step 1 and the nearest neighbor point file obtained in Step 2 into the point cloud encoder to obtain point cloud features at different scales k takes values from 1 to K; Step 4: Input the image patch obtained in Step 1 into the image encoder to obtain the image feature f i k , where k takes values from 1 to K; Step 5: The high-dimensional feature map of the point cloud obtained in Step 3 is input into a multi-layer perceptron to obtain a feature map Then is subjected to nearest neighbor interpolation, and the obtained feature map is concatenated with and input into a multi-layer perceptron to obtain a feature map Then repeat the above operation. Finally, the obtained feature map is concatenated with and input into a multi-layer perceptron to obtain a feature map Then is input into two stacked multi-layer perceptrons to obtain the segmentation result at the point cloud endpoint level Step 6: Input the high-dimensional feature map f obtained in Step 4 i K and the after dimensional transformation into the global knowledge-guided gate module simultaneously, then perform bilinear interpolation on the obtained feature map, and concatenate the obtained feature map with f i K-1 and input the concatenated result into a convolutional block to obtain Input and the after dimensional transformation into the global knowledge-guided gate module simultaneously, and repeat the above operations until Step 7: The segmentation result obtained in Step 5 After dimensional transformation and the one obtained in Step 6 Are simultaneously input into the class knowledge-guided gate module. The class knowledge-guided gate module processes the feature map from the image side After passing through a convolutional layer, a batch normalization layer, a ReLU activation layer, and a Dropout layer, followed by dimensional transformation and transposition, and then performs pixel-wise multiplication with the coarsely segmented result from the point cloud side after dimensional transformation Then passes through a Softmax function layer. The resulting output is transposed and then undergoes pixel-wise multiplication again with the one after dimensional transformation And then performs pixel-wise multiplication again. The resulting output is dimensionally transformed and then concatenated with Finally, the output of the class knowledge-guided gate module is obtained Step 8: The output obtained in Step 7 passes through a convolutional layer to obtain the final image segmentation result Step 9: Calculate the loss function based on the point cloud segmentation obtained in Step 5 and the image segmentation result obtained in Step 8 Step 10: Use the gradient descent algorithm to backpropagate the loss obtained in Step 9 and update the parameters of the entire network; Step 11: Iterate Steps 3 to 10 repeatedly until the training ends.
2. The multi-modal remote sensing data scene segmentation method based on unbalanced knowledge drive according to claim 1, wherein: In step 3, the point cloud encoder sequentially includes a fully connected layer and four local feature aggregation units, and finally the output of the fully connected layer and the output of each local feature aggregation unit can be obtained. The local feature aggregation unit consists of two local spatial encoders, an attention pooling layer, and a random sampling module.
3. The method for multi-modal remote sensing data scene segmentation based on unbalanced knowledge drive according to claim 2, wherein: In Step 3, the local feature aggregation module encodes the coordinate values and features of several nearest points around each point into a one-dimensional vector and then concatenates them into a two-dimensional feature map. The encoding formula used is: where r i n represents the coordinate encoding value of the i-th point and its n-th nearest neighbor point, and MLP represents a multi-layer perceptron, represents concatenation, p i and p i n are the x-y-z coordinates of the i-th point and its n-th nearest neighbor point, respectively.
4. The method for multi-modal remote sensing data scene segmentation based on unbalanced knowledge-driven according to claim 1, wherein: In Step 4, the image encoder is stacked by five convolutional blocks. Each convolutional block consists of two units composed of a convolutional layer with a convolution kernel of 3×3, a batch normalization layer, and a ReLU activation layer. A max pooling layer with a kernel of 2 is passed between every two convolutional blocks.
5. The method for multi-modal remote sensing data scene segmentation based on unbalanced knowledge-driven according to claim 1, wherein: The specific implementation method of Step 5 is as follows; Let K be 5, and input the high-dimensional feature map of the point cloud obtained in step 3 into a multi-layer perceptron to obtain a feature map Then perform nearest neighbor interpolation on and concatenate the resulting feature map with and input the concatenated result into a multi-layer perceptron to obtain a feature map Then perform nearest neighbor interpolation on and concatenate the resulting feature map with and input the concatenated result into a multi-layer perceptron to obtain a feature map Then perform nearest neighbor interpolation on and concatenate the resulting feature map with and input the concatenated result into a multi-layer perceptron to obtain a feature map Then perform nearest neighbor interpolation on and concatenate the resulting feature map with and input the concatenated result into a multi-layer perceptron to obtain a feature map Then input into two stacked multi-layer perceptrons to obtain the segmentation result at the point cloud endpoint level 6. The method for segmenting scenes of multimodal remote sensing data based on unbalanced knowledge-driven according to claim 1, wherein: The specific implementation method of Step 6 is as follows; K is set to 5, and the image high-dimensional feature map f obtained in step 4 is i 5 and dimensionally transformed At the same time, the global knowledge-guided gate module is input, and then the obtained feature map is bilinearly interpolated to obtain the feature map and f i 4 After concatenation, input the convolution block to get Will and dimensionally transformed At the same time, the global knowledge-guided gate module is input, and then the obtained feature map is bilinearly interpolated to obtain the feature map and f i 3 After concatenation, input the convolution block to get Will and dimensionally transformed At the same time, the global knowledge-guided gate module is input, and then the obtained feature map is bilinearly interpolated to obtain the feature map and f i 2 After concatenation, input the convolution block to get Will and dimensionally transformed At the same time, the global knowledge-guided gate module is input, and then the obtained feature map is bilinearly interpolated to obtain the feature map and f i 1 After concatenation, input the convolution block to get 7. The method for scene segmentation of multi-modal remote sensing data based on unbalanced knowledge-driven according to claim 1, characterized in that: The global knowledge-guided gate module described in step 6 takes the feature map from the point cloud and passes it through a max pooling layer and an average pooling layer with an output dimension of 1 respectively. After concatenating the obtained feature maps, they are successively passed through a convolutional layer with a convolutional kernel of 7×7, a batch normalization layer, a ReLU activation layer, and a Sigmoid function layer, and then multiplied pixel-wise with the feature map from the image side After that, it is concatenated with The output of the global knowledge-guided gate module is obtained. The convolutional block consists of two stacked units composed of a convolutional layer with a convolution kernel of 3×3, a batch normalization layer, and a ReLU activation layer.
8. The method for multi-modal remote sensing data scene segmentation based on unbalanced knowledge-driven according to claim 1, wherein: In Step 7, the kernel size of the convolutional block in the category knowledge-guided gate module is 1×1.
9. The method for scene segmentation of multi-modal remote sensing data based on unbalanced knowledge-driven as claimed in claim 1, wherein: In Step 8, the kernel size of the convolutional layer is 1×1.
10. The method for scene segmentation of multi-modal remote sensing data based on unbalanced knowledge-driven according to claim 1, characterized in that: In step 9, the loss function includes the point cloud segmentation loss L pc , the image segmentation loss L i and the similarity loss L pi-SC of the point cloud and image segmentation results: where N is the total number of points in the point cloud, M is the number of finally segmented categories, W and H are the width and height of the image, and are the ground truths of the point cloud and image segmentation results respectively, and KL represents the KL divergence; The total loss function is: L total = L pc + L i + L pi-SC (5).
Citation Information
Patent Citations
Automatic registration method for laser point cloud and sequence panoramic image
CN112767461A
Remote sensing image point cloud joint segmentation method based on cascade cross-modal network
CN114842022A