An Adaptive Fusion Method for Land Cover Classification of Multi-Source Remote Sensing Data

Features are extracted and adaptive fusion is carried out through two-dimensional and three-dimensional networks, the problem of insufficient information fusion in multi-source remote sensing data is solved, and the accuracy and generalization ability of geographic classification are improved.

CN116543191BActive Publication Date: 2025-07-22Chinese People's Liberation Army Cyberspace Force Information Engineering University
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310040195.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-11
Publication Date
2025-07-22
Estimated Expiration
2043-01-11

AI Technical Summary

Technical Problem

In the prior art, the multi-source remote sensing data fusion method ignores the three-dimensional spatial information of point clouds, resulting in insufficient information fusion and low classification accuracy.

Method used

The two-dimensional network and three-dimensional network are used to extract the features of remote sensing images and LiDAR point clouds respectively, and feature alignment is achieved through the sampling and reconstruction module, and non-linear adaptive fusion is used to perform non-linear adaptive fusion, and a classifier is used to classify land objects.

Benefits of technology

Effectively integrate the characteristics of remote sensing images and LiDAR point clouds, avoid information loss, and improve the accuracy and generalization ability of land objects classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116543191B_ABST
    Figure CN116543191B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of remote sensing data processing and application, and particularly relates to a method for classifying ground objects by adaptively fusing multi-source remote sensing data. Remote sensing images and LiDAR point clouds of the same area at the same time are obtained and input into the constructed ground object classification model to obtain the ground object classification result; wherein, the ground object classification model includes a two-dimensional network, a three-dimensional network, a sampling and reconstruction module, a feature fusion module and a classifier; the two-dimensional network is used to extract features from the input remote sensing images to obtain two-dimensional features, and the three-dimensional network is used to extract features from the input LiDAR point clouds to obtain three-dimensional features; the sampling and reconstruction module is used to align the two-dimensional features and the three-dimensional features; the feature fusion module is used to fuse the aligned two-dimensional features and three-dimensional features. The present invention breaks through the problem that most of the existing fusion methods are based on images and the utilization of three-dimensional information is insufficient, provides a new idea for the fusion of heterogeneous data, and has important significance in practical applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing data processing and application, and particularly relates to a method for classifying ground objects by adaptively fusing multi-source remote sensing data. Background Art

[0002] Accurate classification of ground objects based on remote sensing data is one of the important research contents of earth observation. With the improvement of the spectral and temporal resolutions of remote sensing images, a large number of ground object classification methods based on deep learning have been successively proposed. Applying convolutional neural networks to remote sensing images can significantly improve the ability to extract deep features. However, due to the lack of rich and diverse information in a single data source, there are still situations where it is difficult to accurately classify certain ground object categories. Jointly using multi-source remote sensing data is an important solution to break through this bottleneck. The Light Detection And Ranging (LiDAR) technology has the characteristics of being fast, active, and highly penetrable. The point cloud data obtained by it has a stable structure and can objectively and truly express the complex geometric information of the scene, becoming an important data source for high-precision ground three-dimensional information. At present, there have been studies on classifying ground objects by fusing the spectral and spatial characteristics of multi-source remote sensing data. However, limited by the differences in different sensors and data structures, there are still many challenges in the research on fusing multi-modal remote sensing data.

[0003] The research on the fusion classification of images and point clouds is mainly divided into three categories: early fusion at the input layer, intermediate fusion at the feature layer, and late fusion at the decision layer. Interacting with multi-modal data at the feature layer is a more reasonable and flexible fusion strategy and is also the most commonly used fusion method at present. Most of them are based on images. By converting the point cloud into a Digital Surface Model (DSM), spatial features are extracted from the image and the DSM, and then a classifier is used to jointly classify the superimposed spatial and spectral features. In recent years, research has made use of the powerful feature extraction ability of deep learning models in massive data to deeply generalize and extract features and reconstruct and couple the data structure information at the deep feature layer, so that multi-source data can achieve considerable interpretation accuracy at the feature-level hierarchy. However, most of them are also based on two-dimensional images, ignoring the three-dimensional spatial information of the point cloud. The unique three-dimensional advantage structure information is severely lost due to being projected onto a two-dimensional plane before feature input. At the same time, due to the high information coupling and data structure differences between multi-modal data, the existing fusion method based on feature concatenation will lead to insufficient information fusion and even damage the feature learning process of a single modality, and the classification accuracy still needs to be improved. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for classifying ground objects by adaptively fusing multi-source remote sensing data to solve the problem of low accuracy of ground object classification in the prior art.

[0005] To solve the above technical problems, the present invention provides a method for classifying ground objects by adaptively fusing multi-source remote sensing data. Remote sensing images and LiDAR point clouds of the same area at the same time are obtained and input into a constructed ground object classification model to obtain a ground object classification result. Among them, the ground object classification model includes a two-dimensional network, a three-dimensional network, a sampling reconstruction module, a feature fusion module, and a classifier. The two-dimensional network is used to extract features from the input remote sensing images to obtain two-dimensional features, and the three-dimensional network is used to extract features from the input LiDAR point clouds to obtain three-dimensional features. The sampling reconstruction module is used to align the two-dimensional features and three-dimensional features. The feature fusion module is used to perform fusion processing on the aligned two-dimensional features and three-dimensional features. The classifier is used to obtain the ground object classification results of each point cloud in the LiDAR point cloud based on the results of the feature fusion processing.

[0006] The beneficial effects are as follows: The present invention can fuse the input remote sensing images and LiDAR point clouds, and use the two-dimensional network and the three-dimensional network to extract the corresponding spectral-spatial features specific to the remote sensing images and the geometric features of the LiDAR point clouds, ensuring that the features learned by the network are not limited to the common information learned by the modalities and avoiding the loss of unique information. Furthermore, through a non-linear adaptive fusion method, the full fusion of heterogeneous features is effectively achieved, breaking through the problem that most existing fusion methods are based on images and do not utilize three-dimensional information sufficiently, providing a new idea for heterogeneous data fusion and having important significance in practical applications.

[0007] Furthermore, the feature fusion module is used to perform feature fusion processing by the following method: The aligned two-dimensional features and three-dimensional features are summed element-wise to obtain an aggregated feature; pointwise convolution is used to achieve cross-channel information fusion for each point, and local channel features and global channel features are calculated through a bottleneck structure; the attention weights of the two-dimensional features and the attention weights of the three-dimensional features are calculated based on the local channel features and the global channel features; based on the attention weights of the two-dimensional features and the attention weights of the three-dimensional features, as well as the two-dimensional features and the three-dimensional features, the feature fusion processing result is obtained.

[0008] The beneficial effects are as follows: This adaptive fusion processing method samples and reconstructs two-dimensional semantic features onto a three-dimensional point set to achieve heterogeneous feature alignment. The adaptive feature fusion method enables the network to adaptively and dynamically optimize the fusion process of heterogeneous features during training, thereby having stronger generalization ability.

[0009] Furthermore, the local channel features and the global channel features are respectively expressed as:

[0010]

[0011]

[0012] Wherein, L(F) and G(F) respectively represent local channel features and global channel features; g(·) represents global average pooling operation; δ(·) represents ReLU activation function; represents batch normalization processing; PConv i (·) represents pointwise convolution operation with different input and output channel numbers for each layer, and i = 1, 2, 3, 4.

[0013] Furthermore, the attention weight of the two-dimensional feature is expressed as:

[0014]

[0015] Wherein, M(F) represents the attention weight of the two-dimensional feature, and 1 - M(F) represents the attention weight of the three-dimensional feature; represents element-wise addition operation; L(F) and G(F) respectively represent local channel features and global channel features; σ(·) represents Sigmoid activation function.

[0016] Furthermore, the result of feature fusion processing is:

[0017]

[0018] Wherein, represents element-wise multiplication operation; X 2D and X 3D respectively represent two-dimensional feature and three-dimensional feature; Z represents the fused feature map.

[0019] Furthermore, when aligning the two-dimensional feature and the three-dimensional feature, for any three-dimensional point t i (x i , y i , z i ) on the LiDAR point cloud, its pixel position t s (x s , y s ) on the remote sensing image is:

[0020] x s = INT((x i - x0) / dp)

[0021] y s = INT((y i - y0) / dp)

[0022] Wherein, x0 and y0 represent the minimum geometric coordinates in the plane direction of the image area; dp represents the pixel resolution; INT(·) represents the integer function.

[0023] The beneficial effects are as follows: The alignment of features is achieved by sampling the two-dimensional features of the image onto the three-dimensional point set, enabling the fusion of two-dimensional and three-dimensional features.

[0024] Furthermore, when training the ground object classification model, if there are two-dimensional labels of remote sensing images, the two-dimensional labels are used as auxiliary information to enhance the training of the ground object classification model. And for the two-dimensional supervised classification loss function of the three-dimensional point cloud divided into C categories, it is expressed as:

[0025]

[0026] In the formula, L Seg 2D (t s , t s ′ 2D ) represents the two-dimensional supervised classification loss function; t s ′ 2D (n,c) represents the true two-dimensional label of the pixel where the sampled point is located; represents the probability that the ground object classification model predicts that point t s belongs to category c; N represents the total number of points.

[0027] The beneficial effects are as follows: There are two-dimensional labels of remote sensing images, which are applied to the supervised classification of the image semantic segmentation network and used as auxiliary information to enhance the training of the three-dimensional network, which can improve the classification accuracy.

[0028] Furthermore, when training the ground object classification model, for the three-dimensional labels of the LiDAR point cloud, the corresponding three-dimensional supervised classification loss function is expressed as:

[0029]

[0030] In the formula, L Seg 3D (t, t′ 3D ) represents the three-dimensional supervised classification loss function; t′ 3D (n,c) represents the true three-dimensional label of point t obtained by sampling; P t (n,c) represents the probability that the ground object classification model predicts that point t belongs to category c; N represents the total number of points.

[0031] Furthermore, the two-dimensional network is an FCN-8s network.

[0032] The beneficial effects are as follows: FCN-8s can accept image inputs of any size and achieve per-pixel image classification through fully convolutional, upsampling, and skip structures. Among them, the skip structure realizes the fusion of feature maps at different levels, and can better balance the low-level detailed local features and high-level semantic features of the image. Therefore, using FCN-8s as a two-dimensional network for feature extraction results in better classification effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1 is the network structure diagram of the multi-source remote sensing data adaptive fusion land cover classification of the present invention;

[0034] Figure 2 is the structure diagram of the adaptive fusion module of the present invention;

[0035] Figure 3 is the structure diagram of the three-dimensional network of the present invention;

[0036] Figure 4(a) is the training area diagram in the method embodiment of the present invention;

[0037] Figure 4(b) is the test area diagram in the method embodiment of the present invention;

[0038] Figure 5(a) is the experimental result diagram of classifying the test area using only the three-dimensional network DGCNN method;

[0039] Figure 5(b) is the experimental result diagram of classifying the test area using the method of the present invention;

[0040] Figure 5(c) is the standard segmentation result diagram of the test area. DETAILED DESCRIPTION OF THE INVENTION

[0041] The present invention proposes a multi-source remote sensing data adaptive fusion land cover classification method using an independent branch network structure. This method constructs a multi-source remote sensing data adaptive fusion network with independent branches for typical land cover classification. Among them, the two- and three-dimensional independent branch networks can fully extract the corresponding spatial geometric features and semantic features of the two modalities of images and point clouds. The step of cross-modal sampling can obtain point-by-point multi-source features, and a constructed non-linear feature fusion method based on the attention mechanism can realize the adaptive fusion of two- and three-dimensional semantic features. To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0042] Method Embodiment:

[0043] The present invention constructs a multi-source remote sensing data adaptive fusion network with independent branches for typical land cover classification. Hereinafter, this network model is called the land cover classification model, and its network structure diagram is as Figure 1As shown in the figure, the ground object classification model includes a two-dimensional network, a three-dimensional network, a sampling reconstruction module, a feature fusion module, and a classifier. The specific introduction of each part is as follows:

[0044] 1) The two-dimensional network is used to extract features from the input remote sensing image to obtain two-dimensional features. Due to the characteristics of same object with different spectra and different objects with the same spectrum commonly existing in remote sensing images, the present invention adopts an improved multi-scale feature fusion fully convolutional neural network (Fully Convolutional Networks, FCN) for feature extraction of remote sensing images. The basic FCN network can accept image inputs of any size and uses full convolution, upsampling, and skip structures to achieve pixel-by-pixel image classification. Among them, the skip structure realizes the fusion of feature maps at different levels, and can better balance the low-level detailed local features and high-level semantic features of the image. In this embodiment, the FCN-8s network with the best classification effect among the three different resolution fusion methods is used as the two-dimensional semantic feature extraction network. Of course, as other implementation manners, other networks can also be selected.

[0045] 2) The three-dimensional network is used to extract features from the input LiDAR point cloud to obtain three-dimensional features. For large-scale airborne LiDAR point clouds, recently many studies have explored local structures through network improvements to learn their multi-level feature representations. Similarly, based on existing research, the present invention adopts a neural network integrating geometric convolution for feature extraction of airborne LiDAR point clouds. The specific structure of this network can refer to the introduction in "Ground Object Classification of Airborne LiDAR Point Cloud with Multi-Feature Fusion and Geometric Convolution" published in the Journal of Image and Graphics by authors such as Dai Mofan and Xing Shuai in February 2022. The structure is as Figure 3 shown. The model mainly consists of two modules, namely the APD module and the FGC module. The APD module performs processing such as point cloud partitioning on large-area and low-density airborne LiDAR point cloud data and extracts and fuses the input dimension of shallow features. The FGC module includes two branches. The top branch extracts local and global information of the geometric structure of the airborne point cloud based on 3D coordinates, and the bottom branch uses the point-by-point depth semantic feature extraction module of PointNet++ to encode the shallow features. Finally, it aggregates with the geometric depth features obtained by the top branch and performs spatial upsampling to obtain point-by-point multi-scale depth features, that is, three-dimensional features. This network encodes the spatial geometric structure of points through hierarchical convolution and aggregates with global information to be able to extract multi-scale point-by-point depth features, realizing the acquisition of the complex geometric structure of large-area point clouds. At the same time, processing methods such as point cloud partitioning and class balancing are adopted to enhance the applicability of the model to airborne point clouds, retain the original three-dimensional spatial structure at the data input level, and directly output the point-by-point ground object extraction results. Of course, as other implementation manners, other networks in the prior art can also be selected.

[0046] 3) The sampling and reconstruction module is used to align two-dimensional features and three-dimensional features. First, before the two-dimensional network and the three-dimensional network output probabilities, the semantic feature outputs of the two branch networks are kept in the same dimension, and the semantic features are aligned by establishing the connection of the last linear layer. Second, considering that although independent branch networks are used for multi-modal feature extraction, since the point cloud and the image are in different metric spaces, it is still difficult to directly fuse 3D point clouds and 2D images: the three-dimensional coordinates of LiDAR point clouds are accurate and contain more structural information, but the point density is low; remote sensing images have a wide coverage range and rich spectral information, but the ground object resolution is limited and the geometric structure information is insufficient, such as the side of a building, etc. Therefore, when fusing the features of the two types of data, considering that the three-dimensional information and three-dimensional features have good compatibility with the two-dimensional, especially the three-dimensional coordinates of vegetation, the alignment of features is achieved by sampling the two-dimensional features of the image onto the three-dimensional point set. Specifically, sampling the point cloud on the image can obtain a three-dimensional point cloud set Any three-dimensional point t can be obtained i The pixel position t where it is located s (x s , y s ) is shown as follows:

[0047] x s = INT((x i - x0) / dp) y s = INT((y i - y0) / dp)

[0048] In the formula, x0 and y0 represent the minimum geometric coordinates in the plane direction of the area where the image is located; dp represents the pixel resolution; INT(·) represents the rounding function.

[0049] 4) The feature fusion module is used to fuse the aligned two-dimensional features and three-dimensional features. On the feature fusion strategy, simple splicing or adding fusion methods are only fixed linear aggregations of feature maps, and when the scene changes, the generalization ability of the model cannot be guaranteed. Inspired by the visual perception of the human visual system, the attention mechanism can adaptively allocate corresponding weights according to the importance of the target. Therefore, the weights of the semantic features extracted from the point cloud and the image should also be redistributed according to their contributions to the classification performance. Inspired by the attention feature fusion method (Attention Feature Fusion, AFF) in images, the present invention proposes an adaptive fusion method for heterogeneous features, which can dynamically learn and optimize the fusion process of two-dimensional and three-dimensional features of the point cloud in a non-linear manner according to the three-dimensional labels of the point cloud. The specific structure of the feature fusion module is as Figure 2 shown, and the process is as follows:

[0050] ① Given the two-dimensional and three-dimensional features X of the point cloud to be fused 2D , The aggregated feature X is obtained through an element summation operation.

[0051] ② Use pointwise convolution to achieve cross-channel information fusion for each point, and calculate local channel features through a bottleneck structure and global channel features

[0052]

[0053] In the formula, L(F) and G(F) represent local channel features and global channel features respectively; g(·) represents global average pooling operation; δ(·) represents ReLU activation function; represents batch normalization (Batch Normalization, BN); PConv i (·) represents pointwise convolution operations with different input and output channel numbers for each layer, i = 1, 2, 3, 4.

[0054] ③ Both features have the same dimension as the initial feature, which can retain and highlight the subtle details in the underlying features, and then calculate the attention weights M(F) of the two types of features:

[0055]

[0056] In the formula, represents element-wise addition operation; σ(·) represents Sigmoid activation function.

[0057] ④ The fused feature map is expressed as:

[0058]

[0059] In the formula, represents element-wise multiplication operation. Both the fusion weight M(F) and 1 - M(F) are composed of real numbers between 0 and 1. At this time, the network can perform soft selection or weighted average of the weights of X 2D and X 3D , and achieve dynamic optimization during the training process to ensure effective fusion of heterogeneous features, thereby obtaining a soft label that is more robust to data noise and enabling the model to have stronger generalization ability.

[0060] 5) The classifier is used to obtain the ground object classification result of each point cloud in the LiDAR point cloud based on the feature fusion processing result.

[0061] So far, the structure of each part of the object classification model has been introduced. The object classification model is trained using training data. After the training is completed, the trained object classification model can be used to classify objects. Assume that the experimental data is the Vaihingen multi-source data set provided by ISPRS, including the original orthophoto remote sensing image data with a resolution of 0.09m and an average point cloud density of 6.7pts / m 2 Airborne LiDAR point cloud data. Before the test, the two pieces of data with both point cloud and image were registered to obtain two training images and test images with sizes of 2006×3007 and 1919×2569, respectively, as well as the corresponding training point cloud containing 348702 points and the test point cloud containing 174145 points. The test area is shown in Figure 4(a) and Figure 4(b). At the same time, since the category labels of the original point cloud and image data are not unified, they need to be classified and aligned.

[0062] Step 1: Build Figure 1 The feature classification model shown.

[0063] Step 2: Use the training data to train the land object classification model. After the training is completed, the parameter values obtained by the training are assigned to the network to obtain a trained land object classification model.

[0064] When the 2D labels of the image exist, they can be used for supervised classification of the image semantic segmentation network as auxiliary information to enhance the 3D network training. The two-dimensional supervised classification loss function is expressed as:

[0065]

[0066] Where, L Seg 2D (t s ,t s ' 2D ) represents the two-dimensional supervised classification loss function; t s ' 2D (n,c) Indicates the true two-dimensional label of the pixel where the point is located after sampling; Represents the predicted point t of the ground feature classification model s The probability of belonging to category c; N represents the total number of points.

[0067] For the training point cloud obtained after sampling, the three-dimensional supervised classification loss function can be obtained through three-dimensional supervised training under the three-dimensional branch:

[0068]

[0069] Where, L Seg 3D(t, t′ 3D ) represents a three-dimensional supervised classification loss function; t′ 3D (n,c) represents the true three-dimensional label of the sampled point t; P t (n,c) represents the probability that the ground object classification model predicts that the point t belongs to the category c; N represents the total number of points.

[0070] Step 3: Obtain the remote sensing image and LiDAR point cloud of the same area at the same time, and input them into the constructed ground object classification model to obtain the ground object classification result.

[0071] The effects of the present invention are verified and described through the following simulation experiments.

[0072] 1) Simulation conditions: The hardware uses an Intel Core i9-7900 central processing unit, an Nvidia GeForce RTX3090Ti graphics processing unit, and 128GB of memory. The software uses the Pytorch library to implement the present invention.

[0073] 2) Parameter settings: The input batch is set to 16, the number of iterations is set to 200, the initial learning rate is 0.005, the learning rate decay coefficient is 0.5, the decay step size is 20000, and the optimization methods of stochastic gradient descent and L2 regularization are selected.

[0074] 3) Simulation Results: Training is carried out in the training area of the ISPRS multi-source dataset and testing is carried out in the testing area. Precision (Accuracy), Overall Accuracy (OA), and mean Intersection over Union (mIoU) are used as evaluation metrics. A comparative experiment is conducted by comparing the fusion method of the present invention with four 3D semantic segmentation networks, namely PointNet++, PointSIFT, DGCNN, and RandLA-Net. Among them, Fig. 5(a) shows the experimental results of classifying the testing area using only the 3D network DGCNN method; Fig. 5(b) shows the experimental results of classifying the testing area using the method of the present invention; Fig. 5(c) shows the standard segmentation results of the testing area. Table 1 is a comparison table of the final classification results of various methods. At the same time, the adaptive fusion method of the present invention is compared with two other linear fusion methods, and Table 2 shows the comparison results of the three fusion methods. The experimental results show that: ① Compared with the method of classifying point clouds alone, the present invention significantly enriches the available information of the network by introducing image data features through an independent branch network, and significantly improves the extraction ability of vegetation and buildings; ② For typical ground objects, the present invention can obtain better classification results than other advanced point cloud classification methods, and significantly improves the classification accuracy in areas with complex point cloud distributions with the support of images; ③ Compared with the above methods, the present invention effectively alleviates the phenomenon of mixed misclassification existing in other methods in complex scenarios, divides detailed objects more precisely and accurately, and well preserves the boundary information of ground objects.

[0075] Table 1

[0076]

[0077] Table 2

[0078]

[0079] In summary, the present invention has the following characteristics: 1) The present invention designs a two-stream architecture that can simultaneously input images and point clouds and has 2D and 3D branches, which are used to extract the spectral-spatial features specific to remote sensing images and the geometric features of LiDAR point clouds respectively, ensuring that the features learned by the network are not limited to the common information learned by the modality and avoiding the loss of unique information. 2) The present invention designs a non-linear feature fusion method based on attention. On the basis of aligning 2D and 3D semantic features, the 2D semantic features are sampled and reconstructed onto the 3D point set to achieve cross-source feature alignment. The adaptive feature fusion method enables the network to adaptively and dynamically optimize the fusion process of cross-source features during training, thereby having stronger generalization ability.

[0080] Specific embodiments are given above, but the present invention is not limited to the described embodiments. The basic idea of the present invention lies in the above basic solution. For those of ordinary skill in the art, according to the teachings of the present invention, it does not require creative labor to design various deformed models, formulas, and parameters. Changes, modifications, substitutions, and variations made to the embodiments without departing from the principles and spirit of the present invention still fall within the protection scope of the present invention.

Claims

1. An adaptive fusion method for multi-source remote sensing data for land cover classification, characterized in that, Obtain remote sensing images and LiDAR point clouds in the same area at the same time, and input them into the constructed ground object classification model to obtain the ground object classification results; Among them, the ground object classification model includes a two-dimensional network, a three-dimensional network, a sampling reconstruction module, a feature fusion module, and a classifier; the two-dimensional network is used to extract features from the input remote sensing images to obtain two-dimensional features, and the three-dimensional network is used to extract features from the input LiDAR point clouds to obtain three-dimensional features; the sampling reconstruction module is used to align the two-dimensional features and three-dimensional features; the feature fusion module is used to perform fusion processing on the aligned two-dimensional features and three-dimensional features; the classifier is used to obtain the ground object classification results of each point cloud in the LiDAR point cloud according to the feature fusion processing results; When aligning two-dimensional features and three-dimensional features, any three-dimensional point t on the LiDAR point cloud i (x i , y i , z i ) The pixel position t on the remote sensing image s (x s , y s ) is: x s = INT((x i - x0) / dp) y s = INT((y i - y0) / dp) In the formula, x0 and y0 represent the minimum geometric coordinates in the plane direction of the image area; dp represents the pixel resolution; INT(·) represents the rounding function; When training a ground object classification model, if there are two-dimensional labels of remote sensing images, the two-dimensional labels are used as auxiliary information to enhance the training of the ground object classification model, and for the two-dimensional supervised classification loss function \(L\) of the three-dimensional point cloud divided into \(C\) categories Seg 2D (t s ,t s ′ 2D ) is expressed as: where t s ′ 2D (n,c) represents the true two-dimensional label of the pixel where the sampled point t s is located; represents the probability that the point t s predicted by the ground object classification model belongs to the category c; N represents the total number of points; for the three-dimensional label of the LiDAR point cloud, the corresponding three-dimensional supervised classification loss function L Seg 3D (t, t′ 3D ) is expressed as: t′ 3D (n,c) represents the true 3D label of the point t obtained by sampling; P t (n,c) represents the probability that the ground object classification model predicts that the point t belongs to the category c.

2. The multi-source remote sensing data adaptive fusion land cover classification method according to claim 1, wherein The feature fusion module is used to perform feature fusion processing by the following method: Perform element-wise summation on the aligned two-dimensional features and three-dimensional features to obtain aggregated features; Use pointwise convolution to achieve cross-channel information fusion for each point, and calculate local channel features and global channel features through a bottleneck structure; Calculate the attention weights of the two-dimensional features and the attention weights of the three-dimensional features based on the local channel features and the global channel features; Obtain the feature fusion processing results according to the attention weights of the two-dimensional features and the attention weights of the three-dimensional features, as well as the two-dimensional features and the three-dimensional features.

3. The multi-source remote sensing data adaptive fusion ground object classification method according to claim 2, wherein The local channel features and the global channel features are respectively expressed as: In the formula, L(F) and G(F) respectively represent the local channel features and the global channel features; g(·) represents the global average pooling operation; δ(·) represents the ReLU activation function; represents batch normalization processing; PConv i (·) represents the pointwise convolution operation with different input and output channel numbers for each layer, where i = 1, 2, 3, 4.

4. The multi-source remote sensing data adaptive fusion ground object classification method according to claim 2, wherein The attention weight of the two-dimensional feature is expressed as: Wherein, M(F) represents the attention weight of two-dimensional features, and 1 - M(F) represents the attention weight of three-dimensional features; represents the element-wise addition operation; L(F) and G(F) respectively represent local channel features and global channel features; σ(·) represents the Sigmoid activation function.

5. The multi-source remote sensing data adaptive fusion ground object classification method according to claim 4, wherein The feature fusion processing result is: In the formula, represents an element-wise multiplication operation; X 2D and X 3D represent two-dimensional features and three-dimensional features respectively; Z represents the fused feature map.

6. The multi-source remote sensing data adaptive fusion land cover classification method according to any one of claims 1 to 5, characterized in that The two-dimensional network is an FCN-8s network.