Slope natural earth surface point cloud semantic segmentation method based on multi-modal data

Through the semantic segmentation method of slope natural surface point clouds of multimodal data, combined with feature extraction and encoding and decoding of point clouds and remote sensing data, the accuracy of semantic segmentation of slope natural surface point clouds in the existing technology is solved, and high-precision semantic segmentation and feature recognition are achieved.

CN120339604APending Publication Date: 2025-07-18WUHAN UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510320294.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing semantic segmentation method of surface point clouds is difficult to accurately and efficiently segment natural land objects with large scale ranges, complex local geometric features and highly random distributions. Especially in natural land objects and artificial structural characteristics are difficult to distinguish between natural land objects and artificial structures.

Method used

The semantic segmentation method of slope natural surface point cloud based on multimodal data is adopted. By obtaining point cloud data and remote sensing image data for coordinate matching, point cloud and remote sensing features are extracted, and feature encoding and decoding are performed through encoder and decoder, and semantic segmentation is combined with the fully connected layer of the embedded KAN network layer to generate a semantic segmentation result distribution map.

Benefits of technology

Without preprocessing or postprocessing, high-precision semantic segmentation of complex local feature surface point clouds is realized, improving the ability to identify and analyze surface geometry and complex local features, and optimizing the calculation and parameter efficiency of the output layer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339604A_ABST
    Figure CN120339604A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a slope natural earth surface point cloud semantic segmentation method based on multi-modal data, and relates to the technical field of remote sensing recognition, and the method comprises the steps: obtaining point cloud data and remote sensing image data of a research region, and carrying out the feature extraction based on the point cloud data after coordinate matching, and obtaining point cloud features; extracting first remote sensing features based on the remote sensing image data features after coordinate matching, and processing the first remote sensing features to obtain second remote sensing features; and carrying out feature coding through an encoder to obtain fusion features, carrying out feature decoding on the fusion features through a decoder to obtain a decoding result, and outputting a semantic segmentation result distribution diagram of the research area based on the full connection layer of the embedded KAN network layer. According to the method, the earth surface point cloud with the complex local features can be processed with high precision under the condition that preprocessing / post-processing is not needed, the geometric morphology characteristics of the earth surface can be comprehensively captured through the point cloud information, and the recognition and analysis capability of the complex local features of the earth surface can be improved through the spectral information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of remote sensing recognition, and in particular, to a method for semantic segmentation of natural surface point clouds of slopes based on multi-modal data. Background Art

[0002] A slope is a geological body on the earth's surface formed naturally or artificially with a lateral free face. Currently, different natural ground objects on the slope surface can be identified and distinguished through semantic segmentation of natural surface point clouds, which can provide data support for fields such as geological engineering, environmental monitoring, and natural disaster management, and convert large-scale natural surface point cloud data into meaningful and interpretable information, enabling the relevant field staff to more accurately analyze the topographic conditions of the slope, identify potential landslide and erosion areas, and implement corresponding preventive measures.

[0003] Currently, semantic segmentation of surface point clouds is generally carried out mainly through traditional methods such as filtering or manual visual interpretation. These methods can often accurately distinguish indoor point clouds and urban street point clouds because indoor point clouds and urban street point clouds are generally composed of artificial structures, such as buildings, furniture, vehicles, etc., and their structures are usually more regular and uniform, and the geometric features of the point clouds are obvious.

[0004] However, natural surface point clouds usually have a large research scale and include complex natural ground objects, generally with complex local geometric features and highly randomly distributed natural vegetation. In some cases, there are often a small number of residential gathering points in natural surface point clouds, with a certain number of point clouds with artificial structure features. Existing surface point cloud semantic segmentation methods are difficult to accurately and efficiently perform semantic segmentation on natural ground objects with large-scale ranges, complex local geometric features, and highly randomly distributed, and it is difficult to output accurate ground object type results.

[0005] Therefore, there is currently a lack of a method that can accurately perform semantic segmentation on natural surface point clouds of slopes. Summary of the Invention

[0006] The embodiments of the present application provide a method for semantic segmentation of natural surface point clouds of slopes based on multi-modal data to solve the defects of the above-mentioned related technologies, and the technical solutions are as follows:

[0007] In a first aspect, the embodiments of the present application provide a method for semantic segmentation of natural surface point clouds of slopes based on multi-modal data, including:

[0008] Obtain the point cloud data and remote sensing image data of the research area, and perform coordinate matching on the point cloud data and the remote sensing image data;

[0009] Extract features from the point cloud data after coordinate matching to obtain point cloud features; extract features from the remote sensing image data after coordinate matching to obtain the first remote sensing feature, and perform feature enhancement processing and feature optimization processing on the first remote sensing feature to obtain the second remote sensing feature;

[0010] Based on the point cloud features and the second remote sensing feature, perform feature encoding through an encoder to obtain a fused feature;

[0011] Based on the fused feature, perform feature decoding through a decoder to output a decoding result;

[0012] Input the decoding result into the fully connected layer embedded with the KAN network layer, output the probability of the ground object type corresponding to each pixel unit, determine the ground object type corresponding to the pixel unit, and generate the semantic segmentation result distribution map of the research area.

[0013] In an alternative solution of the first aspect, the coordinate matching of the point cloud data and the remote sensing image data includes:

[0014] According to the central point coordinates corresponding to each point cloud area in the point cloud data, determine the remote sensing area within a preset range centered on the central point coordinates, and extract the remote sensing image data within the corresponding remote sensing area;

[0015] Among them, the range of the remote sensing area covers the corresponding point cloud area.

[0016] In an alternative solution of the first aspect, the feature extraction based on the remote sensing image data after coordinate matching to obtain the first remote sensing feature, and performing feature enhancement processing and feature optimization processing on the first remote sensing feature to obtain the second remote sensing feature includes:

[0017] Through multiple convolutional layers, respectively extract features from the remote sensing image data within each remote sensing area to obtain the first remote sensing feature corresponding to each remote sensing area;

[0018] Input the first remote sensing feature into the SA module, determine the channel attention weight through the channel attention module in the SA module, determine the spatial attention weight through the spatial attention module in the SA module, and perform feature enhancement through feature rearrangement operation by combining the channel attention weight and the spatial attention weight, and output the enhanced feature;

[0019] Input the enhanced feature into the KAN network, and perform numerical correction on the enhanced feature through the KAN network to output the second remote sensing feature corresponding to each remote sensing area.

[0020] In an alternative solution of the first aspect, the encoder includes a plurality of encoding layers, and each encoding layer includes a point cloud feature extraction layer and L N convolutional layers; performing the step of feature extraction on the point cloud data after coordinate matching through the plurality of encoding layers in the encoder, specifically including:

[0021] Grouping the point cloud data through the point cloud feature extraction layer in the l-th encoding layer, and extracting the point cloud features corresponding to each group of point clouds. The point cloud features of each l-th layer obtained are:

[0022]

[0023] After obtaining the point cloud features of each l-th layer, through L N convolutional layers, performing the step of feature extraction on the remote sensing image data after coordinate matching, and the steps of feature enhancement processing and feature optimization processing to obtain the second remote sensing feature, and extracting the second remote sensing feature matching the coordinates of each group of point clouds, including:

[0024]

[0025] Performing feature encoding on the point cloud features and the second remote sensing features to obtain the fused feature of the l-th encoding layer, and applying the formula:

[0026]

[0027]

[0028] where N (l) is the number of point clouds in the l-th encoding layer, d (l) is the coordinate dimension of the point clouds in the l-th encoding layer, is the dimension of the point cloud features in the l-th encoding layer, H is the height of the feature map corresponding to the second remote sensing feature, W is the width of the feature map corresponding to the second remote sensing feature, is the dimension of the remote sensing feature output by the l N -th convolutional layer, l N is the ordinal number of the convolutional layer, L N is the total number of convolutional layers, enc is the point cloud feature of the l-th encoding layer, is the second remote sensing feature, and the number of point cloud features in each layer is the same as the number of corresponding second remote sensing features, is the dimension of the second remote sensing feature, is the fused feature output by the l-th encoding layer.

[0029] In an alternative embodiment of the first aspect, the decoder includes a plurality of decoding layers;

[0030] Input the fused features into the decoder, and perform feature decoding based on the fused features through a plurality of decoding layers respectively, and output the decoding result, including:

[0031] Perform upsampling on the input features through the l-th decoding layer in the decoder, and output the upsampled point cloud feature map;

[0032] Concatenate the upsampled point cloud feature map with the fused features output by the (L - l + 1)-th encoding layer, and output the concatenated features;

[0033] Perform feature enhancement on the concatenated features through the CBAM module, and output the decoding result of the corresponding decoding layer;

[0034] Input the decoding result into the (l + 1)-th decoding layer, and perform the steps of upsampling the input features and subsequent steps until the L-th decoding layer, and output the final decoding result;

[0035] Wherein, the number of layers of the decoding layer and the total number of encoding layers are both L.

[0036] In an alternative embodiment of the first aspect, inputting the decoding result into the fully connected layer embedded with the KAN network layer, and outputting the probability of the ground object type corresponding to each pixel unit, including:

[0037] The processing process of the fully connected layer of the embedded KAN network layer for the input decoding result applies the formula:

[0038]

[0039] The fully connected layer of the embedded KAN network layer outputs the probability of the corresponding ground object type based on each input decoding result, determines the corresponding pixel unit according to the position coordinates of the point cloud features and remote sensing features corresponding to the decoding result, and obtains the probability of the ground object type corresponding to each pixel unit in the research area;

[0040] Wherein, x is the decoding result of the L-th decoding layer, W l is the weight matrix, b l is the bias vector, h l is the feature output of the fully connected layer of the embedded KAN network layer, L is the total number of layers of the fully connected layer of the embedded KAN network layer, the KAN network layer is located at the L-th decoding layer, ReLU is the activation function, and ф is a learnable univariate non-linear function.

[0041] In a second aspect, an embodiment of the present application further provides a slope natural ground surface point cloud semantic segmentation device based on multi-modal data, including:

[0042] A data acquisition module, configured to acquire point cloud data and remote sensing image data of a research area, and perform coordinate matching on the point cloud data and the remote sensing image data;

[0043] A feature extraction module, configured to perform feature extraction based on the point cloud data after coordinate matching to obtain point cloud features; the feature extraction module is further configured to perform feature extraction based on the remote sensing image data after coordinate matching, extract first remote sensing features, and perform feature enhancement processing and feature optimization processing on the first remote sensing features to obtain second remote sensing features;

[0044] A feature fusion module, configured to perform feature encoding on the point cloud features and the second remote sensing features through an encoder to obtain fusion features;

[0045] A feature decoding module, configured to perform feature decoding based on the fusion features through a decoder and output a decoding result;

[0046] A result output module, configured to input the decoding result into a fully connected layer embedded with a KAN network layer, output the probability of the ground object type corresponding to each pixel unit, determine the ground object type corresponding to the pixel unit, and generate a semantic segmentation result distribution map of the research area.

[0047] In an alternative solution of the second aspect, the feature extraction module is further configured to perform feature extraction based on the remote sensing image data after coordinate matching, extract first remote sensing features, and perform feature enhancement processing and feature optimization processing on the first remote sensing features to obtain second remote sensing features, including:

[0048] Performing feature extraction on the remote sensing image data in each remote sensing area through multiple convolutional layers in the feature extraction module to extract first remote sensing features corresponding to each remote sensing area;

[0049] Inputting the first remote sensing features into an SA module, determining channel attention weights through a channel attention module in the SA module of the feature extraction module, determining spatial attention weights through a spatial attention module in the SA module, and performing feature enhancement through a feature rearrangement operation in combination with the channel attention weights and the spatial attention weights, and outputting enhanced features;

[0050] Inputting the enhanced features into a KAN network, and performing numerical correction on the enhanced features through the KAN network of the feature extraction module to output second remote sensing features corresponding to each remote sensing area.

[0051] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method provided by the first aspect or any implementation manner of the first aspect of the embodiments of the present application is implemented.

[0052] In a fourth aspect, the present application further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method provided by the first aspect or any implementation manner of the first aspect of the embodiments of the present application is implemented.

[0053] The beneficial effects brought by the technical solutions provided by some embodiments of the present application at least include:

[0054] A method for semantic segmentation of natural surface point clouds of slopes based on multi-modal data provided by an embodiment of the present application can accurately process surface point clouds with complex local features without any preprocessing / postprocessing. First, by combining multi-modal data of point clouds and remote sensing data to enrich the data dimension, the model can understand the surface situation from both geometric and spectral dimensions, improving the model's feature understanding ability. It can not only comprehensively capture the geometric morphological characteristics of the surface through point cloud information, but also improve the recognition and analysis ability of complex local features of the surface through spectral information. By adding a KAN network layer to the fully connected layer, the efficient feature extraction ability of the traditional Multilayer Perceptron Layer (MLP) can be utilized in the early stage, and a more optimized and accurate output can be provided by the KAN layer at the decision stage of the final network's full connection. Finally, without sacrificing the complex feature expressions learned by the deep network in the early stage, the calculation and parameter efficiency of the output layer are optimized. Description of the Drawings

[0055] In order to more clearly illustrate the technical solutions in the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0056] Figure 1 is one of the flow diagrams of a method for semantic segmentation of natural surface point clouds of slopes based on multi-modal data provided by an embodiment of the present application;

[0057] Figure 2 is the second of the flow diagrams of a method for semantic segmentation of natural surface point clouds of slopes based on multi-modal data provided by an embodiment of the present application;

[0058] Figure 3It is a schematic diagram of coordinate matching for a method of semantic segmentation of natural ground surface point clouds of slopes based on multi-modal data provided by an embodiment of the present application;

[0059] Figure 4 It is a schematic diagram of image feature extraction for a method of semantic segmentation of natural ground surface point clouds of slopes based on multi-modal data provided by an embodiment of the present application;

[0060] Figure 5 It is a schematic diagram of feature fusion for a method of semantic segmentation of natural ground surface point clouds of slopes based on multi-modal data provided by an embodiment of the present application;

[0061] Figure 6 It is a schematic diagram of feature decoding for a method of semantic segmentation of natural ground surface point clouds of slopes based on multi-modal data provided by an embodiment of the present application;

[0062] Figure 7 It is a schematic diagram of the overall structure of the fully connected layer of the embedded KAN network layer for a method of semantic segmentation of natural ground surface point clouds of slopes based on multi-modal data provided by an embodiment of the present application;

[0063] Figure 8 It is a comparison chart of the decision optimization accuracy of the fully connected layer of the KAN network layer for a method of semantic segmentation of natural ground surface point clouds of slopes based on multi-modal data provided by an embodiment of the present application;

[0064] Figure 9 It is a comparison chart of the decision optimization efficiency of the fully connected layer of the KAN network layer for a method of semantic segmentation of natural ground surface point clouds of slopes based on multi-modal data provided by an embodiment of the present application;

[0065] Figure 10 It is a schematic diagram of the semantic segmentation result for a method of semantic segmentation of natural ground surface point clouds of slopes based on multi-modal data provided by an embodiment of the present application;

[0066] Figure 11 It is a schematic diagram of the structure of a device for semantic segmentation of natural ground surface point clouds of slopes based on multi-modal data provided by an embodiment of the present application;

[0067] Figure 12 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0068] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.

[0069] As used in the specification, claims and the above drawings of this application, the terms "comprising" and "having", and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or modules is not limited to the listed steps or modules, but may optionally further include steps or modules not listed, or may optionally further include other steps or modules inherent to these processes, methods, products or devices.

[0070] It should be noted that the terms "first" and "second" in this application are only used to distinguish similar objects and do not represent a specific order for the objects. Understandably, "first" and "second" can be interchanged in their specific order or sequence under allowable circumstances. It should be understood that the objects distinguished by "first" and "second" can be interchanged appropriately so that the embodiments of this application described here can be implemented in an order other than those described or illustrated here.

[0071] It should be noted that a method for semantic segmentation of natural surface point clouds of slopes based on multi-modal data provided in this application makes up for the defects of point cloud deep learning models in the related art for surface semantic segmentation. The network provided in this application mainly includes a point cloud feature extraction branch and an image feature extraction branch. Through multiple comparative experiments, it can be determined that the model provided in this application has a significant improvement in the accuracy of surface point cloud semantic segmentation compared to various point cloud models in the related art.

[0072] Understandably, "semantic segmentation" refers to the process of assigning each pixel or pixel unit in an image to a specific category or label, with the aim of identifying and distinguishing different ground object types. For example, semantic segmentation can assign a class label to each pixel in an image, and each pixel in the output annotation map is marked with the ground object category to which it belongs. For example, 4 class labels can be set. The first class is individual vegetation, the second class is vegetation group, the third class is artificial objects (including houses, vehicles, fences, etc.), and the fourth class is the ground surface. The embodiments of this application are not limited thereto.

[0073] The following will describe this application in detail with specific embodiments.

[0074] Next, in combination with Figure 1-2 , a method for semantic segmentation of natural surface point clouds of slopes based on multi-modal data provided by an embodiment of this application will be introduced. Specifically, please refer to Figure 1-2 , Figure 1-2 shows a schematic flowchart of a method for semantic segmentation of natural surface point clouds of slopes based on multi-modal data provided by an embodiment of this application. As Figure 1 shown, the method includes the following steps:

[0075] S101. Obtain the point cloud data and remote sensing image data of the research area, and perform coordinate matching on the point cloud data and the remote sensing image data;

[0076] S102. Extract features based on the point cloud data after coordinate matching to obtain point cloud features; extract features based on the remote sensing image data after coordinate matching to obtain the first remote sensing features, and perform feature enhancement processing and feature optimization processing on the first remote sensing features to obtain the second remote sensing features;

[0077] S103. Perform feature encoding on the point cloud features and the second remote sensing features by an encoder to obtain fused features;

[0078] S104. Perform feature decoding on the fused features by a decoder to output a decoding result;

[0079] S105. Input the decoding result into a fully connected layer embedded with a KAN network layer, output the probability of the ground object type corresponding to each pixel unit, determine the ground object type corresponding to the pixel unit, and generate the semantic segmentation result distribution map of the research area.

[0080] Specifically, in S101, after determining the point cloud data and the remote sensing image data, the point cloud data and the remote sensing image data can be aligned. Specifically, the remote sensing area within a preset range centered on the center point coordinates can be determined according to the center point coordinates corresponding to each point cloud area in the point cloud data, and the remote sensing image data within the corresponding remote sensing area can be extracted;

[0081] Among them, the range of the remote sensing area covers the corresponding point cloud area.

[0082] In some embodiments, the input point cloud is where N is the number of points, d is the d-dimensional coordinates of the points, and the input NDVI data is N is the number of points, H and W are the height and width of the image respectively. Different from the RGB image, the number of channels of the NDVI data is 1.

[0083] For example Figure 3 as shown, the point cloud area is like Figure 3 the red rectangular box shown on the left. The center point can determine the center point coordinates corresponding to each point cloud area. The point cloud area can correspond to all the point cloud data within a radius of 5 m around the center point. The remote sensing area with a size of 20 m × 20 m around the center point can be selected according to the center point of the point cloud area, such as Figure 3 the rectangular box shown on the right. Among them, the gray value of the color block in the remote sensing area corresponds to the NDVI value. The black part indicates that there is no vegetation in the corresponding pixel, and the gray part indicates that there is vegetation in the corresponding pixel. The embodiments of the present application do not limit this.

[0084] During the training process of the encoder model, the same number of center points as the model batch size m can be randomly selected from the point cloud data. The number of center points selected in each round of training is equal to the batch size. After n rounds of epoch training, the model can select m×n independent regions for learning. The coordinates of the center points can also be used to perform corresponding raster cropping on the NDVI data. To ensure that the cropped NDVI regions can completely cover the point cloud data, 4 raster regions around the center points can be selected for cropping, with a cropping size of 20 meters × 20 meters, so as to ensure that the cropped range can completely contain all relevant point clouds. Specifically, the point cloud region in each training batch size batch corresponds to an NDVI cropping region. This method of pairing regions of multi-modal data can ensure that the spatial features of the point cloud and NDVI data can be combined consistently and effectively during the model training process.

[0085] Specifically, in S102, in the point cloud feature extraction branch, the point cloud feature extraction branch can be established by using the PointNet++ network as the basic network architecture, and the point cloud features at each point cloud position can be extracted.

[0086] It should be noted that the remote sensing image data in this application is NDVI image. The image feature extraction branch of S102 extracts the Normalized Difference Vegetation Index (NDVI). NDVI is an index widely used in the fields of remote sensing and agriculture, and is usually used to evaluate the health status and coverage of vegetation. NDVI is calculated by using the data of the near-infrared (NIR) band and the red light band in the multi-spectral images obtained by remote sensing platforms such as satellites or drones.

[0087] In some embodiments, as Figure 4 shown, in the image feature extraction branch of S102, the remote sensing image data in each remote sensing region can be respectively subjected to feature extraction through multiple convolutional layers, and the first remote sensing features corresponding to each remote sensing region can be extracted;

[0088] The first remote sensing features are input into the SA module. The channel attention weights are determined through the channel attention module in the SA module, the spatial attention weights are determined through the spatial attention module in the SA module, and feature enhancement is performed by combining the channel attention weights and the spatial attention weights through a feature rearrangement operation, and the enhanced features are output;

[0089] The enhanced features are input into the KAN network, and the enhanced features are numerically corrected through the KAN network, and the second remote sensing features corresponding to each remote sensing region are output.

[0090] Exemplarily, a convolutional neural network (CNN) can be introduced as the basic framework, and 4 convolutional layers are used for feature extraction. The Shuffle Attention (SA) module is incorporated to improve the controllability and stability of NDVI features in multimodal fusion, and the SA module is introduced in the fourth convolutional layer of the CNN network.

[0091] It can be understood that the spectral boundary features of NDVI data are enhanced by the attention mechanism SA module in the last layer, improving the accuracy and robustness of the semantic segmentation results. The KAN layer after the SA module enables the model to retain more spatial and context information at the decision-making level through univariate function decomposition. Compared with traditional MLP, this method helps to maintain the high-level spatial features extracted from the convolutional layer, avoiding information loss during the fully connected process, in order to effectively improve the model performance while maintaining the running efficiency of the entire network.

[0092] Exemplarily, as Figure 4 shown, Figure 4 FIG. is a schematic flow diagram of image feature extraction provided by this application. By adding the SA module to the last layer of the 4-layer CNN convolutional network, it can be ensured that the model can comprehensively consider all key feature information and optimize the decision-making process. At the same time, the spectral boundary features of NDVI data are enhanced by the SA attention mechanism in the last layer, thereby improving the accuracy and robustness of the segmentation results. This deployment strategy can maximize the utilization efficiency of computing resources, avoid complex attention calculations at each level, and thus balance performance and computational cost. The SA module adjusts the feature extraction range of the CNN convolution, enhancing the difference of features and improving the model's discrimination ability for different regions and features.

[0093] It can be understood that the Shuffle Attention (SA) module is a self-attention mechanism module designed to enhance feature representation through feature rearrangement and attention mechanism, combining the advantages of channel attention and spatial attention, and introducing feature rearrangement operations to better capture global information and local details.

[0094] In some embodiments, the number of channels c of the input features of the SA module is 512, which is divided into 64 groups (G = 64), so the number of channels in each group is 8. The SA module improves the feature expression ability through grouping and shuffling. To ensure effective shuffling of features within each group, G is an integer divisor of the number of channels. Therefore, choosing 64 as a reasonable divisor of 512 ensures uniform channel division. Two training batches of NDVI features can be randomly selected. Further, the range of the extracted NDVI feature values is transformed from the original -0.05 to 0.2 to 0 to 0.4 through the KAN layer. The KAN layer corrects the NDVI feature values, such as eliminating negative values and expanding positive values, so as to improve the controllability and stability of NDVI features in multimodal fusion, avoid adverse weight conflicts or information loss caused by negative features in the feature addition of modal fusion, and make the multimodal fusion process smoother and more efficient.

[0095] Exemplarily, as Figure 4 shown, before the SA module, convolutional features are obtained through four 1DConv layers in sequence, and the convolutional features are input into the SA module. After the SA module, there is also a maxpool layer. Figure 4 where k is the number of remote sensing image grids. The size of the feature map is reduced through the maxpool layer, and then the corrected feature map is output through the KAN network. n is the number of features of the output layer during the encoding process.

[0096] In some embodiments, as Figure 5 shown, in S103, feature encoding is performed based on the point cloud features and the second remote sensing features through an encoder to obtain fused features, specifically including:

[0097] The encoder includes multiple encoding layers, and each encoding layer includes a point cloud feature extraction layer and L N convolutional layers;

[0098] Feature encoding is performed based on the point cloud features and the second remote sensing features through multiple encoding layers in sequence to obtain fused features, including:

[0099] The point cloud feature extraction layer in the l-th encoding layer groups the point cloud data and extracts the point cloud features corresponding to each group of point clouds. The point cloud features of each l-th layer obtained are:

[0100]

[0101] Exemplarily, Figure 5 in which K, K1, K2, and K3 respectively represent the number of groups of point clouds for each point cloud grouping, which is equivalent to downsampling the point cloud data. One group of point clouds after each grouping is used as a point cloud after downsampling;

[0102] After obtaining the point cloud features of each l layers, through L N convolutional layers perform the steps of feature extraction on the remotely sensed image data after coordinate matching, and the steps of feature enhancement processing and feature optimization processing to obtain the second remotely sensed feature, and extract the second remotely sensed feature matching each group of point cloud coordinates, including:

[0103]

[0104] Perform feature encoding on the point cloud feature and the second remotely sensed feature to obtain the fusion feature of the l-th encoding layer, and apply the formula:

[0105]

[0106] where N (l) is the number of point clouds in the l-th encoding layer, d (l) is the coordinate dimension of the point clouds in the l-th encoding layer, is the dimension of the point cloud feature in the l-th encoding layer, H is the height of the feature map corresponding to the second remotely sensed feature, W is the width of the feature map corresponding to the second remotely sensed feature, is the dimension of the remotely sensed feature output by the l N -th convolutional layer, l N is the ordinal number of the convolutional layer, L N is the total number of convolutional layers, enc is the point cloud feature of the l-th encoding layer, is the second remotely sensed feature, and the number of point cloud features in each layer is the same as the number of the corresponding second remotely sensed feature, is the dimension of the second remotely sensed feature, is the fusion feature output by the l-th encoding layer.

[0107] It can be understood that the model fuses NDVI features at different scales, which enhances the expression ability of each layer of features, not only improves the model's understanding ability of large-scale terrain structures, but also makes it perform excellently in fine-grained vegetation segmentation. At the same time, the way of adding features retains the spatial geometric information of the point cloud and the vegetation information of NDVI without increasing the complexity of the model, and enhances the overall expression ability of the model.

[0108] In some embodiments, as Figure 6 shown, perform feature decoding on the fusion feature through a decoder, and output a decoding result. Among them, the decoder includes multiple decoding layers. Specifically, the fusion feature can be input into the decoder, and feature decoding is respectively performed on the fusion feature through multiple decoding layers to output a decoding result, including:

[0109] The input features are upsampled by the l-th decoding layer in the decoder to output the upsampled point cloud feature map;

[0110] The upsampled point cloud feature map is concatenated with the fused features output by the (L - l + 1)-th encoding layer to output the concatenated features;

[0111] The concatenated features are subjected to feature enhancement processing through the CBAM module to output the decoding result corresponding to the decoding layer;

[0112] The decoding result is input into the (l + 1)-th decoding layer to perform the above steps of upsampling the input features and subsequent steps until the L-th decoding layer, and the final decoding result is output;

[0113] Wherein, the number of layers of the decoding layer and the total number of layers of the encoding layer are both L.

[0114] Specifically, upsampling is performed based on the fused features output by the l-th encoding layer. The number of input points for upsampling in the l-th layer is N (l) , the coordinate dimension of the points in the l-th layer is d (l) , and the output feature dimension of PointNet++ for the points in the l-th layer is The features of the input point cloud in the l-th layer are:

[0115]

[0116] After upsampling, the number of output points changes from N (l) to N (l+1) , while the feature dimension is to The features after upsampling by PointNet++ are concatenated with the output features of the corresponding Encoder layer . L is the total number of layers in the encoding process, and the number of layers in the encoding process and the decoding process is the same. The formula is applied:

[0117]

[0118] After introducing the concatenated features into the CBAM module, the number of features is enhanced through the attention mechanism, and the original number of features remains unchanged. The feature dimension is So the output feature dimension is

[0119]

[0120] Specifically, in each layer of the network during the decoding process, the input features are first upsampled, then concatenated with the output features of the corresponding encoding layer, and then the feature representation is enhanced through the CBAM module. The output features of each layer are passed to the next layer for further processing. After being processed by the CBAM module in each layer, enhanced features are obtained, and these features continue to be passed to the next layer until the last layer of the decoding process.

[0121] In an optional example, the point cloud branch compared the original PointNet++ network, the KAN and PointNet++ network, the CBAM and PointNet++ network, and the complete network. The effectiveness of the module was verified through three main evaluation metrics: the mean intersection over union (IoU) of each point cloud label, the model accuracy (Accuracy), and the model mean loss. The intersection over union (IoU) of the point cloud label evaluates the overlapping degree between the model prediction and the true label. The model accuracy (Accuracy) represents the proportion of points that meet the correct standard in the model. The mean loss reflects the overall deviation of the model and is used to evaluate the robustness and accuracy of the model.

[0122] As shown in Table 1, both the mean intersection over union and the accuracy in the training and testing after adding the CBAM attention mechanism model have been significantly improved. In the point cloud feature extraction branch, compared with the basic PointNet++, after adding the KAN network and the CBAM module, the overall intersection over union of the model has increased by 2.7%, the accuracy has increased by 3.2%, and the mean loss has decreased by 0.033.

[0123] Table 1 Comparison of the effectiveness of the attention mechanism in the point cloud feature extraction branch

[0124]

[0125] In some embodiments, in S105, according to the representation of KAN to define the Kolmogorov - Arnold theorem, any multivariable function can be represented as a combination of a series of univariate functions. The process of inputting the decoding result into the fully - connected layer of the embedded KAN network layer and outputting the probability of the ground object type corresponding to each pixel unit includes:

[0126] During the process of the fully - connected layer of the embedded KAN network layer processing the input decoding result, as Figure 7 shown, from the input layer of the MLP to the processing through multiple hidden layers, and finally the probability of the ground object type corresponding to each pixel unit is output through the KAN layer. The specific application formula is:

[0127]

[0128] The fully connected layer embedded with the KAN network layer outputs the probability of the corresponding ground object type based on the decoding result of each input, determines the corresponding pixel unit according to the position coordinates of the point cloud features and remote sensing features corresponding to the decoding result, and obtains the probability of the corresponding ground object type for each pixel unit in the research area;

[0129] where x is the decoding result of the L-th decoding layer, W l is the weight matrix, b l is the bias vector, h l is the feature output of the fully connected layer of the embedded KAN network layer, L is the total number of layers of the fully connected layer of the embedded KAN network layer, the KAN network layer is located in the L-th decoding layer, ReLU is the activation function, and ф is a learnable univariate non-linear function.

[0130] It should be noted that the decoding layer as a whole can be regarded as a multi-layer perceptron structure, and a KAN network layer is embedded in the last layer of this multi-layer perceptron structure.

[0131] It can be understood that in the KAN network ablation experiment, first, the model accuracies of four structures, namely MLPs, MLPs and KAN, CNN, and CNN and KAN, are compared in the two-dimensional image dataset minst. Figure 8 It is a comparison chart of the decision optimization accuracy of the fully connected layer of the KAN network layer. When the KAN network is added to the fully connected layer of the model decision stage, the model progress accuracy is improved by 0.81% in mlp, and the accuracy is improved by an average of 0.12% in CNN convolution (the CNN network itself has relatively high accuracy when processing image data).

[0132] At the same time, in order to improve the operation efficiency of the KAN network, the original KAN includes piecewise polynomial scaling for each activation function. In the embodiments of the present application, after disabling the piecewise polynomial scaling, the efficiency comparisons of two network structures, namely MLPs and KAN and CNN and KAN, are respectively carried out. Figure 9 It is a comparison chart of the decision optimization efficiency of the fully connected layer of the optimized KAN network layer.

[0133] In an optional embodiment, after disabling the piecewise polynomial scaling, the model is more efficient (the efficiency is improved by 3.8%), and the running time of the model is shorter. This ensures that the network can utilize the efficient feature extraction ability of the traditional MLP in the early stage when processing multi-modal data, and provide more optimized and accurate outputs through the KAN layer in the final decision stage of the network fully connected. Finally, without sacrificing the complex feature expressions learned by the deep network in the early stage, the calculation and parameter efficiency of the output layer are optimized.

[0134] In some embodiments, the point cloud data of a natural slope at a hydropower station in the upper reaches of the Yunnan section of the Lancang River Basin can be selected. The coordinate system is the CGCS2000 geodetic coordinate system, the data collection time is July 2022, and the training dataset and test dataset are selected. There are a total of 4 types of label settings. The first type is single vegetation, the second type is vegetation group, the third type is artificial objects (including houses, vehicles, fences, etc.), and the fourth type is the ground surface. At the same time, in order to further optimize the results generated by the trained model and reduce the number of false positives, probabilistic analysis and optimization can be performed on the annotation results output by the model. The model introduces the lower bound of the 95% confidence interval as the threshold for result optimization. This strategy helps to filter out the results that the model is not very sure about the classification. That is, only when the prediction probability of a certain type of sample is greater than the threshold of that type, it will be labeled as the specified class, and the result less than the threshold will be labeled white as the unlabeled class.

[0135] By downloading the 10-meter resolution image data containing the target area from the official data platform of Sentinel-2 (Copernicus Open Access Hub), which is the same as the point cloud data. After atmospheric correction of the image, the red light (Band 4, wavelength 665nm) and near-infrared light (Band 8, wavelength 842nm) band data with a resolution of 10 meters in the image are extracted, and the NDVI value is calculated for each pixel. The calculation results are saved as a GeoTIFF format NDVI image with a spatial resolution of 10 meters, and the coordinate system is converted to CGCS2000.

[0136] Optionally, for the basic hardware configuration of model training, the GPU can be selected as NVIDIA RTX4090 (24G), and the CPU can be AMD PRO 5965WX with 24 cores and 48 threads. 32 training epochs can be set, the batch size is 16, the initial learning rate is set to 0.001, and a learning rate decay strategy with a decay coefficient of 0.0004 is used. The Adam optimizer is selected to improve the convergence speed and stability of the model. Figure 10 It is a schematic diagram of the natural terrain segmentation result of slope natural ground surface point cloud semantic segmentation based on multi-modal data provided by an embodiment of the present application.

[0137] The following is the device embodiment of the present application, which can be used to execute the method embodiment of the present application. For the details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.

[0138] Next, please refer to Figure 11, which is a schematic structural diagram of a slope natural ground surface point cloud semantic segmentation device provided for an exemplary embodiment of the present application. This device can be implemented as all or part of a terminal through software, hardware, or a combination of both, and can also be integrated as an independent module on a server. A slope natural ground surface point cloud semantic segmentation device in an embodiment of the present application can be applied to a terminal or the cloud. The device 1100 includes a data acquisition module 1101, a feature extraction module 1102, a feature fusion module 1103, a feature decoding module 1104, and a result output module 1105, where:

[0139] The data acquisition module 1101 is configured to acquire point cloud data and remote sensing image data of a research area, and perform coordinate matching on the point cloud data and the remote sensing image data;

[0140] The feature extraction module 1102 is configured to extract features based on the point cloud data after coordinate matching to obtain point cloud features; the feature extraction module is further configured to extract features based on the remote sensing image data after coordinate matching, extract a first remote sensing feature, and perform feature enhancement processing and feature optimization processing on the first remote sensing feature to obtain a second remote sensing feature;

[0141] The feature fusion module 1103 is configured to perform feature encoding on the point cloud features and the second remote sensing feature through an encoder to obtain a fusion feature;

[0142] The feature decoding module 1104 is configured to perform feature decoding on the fusion feature through a decoder and output a decoding result;

[0143] The result output module 1105 is configured to input the decoding result into a fully connected layer embedded with a KAN network layer, output the probability of the ground object type corresponding to each pixel unit, determine the ground object type corresponding to the pixel unit, and generate a semantic segmentation result distribution map of the research area.

[0144] In some embodiments, the feature extraction module 1102 is further configured to extract features based on the remote sensing image data after coordinate matching, extract a first remote sensing feature, and perform feature enhancement processing and feature optimization processing on the first remote sensing feature to obtain a second remote sensing feature, including:

[0145] Extract first remote sensing features corresponding to each remote sensing area by respectively extracting features of the remote sensing image data in each remote sensing area through multiple convolutional layers in the feature extraction module;

[0146] Input the first remote sensing feature into the SA module, determine the channel attention weight through the channel attention module in the SA module of the feature extraction module, determine the spatial attention weight through the spatial attention module in the SA module, and perform feature enhancement by combining the channel attention weight and the spatial attention weight through a feature rearrangement operation to output the enhanced feature;

[0147] Input the enhanced feature into the KAN network, and perform numerical correction on the enhanced feature through the KAN network of the feature extraction module to output the second remote sensing feature corresponding to each remote sensing area.

[0148] It should be noted that when the device 1100 provided in the above embodiment executes a method for semantic segmentation of natural ground point clouds of slopes based on multi-modal data, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the embodiment of the method for semantic segmentation of natural ground point clouds of slopes based on multi-modal data belong to the same concept, and the implementation process is detailed in the method embodiment, which will not be elaborated here.

[0149] The embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method in any of the above embodiments.

[0150] Please refer to Figure 12 , which is a structural block diagram of an electronic device provided by the embodiment of the present application.

[0151] As Figure 12 shown, the electronic device 1200 includes a processor 1201 and a memory 1202.

[0152] In the embodiment of the present application, the processor 1201 is the control center of the computer system, which can be the processor of a physical machine or the processor of a virtual machine. The processor 1201 can include one or more processing cores, such as a 4-core processor or an 8-core processor. The processor 1201 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array).

[0153] The processor 1201 may also include a main processor and a co-processor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the co-processor is a low-power processor for processing data in the standby state.

[0154] The memory 1202 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1202 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments of the present application, the non-transitory computer-readable storage media in the memory 1202 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1201 to implement the method in the embodiments of the present application.

[0155] In some embodiments, the electronic device 1200 further includes: a peripheral device interface 1203 and at least one peripheral device 1204. The processor 1201, the memory 1202, and the peripheral device interface 1203 may be connected by a bus or signal lines. Each peripheral device 1204 may be connected to the peripheral device interface 1203 through a bus, signal lines, or a circuit board. Specifically, the peripheral device 1204 includes: a display screen, a camera, and an audio circuit. The peripheral device interface 1203 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1201 and the memory 1202.

[0156] In some embodiments of the present application, the processor 1201, the memory 1202, and the peripheral device interface 1203 are integrated on the same chip or circuit board; in some other embodiments of the present application, any one or two of the processor 1201, the memory 1202, and the peripheral device interface 1203 may be implemented on a separate chip or circuit board. The embodiments of the present application do not make specific limitations on this.

[0157] The block diagram of the electronic device structure shown in the embodiments of the present application does not constitute a limitation on the electronic device 1200. The electronic device 1200 may include more or fewer components than shown in the figure, or combine some components, or adopt a different component layout.

[0158] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method in any of the foregoing embodiments. Among them, the computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic cards or optical cards, nano-systems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.

[0159] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solutions, or the part that contributes to the related technologies, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical disks, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.

Claims

1. A method for semantic segmentation of natural surface point clouds of slopes based on multi-modal data, characterized in that Including: Obtain the point cloud data and remote sensing image data of the research area, and perform coordinate matching on the point cloud data and the remote sensing image data; Extract features based on the point cloud data after coordinate matching to obtain point cloud features; Extract features based on the remote sensing image data after coordinate matching, extract the first remote sensing feature, and perform feature enhancement processing and feature optimization processing on the first remote sensing feature to obtain the second remote sensing feature; Perform feature encoding on the point cloud features and the second remote sensing feature through an encoder to obtain a fused feature; Perform feature decoding on the fused feature through a decoder and output a decoding result; Input the decoding result into a fully connected layer embedded with a KAN network layer, output the probability of the ground object type corresponding to each pixel unit, determine the ground object type corresponding to the pixel unit, and generate a semantic segmentation result distribution map of the research area.

2. A method for semantic segmentation of natural ground point clouds of slopes based on multi-modal data according to claim 1, characterized in that, The performing coordinate matching on the point cloud data and the remote sensing image data includes: Determine a remote sensing area within a preset range centered on the center point coordinates according to the center point coordinates corresponding to each point cloud area in the point cloud data, and extract the remote sensing image data within the corresponding remote sensing area; Wherein, the range of the remote sensing area covers the corresponding point cloud area.

3. A method for semantic segmentation of natural surface point clouds of slopes based on multi-modal data according to claim 2, characterized in that, The extracting features based on the remote sensing image data after coordinate matching, extracting the first remote sensing feature, and performing feature enhancement processing and feature optimization processing on the first remote sensing feature to obtain the second remote sensing feature includes: Extract features of the remote sensing image data within each remote sensing area through multiple convolutional layers respectively to obtain the first remote sensing feature corresponding to each remote sensing area; Input the first remote sensing feature into an SA module, determine channel attention weights through the channel attention module in the SA module, determine spatial attention weights through the spatial attention module in the SA module, and perform feature enhancement by combining the channel attention weights and the spatial attention weights through a feature rearrangement operation, and output the enhanced feature; Input the enhanced feature into a KAN network, and perform numerical correction on the enhanced feature through the KAN network to output the second remote sensing feature corresponding to each remote sensing area.

4. A method for semantic segmentation of natural surface point clouds of slopes based on multi-modal data according to claim 3, characterized in that, The encoder includes a plurality of encoding layers, and each of the encoding layers includes a point cloud feature extraction layer and L N convolutional layers; performing the step of feature extraction on the point cloud data after coordinate matching through the plurality of encoding layers in the encoder, specifically including: Group the point cloud data through the point cloud feature extraction layer in the l-th encoding layer, and extract the point cloud features corresponding to each group of point clouds. The point cloud features of each l-th layer are: After obtaining the point cloud features of each l layers, through L N performing the step of feature extraction on the remotely sensed image data after coordinate matching, and the steps of feature enhancement processing and feature optimization processing to obtain the second remotely sensed feature, and extracting the second remotely sensed feature matching each group of point cloud coordinates, including: Perform feature encoding on the point cloud features and the second remote sensing feature to obtain the fused feature of the l-th encoding layer, and apply the formula: Among them, N (l) is the number of point clouds in the l-th encoding layer, d (l) is the coordinate dimension of the point clouds in the l-th encoding layer, is the dimension of the point cloud features in the l-th encoding layer, H is the height of the feature map corresponding to the second remote sensing feature, and W is the width of the feature map corresponding to the second remote sensing feature. is the dimension of the remote sensing feature output by the l-th N convolution layer, l N is the ordinal number of the convolution layer, and L N is the total number of convolution layers. The is the point cloud feature of the l-th encoding layer, is the second remote sensing feature. The number of point cloud features in each layer is the same as the number of the corresponding second remote sensing features. is the dimension of the second remote sensing feature, is the fused feature output by the l-th encoding layer.

5. A method for semantic segmentation of natural ground point clouds of slopes based on multi-modal data according to claim 4, characterized in that, The decoder includes multiple decoding layers; Input the fused feature into the decoder, and perform feature decoding on the fused feature through multiple decoding layers respectively to output a decoding result, including: Perform upsampling processing on the input feature through the l-th decoding layer in the decoder to output an upsampled point cloud feature map; Stitch the upsampled point cloud feature map with the fused feature output by the (L - l + 1)-th encoding layer to output a stitched feature; Perform feature enhancement processing on the stitched feature through a CBAM module to output the decoding result corresponding to the decoding layer; Input the decoding result into the (l + 1)-th decoding layer, and perform the steps of upsampling the input features and subsequent steps until the L-th decoding layer, and output the final decoding result; Among them, the number of decoding layers and the total number of encoding layers are both L.

6. A method for semantic segmentation of natural surface point clouds of slopes based on multi-modal data according to claim 5, characterized in that, The step of inputting the decoding result into the fully connected layer of the embedded KAN network layer and outputting the probability of the ground object type corresponding to each pixel unit includes: The processing process of the fully connected layer of the embedded KAN network layer for the input decoding result applies the formula: The fully connected layer of the embedded KAN network layer outputs the probability of the corresponding ground object type based on each input decoding result, determines the corresponding pixel unit according to the position coordinates of the point cloud feature and the remote sensing feature corresponding to the decoding result, and obtains the probability of the ground object type corresponding to each pixel unit in the study area; where x is the decoding result of the L-th decoding layer, W l is the weight matrix, b l is the bias vector, h l is the feature output of the fully connected layer of the embedded KAN network layer, L is the total number of layers of the fully connected layer of the embedded KAN network layer, the KAN network layer is located in the L-th decoding layer, ReLU is the activation function, and ф is a learnable univariate non-linear function.

7. A slope natural ground point cloud semantic segmentation device based on multimodal data, characterized in that Including: A data acquisition module, configured to acquire point cloud data and remote sensing image data of a study area, and perform coordinate matching on the point cloud data and the remote sensing image data; A feature extraction module, configured to extract point cloud features based on the coordinate-matched point cloud data; the feature extraction module is further configured to extract first remote sensing features based on the coordinate-matched remote sensing image data, and perform feature enhancement processing and feature optimization processing on the first remote sensing features to obtain second remote sensing features; A feature fusion module, configured to perform feature encoding on the point cloud features and the second remote sensing features through an encoder to obtain fused features; A feature decoding module, configured to perform feature decoding on the fused features through a decoder and output a decoding result; A result output module, configured to input the decoding result into the fully connected layer of the embedded KAN network layer, output the probability of the ground object type corresponding to each pixel unit, determine the ground object type corresponding to the pixel unit, and generate a semantic segmentation result distribution map of the study area.

8. A device for semantic segmentation of natural ground point clouds of slopes based on multi-modal data according to claim 7, characterized in that, The feature extraction module is further configured to extract first remote sensing features based on the coordinate-matched remote sensing image data, and perform feature enhancement processing and feature optimization processing on the first remote sensing features to obtain second remote sensing features, including: Performing feature extraction on the remote sensing image data in each remote sensing area through a plurality of convolutional layers in the feature extraction module to extract first remote sensing features corresponding to each remote sensing area; Inputting the first remote sensing feature into the SA module, determining channel attention weights through the channel attention module in the SA module of the feature extraction module, determining spatial attention weights through the spatial attention module in the SA module, and performing feature enhancement through feature rearrangement operation in combination with the channel attention weights and the spatial attention weights, and outputting the enhanced feature; Inputting the enhanced feature into the KAN network, and performing numerical correction on the enhanced feature through the KAN network of the feature extraction module to output second remote sensing features corresponding to each remote sensing area.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 6.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Cited By

  • Multi-modal slope state monitoring method based on infrasound and Beidou signals

    CN121114239A