Radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling

By adopting the technology of fusion interactive learning and dynamic sampling in the LiDAR point cloud semantic segmentation method, combining multi-scale context information fusion and self-attention module, the existing methods have solved the shortcomings in accuracy and computing efficiency balance and information utilization, and achieved efficient segmentation and generalization capabilities in complex scenarios.

CN119942124APending Publication Date: 2025-05-06CHONGQING UNIV OF TECH

Patent Information

Application Number
CN202510102243.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When processing three-dimensional point cloud data, the existing LiDAR point cloud semantic segmentation method is difficult to achieve a balance of accuracy and computing efficiency, and fails to make full use of coordinate values, distance and intensity information, resulting in insufficient segmentation accuracy and generalization capabilities in complex urban environments and dynamic scenarios.

Method used

The radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling is adopted. The three-dimensional point cloud data is converted into two-dimensional distance images through spherical projection, and the coordinate value, distance and intensity information is retained, and feature information is fused and enhanced through the multi-scale context information fusion module and the self-attention module, and efficient decoding is carried out in combination with the dynamic sampling module.

Benefits of technology

It achieves an effective balance of accuracy and computing efficiency, makes full use of important information in LiDAR point cloud data, improves the segmentation accuracy and generalization capabilities of the model in complex urban environments and dynamic scenarios, and is suitable for real-time decision-making applications such as autonomous driving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942124A_ABST
    Figure CN119942124A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of laser radar three-dimensional point cloud semantic segmentation, and particularly relates to a radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling. According to the method, effective balance of precision and calculation efficiency can be realized, coordinate values, distance and intensity information in laser radar three-dimensional point cloud data are fully utilized, and the segmentation precision and generalization ability of the model in a complex urban environment and a dynamic scene are improved. Through multi-channel information fusion and enhancement, dynamic sampling and efficient decoding and a balance strategy of precision and calculation efficiency, the method not only improves the segmentation precision and calculation efficiency of the model, but also enhances the understanding ability and generalization ability of the model to complex scenes, and provides powerful technical support for real-time decision application such as automatic driving and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of semantic segmentation of three-dimensional point clouds using laser radar, and in particular relates to a semantic segmentation method of radar point clouds based on fusion of interactive learning and dynamic sampling. Background Art

[0002] In recent years, LiDAR (Light Detection and Ranging) technology has made significant progress in capturing three-dimensional information of real-world objects. With its high precision, high resolution, and ability to work around the clock, LiDAR technology has shown great application potential in real-time decision-making fields such as autonomous driving, robot navigation, and environmental perception. LiDAR 3D point cloud data provides rich spatial information, which enables effective scene parsing and label assignment for different environmental elements. Semantic segmentation of LiDAR 3D point clouds assigns each point in the 3D point cloud data collected by the LiDAR to a predefined semantic category. These semantic categories usually represent different objects, surfaces, or structures in the environment, such as "car", "road", "building", or "tree". Therefore, developing advanced and efficient LiDAR point cloud semantic segmentation models has important academic and application value.

[0003] Traditional convolutional neural networks (CNNs) have achieved remarkable results in image processing, natural language processing and other fields. However, when processing LiDAR point cloud data, traditional CNNs are difficult to apply directly due to its disorder and irregular structure. In order to overcome this challenge, the research focus in related fields has shifted to exploring new LiDAR point cloud semantic segmentation network models. At present, LiDAR point cloud semantic segmentation methods based on deep learning are mainly divided into four categories: original point-based methods, voxel-based methods, range image-based methods and multimodal fusion-based methods.

[0004] Although point-based and voxel-based methods perform well in terms of accuracy, they usually have high model complexity and computational cost. Point-based methods face challenges in computational efficiency when processing large-scale point cloud data, while voxel-based methods may lead to loss of spatial resolution and quantization errors. Although methods based on multimodal fusion have advantages in scene parsing, the data synchronization and alignment between different sensors are still difficult problems faced by current technologies. Given the strict requirements of applications such as autonomous driving on real-time performance and inference speed, distance image-based methods have attracted much attention due to their processing efficiency and relatively simple model structure. Distance image-based methods reduce computational complexity by projecting three-dimensional point cloud data onto a two-dimensional plane and processing it using two-dimensional CNN. However, existing distance image-based methods still have some problems. For example, although models such as KPRNet and Lite-HDSeg have excellent performance, they have high model complexity and large number of parameters, resulting in slow inference speed, which limits their application in real-time autonomous driving. Models such as FIDNet and CENet use a full interpolation decoding module that only contains bilinear interpolation, which affects the effective capture of contextual semantics and leads to insufficient segmentation accuracy. At the same time, this type of method lacks effective preprocessing of the two-dimensional distance image after the original three-dimensional point cloud is projected, such as the fusion and enhancement of coordinate values, distance and intensity information, which further limits the improvement of segmentation performance. In recent years, TransRVNet proposed a multi-residual channel interaction attention network model. This method learns individual feature spaces for each independent channel through channel interaction, thereby refining the two-dimensional distance image after spherical coordinate projection. However, this method mainly focuses on the interaction between independent channels, ignoring the intrinsic connection and interaction between coordinate values, distance and intensity information, resulting in limited spatial perception of the model. This processing method that ignores the intrinsic information association means that the segmentation accuracy and generalization ability of the model in complex scenes still need to be improved.

[0005] The reasons for the existing technical problems are mainly due to the insufficient preprocessing and feature extraction of three-dimensional point cloud data, and the inefficient and inaccurate interaction mode of the model when processing this information. Coordinate values, distance and intensity information are important components of LiDAR point cloud data, and they contain rich spatial information and contextual semantics. However, existing methods often fail to fully integrate and utilize this information, resulting in the model being difficult to achieve an ideal balance between segmentation accuracy and computational efficiency. In addition, the interaction mode of the model when processing this information also lacks sufficient flexibility and robustness, making it difficult to adapt to complex and changing urban environments and dynamic scenes. The difficulty in solving the problem lies in how to achieve effective integration and utilization of this information while ensuring computational efficiency; how to design a model structure that can capture local detail features and grasp global contextual semantics; these are the technical problems that need to be solved in the current field of LiDAR point cloud semantic segmentation.

[0006] Therefore, how to achieve an effective balance between accuracy and computational efficiency, while making full use of the coordinate values, distance and intensity information in the LiDAR point cloud data to improve the segmentation accuracy and generalization ability of the model in complex urban environments and dynamic scenes has become a problem that needs to be solved urgently. Summary of the invention

[0007] In view of the above-mentioned shortcomings of the prior art, the present invention provides a radar point cloud semantic segmentation method based on the fusion of interactive learning and dynamic sampling, which can achieve an effective balance between accuracy and computational efficiency, while making full use of the coordinate values, distance and intensity information in the LiDAR point cloud data, and improve the segmentation accuracy and generalization ability of the model in complex urban environments and dynamic scenes.

[0008] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0009] The radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling includes the following steps:

[0010] S1, obtain laser radar three-dimensional point cloud data;

[0011] S2. Use a spherical projection method to project the three-dimensional point cloud data into a two-dimensional range image; each pixel in the two-dimensional range image contains five channels (x, y, z, d, r) information; where (x, y, z) is the coordinate value information, indicating the spatial position of the pixel; d is the distance information, indicating the distance between the pixel and the laser radar; r is the intensity information, indicating the reflection intensity of the object surface at the pixel;

[0012] S3, through the three branches, the coordinate value and depth information, the coordinate value and intensity information, the depth and intensity information are respectively fused through the multi-scale context information fusion module MCIF to obtain the corresponding preliminary fusion information; and the self-attention module SA is used to capture the intra-branch dependency and feature enhancement of the preliminary fusion information of the three branches to obtain the corresponding preliminary enhancement information;

[0013] S4, aggregate the preliminary enhanced information of the three branches to obtain comprehensive fusion information, feed the comprehensive fusion information into the SA module, capture the global dependency relationship between branches and enhance the features to obtain comprehensive enhanced information;

[0014] S5, feeding the comprehensive enhancement information to the input module, the input module extracts the initial feature map with the same resolution as the original two-dimensional range image, and generates a multi-layer feature map with successively decreasing resolutions in the encoding stage; in the decoding stage, the multi-layer feature map is upsampled layer by layer through the dynamic sampling module, and finally the feature map is restored to the same resolution as the initial feature map; and the final 2D prediction map is obtained;

[0015] S6. Process the 2D prediction image to generate a 3D point cloud prediction.

[0016] Compared with the prior art, the present invention has the following beneficial effects:

[0017] 1. Multi-channel information fusion and enhancement. This method converts three-dimensional point cloud data into two-dimensional range images through spherical projection, and retains the information of five channels: coordinate values ​​(x, y, z), distance (d) and intensity (r). Existing LiDAR-based point cloud semantic segmentation methods lack effective preprocessing of the projected range images and ignore the intrinsic connection between channels of different modalities. Compared with existing methods that only use single or a small amount of channel information, multi-channel fusion in this method can more comprehensively capture the spatial features and contextual information of point cloud data. Furthermore, through the multi-scale contextual information fusion module MCIF and the self-attention module SA, the method effectively fuses coordinate values ​​with depth information, coordinate values ​​with intensity information, and depth with intensity information, and realizes feature enhancement and dependency capture within and between branches, thereby learning the potential relationship between different physical quantities in point cloud data. This enhances the model's ability to understand complex scenes.

[0018] By fusing and interactively extracting information from different modal channels in the range image, the potential relationship between different physical quantities in the point cloud data can be learned, which can reduce the risk of information compression and loss in the process of projecting the original three-dimensional point cloud into the two-dimensional range image.

[0019] 2. Dynamic sampling and efficient decoding. In the decoding stage, this method uses a dynamic sampling module to upsample the multi-layer feature maps layer by layer, rather than the traditional fixed upsampling method. Existing LiDAR-based point cloud semantic segmentation methods interpolate low-resolution features through fixed rules in the decoding stage, fail to fully consider the semantic information in the feature map, and limit the effective capture of contextual semantics. The dynamic sampling in this method can adaptively adjust the sampling strategy according to contextual information and task requirements, thereby improving the flexibility and efficiency of the decoding process. Compared with the existing technology that only uses simple decoding methods such as bilinear interpolation, the dynamic sampling module of this method can more finely restore the detailed information of the feature map and improve the resolution and accuracy of the prediction map.

[0020] 3. Balance between accuracy and computational efficiency. This method achieves an effective balance between computational efficiency and segmentation accuracy by integrating interactive learning and dynamic sampling techniques. On the one hand, multi-channel information fusion and feature enhancement improve the segmentation accuracy of the model; on the other hand, dynamic sampling and efficient decoding reduce the computational complexity of the model and meet the needs of real-time applications. Compared with existing methods with high model complexity and slow inference speed, this method significantly improves computational efficiency while ensuring accuracy, and is more suitable for real-time decision-making scenarios such as autonomous driving.

[0021] 4. Segmentation accuracy and generalization ability in complex scenes. This method makes full use of the coordinate values, distance and intensity information in the LiDAR point cloud data, and improves the model's ability to understand complex urban environments and dynamic scenes through multi-scale context information fusion and self-attention mechanism. Compared with existing methods that ignore the intrinsic connection between this information or only use part of the information, this method can better adapt to changes in different scenes and lighting conditions, and improve the generalization ability and robustness of the model.

[0022] In summary, this method can achieve an effective balance between accuracy and computational efficiency, while making full use of the coordinate values, distance and intensity information in the LiDAR point cloud data to improve the segmentation accuracy and generalization ability of the model in complex urban environments and dynamic scenes. Through multi-channel information fusion and enhancement, dynamic sampling and efficient decoding, and a balance strategy between accuracy and computational efficiency, this method not only improves the segmentation accuracy and computational efficiency of the model, but also enhances the model's ability to understand and generalize complex scenes, providing strong technical support for real-time decision-making applications such as autonomous driving.

[0023] Preferably, in S2, when the three-dimensional point cloud data is projected onto the two-dimensional range image using a spherical projection method, each three-dimensional point p=(x, y, z) having Cartesian coordinates is projected onto the two-dimensional range image by spherical mapping. is converted to the corresponding image coordinate system, the formula is:

[0024]

[0025] Where f = f u +f d represents the vertical field of view of the sensor; the depth d of a point is expressed as The size of the distance image obtained after projection is (H, W, 5), where each pixel contains 5 channels (x, y, z, d, r), and r is the intensity information of the point.

[0026] This setting can achieve 1. Efficient data dimensionality reduction. Through the spherical projection method, each point in the 3D point cloud data is effectively converted to a pixel in the 2D image coordinate system. This conversion not only reduces the dimensionality of the data, but also retains important information in the original point cloud, such as coordinate values, depth, and intensity. The reduced-dimensional data is easier to process and analyze, while reducing the consumption of computing resources and improving the efficiency of subsequent processing.

[0027] 2. It can retain key information. The projection formula takes into account the Cartesian coordinates (x, y, z), depth d, and intensity r of the point, which are crucial for the subsequent semantic segmentation task. The depth information d reflects the distance between the point and the sensor, which helps the model understand the spatial structure of the scene; the intensity information r reflects the reflective properties of the object surface, which helps to distinguish different objects and materials.

[0028] The projected 2D range image has a fixed resolution (H, W) and number of channels (5), which is conducive to the design and implementation of subsequent processing algorithms. For example, in the semantic segmentation task, a deep learning model such as a convolutional neural network can be used to process the 2D range image to achieve accurate segmentation of point cloud data.

[0029] Preferably, in S3, the multi-scale context information fusion module MCIF includes three basic blocks connected in sequence;

[0030] The first basic block includes a 1×1 convolution layer and multiple non-1×1 convolution layers with different kernels, which are used to extract input features. After the multiple non-1×1 convolution layers with different kernels, another 1×1 convolution layer is connected to fuse the multi-scale information of the convolution layers with non-1×1 kernels. After the output residuals of the two 1×1 convolution layers are connected, the input features of the second basic block are obtained.

[0031] The second basic block includes a 1×1 convolution layer and a non-1×1 convolution layer, which are used to extract input features. After the non-1×1 convolution layer, there is also a non-1×1 hole convolution layer. The outputs of the 1×1 convolution layer and the non-1×1 hole convolution layer of the second basic block are connected with the input feature residual of the first basic block to obtain the input features of the third basic block.

[0032] The third basic block has the same structure as the first basic block.

[0033] With this setting, 1. Through convolutional layers with different kernel sizes, MCIF can extract multi-scale information from input features. This information is essential for understanding complex scene structures, capturing objects of different scales, and distinguishing different object categories. The 1×1 convolutional layer is used to fuse this multi-scale information and generate more robust and discriminative feature representations.

[0034] 2. In the second basic block, the non-1×1 hole convolution layer can increase the receptive field of the convolution kernel while keeping the resolution of the feature map unchanged. This helps capture contextual information in a wider range and improves the model's ability to understand the global structure.

[0035] 3. The use of residual connections helps alleviate the gradient vanishing problem in deep neural networks, allowing the network to more easily learn useful feature representations. At the same time, it can also promote the flow of information, making it easier for information from the previous layer to be transferred to the next layer, thereby improving the performance of the entire network.

[0036] 4. Full utilization of multi-scale information. Through the sequential connection and residual connection of the three basic blocks, MCIF can fully utilize the multi-scale information in the input features and generate richer and more discriminative feature representations. This is crucial to improving the performance of tasks such as semantic segmentation.

[0037] In summary, the multi-scale context information fusion module MCIF can effectively improve the performance of the model by extracting and fusing multi-scale information, introducing hole convolution, using residual connection, and making full use of the multi-scale information in the input features, providing strong support for subsequent tasks such as semantic segmentation.

[0038] Preferably, the convolutional layers with different non-1×1 kernels of the first basic block are convolutional layers with 3×3, 5×5 and 7×7 kernels respectively.

[0039] With this setting, the first basic block can capture information of different scales in the input feature map by using convolution kernels of different sizes (3×3, 5×5, 7×7). Small convolution kernels (such as 3×3) can capture local detail features, while large convolution kernels (such as 5×5, 7×7) can capture contextual information of a larger range. This multi-scale feature extraction helps the model better understand complex scene structures and improve segmentation accuracy.

[0040] Convolution kernels of different sizes can learn different feature representations, which can be further integrated and utilized in subsequent processing. By combining features of multiple scales, the model can generate richer and more discriminative feature representations, which is crucial to improving the performance of the model.

[0041] Preferably, the convolution layer of the second basic block that is not 1×1 is a convolution layer with a 3×3 kernel; and the dilated convolution layer of the second basic block that is not 1×1 is a 3×3 dilated convolution layer.

[0042] With this setting, 1. The 3×3 convolution layer can capture local detail information in the input feature map. This local feature extraction is crucial for understanding details such as textures and edges in images. Compared with larger convolution kernels, the 3×3 convolution layer has fewer parameters, which helps reduce the complexity of the model and the risk of overfitting. At the same time, smaller convolution kernels also mean less computation, which helps improve the operating efficiency of the model. The 3×3 convolution layer can be flexibly combined with other convolution layers or operations (such as pooling, activation functions, etc.) to build a more complex network structure. This flexibility enables the model to be customized and optimized according to different task requirements.

[0043] 2. Dilated convolution (also known as expanded convolution) increases the receptive field by inserting holes in the convolution kernel without increasing the size of the convolution kernel. The 3×3 dilated convolution layer can capture a wider range of contextual information while maintaining computational efficiency. Dilated convolution can increase the receptive field without changing the size of the input feature map, which is crucial to maintaining the resolution and detail information of the feature map. This helps the model to better utilize this information in subsequent processing. When the dilated convolution layer is used in combination with convolution layers of other scales (such as 1×1 convolution layers or dilated convolution layers with different expansion rates), multi-scale feature fusion can be achieved. This fusion helps improve the model's ability to understand complex scenes and enhance the discriminative power of features. Compared with traditional downsampling and upsampling operations, dilated convolution can increase the receptive field without losing information. This helps the model maintain higher accuracy and stability in tasks such as segmentation and detection.

[0044] Preferably, in S5, in the encoding stage, 4 layers of feature maps are generated in sequence; the resolutions of the 4 layers of feature maps are 1, 1 / 2, 1 / 4, and 1 / 8 of the initial feature maps, respectively.

[0045] With this setting, the encoding stage generates four layers of feature maps of different resolutions in sequence, which can capture multi-scale information in the input image, optimize computational efficiency and memory usage, integrate contextual information, build feature pyramids, adapt to objects of different scales, and promote information flow and fusion. These effects work together to improve the performance of the model, enabling it to better understand and process complex image data.

[0046] Preferably, in S5, in the decoding stage, the dynamic sampling module is a dynamic upsampler Dysample-S+.

[0047] This setting uses the dynamic upsampler DySample-S+ as the dynamic sampling module in the decoding stage of S5, which can bring multiple effects such as high efficiency, light weight, strong adaptability, improved model performance, and easy integration and deployment. These effects work together in the decoding stage of the model, helping to improve the overall performance and meet the needs of practical applications.

[0048] Preferably, Dysample-S+ comprises a receiving module, an offset generating module, an upsampling module and an output module;

[0049] The receiving module is used to receive the feature map generated from the encoding stage; the offset generation module is used to perform preset operations on the received feature map to generate a new offset; the upsampling module is used to use the generated offset to adjust the original grid position to achieve dynamic upsampling of the feature map; the output module is used to output the feature map that has been restored to the same resolution as the initial feature map after dynamic upsampling, which is used for subsequent decoding and classification tasks.

[0050] With this setup, Dysample-S+ achieves efficient and high-quality feature map upsampling through the collaborative work of various modules. Compared with traditional upsampling methods, Dysample-S+ shows significant advantages in computational efficiency, feature map quality, and task adaptability. Due to its high efficiency, flexibility, and strong adaptability, Dysample-S+ has broad application prospects in image processing, computer vision, and other fields. It can be used in various intensive prediction tasks to improve the performance and accuracy of the model.

[0051] Preferably, when the offset generation module performs a preset operation on the received feature map, the preset operation includes pixel shuffling-linear projection and linear projection-pixel shuffling.

[0052] Such a setting,1. Flexibility. The two combinations of pixel shuffle-linear projection and linear projection-pixel shuffle provide different feature processing orders, allowing the offset generation module to be flexibly selected and adjusted according to different input feature maps and task requirements.

[0053] 2. Efficiency. Through reasonable combination operations, the offset generation module can efficiently generate high-quality offsets, providing strong support for the subsequent upsampling process. At the same time, these two combinations also avoid complex calculation processes and improve overall processing efficiency.

[0054] 3. Accuracy. The generated offsets can accurately guide the upsampling process, ensuring that the feature maps can be upsampled to higher resolutions as expected. This helps improve the accuracy and performance of subsequent decoding and classification tasks.

[0055] Preferably, when the offset generation module generates a new offset, the processing process is as shown in the following formula:

[0056]

[0057] in, Indicates the generated offset; Represents input features; and Indicates the preset two-pixel reorganization; linear indicates linear transformation.

[0058] Such an upsampling process helps to improve the upsampling accuracy and the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to make the purpose, technical solution and advantages of the invention more clear, the present invention will be further described in detail below with reference to the accompanying drawings, in which:

[0060] Figure 1 A flowchart of the method is shown in FIG.

[0061] Figure 2 Schematic diagram of the process of semantic segmentation of laser radar three-dimensional point cloud data in Example 1;

[0062] Figure 3 Schematic diagram of the structure of the multi-scale context information fusion module MCIF in Example 1;

[0063] Figure 4 This is a schematic diagram of the qualitative analysis of the SemanticKITTI validation set in Example 2;

[0064] Figure 5 This is a schematic diagram of the qualitative analysis of the SemanticKITTI validation set in Example 2. DETAILED DESCRIPTION

[0065] The following is a further detailed description through specific implementation methods:

[0066] Embodiment 1

[0067] like Figure 1 , Figure 2 As shown, this embodiment discloses a radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling. It should be noted that Figure 2 In this paper, each step in this method is embodied in the form of a module and an intelligent model is used. Figure 2 The model in can automatically implement this method.

[0068] This includes the following steps:

[0069] S1, obtain laser radar three-dimensional point cloud data;

[0070] S2. Use the spherical projection method to project the three-dimensional point cloud data into a two-dimensional range image; each pixel in the two-dimensional range image contains five channels of (x, y, z, d, r) information; wherein (x, y, z) is the coordinate value information, indicating the spatial position of the pixel; d is the distance information, indicating the distance between the pixel and the laser radar; r is the intensity information, indicating the reflection intensity of the object surface at the pixel.

[0071] In a specific implementation, when the spherical projection method is used to project the 3D point cloud data onto the 2D range image, each 3D point p = (x, y, z) with Cartesian coordinates is mapped onto the spherical surface. is converted to the corresponding image coordinate system, the formula is:

[0072]

[0073] Where f = f u +f d represents the vertical field of view of the sensor; the depth d of a point is expressed as The size of the distance image obtained after projection is (H, W, 5), where each pixel contains 5 channels (x, y, z, d, r), and r is the intensity information of the point.

[0074] Through the spherical projection method, each point in the 3D point cloud data is effectively converted to a pixel in the 2D image coordinate system. This conversion not only reduces the dimension of the data, but also retains important information in the original point cloud, such as coordinate values, depth, and intensity. The reduced-dimensional data is easier to process and analyze, while reducing the consumption of computing resources and improving the efficiency of subsequent processing. Key information can also be retained. The projection formula takes into account the Cartesian coordinates (x, y, z), depth d, and intensity r of the point, which are crucial for the subsequent semantic segmentation task. The depth information d reflects the distance between the point and the sensor, which helps the model understand the spatial structure of the scene; the intensity information r reflects the reflective properties of the object surface, which helps to distinguish different objects and materials. In addition, the projected 2D distance image has a fixed resolution (H, W) and number of channels (5), which is conducive to the design and implementation of subsequent processing algorithms. For example, in the semantic segmentation task, deep learning models such as convolutional neural networks can be used to process the 2D distance image to achieve accurate segmentation of point cloud data.

[0075] On this basis, the point cloud segmentation problem is transformed into an image segmentation problem.

[0076] S3. Through the three branches, the coordinate value and depth information, the coordinate value and intensity information, and the depth and intensity information are fused with feature information through the multi-scale context information fusion module MCIF to obtain the corresponding preliminary fusion information; and the self-attention module SA is used to capture the intra-branch dependency and enhance the features of the preliminary fusion information of the three branches to obtain the corresponding preliminary enhanced information.

[0077] In the two-dimensional range image, the five channels contained in each pixel present information of different modalities, and these channels play their own unique roles in the learning process of feature extraction. The coordinate value channel provides an accurate spatial position for each pixel, the depth channel reflects the actual distance between each point and the laser radar, and the reflection intensity channel reveals the material characteristics of the object surface. In addition, according to the projection formula of converting the original three-dimensional point cloud data into a two-dimensional range image, it can be known that the coordinate and depth channels are closely related to each other. Their synergy enables the segmentation model to more accurately identify and distinguish object features in complex scenes. Specifically, the combination of distance and coordinates can help the model more accurately understand the relative position relationship between the point and the sensor, and enhance the insight into the shape and depth structure of the object. The combination of intensity and coordinates helps to identify the reflection characteristics of different object surfaces and improve the accuracy of object classification through subtle differences in surface materials. Therefore, in order to capture the potential physical quantity relationship between different modal channels and reduce the noise interference of a single channel, the present invention constructs a multi-channel interactive feature extraction module (Multi-Channel Interactive Feature Extraction Module, MCIF) to extract multi-scale features from the combined channel.

[0078] like Figure 3 As shown, the multi-scale context information fusion module MCIF includes three basic blocks connected in sequence;

[0079] The first basic block includes one 1×1 convolutional layer and multiple non-1×1 convolutional layers with different kernels, which are used to extract input features respectively; after the multiple non-1×1 convolutional layers with different kernels, there is another 1×1 convolutional layer connected to fuse the multi-scale information of the convolutional layers with non-1×1 kernels; after the output residuals of the two 1×1 convolutional layers are connected, the input features of the second basic block are obtained.

[0080] In the specific implementation, the convolution layers of the first basic block with different non-1×1 kernels are convolution layers with 3×3, 5×5, and 7×7 kernels. By using convolution kernels of different sizes (3×3, 5×5, 7×7), the first basic block can capture information of different scales in the input feature map. Small convolution kernels (such as 3×3) can capture local detail features, while large convolution kernels (such as 5×5, 7×7) can capture contextual information in a wider range. This multi-scale feature extraction helps the model better understand the complex scene structure and improve the accuracy of segmentation. Convolution kernels of different sizes can learn different feature representations, which can be further fused and utilized in subsequent processing. By combining features of multiple scales, the model can generate richer and more discriminative feature representations, which is crucial to improving the performance of the model.

[0081] It should be noted that the structure of the third basic block is the same as that of the first basic block.

[0082] The second basic block includes a 1×1 convolution layer and a non-1×1 convolution layer, which are used to extract input features. After the non-1×1 convolution layer, there is also a non-1×1 dilated convolution layer. The outputs of the 1×1 convolution layer and the non-1×1 dilated convolution layer of the second basic block are connected with the residual input features of the first basic block to obtain the input features of the third basic block. In specific implementation, the non-1×1 convolution layer of the second basic block is a convolution layer with a 3×3 kernel; the non-1×1 dilated convolution layer of the second basic block is a 3×3 dilated convolution layer.

[0083] In this way, the 3×3 convolution layer is able to capture local detail information in the input feature map. This local feature extraction is crucial for understanding details such as textures and edges in images. Compared with larger convolution kernels, the 3×3 convolution layer has a smaller number of parameters, which helps reduce the complexity of the model and the risk of overfitting. At the same time, smaller convolution kernels also mean less computation, which helps improve the operating efficiency of the model. The 3×3 convolution layer can be flexibly combined with other convolution layers or operations (such as pooling, activation functions, etc.) to build more complex network structures. This flexibility enables the model to be customized and optimized according to different task requirements. In addition, the dilated convolution (also known as the expanded convolution) increases the receptive field by inserting holes in the convolution kernel without increasing the size of the convolution kernel. The 3×3 dilated convolution layer is able to capture a wider range of contextual information while maintaining computational efficiency. The dilated convolution can increase the receptive field without changing the size of the input feature map, which is crucial to maintaining the resolution and detail information of the feature map. This helps the model better utilize this information in subsequent processing. When the dilated convolution layer is used in combination with convolution layers of other scales (such as 1×1 convolution layers or dilated convolution layers with different expansion rates), multi-scale feature fusion can be achieved. This fusion helps improve the model's ability to understand complex scenes and enhance the discriminative power of features. Compared with traditional downsampling and upsampling operations, dilated convolution can increase the receptive field without losing information. This helps the model maintain higher accuracy and stability in tasks such as segmentation and detection.

[0084] Through convolutional layers with different kernel sizes, MCIF can extract multi-scale information from input features. This information is essential for understanding complex scene structures, capturing objects of different scales, and distinguishing different object categories. The 1×1 convolutional layer is used to fuse this multi-scale information to generate more robust and discriminative feature representations. In the second basic block, the non-1×1 hole convolution layer can increase the receptive field of the convolution kernel while keeping the resolution of the feature map unchanged. This helps capture contextual information in a larger range and improves the model's ability to understand the global structure. In addition, the use of residual connections helps alleviate the gradient vanishing problem in deep neural networks, making it easier for the network to learn useful feature representations. At the same time, it can also promote the flow of information, making it easier for information from the previous layer to be passed to the next layer, thereby improving the performance of the entire network. Through the sequential connection and residual connection of the three basic blocks, MCIF can make full use of the multi-scale information in the input features and generate richer and more discriminative feature representations. This is essential for improving the performance of tasks such as semantic segmentation.

[0085] S4: Aggregate the preliminary enhanced information of the three branches to obtain comprehensive fusion information, feed the comprehensive fusion information into the SA module, capture the global dependency relationship between branches and enhance the features to obtain comprehensive enhanced information.

[0086] In order to further refine the multi-scale features extracted from the two-dimensional range image, the present invention introduces the spatial channel attention module (SA) in SA-Net to refine the extracted fused channel interaction features to effectively model the dependencies between different combined channels.

[0087] S5. Feed the comprehensive enhancement information to the input module, which extracts an initial feature map with the same resolution as the original two-dimensional range image, and generates a multi-layer feature map with successively decreasing resolutions in the encoding stage; in the decoding stage, the multi-layer feature map is upsampled layer by layer through the dynamic sampling module, and finally the feature map is restored to the same resolution as the initial feature map; and the final 2D prediction map is obtained.

[0088] In specific implementation, in the encoding stage, 4 layers of feature maps are generated in sequence; the resolutions of the 4 layers of feature maps are 1, 1 / 2, 1 / 4, and 1 / 8 of the initial feature maps, respectively. In this way, the encoding stage generates 4 layers of feature maps with different resolutions in sequence, which can capture the multi-scale information in the input image, optimize the computational efficiency and memory usage, integrate contextual information, build a feature pyramid, adapt to objects of different scales, and promote information flow and fusion. These effects work together to improve the performance of the model, enabling it to better understand and process complex image data.

[0089] In the decoding stage, the dynamic sampling module is a dynamic upsampler Dysample-S+. In this way, using the dynamic upsampler DySample-S+ as a dynamic sampling module can bring multiple effects such as high efficiency, light weight, strong adaptability, improved model performance, and easy integration and deployment. These effects work together in the decoding stage of the model, which helps to improve the overall performance and meet the needs of practical applications.

[0090] In semantic segmentation, commonly used feature upsampling methods include nearest neighbor interpolation and bilinear interpolation. These methods interpolate low-resolution features through fixed rules, but they fail to fully consider the semantic information in the feature map, limiting the effective capture of contextual semantics. The nearest neighbor interpolation simply copies the nearest pixel value in the low-resolution feature map to the high-resolution position. Although it has high computational efficiency, it is easy to cause pixel block effects in the upsampled image due to the lack of smooth transition. In particular, it shows obvious limitations in the segmentation tasks of object boundaries or fine structures. Bilinear interpolation generates new pixel values ​​by weighted average of adjacent pixels. Although it can achieve a certain degree of smooth transition, its fixed calculation rules rely on geometric continuity rather than semantic information. Therefore, when processing complex scenes or irregular object boundaries, it is easy to introduce blurring effects, resulting in blurred category boundaries and loss of small object details. To this end, the present invention adopts a lightweight and more efficient dynamic upsampler Dysample-S+ in the encoding and decoding stage to replace the traditional nearest neighbor interpolation and bilinear interpolation, which significantly improves the segmentation accuracy and boundary detail retention effect.

[0091] Dysample-S+ includes a receiving module, an offset generation module, an upsampling module and an output module; the receiving module is used to receive the feature map generated from the encoding stage; the offset generation module is used to perform preset operations on the received feature map to generate a new offset; the upsampling module is used to use the generated offset to adjust the original grid position to achieve dynamic upsampling of the feature map; the output module is used to output the feature map restored to the same resolution as the initial feature map after dynamic upsampling, which is used for subsequent decoding and classification tasks.

[0092] Dysample-S+ achieves efficient and high-quality feature map upsampling through the collaborative work of various modules. Compared with traditional upsampling methods, Dysample-S+ shows significant advantages in computational efficiency, feature map quality and task adaptability. Due to its high efficiency, flexibility and strong adaptability, Dysample-S+ has broad application prospects in image processing, computer vision and other fields. It can be used in various intensive prediction tasks to improve the performance and accuracy of the model.

[0093] Among them, when the offset generation module performs a preset operation on the received feature map, the preset operation includes pixel shuffling-linear projection and linear projection-pixel shuffling. In this way, the two combinations of pixel shuffling-linear projection and linear projection-pixel shuffling provide different feature processing orders, so that the offset generation module can be flexibly selected and adjusted according to different input feature maps and task requirements. In addition, through reasonable combination operations, the offset generation module can efficiently generate high-quality offsets, providing strong support for the subsequent upsampling process. At the same time, these two combinations also avoid complex calculation processes and improve overall processing efficiency. In addition, the generated offset can accurately guide the upsampling process to ensure that the feature map can be upsampled to a higher resolution as expected. This helps to improve the accuracy and performance of subsequent decoding and classification tasks.

[0094] When the offset generation module generates a new offset, the processing process is as follows:

[0095]

[0096] in, Indicates the generated offset; Represents input features; and Indicates the preset two-pixel reorganization; linear indicates linear transformation.

[0097] Such an upsampling process helps to improve the upsampling accuracy and the robustness of the model.

[0098] S6. Process the 2D prediction image to generate a 3D point cloud prediction.

[0099] This method converts 3D point cloud data into 2D range images through spherical projection, and retains the information of five channels: coordinate values ​​(x, y, z), distance (d) and intensity (r). Existing LiDAR-based point cloud semantic segmentation methods lack effective preprocessing of the projected range images and ignore the intrinsic connection between different modal channels. Compared with existing methods that only use single or a small amount of channel information, multi-channel fusion in this method can more comprehensively capture the spatial features and contextual information of point cloud data. Furthermore, through the multi-scale context information fusion module MCIF and the self-attention module SA, the method effectively fuses coordinate values ​​with depth information, coordinate values ​​with intensity information, and depth with intensity information, realizes feature enhancement and dependency capture within and between branches, and thus learns the potential relationship between different physical quantities in point cloud data. This enhances the model's ability to understand complex scenes. By fusing and interactively extracting the information of different modal channels in the range image, the potential relationship between different physical quantities in the point cloud data can be learned, which can reduce the risk of information compression and loss in the process of projecting the original 3D point cloud into the 2D range image.

[0100] In the decoding stage, this method uses a dynamic sampling module to upsample the multi-layer feature map layer by layer instead of the traditional fixed upsampling method. The existing LiDAR-based point cloud semantic segmentation method interpolates low-resolution features through fixed rules in the decoding stage, fails to fully consider the semantic information in the feature map, and limits the effective capture of contextual semantics. The dynamic sampling in this method can adaptively adjust the sampling strategy according to the context information and task requirements, thereby improving the flexibility and efficiency of the decoding process. Compared with the existing technology that only uses simple decoding methods such as bilinear interpolation, the dynamic sampling module of this method can more finely restore the detailed information of the feature map and improve the resolution and accuracy of the prediction map. In addition, this method achieves an effective balance between computational efficiency and segmentation accuracy by integrating interactive learning and dynamic sampling technology. On the one hand, multi-channel information fusion and feature enhancement improve the segmentation accuracy of the model; on the other hand, dynamic sampling and efficient decoding reduce the computational complexity of the model and meet the needs of real-time applications. Compared with the existing methods with high model complexity and slow inference speed, this method significantly improves computational efficiency while ensuring accuracy, and is more suitable for real-time decision-making scenarios such as autonomous driving. In addition, this method makes full use of the coordinate values, distance and intensity information in the LiDAR point cloud data, and improves the model's ability to understand complex urban environments and dynamic scenes through multi-scale context information fusion and self-attention mechanism. Compared with existing methods that ignore the intrinsic connection between this information or only use part of the information, this method can better adapt to changes in different scenes and lighting conditions, and improve the generalization ability and robustness of the model.

[0101] This method can achieve an effective balance between accuracy and computational efficiency, while making full use of the coordinate values, distance and intensity information in the LiDAR point cloud data to improve the segmentation accuracy and generalization ability of the model in complex urban environments and dynamic scenes. Through multi-channel information fusion and enhancement, dynamic sampling and efficient decoding, and a balance strategy between accuracy and computational efficiency, this method not only improves the segmentation accuracy and computational efficiency of the model, but also enhances the model's ability to understand and generalize complex scenes, providing strong technical support for real-time decision-making applications such as autonomous driving.

[0102] Embodiment 2

[0103] In order to better illustrate the effect of this method, the following verification is carried out.

[0104] In this verification experiment, three public benchmark datasets, SemanticKITTI, SemanticPOSS, and Nuscenes, were used to conduct experiments and analyze the results to verify the effectiveness of the algorithm. SemanticKITTI is an annotated dataset for semantic segmentation of point clouds in autonomous driving scenes, including 43,551 full 3D scan data from 22 LiDAR scan sequences, of which sequences 00 to 10 (a total of 19130 scans) are used for training, sequences 11 to 21 (a total of 20351 scans) are used for testing, and sequence 08 (4071 scans) are used for verification. SemanticPOSS is a smaller, sparser, and more challenging benchmark collected by Peking University. It consists of 2988 different complex LiDAR scenes, each with a large number of sparse dynamic instances (such as pedestrians and bicycles). SemanticPOSS is divided into 6 parts, of which part 2 is used as a test set and the other parts are used as training sets. NuScenes consists of 1,000 20-second scenes in Boston and Singapore, including various urban scenes, lighting, and weather conditions. In addition, there are 16 annotated semantic classes and the dataset is split into 28,130 training point cloud scans and 6,019 validation point cloud scans.

[0105] Mean Intersection over Union (MIoU) and parameters are used as the measurement indicators of accuracy and model complexity. These two evaluation indicators are the main standard metrics in the current field of point cloud semantic segmentation.

[0106] The experimental environment is tested on a computer equipped with two NVIDIA RTX 4080 GPUs. During training, data augmentation is performed by random rotation, random point dropout, and adding random noise to the X, Y, and Z values, and the weight decay is set to 1e-4. The learning rate is dynamically adjusted using a cosine annealing learning rate scheduler during training. For SemanticKITTI, 100 epochs are trained with an initial learning rate of 1e-2. For SemanticPOSS and Nuscenes, the network is trained for 50 epochs, and the minimum and maximum learning rates are set to 1e-5 and 2e-3.

[0107] Table 1 Comparison of mIoU (%) with other range image based LiDAR segmentation methods on the SemanticKITTI test set. The best results are marked in bold and the suboptimal results are underlined.

[0108] Table 1

[0109]

[0110] Table 1 shows the comparison of the present invention (i.e., Range-FDSeg in the table) with representative models on the SemanticKITTI test set (sequences 11-21). In terms of mIoU scores, the present method achieves advanced performance at different resolutions. Specifically, at an input resolution of 64×512px, Range-FDSeg performs optimally in 9 categories and achieves suboptimal performance in 5 categories. In addition, Range-FDSeg achieves the highest mIoU (64.1%) at an input resolution of 64×1024px, and performs particularly well in the segmentation of categories such as bicycles, trucks, other vehicles, people, and cyclists. This excellent performance is attributed to the proposed multimodal channel fusion strategy, which can effectively capture and distinguish details in challenging scenes. It is worth noting that Range-FDSeg not only performs well at an input resolution of 64×1024px, but also outperforms other methods at higher resolutions. Compared with the more advanced methods RangeVit and TORNA-HiRes, Range-FDSeg achieves the best results (64.1mIoU vs.64.0mIoU vs.63.1mIoU) at a lower resolution (64×1024px vs.64×2048px vs.128×2048px) and with smaller parameters (7M vs.21M).

[0111] In order to more effectively illustrate the improvement of the proposed method over the baseline model, a qualitative comparison is performed by visualizing the wrong prediction points. Figure 4As shown in the figure, the present invention compares FIDNet and Range-FDSeg on the SemanticKITTI validation set. Among them, (a) original point cloud. (b) true value. (c) FIDNet, the number of error points is 13062. (d) the present invention method, the number of error points is 10662. Compared with the baseline model, the method significantly reduces the number of incorrectly predicted points, and the prediction results are closer to the true annotations. The error visualization diagram shows that the present invention has a significant improvement over the original model, verifying the effectiveness and superiority of the method.

[0112] In order to quantitatively evaluate the effectiveness of different components, a series of ablation experiments were conducted on the SemanticKITTI validation set. In the experiment, FIDNet was selected as the benchmark for fair comparison, and the resolution of the input image was set to 64×512px. Subsequently, the FIL module and Dysample-S+ dynamic upsampler were added to the baseline method. Then, the boundary loss, auxiliary loss, and kernel size were gradually introduced to evaluate their contribution to the model performance.

[0113] Table 2 Ablation study evaluation on SemanticKITTI validation set

[0114]

[0115] Table 2 shows the parameters and mIoU scores of the model after adding each component. The first row is the result of the FIDNet baseline model. The second row of results shows that after adding the FIL module, the model accuracy is improved by 1.0%, indicating that the FIL module effectively reduces the risk of information compression and loss in the prediction process. The third row reports the performance after the introduction of the dynamic upsampler Dysample-S+, which is about 0.4% higher than the previous row, proving that Dysample-S+ can effectively preserve boundary details and only adds a very small amount of parameters. Rows 4 to 5 show the effectiveness of additional loss functions. In particular, the boundary loss brings a performance improvement of more than 1.8%, while the auxiliary loss increases the model's mIoU score by 2.3%. The last row shows the improvement in model performance after replacing the 1×1 convolution in the input module with a 3×3 convolution.

[0116] Figure 5 The visualization results of 2D image prediction of the baseline model and after adding the designed FIL module are shown. (a) Distance image. (b) True value. (c) Baseline model. (d) Baseline model with FIL module added. The red box highlights the prediction differences between different methods and the ground truth of certain categories. It can be clearly seen that the FIL module significantly enhances the segmentation performance under various input frames.

[0117] Experimental results on the SemanticKITTI, SemanticPOSS and NuScenes datasets show that the three-dimensional point cloud semantic segmentation method proposed in the present invention has achieved excellent results in segmentation performance and optimized the segmentation accuracy of multiple categories in dynamic and complex scenes.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit the technical solution. Those skilled in the art should understand that those modifications or equivalent substitutions of the technical solution of the present invention that do not depart from the purpose and scope of the technical solution should be included in the scope of the claims of the present invention.

Claims

1. A radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling, characterized in that: The following steps are involved: S1, obtain laser radar three-dimensional point cloud data; S2. Use a spherical projection method to project the three-dimensional point cloud data into a two-dimensional range image; each pixel in the two-dimensional range image contains five channels (x, y, z, d, r) information; where (x, y, z) is the coordinate value information, indicating the spatial position of the pixel; d is the distance information, indicating the distance between the pixel and the laser radar; r is the intensity information, indicating the reflection intensity of the object surface at the pixel; S3, through the three branches, the coordinate value and depth information, the coordinate value and intensity information, the depth and intensity information are respectively fused through the multi-scale context information fusion module MCIF to obtain the corresponding preliminary fusion information; and the self-attention module SA is used to capture the intra-branch dependency and feature enhancement of the preliminary fusion information of the three branches to obtain the corresponding preliminary enhancement information; S4, aggregate the preliminary enhanced information of the three branches to obtain comprehensive fusion information, feed the comprehensive fusion information into the SA module, capture the global dependency relationship between branches and enhance the features to obtain comprehensive enhanced information; S5, feeding the comprehensive enhancement information to the input module, the input module extracts the initial feature map with the same resolution as the original two-dimensional range image, and generates a multi-layer feature map with successively decreasing resolutions in the encoding stage; in the decoding stage, the multi-layer feature map is upsampled layer by layer through the dynamic sampling module, and finally the feature map is restored to the same resolution as the initial feature map; and the final 2D prediction map is obtained; S6. Process the 2D prediction image to generate a 3D point cloud prediction.

2. The radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling as claimed in claim 1, characterized in that: In S2, when the 3D point cloud data is projected onto the 2D range image using the spherical projection method, each 3D point p = (x, y, z) with Cartesian coordinates is mapped onto the spherical surface. is converted to the corresponding image coordinate system, the formula is: Where f = f u +f d represents the vertical field of view of the sensor; the depth d of a point is expressed as The size of the distance image obtained after projection is (H, W, 5), where each pixel contains 5 channels (x, y, z, d, r), and r is the intensity information of the point.

3. The radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling as claimed in claim 1, characterized in that: In S3, the multi-scale context information fusion module MCIF includes three basic blocks connected sequentially; The first basic block includes a 1×1 convolution layer and multiple non-1×1 convolution layers with different kernels, which are used to extract input features. After the multiple non-1×1 convolution layers with different kernels, another 1×1 convolution layer is connected to fuse the multi-scale information of the convolution layers with non-1×1 kernels. After the output residuals of the two 1×1 convolution layers are connected, the input features of the second basic block are obtained. The second basic block includes a 1×1 convolution layer and a non-1×1 convolution layer, which are used to extract input features. After the non-1×1 convolution layer, there is also a non-1×1 hole convolution layer. The outputs of the 1×1 convolution layer and the non-1×1 hole convolution layer of the second basic block are connected with the input feature residual of the first basic block to obtain the input features of the third basic block. The third basic block has the same structure as the first basic block.

4. The radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling as claimed in claim 3, characterized in that: The convolutional layers with different non-1×1 kernels of the first basic block are 3×3, 5×5, and 7×7 kernel convolutional layers.

5. The radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling as claimed in claim 3, characterized in that: The second basic block’s non-1×1 convolution layer is a 3×3 kernel convolution layer; the second basic block’s non-1×1 atrous convolution layer is a 3×3 atrous convolution layer.

6. The radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling as claimed in claim 1, characterized in that: In S5, during the encoding stage, 4 layers of feature maps are generated in sequence; the resolutions of the 4 layers of feature maps are 1, 1 / 2, 1 / 4, and 1 / 8 of the initial feature maps, respectively.

7. The radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling as claimed in claim 1, characterized in that: In S5, in the decoding stage, the dynamic sampling module is a dynamic upsampler Dysample-S+.

8. The radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling as claimed in claim 7, characterized in that: Dysample-S+ includes a receiving module, an offset generation module, an upsampling module and an output module; The receiving module is used to receive the feature map generated from the encoding stage; the offset generating module is used to perform a preset operation on the received feature map to generate a new offset; The upsampling module is used to adjust the original grid position using the generated offset to achieve dynamic upsampling of the feature map; The output module is used to output a feature map that has been restored to the same resolution as the initial feature map after dynamic upsampling for subsequent decoding and classification tasks.

9. The radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling as claimed in claim 8, characterized in that: When the offset generation module performs a preset operation on the received feature map, the preset operation includes pixel shuffling-linear projection and linear projection-pixel shuffling.

10. The radar point cloud semantic segmentation method based on fusion interactive learning and dynamic sampling as claimed in claim 9, characterized in that: When the offset generation module generates a new offset, the processing process is as follows: in, Represents the generated offset; x represents the input feature; shuffle1(x) and shuffle2(x) represent the preset two pixel reorganizations; linear represents linear transformation.

Citation Information

Patent Citations

  • Remote sensing semantic segmentation method fusing optical image and laser radar point cloud

    CN116246074A

  • Urban road environment semantic segmentation method and system based on multi-view fusion

    CN117095162A

  • Laser radar point cloud semantic segmentation method based on feature interaction fusion

    CN118351314A

  • Method for semantic segmentation using correlations and regional associations of multi-scale features, and computer program recorded on record-medium for executing method thereof

    KR102546206B1

Cited By

  • Semantic segmentation method and system based on 3D point cloud

    CN121353677A