Point cloud learning feature representation method and device based on neighborhood geometry embedding

By combining multi-scale neighborhood construction and self-attention mechanism, the problem of insufficient neighborhood scale adaptation in autonomous driving point cloud processing is solved, more accurate point cloud feature expression and environmental information capture are achieved, and the target detection and recognition accuracy of the autonomous driving system is improved.

CN120599434BActive Publication Date: 2025-10-03NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511103591.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-07
Publication Date
2025-10-03
Estimated Expiration
2045-08-07

AI Technical Summary

Technical Problem

Existing point cloud processing methods for autonomous driving have insufficient neighborhood scale adaptation when faced with objects of different sizes and complex scenes, resulting in low target recognition accuracy and poor feature aggregation robustness, making it difficult to provide reliable point cloud feature expression in dynamic environments.

Method used

A point cloud learning feature representation method based on neighborhood geometric embedding is adopted. Multi-scale neighborhoods are constructed through a multi-branch network. The neighboring point response map is learned in combination with the self-attention mechanism. The optimal neighborhood is dynamically selected and features of different scales are integrated to enhance the flexibility and accuracy of feature extraction.

Benefits of technology

It improves the adaptability and accuracy of point cloud feature expression, enhances the accuracy of target detection, recognition and tracking tasks, and can capture environmental information more comprehensively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599434B_ABST
    Figure CN120599434B_ABST
Patent Text Reader

Abstract

This application relates to a method and apparatus for learning feature representation for point clouds based on neighborhood geometry embedding. The method comprises: acquiring 3D point cloud data of the three-dimensional spatial environment surrounding a vehicle; performing feature extraction on the 3D point cloud data to obtain an initial point feature representation; constructing a multi-scale neighborhood of a target point using a pre-set neighborhood determination strategy; learning each neighboring point response map from the initial point feature representation in each neighborhood within the multi-scale neighborhood according to a neighborhood point feature aggregation model to obtain multi-scale features; and performing feature fusion on multiple multi-scale features to obtain a final point cloud learning feature representation. This method can provide reliable point cloud feature representation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of autonomous driving technology, and in particular to a point cloud learning feature representation method and device based on neighborhood geometry embedding. Background Art

[0002] In the field of autonomous driving, point cloud data, with its high-precision three-dimensional description of the vehicle's surroundings, has become a core data format for environmental perception. LiDAR and other sensors collect massive amounts of point cloud data in real time, allowing the autonomous driving system to construct a three-dimensional scene containing rich information about roads, obstacles, pedestrians, and other vehicles. This is the cornerstone for safe and reliable driving decisions. From detecting dense pedestrians and vehicles on complex urban streets to accurately identifying vehicles and road conditions at long distances on highways, the effective processing of point cloud data is directly related to the safety and intelligence level of autonomous driving.

[0003] Traditional point cloud processing methods for autonomous driving, such as PointNet and PointNet++, rely on predefined neighborhoods or fixed radius strategies to extract local geometric features, exposing significant flaws. First, the size of objects in autonomous driving scenarios varies greatly, from tiny traffic signs to massive trucks. A single neighborhood scale is difficult to adapt. For small objects, a large neighborhood can easily miss key details. For large objects, a small neighborhood fails to capture the entire structure, introducing inaccurate information or noise that severely impacts object recognition accuracy. Second, existing methods often employ simple, uniform feature aggregation within a neighborhood, such as max pooling and average pooling. These methods ignore the complex geometric relationships between points and fail to fully exploit the spatial structure inherent in point cloud data. This leads to inadequate representation of object geometry and reduces the system's ability to understand complex scenes. To address these challenges, current technologies are attempting to incorporate graph convolutions (such as DGCNN) or attention mechanisms (such as PointTransformer). However, in the dynamic and changeable autonomous driving environment, they are still unable to adaptively determine the optimal neighborhood scale based on different objects and scenes, and lack flexibility when facing complex road conditions and diverse targets. At the same time, the feature aggregation process is not robust enough and is easily affected by interference factors such as occlusion and noise. It is difficult to stably and accurately provide reliable point cloud feature expressions for autonomous driving decisions under various extreme weather and complex lighting conditions, which limits the widespread application and safety improvement of autonomous driving systems. Summary of the Invention

[0004] Based on this, it is necessary to provide a point cloud learning feature representation method and device based on neighborhood geometric embedding that can provide reliable point cloud feature expression to address the above technical problems.

[0005] A point cloud learning feature representation method based on neighborhood geometric embedding, the method comprising:

[0006] Acquire 3D point cloud data of the three-dimensional space surrounding the vehicle; perform feature extraction on the 3D point cloud data to obtain the initial point feature representation; and construct a multi-scale neighborhood of the target point using a pre-set neighborhood determination strategy.

[0007] In each neighborhood within the multi-scale neighborhood, the neighborhood point feature aggregation model is used to learn each neighboring point response map from the initial point feature representation to obtain multi-scale features.

[0008] Feature fusion is performed on multiple multi-scale features to obtain the final point cloud learning feature representation.

[0009] A point cloud learning feature representation device based on neighborhood geometry embedding, the device comprising:

[0010] The multi-scale neighborhood construction module is used to obtain 3D point cloud data of the three-dimensional spatial environment around the vehicle; extract features from the 3D point cloud data to obtain the initial point feature representation; and construct a multi-scale neighborhood of the target point using a pre-set neighborhood determination strategy;

[0011] The feature representation learning module is used to learn each neighboring point response map from the initial point feature representation in each neighborhood within the multi-scale neighborhood according to the neighborhood point feature aggregation model to obtain multi-scale features;

[0012] The feature representation fusion module is used to fuse multiple multi-scale features to obtain the final point cloud learning feature representation.

[0013] The above-mentioned point cloud feature representation method and device based on neighborhood geometric embedding first designs a multi-branch network to extract neighborhood features at different scales. When processing small objects, a small-scale neighborhood focuses on key details, while when processing large objects, a large-scale neighborhood covers the entire structure. This avoids information loss or noise, and more comprehensively and accurately captures the features of different objects. The optimal neighborhood is then dynamically selected using regression fusion coefficients, avoiding the limitations of manually preset scales. Adaptive neighborhood fusion automatically adjusts the neighborhood scale based on the specific point cloud data characteristics and scenario requirements, allowing the model to find the most suitable neighborhood for feature extraction in different situations. This improves the accuracy and flexibility of feature extraction and enhances the adaptability of point cloud feature representation to complex environments. Finally, a self-attention mechanism is used to model the geometric correlation between neighborhood points and the target point, learning a response map for each neighboring point. Unlike traditional equal aggregation, this method considers the geometric relationship between points, highlighting neighbors important to the target point, suppressing the influence of irrelevant or interfering points, and improving the discriminative performance of feature aggregation. Multiple multi-scale features are then fused to integrate feature information extracted from neighborhoods of different scales. Features at different scales contain information at different levels and aspects of an object. Small-scale features focus on details, while large-scale features reflect overall structure. Fusion of these features leverages the strengths of each scale, making feature representations learned from point clouds more comprehensive and rich. This provides autonomous driving systems with more accurate and detailed environmental information, helping to improve the accuracy of tasks such as object detection, recognition, and tracking. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 1 is a flow chart of a point cloud learning feature representation method based on neighborhood geometry embedding in one embodiment;

[0015] Figure 2 This is a diagram of the overall architecture of NG-Net in one embodiment;

[0016] Figure 3 A schematic diagram illustrating three strategies for neighborhood determination in one embodiment; Figure 3 (a) is a schematic diagram of selecting 5 close points as the neighborhood. Figure 3 (b) is a schematic diagram of selecting points inside the sphere as the neighborhood. Figure 3 (c) Schematic diagram of sampling 5 points to form a neighborhood based on the determination of the sphere;

[0017] Figure 4 Schematic diagram of an NGL module for point feature representation learning in another embodiment;

[0018] Figure 5 A detailed structural diagram of the NPFA module in one embodiment;

[0019] Figure 6 A diagram illustrating a partial segmentation result in one embodiment;

[0020] Figure 7 This is a semantic segmentation result diagram in one embodiment; Figure 7 (a) is a diagram showing the visualization of the mIoU of 66.42% achieved in the semantic segmentation task. Figure 7 (b) A visualization diagram showing the results of achieving 66.42% mIoU and 85.2% overall accuracy (OA) in the semantic segmentation task;

[0021] Figure 8 1 is a structural block diagram of a point cloud learning feature representation device based on neighborhood geometric embedding in one embodiment;

[0022] Figure 9 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0023] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0024] In one embodiment, Figure 1 As shown in FIG, a point cloud learning feature representation method based on neighborhood geometric embedding is provided, which includes the following steps:

[0025] Step 102: Acquire 3D point cloud data of the three-dimensional space environment around the vehicle; perform feature extraction on the 3D point cloud data to obtain an initial point feature representation; and construct a multi-scale neighborhood of the target point using a pre-set neighborhood determination strategy.

[0026] NG-Net (this application) improves on this approach in two aspects: perceiving multiple neighborhoods and learning to adaptively decide preferred neighborhoods, and modeling the correlation with the target point instead of treating each neighboring point equally.

[0027] The overall structure of NG-Net is as follows Figure 2 As shown, given an input point cloud, it is first calibrated. Then, our Neighbor Geometry Learning module (NGL module) is applied to extract point features. For segmentation tasks, point-by-point features are extracted to learn segmentation labels. For classification tasks, point features are pooled into a global representation, and the final classification score is regressed.

[0028] Before entering the NGL-module, NG-Net (this application) first passes a MLP network processing 3D point cloud , which takes 3D coordinates as input and outputs the initial point feature representation .

[0029] (1)

[0030] in Represents the parameters of the MLP network, d represents the feature dimension, Indicates a point feature representation.

[0031] The neighborhood represents the receptive field of the target point feature. There are three common neighborhood determination strategies.

[0032] 1) Predefined points. Figure 3 As shown, a predefined number of points are selected near the target point to build a neighborhood based on the spatial distance.

[0033] 2) Predefined range: A radius is given with the target point as the center, and the points inside the sphere constitute the neighborhood.

[0034] 3) Fusion of the two aforementioned strategies. Within a given sphere, adjacent points are sampled to a given number. If there are not enough points within the sphere, the target point is repeated. This precise neighborhood provides complete detail and reduces interference from irrelevant points.

[0035] Figure 3 Illustration of the three strategies determined for the neighborhood. The target point is shown in red. Figure 3 In (a), five close points are selected as the neighborhood. Figure 3 In (b), select a point inside the sphere. Figure 3 In (c), based on the determination of the sphere, 5 points are sampled to form a neighborhood.

[0036] Using the above-mentioned domain determination strategy, the target point , build L Different radii Neighborhood , ensuring coverage of geometric information from local to global.

[0037] Objects in autonomous driving scenarios vary greatly in size, and traditional single-scale neighborhood methods are flawed. By constructing a multi-scale neighborhood using a pre-defined neighborhood determination strategy, we can cover objects of varying sizes. For example, by designing a multi-branch network to extract features from neighborhoods of different scales, we can use a small-scale neighborhood to focus on key details when processing small objects, while using a large-scale neighborhood to cover the entire structure when processing large objects. This avoids information loss or the introduction of noise, and allows for a more comprehensive and accurate capture of the characteristics of different objects.

[0038] Traditional methods, which predefine neighborhoods, struggle to adapt to complex scenarios. However, this approach dynamically selects the optimal neighborhood through regression fusion coefficients, avoiding the limitations of manually pre-set scales. Adaptive neighborhood fusion automatically adjusts the neighborhood scale based on the specific characteristics of point cloud data and scenario requirements, enabling the model to find the most suitable neighborhood for feature extraction in different situations. This improves the accuracy and flexibility of feature extraction and enhances the adaptability of point cloud feature representation to complex environments.

[0039] Step 104 : In each neighborhood within the multi-scale neighborhood, a response map of each neighboring point is learned from the initial point feature representation according to the neighborhood point feature aggregation model to obtain a multi-scale feature.

[0040] like Figure 4 The figure shows the NGL-module schematic diagram of point feature representation learning. In the NGL module, in order to learn the response coefficient of each neighboring point to determine its contribution to the target point feature, a self-attention-based network is designed to learn the response map. The network proposes a new neighboring point feature aggregation module, called NPFA, which is located in the NGL module as shown in the figure. Figure 5 shown.

[0041] The neighborhood point feature aggregation model mainly consists of three parts:

[0042] 1. Geometric perception: Enhance features through self-attention mechanism to calculate query Q, key K, value V:

[0043] (2)

[0044] Generating attention maps , after weighted aggregation, enhanced features are obtained .

[0045] Then, V is weighted averaged based on the self-attention map M to form the relevant information S, that is, . According to the residual principle, the input feature map is Enhancement: The entire feature map is embedded into each point representation. To improve the interactivity between feature channels at each point, a U-Net structure is used in the channel dimension. An MLP network is used to compress the number of channels and then restore them to their original size.

[0046] 2. Channel association: given feature map , first take the average of all neighboring points to produce a global representation, that is ,in represents the averaging operation. Then, in A fully connected network is used, which consists of 3 fully connected layers to learn a channel graph representing the key channels ,in , represents a fully connected network parameterized by ν. The feature map is further updated based on the channel map, i.e. , where key channels are further emphasized.

[0047] 3. Response map generation: Use MLP to generate the response map C, and retain the key point features through Top-K truncation.

[0048] (3)

[0049] Among them, Is a parameter The MLP network of Top-K response parameters is , where the truncation function The K largest elements are retained and the remaining parameters are set to 0. This means that irrelevant neighboring points are disabled. Finally, the activated point features are aggregated by weighted averaging, that is:

[0050] (4)

[0051] A self-attention mechanism is used to model the geometric correlation between neighboring points and the target point, learning a response map for each neighboring point. Unlike traditional equal aggregation, this method considers the geometric relationship between points, highlighting neighboring points that are important to the target point, suppressing the influence of irrelevant or interfering points, and improving the discriminability of feature aggregation. For example, when identifying a vehicle, the geometric correlation between points can be used to more accurately aggregate the vehicle's point cloud features, eliminating interference from the surrounding point cloud, making the final feature representation more representative and discriminative.

[0052] Step 106: perform feature fusion on multiple multi-scale features to obtain the final point cloud learning feature representation.

[0053] Multi-scale features After concatenation, input the fully connected layer, where are connected in series ,in, Regress to an L-dimensional parameter through a fully connected network , p is used to express preference. And based on p, these feature representations are fused. In form,

[0054] (5)

[0055] (6)

[0056] in represents a fully connected network whose output is normalized to p through a softmax operation, is the first i elements, implicitly indicating the neighborhood selection and guiding the feature fusion process, the feature is represented as , which can be used for different tasks.

[0057] Multiple multi-scale features are fused to integrate feature information extracted from neighborhoods at different scales. Features at different scales contain information at different levels and aspects of an object. Small-scale features focus on details, while large-scale features reflect overall structure. Fusion of these features leverages the strengths of each scale, making the feature representation learned from point clouds more comprehensive and rich. This provides autonomous driving systems with more accurate and detailed environmental information, helping to improve the accuracy of tasks such as object detection, recognition, and tracking.

[0058] The above-mentioned point cloud feature representation method and device based on neighborhood geometric embedding first designs a multi-branch network to extract neighborhood features at different scales. When processing small objects, a small-scale neighborhood focuses on key details, while when processing large objects, a large-scale neighborhood covers the entire structure. This avoids information loss or noise, and more comprehensively and accurately captures the features of different objects. The optimal neighborhood is then dynamically selected using regression fusion coefficients, avoiding the limitations of manually preset scales. Adaptive neighborhood fusion automatically adjusts the neighborhood scale based on the specific point cloud data characteristics and scenario requirements, allowing the model to find the most suitable neighborhood for feature extraction in different situations. This improves the accuracy and flexibility of feature extraction and enhances the adaptability of point cloud feature representation to complex environments. Finally, a self-attention mechanism is used to model the geometric correlation between neighborhood points and the target point, learning a response map for each neighboring point. Unlike traditional equal aggregation, this method considers the geometric relationship between points, highlighting neighbors important to the target point, suppressing the influence of irrelevant or interfering points, and improving the discriminative performance of feature aggregation. Multiple multi-scale features are then fused to integrate feature information extracted from neighborhoods of different scales. Features at different scales contain information at different levels and aspects of an object. Small-scale features focus on details, while large-scale features reflect overall structure. Fusion of these features leverages the strengths of each scale, making feature representations learned from point clouds more comprehensive and rich. This provides autonomous driving systems with more accurate and detailed environmental information, helping to improve the accuracy of tasks such as object detection, recognition, and tracking.

[0059] In one embodiment, feature extraction is performed on 3D point cloud data to obtain an initial point feature representation, including:

[0060] Perform feature extraction on 3D point cloud data and obtain the initial point feature representation as

[0061] ;

[0062] in, Represents the parameters of the MLP network.

[0063] In one embodiment, a pre-set neighborhood determination strategy includes a predefined point number strategy, a predefined range strategy, and a fusion strategy; the predefined point number strategy is to select a predefined number of points near the target point based on the spatial distance to construct a neighborhood; the predefined range strategy is to give a radius with the target point as the center, and the points inside the sphere constitute the neighborhood; the fusion strategy is to sample adjacent points to a given number within a given sphere, and if there are not enough points inside the sphere, the target point is repeated.

[0064] In one embodiment, the neighborhood point feature aggregation model includes a geometric perception module, a channel association module, and a response map generation module; in each neighborhood within the multi-scale neighborhood, each neighborhood point response map is learned from the initial point feature representation according to the neighborhood point feature aggregation model to obtain multi-scale features, including:

[0065] In the geometric perception module, the initial point feature representation is enhanced through the self-attention mechanism to obtain the enhanced features;

[0066] In the channel association module, the enhanced features are channel-associated to obtain key features;

[0067] In the response graph generation module, a response graph is generated based on key features and key point features are retained through Top-K truncation.

[0068] In one embodiment, the geometric perception module performs feature enhancement on the initial point feature representation through a self-attention mechanism to obtain enhanced features, including:

[0069] Calculate the query Q, key K, and value V, generate a self-attention map based on the query Q and key K, and obtain the feature map through weighted aggregation ; Based on the self-attention map M, weighted average V is performed to form relevant information S, that is, ; According to the residual principle, the input initial point feature representation is performed Enhance, get the enhanced features, Represents the initial point feature representation.

[0070] In one embodiment, the channel correlation module performs channel correlation on the enhanced features to obtain key features, including:

[0071] Given enhanced features , first take the average of all neighboring points to produce a global representation, that is ,in represents an average operation;

[0072] exist A fully connected network is used to learn a channel graph s representing the key channels, where , represents a fully connected network parameterized by ν;

[0073] Based on the channel map, the enhanced features are further updated to obtain the key features .

[0074] In one embodiment, generating a response map based on key features in a response map generation module and retaining key point features through Top-K truncation includes:

[0075] The response graph C generated by MLP is

[0076] ;

[0077] in, Is a parameter MLP network, Indicates key features;

[0078] Top-K response parameters, i.e. , where the truncation function Keep the K largest elements and set the remaining parameters to 0. The key point features are retained through Top-K truncation, that is:

[0079] ;

[0080] in, Indicates a point The feature representation of Indicates the number of points, Indicates the assignment to the point The weight of .

[0081] In one embodiment, multiple multi-scale features are fused to obtain a final point cloud learning feature representation, including:

[0082] Multi-scale features After concatenation, input the fully connected layer, where are connected in series ,in Regress to an L-dimensional parameter through a fully connected network , p is used to express preference, and then these feature representations are fused based on p to obtain the final point cloud learning feature representation as ;

[0083] ;

[0084] in, represents a fully connected network whose output is normalized by the softmax operation to , yes No.i elements, Indicates the dimension.

[0085] In a specific embodiment, in the segmentation experiment on the ShapeNetPart dataset, a large-scale dataset containing 16,881 three-dimensional shapes was used, of which 14,006 samples were used for training and 2,874 samples were used for testing, covering 16 different categories of objects, each of which contained 2-6 separable parts. In the data preprocessing stage, the present application uniformly samples each three-dimensional object, extracts 2,048 spatial points, and uses only its three-dimensional coordinates (x, y, z) as input features without relying on additional information such as normal vectors. Experimental results show that the present application achieved an excellent performance of 86.4% in the mean intersection over union (mIoU) indicator, which is significantly better than current mainstream methods, including PointNet++ (85.1%) and DGCNN (85.2%). Especially when dealing with objects with complex geometric structures, such as Figure 6 As shown in the figure, this application demonstrates excellent segmentation capabilities; for objects such as tables that consist of a flat surface and supporting structures, it can accurately segment different parts such as the tabletop and legs. This excellent performance is mainly due to the multi-scale neighborhood feature extraction mechanism and attention-weighted feature aggregation strategy proposed in this application. This enables the network to adaptively capture local geometric features at different scales and pay more attention to key areas, thereby achieving accurate segmentation of complex geometric structures.

[0086] In the semantic segmentation experiment on the S3DIS dataset, a large-scale indoor scene point cloud dataset released by Stanford University was used. This dataset contains 272 complete room scans in 6 different areas, and each point is annotated with 13 semantic categories (such as walls, floors, tables and chairs, doors and windows, etc.). In the experimental setup, each room was divided into 1m×1m blocks, 4,096 points were sampled from each block as input, and a 9-dimensional feature vector (containing XYZ coordinates, RGB color information, and normalized spatial position coordinates) was used as the initial feature representation. This method achieved 66.42% mIoU and 85.2% overall accuracy (OA) in the semantic segmentation task, significantly outperforming baseline methods such as PointNet (47.6% mIoU) and DGCNN (56.1% mIoU); Figure 7 (a) is a diagram showing the visualization of the mIoU of 66.42% achieved in the semantic segmentation task. Figure 7(b) A visualization of the semantic segmentation results, achieving a mean Intersection Over Union (MIoU) of 66.42% and an overall accuracy (OA) of 85.2%. This method accurately identifies various elements in complex indoor scenes: it maintains the integrity of segmentation boundaries for large planar structures (such as walls and ceilings); accurately distinguishes between different instances of furniture objects (such as tables, chairs, and bookshelves); and accurately segments slender structures (such as door frames and lamps). This excellent performance stems from the NPFA module's enhanced representation of local geometric features and the multi-scale fusion mechanism's adaptive processing of objects of different sizes, enabling the network to simultaneously capture both the global structure and local details of the scene.

[0087] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0088] In one embodiment, Figure 8 As shown, a point cloud learning feature representation device based on neighborhood geometric embedding is provided, comprising: a multi-scale neighborhood construction module 802, a feature representation learning module 804 and a feature representation fusion module 806, wherein:

[0089] The multi-scale neighborhood construction module 802 is used to obtain 3D point cloud data of the three-dimensional space environment around the vehicle; perform feature extraction on the 3D point cloud data to obtain an initial point feature representation; and construct a multi-scale neighborhood of the target point using a pre-set neighborhood determination strategy;

[0090] A feature representation learning module 804 is configured to learn, in each neighborhood within the multi-scale neighborhood, a response map of each neighboring point from the initial point feature representation according to a neighborhood point feature aggregation model to obtain a multi-scale feature;

[0091] The feature representation fusion module 806 is used to perform feature fusion on multiple multi-scale features to obtain the final point cloud learning feature representation.

[0092] Regarding the specific limitations of the point cloud learning feature representation device based on neighborhood geometry embedding, please refer to the limitations of the point cloud learning feature representation method based on neighborhood geometry embedding above, which will not be repeated here. The various modules in the above-mentioned point cloud learning feature representation device based on neighborhood geometry embedding can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0093] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 9 As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a point cloud learning feature representation method based on neighborhood geometry embedding is implemented. The display screen of the computer device can be a liquid crystal display or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.

[0094] Those skilled in the art will understand that Figure 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0095] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0096] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0097] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A point cloud learning feature representation method based on neighborhood geometry embedding, characterized in that: The method comprises: Acquire 3D point cloud data of the three-dimensional space environment around the vehicle; perform feature extraction on the 3D point cloud data to obtain an initial point feature representation; and construct a multi-scale neighborhood of the target point using a pre-set neighborhood determination strategy; In each neighborhood within the multi-scale neighborhood, learning each neighboring point response map from the initial point feature representation according to a neighborhood point feature aggregation model to obtain a multi-scale feature; Perform feature fusion on multiple multi-scale features to obtain the final point cloud learning feature representation; The neighborhood point feature aggregation model includes a geometric perception module, a channel association module, and a response map generation module. In each neighborhood within the multi-scale neighborhood, each neighborhood point response map is learned from the initial point feature representation according to the neighborhood point feature aggregation model to obtain multi-scale features, including: In the geometric perception module, feature enhancement is performed on the initial point feature representation through a self-attention mechanism to obtain enhanced features; Performing channel association on the enhanced features in the channel association module to obtain key features; The response graph generation module generates a response graph according to key features and retains key point features through Top-K truncation; In the geometric perception module, the initial point feature representation is enhanced through a self-attention mechanism to obtain enhanced features, including: Calculate the query Q, key K, and value V, generate a self-attention map based on the query Q and key K, and obtain the feature map through weighted aggregation ; Based on the self-attention map M, weighted average V is performed to form relevant information S, that is, ; According to the residual principle, the input initial point feature representation is performed Enhance, get the enhanced features, represents the initial point feature representation; The channel association module performs channel association on the enhanced features to obtain key features, including: Given the enhanced feature Γ, take the average of all neighboring points to generate a global representation, that is, ,in represents an average operation; exist φ A fully connected network is used to learn a channel graph s representing the key channels, where , represents a fully connected network parameterized by ν; Based on the channel map, the enhanced features are further updated to obtain the key features .

2. The method according to claim 1, characterized in that Performing feature extraction on the 3D point cloud data to obtain an initial point feature representation includes: Feature extraction is performed on the 3D point cloud data to obtain the initial point feature representation: in, represents the parameters of the MLP network, Represents 3D point cloud data.

3. The method according to claim 1, characterized in that The pre-set neighborhood determination strategy includes a predefined point number strategy, a predefined range strategy and a fusion strategy; the predefined point number strategy is to select a predefined number of points near the target point according to the spatial distance to construct a neighborhood; the predefined range strategy is to give a radius with the target point as the center, and the points inside the sphere constitute the neighborhood; the fusion strategy is to sample adjacent points to a given number within a given sphere, and if there are not enough points inside the sphere, the target point is repeated.

4. The method according to claim 1, wherein The response graph generation module generates a response graph based on key features and retains key point features through Top-K truncation, including: The response graph C generated by MLP is: in, Is a parameter MLP network, represents the key feature, Γ represents the enhanced feature; Top-K response parameters, i.e. , where the truncation function Keep the K largest elements and set the remaining parameters to 0. The key point features are retained through Top-K truncation, that is: in, Indicates a point The feature representation of Indicates the number of points, Indicates the assignment to the point The weight of .

5. The method according to claim 4, characterized in that Perform feature fusion on multiple multi-scale features to obtain the final point cloud learning feature representation, including: Multi-scale features After concatenation, input the fully connected layer, where are connected in series ,in, Regress to an L-dimensional parameter through a fully connected network , p represents the preference, and then these feature representations are fused based on p to obtain the final point cloud learning feature representation: in, represents a fully connected network whose output is normalized to p through a softmax operation, is the first i elements, L Indicates the dimension.

6. A point cloud learning feature representation device based on neighborhood geometry embedding, characterized in that: The device comprises: A multi-scale neighborhood construction module is used to obtain 3D point cloud data of the three-dimensional spatial environment around the vehicle; perform feature extraction on the 3D point cloud data to obtain an initial point feature representation; and construct a multi-scale neighborhood of the target point using a pre-set neighborhood determination strategy; A feature representation learning module is configured to learn, in each neighborhood within the multi-scale neighborhood, a response map of each neighboring point from an initial point feature representation according to a neighborhood point feature aggregation model to obtain multi-scale features; the neighborhood point feature aggregation model includes a geometric perception module, a channel association module, and a response map generation module; and, in each neighborhood within the multi-scale neighborhood, learn, in each neighborhood within the multi-scale neighborhood, a response map of each neighboring point from an initial point feature representation according to the neighborhood point feature aggregation model to obtain multi-scale features, including: In the geometric perception module, feature enhancement is performed on the initial point feature representation through a self-attention mechanism to obtain enhanced features; Performing channel association on the enhanced features in the channel association module to obtain key features; The response graph generation module generates a response graph according to key features and retains key point features through Top-K truncation; In the geometric perception module, the initial point feature representation is enhanced through a self-attention mechanism to obtain enhanced features, including: Calculate the query Q, key K, and value V, generate a self-attention map based on the query Q and key K, and obtain the feature map through weighted aggregation ; Based on the self-attention map M, weighted average V is performed to form relevant information S, that is, ; According to the residual principle, the input initial point feature representation is performed Enhance, get the enhanced features, represents the initial point feature representation; The channel association module performs channel association on the enhanced features to obtain key features, including: Given the enhanced feature Γ, take the average of all neighboring points to generate a global representation, that is, ,in represents an average operation; exist φ A fully connected network is used to learn a channel graph s representing the key channels, where , represents a fully connected network parameterized by ν; Based on the channel map, the enhanced features are further updated to obtain the key features ; The feature representation fusion module is used to fuse multiple multi-scale features to obtain the final point cloud learning feature representation.

Citation Information

Patent Citations

  • Indoor point cloud semantic segmentation method based on self-adaption and multi-scale feature fusion

    CN119131391A

  • Three-dimensional point cloud data segmentation method and system based on multi-scale point features

    CN119648536A