A 3D point cloud segmentation method, device, electronic device, and storage medium
Through the BALR module and cross-layer cross attention network in the BALR-NET model, the edge blur problem caused by the abstract aggregation module does not consider the diversity of neighboring points, and a clearer three-dimensional point cloud segmentation effect is achieved.
Patent Information
- Application Number
- CN202510592634.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-09
AI Technical Summary
In the prior art, the abstract aggregation module does not fully consider the diversity between neighboring points when performing semantic segmentation tasks, resulting in blurred object edges in semantic label prediction.
Using the BALR-NET model, the BALR module and a cross-layer cross-attention network are introduced. Through feature correction and feature comparison and extraction, the edge perception ability of local features is enhanced, the feature isolation state is broken, and the edge segmentation clarity is improved.
It effectively reduces the ambiguity of the object edge in semantic prediction, improves the segmentation accuracy of three-dimensional point cloud data, and performs excellently in object edge segmentation.
Smart Images

Figure CN120107606B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of point cloud segmentation, and in particular, to a three-dimensional point cloud segmentation method, device, electronic device, and storage medium. Background Art
[0002] With the technological breakthroughs of point cloud acquisition devices such as lidar and RGB-D cameras, point cloud processing methods do not need to rely on expensive mesh reconstruction or denoising processes, enabling point cloud data to be more and more widely used in multiple application fields of three-dimensional scene understanding, such as robotics, autonomous driving, urban planning, infrastructure maintenance, and digital cultural heritage protection. In these applications, point cloud semantic segmentation, as the core link for understanding and parsing three-dimensional scenes, serves to divide the original point cloud data into several subsets with different semantic labels, and has thus received great attention and research in recent years. Among them, the segmentation of edges between three-dimensional objects has always been the focus of semantic segmentation tasks, and some researchers have made contributions in this direction in recent years. However, without exception, they need to introduce edge supervision in the model, which undoubtedly greatly increases the annotation work of point cloud data and is not conducive to practical engineering applications. Analyzing the process of point cloud feature extraction, due to the irregularity, sparsity, and disorder of point cloud data, directly operating on point clouds, extracting features, and regressing semantic labels have long been a technical challenge.
[0003] The existing technology has proposed an abstract aggregation module (SA), which is an efficient local feature extraction method. Since the SA module was proposed, it has become a key component for many models to learn local region features and is used in combination with other modules to complete semantic segmentation tasks. However, when the SA module performs max pooling to aggregate all point features to the central representative point, it does not fully consider the possible diversity among neighboring points. For example, if the central representative point is located at the edge of an object, its neighborhood may contain points of neighboring objects. In this case, the SA module directly aggregates the features of these points to the central representative point, which may have a negative impact on semantic label prediction, resulting in the blurring of object edges in semantic prediction.
[0004] Therefore, there is an urgent need to propose a three-dimensional point cloud segmentation method, device, electronic device, and storage medium to solve the technical problem in the existing technology that when the abstract aggregation module performs semantic segmentation tasks, it does not fully consider the possible diversity among neighboring points, which may have a negative impact on semantic label prediction, resulting in the blurring of object edges in semantic prediction. Summary of the Invention
[0005] In view of this, it is necessary to provide a three-dimensional point cloud segmentation method, device, electronic device and storage medium to solve the technical problem that in the prior art, when the abstract aggregation module performs semantic segmentation tasks, it does not fully consider the possible diversity among neighboring points, which may have a negative impact on semantic label prediction, resulting in blurred object edges in semantic prediction.
[0006] To solve the above problems, the present invention provides a three-dimensional point cloud segmentation method, including:
[0007] Obtain three-dimensional point cloud data;
[0008] Extract features from the three-dimensional point cloud data to obtain a first set of feature points;
[0009] Modify the features in the first set of feature points to obtain a second set of feature points;
[0010] Respectively perform feature correction extraction and feature comparison extraction on the second set of feature points to obtain a third set of feature points and a fourth set of feature points;
[0011] Obtain a three-dimensional point cloud segmentation result according to the fifth set of feature points.
[0012] In a possible implementation manner, the modifying the features in the first set of feature points to obtain a second set of feature points includes:
[0013] Sample the first set of feature points to obtain a set of center points;
[0014] Group the first set of feature points with the center points in the set of center points as the centers to obtain neighborhoods;
[0015] Obtain a local feature set according to the features of the neighboring points in the neighborhood;
[0016] Based on the BALR module and the center points, modify the features of the neighboring points in the local feature set to obtain a target local feature set;
[0017] Extract features from the target local feature set to obtain a second set of feature points.
[0018] In a possible implementation manner, after obtaining the local feature set according to the features of the neighboring points in the neighborhood, it further includes:
[0019] Determine whether there are features of neighboring points in the local feature set that are different from the features of the center points;
[0020] If so, based on the BALR module and the center points, modify the features of the neighboring points in the local feature set to obtain a target local feature set.
[0021] In a possible implementation, the domain points in the local feature set include neighborhood point features and neighborhood point spatial coordinates; the central point includes central point features and central point spatial coordinates; the modification of the features of the domain points in the local feature set based on the BALR module and the central point to obtain the target local feature set includes:
[0022] Calculate the feature difference between the neighborhood point features and the central point features;
[0023] Calculate the coordinate difference between the neighborhood point spatial coordinates and the central point spatial coordinates;
[0024] Encode the feature difference and the coordinate difference respectively to obtain the semantic weight and the spatial weight;
[0025] Modify the domain points in the local feature set according to the semantic weight and the spatial weight to obtain the target local feature set.
[0026] In a possible implementation, feature comparison and extraction are performed on the second feature point set to obtain a fourth feature point set, including:
[0027] Feature comparison and extraction are performed on the second feature point set to obtain a fourth feature point set, including:
[0028] Perform feature comparison on the second feature point set based on the cross-layer cross-attention network to obtain feature connections;
[0029] Determine the weight between the second feature point set and the representative point set according to the feature connection;
[0030] Obtain the fourth feature point set according to the feature connection and the weight.
[0031] In a possible implementation, the performing feature comparison on the second feature point set based on the cross-layer cross-attention network to obtain feature connections includes:
[0032] Select the second feature point set based on the farthest point sampling method to obtain a representative point set;
[0033] Process the second feature point set and the representative point set based on the cross-layer cross-attention network to obtain feature connections.
[0034] In a possible implementation, the performing feature extraction on the target local feature set to obtain a second feature point set includes:
[0035] Perform feature extraction on the target local feature set to obtain a domain aggregation feature centered on the central point;
[0036] Obtain the corrected features output by the BALR module according to the domain aggregation features;
[0037] Concatenate the corrected features with the set of spatial coordinates of neighborhood points in the local feature set to obtain a second feature point set.
[0038] On the other hand, the present invention also provides a three-dimensional point cloud segmentation device, including:
[0039] A data acquisition module for acquiring three-dimensional point cloud data;
[0040] A point set sampling module for extracting features from the three-dimensional point cloud data to obtain a first feature point set;
[0041] A first feature extraction module for correcting the features in the first feature point set to obtain a second feature point set;
[0042] A second feature extraction module for respectively performing feature correction extraction and feature comparison extraction on the second feature point set to obtain a third feature point set and a fourth feature point set;
[0043] A segmentation result determination module for obtaining a three-dimensional point cloud segmentation result according to the fifth feature point set.
[0044] On the other hand, an embodiment of the present invention discloses an electronic device, including: a processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, each step of the above-mentioned three-dimensional point cloud segmentation method embodiment is implemented.
[0045] On the other hand, an embodiment of the present invention discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, each step of the above-mentioned three-dimensional point cloud segmentation method embodiment is implemented.
[0046] The beneficial effects of the present invention are as follows: three-dimensional point cloud data is acquired, features are extracted from the three-dimensional point cloud data to obtain a first feature point set; the features in the first feature point set are corrected to obtain a second feature point set; feature correction extraction and feature comparison extraction are respectively performed on the second feature point set to obtain a third feature point set and a fourth feature point set; a three-dimensional point cloud segmentation result is obtained according to the fifth feature point set, so that the three-dimensional point cloud data can be corrected, and feature correction extraction and feature comparison extraction can also be performed, so that the extracted features can be filtered, and the fuzziness of the object edge in semantic prediction is reduced. Description of the Drawings
[0047] Figure 1Schematic diagram of a process of an embodiment of the 3D point cloud segmentation method provided by the present invention;
[0048] Figure 2 Schematic diagram of a structure of an embodiment of the BALR-NET model provided by the present invention;
[0049] Figure 3 For the present invention Figure 1 Schematic diagram of a process of an embodiment of step S103 in the present invention;
[0050] Figure 4 For the present invention Figure 3 Schematic diagram of a process of an embodiment of step S304 in the present invention;
[0051] Figure 5 For the present invention Figure 1 Schematic diagram of a process of an embodiment of step S104 in the present invention;
[0052] Figure 6 Schematic diagram of a structure of an embodiment of the 3D point cloud segmentation device provided by the present invention;
[0053] Figure 7 Schematic diagram of a structure of an embodiment of the electronic device provided by the present invention. Detailed implementation manners
[0054] The following will specifically describe the preferred embodiments of the present invention in conjunction with the accompanying drawings. The accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.
[0055] As Figure 1 shown, a specific embodiment of the present invention discloses a 3D point cloud segmentation method, including:
[0056] S101. Obtain 3D point cloud data;
[0057] S102. Extract features from the 3D point cloud data to obtain a first feature point set;
[0058] S103. Correct the features in the first feature point set to obtain a second feature point set;
[0059] S104. Respectively perform feature correction extraction and feature comparison extraction on the second feature point set to obtain a third feature point set and a fourth feature point set;
[0060] S105. Obtain the 3D point cloud segmentation result according to the fifth feature point set.
[0061] It should be understood that the method for obtaining the three-dimensional point cloud data in step S101 may be to obtain a three-dimensional point cloud data set according to a radar detection device, or to call a historically stored three-dimensional point cloud data set from a storage medium.
[0062] In a specific embodiment of the present invention, a BALR-NET model is provided. The BALR-NET model introduces a BALR module with edge perception function to replace the SA module for local feature extraction. In addition, in order to break through the isolated state of the local features of the last layer of the PointNet++ model encoding and enhance the information exchange between the local features of the last layer, the BALR-NET model designs a cross-layer cross-attention network. This network establishes an association between the feature of the representative point of the next layer and the groups with a relatively large spatial distance from the previous layer that generates this feature. The input of the BALR-NET model is three-dimensional point cloud data , such as Figure 2 shown, the input of the th layer of the encoder of the BALR-NET model is , represents the number of input points, represents the feature dimension of the input points, including 3D spatial information and dimensional semantic information. After farthest point sampling (Farthest Point Sampling, FPS) ( Figure 2 Sample in ), the point set containing points is , where represents the Euclidean space position of point , represents the dimensional feature of the seed point . The output of the th layer of the encoder is , represents the feature dimension of the output points. At the same time, considering that the semantic information gradually enriches with the increase of the network depth, while the fine-grained spatial information gradually disappears, the BALR-NET model uses the BALR module with edge perception ability to perform local feature extraction on the downsampled point cloud in the second and third layers of the encoding. In particular, in order to break the isolated state of the local features of the last layer (the fourth layer) of the PointNet++ model B encoding, the BALR-NET model incorporates a cross-layer cross-attention network into the BALR module in the third layer to perform local feature extraction on the downsampled point cloud.
[0063] The first layer of the encoder of the BALR-NET model is the SA module, and the input is , and the point set after farthest point sampling (FPS) is , the SA module is used to extract features to obtain the first feature point set . Among them, the process of the SA module extracting features from the three-dimensional point cloud data can be set according to the actual situation, and the embodiments of the present invention do not limit this here. The second layer is the BALR module, which processes The point set after farthest point sampling (FPS) is . The BALR module is used to correct the features in it to obtain the second feature point set . The third layer is the BALR module and the cross-layer cross-attention network, which process The point set after farthest point sampling (FPS) is . The BALR module and the cross-layer cross-attention network CLCA (i.e., CLCA-BALR) are respectively used to extract features to obtain the third feature point set and the fourth feature point set, and splice the third feature point set and the fourth feature point set to obtain the fifth feature point set . Specifically, the BALR module can be used to extract features from the second feature point set to obtain the third feature point set, and the cross-layer cross-attention network can also be used to extract features from the second feature point set to obtain the fourth feature point set. Then, the fourth feature point set and the fifth feature point set are spliced and combined to obtain the fifth feature point set . The fourth layer is the SA module, which processes the fifth feature point set The point set after farthest point sampling (FPS) is . The SA module is used to extract features to obtain . The output of the fourth layer . After the feature extraction in the first, second, and third layers, the point sets obtained by feature extraction ( , and ) can be copied and sent to the decoder (Decoder) through skip concatenation. , , and are merged and then decoded by the decoder (Decoder), and then segmented (Segmentation Head), and the three-dimensional point cloud segmentation result can be obtained.
[0064] Compared with the prior art, the present embodiment provides for extracting feature points from three-dimensional point cloud data to obtain a first set of feature points; correcting the features in the first set of feature points to obtain a second set of feature points; respectively extracting features from the second set of feature points based on the BALR module and the cross-layer cross-attention network to obtain a third set of feature points and a fourth set of feature points; obtaining a three-dimensional point cloud segmentation result according to a fifth set of feature points, so that three-dimensional feature extraction can be performed on the three-dimensional point cloud data according to the BALR module and the cross-layer cross-attention network, and the extracted features can be filtered, reducing the fuzziness of object edges in semantic prediction.
[0065] In some embodiments of the present invention, as Figure 3 shown, step S103 includes:
[0066] S301. Sampling the first set of feature points to obtain a set of center points;
[0067] S302. Grouping the first set of feature points with the center points in the set of center points as the centers to obtain neighborhoods;
[0068] S303. Obtaining a local feature set according to the features of the domain points in the neighborhood;
[0069] S304. Correcting the features of the domain points in the local feature set based on the BALR module and the center points to obtain a target local feature set;
[0070] S305. Extracting features from the target local feature set to obtain a second set of feature points.
[0071] In a specific embodiment of the present invention, the input of the second layer is the first set of feature points output by the first layer , and then it can be sampled by farthest point sampling (FPS) for to obtain a set of center points , which includes center points , i denotes the i th center point. Grouping with the point as the center to form neighborhoods, that is, the neighborhoods corresponding to each center point i , and then constructing a local feature set with the features of the domain points in the neighborhood, as shown in formula (1):
[0072] (1)
[0073] In the formula, are respectively the features and spatial coordinates of the th domain point of the center point , is the number of points in the neighborhood.
[0074] The SA module extracts features from all domain points in the local feature set, as shown in formula (2):
[0075] (2)
[0076] In the formula, represents a shared multi-layer perceptron, represents max pooling, is the local domain feature centered on
[0077] In the SA module, if the center point is located on the boundary between objects, then the set constructed with the center point may contain features of points with different semantics from These features will directly affect resulting in blurred edge segmentation. To solve this problem, the BALR-NET model proposes the BALR module to adjust . Specifically, when the semantics of the th domain point is different from that of the center point, the BALR module will correct and update its feature to obtain the target local feature set. Then, the target local feature set is brought into formula (2) for feature extraction, and the second feature point set can be obtained.
[0078] In some embodiments of the present invention, after step S303, it further includes:
[0079] Determine whether there are features of domain points in the local feature set that are different from the features of the center point;
[0080] If so, correct the features of the domain points in the local feature set based on the BALR module and the center point to obtain the target local feature set.
[0081] In a specific embodiment of the present invention, after obtaining the local feature set, it can be determined whether there are features of domain points in the local feature set that are different from the features of the center point; if so, step S304 can be performed; if not, it means that the difference between the features of the domain points and the features of the center point is relatively small, and step S304 is also performed, but the result after correction is not much different from that before correction, approaching no correction.
[0082] In some embodiments of the present invention, the neighborhood points in the local feature set include neighborhood point features and neighborhood point spatial coordinates; the center point includes center point features and center point spatial coordinates; asFigure 4 As shown in Figure 4 , step S304 includes:
[0083] S401. Calculate the feature difference between the neighborhood point features and the center point features;
[0084] S402. Calculate the coordinate difference between the neighborhood point spatial coordinates and the center point spatial coordinates;
[0085] S403. Encode the feature difference and the coordinate difference respectively to obtain the semantic weight and the spatial weight;
[0086] S404. Correct the neighborhood points in the local feature set according to the semantic weight and the spatial weight to obtain the target local feature set.
[0087] In a specific embodiment of the present invention, for each neighborhood point in the local feature set the BALR module first calculates the feature difference between the neighborhood point feature and the center point feature as well as the coordinate difference between the neighborhood point spatial coordinate and the center point spatial coordinate of the center point as shown in formula (3) and (4): of the coordinate difference , and encodes the feature difference and the coordinate difference respectively to obtain the semantic weight and the spatial weight of each neighborhood point
[0088] (3)
[0089] (4)
[0090] In the formula, is a non - linear mapping function that maps the value of to the range of 0 to 1, is the convolution function, is the batch normalization, represents the non - linear activation function. , The values of are all between 0 and 1. The greater the feature difference, the closer is to 1, and the closer the distance, the closer
[0091] Then calculate the corrected neighborhood point feature as shown in formula (5):
[0092] (5)
[0093] In the formula, represents matrix multiplication, , update the center point S i the domain set of, to obtain the target local feature set .
[0094] In some embodiments of the present invention, step S305 includes:
[0095] Extract features from the target local feature set to obtain domain aggregation features centered on the center point;
[0096] According to the domain aggregation features, obtain the corrected features output by the BALR module;
[0097] Concatenate the corrected features with the set of neighborhood point spatial coordinates in the local feature set to obtain the second feature point set.
[0098] In a specific embodiment of the present invention, the corrected features in the target local feature set can be substituted into formula (2) to extract domain features, and domain aggregation features centered on the center point are obtained, then the BALR module can output corrected features according to the domain aggregation features , and then the corrected features are concatenated with the set of neighborhood point spatial coordinates in the local feature set, and the second feature point set output by the second layer can be obtained , where the domain aggregation features are shown in formula (6):
[0099] (6)
[0100] In some embodiments of the present invention, as Figure 5 shown, step S104 includes:
[0101] S501. Perform feature comparison on the second feature point set based on the cross-layer cross-attention network to obtain feature connections;
[0102] S502. Determine the weights between the second feature point set and the representative point set according to the feature connections;
[0103] S503. Obtain the fourth feature point set according to the feature connections and weights.
[0104] In some embodiments of the present invention, performing feature comparison on the second feature point set based on the cross-layer cross-attention network to obtain feature connections includes:
[0105] Select the second feature point set based on the farthest point sampling method to obtain a representative point set;
[0106] Process the second feature point set and the representative point set based on the cross-layer cross-attention network to obtain feature connections.
[0107] In a specific embodiment of the present invention, in order to expand the receptive field and strengthen the information exchange between all local regions, the BALR-NET model incorporates a cross-layer cross-attention network into the BALR module of the third layer. This network establishes the connection between the representative point features of the fourth layer and the features of the third layer, and then embeds the representative point features of the fourth layer into the center point features of the third layer local domain, so that when the fourth layer performs local feature extraction, it not only aggregates the point features of this domain, but also obtains the features of other domains, thereby strengthening the information exchange between local regions. Assume that the representative point set after sampling in the third layer is , the point set input to the fourth layer , and the representative point set after sampling is . The specific calculation process of the cross-layer cross-attention network is as follows.
[0108] First, in the third layer , use the farthest point sampling to select the representative point set , , the representative point set is the same as the fourth layer representative point set in the Euclidean space. Then, establish the connection between the point set and the points in the third layer to obtain the feature connection ( , , ). , , are calculated as shown in formula (7):
[0109] (7)
[0110] Among them, represents the linear layer, represents 's semantic feature, represents 's semantic feature.
[0111] Then calculate the weight between the fourth layer representative point and the second feature point set of the third layer as shown in formula (8):
[0112] (8)
[0113] In the formula, represents the summation function, represents matrix multiplication, 。
[0114] Then, according to the feature connection and weight, the established representative point set can be calculated The fourth feature point set after the feature connection with the second feature point set of the third layer, as shown in Formula (9):
[0115] (9)
[0116] In the formula, represents the convolution operation, represents the addition, is the fifth feature output by the cross-layer cross-attention network
[0117] Finally, the fourth feature point set is mixed with the third feature point set of the BALR module of the third layer to obtain the feature output by the CLCA-BALR module of the third layer , concatenate and the coordinates of the representative points of the third layer to obtain the fifth feature point set output by the third layer encoding , as shown in Formula (10):
[0118] (10)
[0119] In the formula, represents the convolution operation, represents the concatenation operation, is the output of the BALR module
[0120] The embodiment of the present invention proposes a novel point cloud segmentation model (BALR-NET), which shows excellent performance in semantic segmentation and component segmentation tasks, especially in dealing with the edge segmentation between objects, effectively improving the clarity of edge segmentation. Aiming at the problem of blurred edge segmentation between objects in semantic prediction, an edge perception mechanism based on the semantic distance and Euclidean distance between neighboring points is innovatively introduced, and an edge perception local feature expression module (BALR) is proposed to enhance the feature representation of edge points. Aiming at the problem of feature isolation in the feature encoding process of the traditional abstract aggregation module (SA), a cross-layer cross-attention network (CLCA) with the representative points of the next layer as the query is designed. This network breaks the isolation state between features by establishing long-range dependencies between the upper and lower layers
[0121] To better implement the three-dimensional point cloud segmentation method in the embodiment of the present invention, correspondingly, the embodiment of the present invention also provides a three-dimensional point cloud segmentation device, as Figure 6 shown, the three-dimensional point cloud segmentation device 600 includes:
[0122] A data acquisition module 601 for acquiring three-dimensional point cloud data;
[0123] A point set sampling module 602 for extracting features from the three-dimensional point cloud data to obtain a first feature point set;
[0124] A first feature extraction module 603 for correcting the features in the first feature point set to obtain a second feature point set;
[0125] A second feature extraction module 604 for respectively performing feature correction extraction and feature comparison extraction on the second feature point set to obtain a third feature point set and a fourth feature point set;
[0126] A segmentation result determination module 605 for obtaining a three-dimensional point cloud segmentation result according to a fifth feature point set.
[0127] The three-dimensional point cloud segmentation device 600 provided in the above embodiment can implement the technical solutions described in the above three-dimensional point cloud segmentation method embodiment. The specific implementation principles of the above modules or units can be referred to the corresponding content in the above three-dimensional point cloud segmentation method embodiment, which will not be elaborated here.
[0128] As Figure 7 shown, the present invention also correspondingly provides an electronic device 700. The electronic device 700 includes a processor 701, a memory 702, and a display 703. Figure 7 Only some components of the electronic device 700 are shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented.
[0129] The memory 702 can be an internal storage unit of the electronic device 700 in some embodiments, such as the hard disk or memory of the electronic device 700. The memory 702 can also be an external storage device of the electronic device 700 in other embodiments, such as a plug-in hard disk equipped on the electronic device 700, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.
[0130] Furthermore, the memory 702 can also include both the internal storage unit and the external storage device of the electronic device 700. The memory 702 is used to store the application software installed on the electronic device 700 and various types of data.
[0131] The processor 701 can be a Central Processing Unit (CPU), a microprocessor, or other data processing chips in some embodiments, for running the program code stored in the memory 702 or processing data, such as the three-dimensional point cloud segmentation method in the present invention.
[0132] In some embodiments, the display 703 can be an LED display, a liquid crystal display, a touch liquid crystal display, an OLED (Organic Light-Emitting Diode) touch device, etc. The display 703 is used to display information of the electronic device 700 and to display a visual user interface. The components 701-703 of the electronic device 700 communicate with each other via a system bus.
[0133] In some embodiments of the present invention, when the processor 701 executes the 3D point cloud segmentation program in the memory 702, the following steps can be achieved:
[0134] Obtain 3D point cloud data;
[0135] Extract features from the 3D point cloud data to obtain a first set of feature points;
[0136] Correct the features in the first set of feature points to obtain a second set of feature points;
[0137] Respectively perform feature correction extraction and feature comparison extraction on the second set of feature points to obtain a third set of feature points and a fourth set of feature points;
[0138] Obtain the 3D point cloud segmentation result according to the fifth set of feature points.
[0139] It should be understood that when the processor 701 executes the 3D point cloud segmentation program in the memory 702, in addition to the above functions, other functions can also be achieved. For details, please refer to the description of the corresponding method embodiments above.
[0140] Furthermore, the type of the electronic device 700 mentioned in the embodiments of the present invention is not specifically limited. The electronic device 700 can be a mobile phone, a tablet computer, a personal digital assistant (PDA), a wearable device, a laptop computer, etc. Exemplary embodiments of the portable electronic device include, but are not limited to, portable electronic devices equipped with IOS, android, microsoft, or other operating systems. The above portable electronic devices can also be other portable electronic devices, such as a laptop computer with a touch-sensitive surface (e.g., a touch panel). It should also be understood that in some other embodiments of the present invention, the electronic device 700 can also not be a portable electronic device, but a desktop computer with a touch-sensitive surface (e.g., a touch panel).
[0141] Accordingly, an embodiment of the present application further provides a computer-readable storage medium, which is used to store computer-readable programs or instructions. When the programs or instructions are executed by a processor, the method steps or functions of the 3D point cloud segmentation method provided in the above method embodiments can be implemented.
[0142] Those skilled in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware (such as a processor, a controller, etc.) through a computer program, and the computer program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.
[0143] The 3D point cloud segmentation method and device provided by the present invention have been introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A three-dimensional point cloud segmentation method, characterized in that Including: Obtaining three-dimensional point cloud data; Performing feature extraction on the three-dimensional point cloud data to obtain a first set of feature points; Correcting the features in the first set of feature points to obtain a second set of feature points; Respectively performing feature correction extraction and feature comparison extraction on the second set of feature points to obtain a third set of feature points and a fourth set of feature points, and splicing the third set of feature points and the fourth set of feature points to obtain a fifth set of feature points; Obtaining a three-dimensional point cloud segmentation result according to the fifth set of feature points; The correcting the features in the first set of feature points to obtain a second set of feature points includes: Sampling the first set of feature points to obtain a set of center points; Grouping the first set of feature points with the center points in the set of center points as the center to obtain neighborhoods; Obtaining a local feature set according to the features of the neighborhood points in the neighborhood; Based on the BALR module and the center point, correcting the features of the neighborhood points in the local feature set to obtain a target local feature set; Performing feature extraction on the target local feature set to obtain a second set of feature points; The neighborhood points in the local feature set include neighborhood point features and neighborhood point spatial coordinates; the center point includes center point features and center point spatial coordinates; the correcting the features of the neighborhood points in the local feature set based on the BALR module and the center point to obtain a target local feature set includes: Calculating the feature difference between the neighborhood point features and the center point features; Calculating the coordinate difference between the neighborhood point spatial coordinates and the center point spatial coordinates; Encoding the feature difference and the coordinate difference respectively to obtain a semantic weight and a spatial weight; Correcting the neighborhood points in the local feature set according to the semantic weight and the spatial weight to obtain a target local feature set.
2. The three-dimensional point cloud segmentation method according to claim 1, characterized in that After obtaining the local feature set according to the features of the neighborhood points in the neighborhood, it further includes: Judging whether there are neighborhood point features in the local feature set that are different from the center point features; If so, based on the BALR module and the center point, correcting the features of the neighborhood points in the local feature set to obtain a target local feature set.
3. The three-dimensional point cloud segmentation method according to claim 1, wherein Performing feature comparison extraction on the second set of feature points to obtain a fourth set of feature points, including: Performing feature comparison on the second set of feature points based on a cross-layer cross-attention network to obtain feature connections; Determining the weights between the second set of feature points and the representative point set according to the feature connections; Obtaining a fourth set of feature points according to the feature connections and the weights.
4. The 3D point cloud segmentation method according to claim 3, wherein The performing feature comparison on the second set of feature points based on a cross-layer cross-attention network to obtain feature connections includes: Selecting the second set of feature points based on the farthest point sampling method to obtain a representative point set; Processing the second set of feature points and the representative point set based on a cross-layer cross-attention network to obtain feature connections.
5. The three-dimensional point cloud segmentation method according to claim 1, wherein The performing feature extraction on the target local feature set to obtain a second set of feature points includes: Performing feature extraction on the target local feature set to obtain a neighborhood aggregation feature centered on the center point; Obtain the corrected features output by the BALR module according to the domain aggregation features; Concatenate the corrected features with the set of neighborhood point spatial coordinates in the local feature set to obtain a second feature point set.
6. A three-dimensional point cloud segmentation device, characterized in that, Including: A data acquisition module for acquiring three-dimensional point cloud data; A point set sampling module for extracting features from the three-dimensional point cloud data to obtain a first feature point set; A first feature extraction module for correcting the features in the first feature point set to obtain a second feature point set; A second feature extraction module for respectively performing feature correction extraction and feature comparison extraction on the second feature point set to obtain a third feature point set and a fourth feature point set, and concatenating the third feature point set and the fourth feature point set to obtain a fifth feature point set; A segmentation result determination module for obtaining a three-dimensional point cloud segmentation result according to the fifth feature point set; The step of correcting the features in the first feature point set to obtain a second feature point set includes: Sampling the first feature point set to obtain a center point set; Grouping the first feature point set with the center point in the center point set as the center to obtain neighborhoods; Obtain a local feature set according to the features of the neighborhood points in the neighborhood; Based on the BALR module and the center point, correct the features of the neighborhood points in the local feature set to obtain a target local feature set; Extract features from the target local feature set to obtain a second feature point set; The neighborhood points in the local feature set include neighborhood point features and neighborhood point spatial coordinates; the center point includes center point features and center point spatial coordinates; the step of correcting the features of the neighborhood points in the local feature set based on the BALR module and the center point to obtain a target local feature set includes: Calculate the feature difference between the neighborhood point features and the center point features; Calculate the coordinate difference between the neighborhood point spatial coordinates and the center point spatial coordinates; Encode the feature difference and the coordinate difference respectively to obtain a semantic weight and a spatial weight; Correct the neighborhood points in the local feature set according to the semantic weight and the spatial weight to obtain a target local feature set.
7. An electronic device, characterized in that, Including: A processor, a memory, and a computer program stored on the memory and capable of running on the processor. When the computer program is executed by the processor, the steps of the three-dimensional point cloud segmentation method according to any one of claims 1-5 are implemented.
8. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium. When the computer program is executed by the processor, the steps of the three-dimensional point cloud segmentation method according to any one of claims 1-5 are implemented.
Citation Information
Patent Citations
Point cloud segmentation method based on feature deviation value and attention mechanism
CN116958956A