Feature Fusion Point Cloud Salient Object Detection Method, Device, Equipment and Medium
Through the point cloud significance object detection method of feature fusion, the multi-level feature extraction and fusion of point cloud maps is used to solve the problem of difficulty in significance object detection in point cloud scenarios, and accurate significance object recognition and data set production are achieved.
Patent Information
- Application Number
- CN202210254638.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-15
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-03-15
AI Technical Summary
In a point cloud scenario, whether the same object is significantly changed due to the change of observation point points, resulting in the problem of not determining the significance target of the point cloud scenario and the production of the point cloud significance target data set.
A point cloud significance object detection method for feature fusion is proposed. Through the object detection model, including feature fusion module, point cloud perception module and significance perception module, the point cloud map is acquired and multi-level feature extraction is performed, feature fusion and prediction is generated, and the prediction results of the point cloud map are generated.
The goal of accurately determining the significance of point cloud map is achieved, the difficulty of significance target detection in point cloud scenarios is solved, and the production efficiency of point cloud significance target data set is improved.
Smart Images

Figure CN114612652B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and in particular, to a method, device, equipment and medium for point cloud saliency object detection with feature fusion. Background Art
[0002] Visual saliency refers to the local area in the field of view that can attract human visual attention when a human observes a certain area, and this local area is called the salient area. Saliency detection is mainly used to highlight the salient areas in images or videos. Generally speaking, saliency detection is widely used in fields such as image segmentation, object detection and recognition, image retrieval, and image compression. Conducting relevant research work has very important practical significance.
[0003] At present, there is still a certain gap in the work of point cloud saliency detection. Compared with traditional saliency object detection, there is a "saliency conflict" in the point cloud scenario, that is, due to the influence of factors such as occlusion, the saliency of the same object changes due to the change of the observation viewpoint.
[0004] Due to the change of the observation viewpoint, the saliency of the same object changes, resulting in problems such as the inability to determine the saliency object in the point cloud scenario and the inability to create a point cloud saliency object dataset. Summary of the Invention
[0005] The main purpose of the present invention is to propose a method, device, equipment and medium for point cloud saliency object detection with feature fusion, aiming to accurately determine the saliency object of the point cloud.
[0006] To achieve the above object, the present invention provides a method for point cloud saliency object detection with feature fusion. The method for point cloud saliency object detection with feature fusion is applied to an object detection model. The object detection model includes a feature fusion module, a point cloud perception module, and a saliency perception module. The method for point cloud saliency object detection with feature fusion includes the following steps:
[0007] Obtain a point cloud map, and perform multi-level feature extraction on the point cloud map through a preset backbone network model to obtain multi-layer features corresponding to different sampling scales of the point cloud map;
[0008] Based on the topmost feature of the multi-layer features, perform preliminary fusion features through the point cloud perception module to obtain semantic features corresponding to the topmost feature;
[0009] Based on the multi-layer features, perform feature aggregation through the feature fusion module to obtain compact features corresponding to the multi-layer features;
[0010] Based on the compact features, perform hierarchical fusion through the point cloud perception module to obtain multi-scale features corresponding to the compact features;
[0011] Based on the semantic features and the multi-scale features, perform prediction through the saliency perception module to generate a prediction result corresponding to the point cloud map.
[0012] Preferably, before the step of obtaining the point cloud map and performing feature extraction on the point cloud map through a preset backbone network model to obtain multi-layer features corresponding to the point cloud map, the point cloud saliency object detection method based on feature fusion further includes:
[0013] Obtain sample point cloud maps of multiple different scenarios, and construct a training set for training the initial model with the sample point cloud maps;
[0014] Obtain a saliency point cloud map of the marked significant object corresponding to each sample point cloud map;
[0015] Use the sample point cloud maps in the training set as the input of the initial model, use the saliency point cloud map as the output of the initial model, and perform iterative training on the initialized model to obtain the target detection model.
[0016] Preferably, the step of obtaining the point cloud map and performing multi-level feature extraction on the point cloud map through a preset backbone network model to obtain multi-layer features of different sampling scales corresponding to the point cloud map includes:
[0017] Obtain the point cloud map and input the point cloud map into a preset backbone network model;
[0018] Perform multi-level feature extraction on the point cloud map through the preset backbone network model to obtain the first-layer feature, second-layer feature, third-layer feature, and fourth-layer feature of different sampling scales corresponding to the point cloud map.
[0019] Preferably, the multi-layer features include the first-layer feature, second-layer feature, third-layer feature, and fourth-layer feature, the fourth-layer feature is the topmost feature of the multi-layer features, and the step of obtaining semantic features corresponding to the topmost feature by performing preliminary fusion features on the topmost feature of the multi-layer features through the point cloud perception module includes:
[0020] Input the fourth-layer feature into the point cloud perception module, and branch the fourth-layer feature to obtain the branched fourth-layer feature;
[0021] Find the first nearest neighbor points of each point of the branched fourth-layer feature according to the preset nearest neighbor algorithm;
[0022] Obtain the first coordinate information of each point of each fourth - layer feature and each first nearest - neighbor point, and calculate the first relative position between each point of the fourth - layer feature and the first nearest - neighbor point through the first coordinate information;
[0023] According to a preset splicing formula, fuse and splice the first nearest - neighbor point and the first relative position to obtain the first local area corresponding to the fourth - layer feature;
[0024] According to a preset fusion formula, preliminarily fuse the fourth - layer feature and the first local area to obtain the semantic feature corresponding to the fourth - layer feature.
[0025] Preferably, the multi - layer features include the first - layer feature, the second - layer feature, the third - layer feature, and the fourth - layer feature. The step of aggregating features through the feature fusion module based on the multi - layer features to obtain the compact feature corresponding to the multi - layer features includes:
[0026] Input the first - layer feature, the second - layer feature, the third - layer feature, and the fourth - layer feature into the feature fusion module;
[0027] Upsample the fourth - layer feature, and splice and aggregate the upsampled fourth - layer feature and the third - layer feature through a preset multi - layer perceptron to obtain a third - layer splicing result;
[0028] Upsample the third - layer splicing result, and splice and aggregate the upsampled third - layer splicing result and the second - layer feature through a preset multi - layer perceptron to obtain a second - layer splicing result;
[0029] Upsample the second - layer splicing result, and splice and aggregate the upsampled second - layer splicing result and the first - layer feature through a preset multi - layer perceptron to obtain the compact feature corresponding to the first - layer feature, the second - layer feature, the third - layer feature, and the fourth - layer feature.
[0030] Preferably, the step of performing hierarchical fusion through the point - cloud perception module based on the compact feature to obtain the multi - scale feature corresponding to the compact feature includes:
[0031] Input the compact feature into the point - cloud perception module, and branch the compact feature to obtain a branched compact feature;
[0032] Find the second nearest - neighbor points of each point of the branched compact feature according to a preset nearest - neighbor algorithm;
[0033] Obtain the second coordinate information of each point of each compact feature and each second nearest - neighbor point, and calculate the second relative position between each point of the compact feature and the second nearest - neighbor point through the second coordinate information;
[0034] According to a preset splicing formula, the second nearest neighbor point and the second relative position are fused and spliced to obtain a second local region corresponding to the compact feature;
[0035] According to a preset fusion formula, the compact feature and the second local region are preliminarily fused to obtain a multi-scale feature corresponding to the compact feature.
[0036] Preferably, the step of predicting a prediction result corresponding to the point cloud map through the saliency perception module based on the semantic feature and the multi-scale feature includes:
[0037] Input the semantic feature and the multi-scale feature into the saliency perception module;
[0038] Upsample the semantic feature and process the upsampled semantic feature through a preset multi-layer perceptron to obtain a processed semantic feature;
[0039] According to a preset splicing algorithm, the processed semantic feature and the multi-scale feature are fused and spliced to obtain a splicing result corresponding to the processed semantic feature and the multi-scale feature;
[0040] Predict the splicing result through a preset multi-layer perceptron to obtain a prediction result corresponding to the splicing result.
[0041] In addition, to achieve the above object, the present invention also provides a point cloud saliency object detection device for feature fusion, and the point cloud saliency object detection device for feature fusion includes:
[0042] An acquisition module, configured to acquire a point cloud map and perform multi-level feature extraction on the point cloud map through a preset backbone network model to obtain multi-layer features of different sampling scales corresponding to the point cloud map;
[0043] A fusion module, configured to perform preliminary fusion of features through the point cloud perception module based on the topmost feature of the multi-layer features to obtain a semantic feature corresponding to the topmost feature;
[0044] An aggregation module, configured to perform feature aggregation through the feature fusion module based on the multi-layer features to obtain a compact feature corresponding to the multi-layer features;
[0045] A grading module, configured to perform hierarchical fusion through the point cloud perception module based on the compact feature to obtain a multi-scale feature corresponding to the compact feature;
[0046] A prediction module, configured to perform prediction through the saliency perception module based on the semantic features and the multi-scale features, and generate a prediction result corresponding to the point cloud map.
[0047] In addition, to achieve the above object, the present invention further provides a device, which is a feature fusion point cloud saliency object detection device. The feature fusion point cloud saliency object detection device includes: a memory, a processor, and a feature fusion point cloud saliency object detection program stored on the memory and executable on the processor. When the feature fusion point cloud saliency object detection program is executed by the processor, the steps of the feature fusion point cloud saliency object detection method described above are implemented.
[0048] In addition, to achieve the above object, the present invention further provides a medium, which is a computer-readable storage medium. A feature fusion point cloud saliency object detection program is stored on the computer-readable storage medium. When the feature fusion point cloud saliency object detection program is executed by a processor, the steps of the feature fusion point cloud saliency object detection method described above are implemented.
[0049] The present invention provides a feature fusion point cloud saliency object detection method, device, equipment and medium. The feature fusion point cloud saliency object detection method is applied to an object detection model, and the object detection model includes a feature fusion module, a point cloud perception module and a saliency perception module. The feature fusion point cloud saliency object detection method includes: obtaining a point cloud map, and performing multi-level feature extraction on the point cloud map through a preset backbone network model to obtain multi-layer features of different sampling scales corresponding to the point cloud map; based on the topmost feature of the multi-layer features, performing preliminary fusion features through the point cloud perception module to obtain semantic features corresponding to the topmost feature; based on the multi-layer features, performing feature aggregation through the feature fusion module to obtain compact features corresponding to the multi-layer features; based on the compact features, performing hierarchical fusion through the point cloud perception module to obtain multi-scale features corresponding to the compact features; based on the semantic features and the multi-scale features, performing prediction through the saliency perception module to generate a prediction result corresponding to the point cloud map. The present invention obtains a point cloud map, performs multi-level feature extraction on the point cloud map through a preset backbone network to obtain multi-layer features of different sampling scales; inputs the topmost feature of the multi-layer features into the point cloud perception module for preliminary fusion to obtain corresponding semantic features; inputs the multi-layer features into the feature fusion module for feature aggregation to obtain corresponding compact features; inputs the compact features into the point cloud perception module for hierarchical fusion to obtain corresponding multi-scale features; inputs the semantic features and the multi-scale features into the saliency perception module for prediction to obtain the prediction result of the point cloud map, so as to accurately determine the salient object of the point cloud map. Description of the Drawings
[0050] Figure 1 It is a schematic diagram of the device structure of the hardware operating environment involved in the solution of the embodiment of the present invention;
[0051] Figure 2 It is a schematic flowchart of the first embodiment of the point cloud saliency object detection method for feature fusion of the present invention;
[0052] Figure 3 It is a schematic diagram of the structure of the overall model of the first embodiment of the point cloud saliency object detection method for feature fusion of the present invention;
[0053] Figure 4 It is a schematic flowchart of the sub - process of the first embodiment of the point cloud saliency object detection method for feature fusion of the present invention;
[0054] Figure 5 It is a schematic diagram of the structure of the object detection model of the first embodiment of the point cloud saliency object detection method for feature fusion of the present invention;
[0055] Figure 6 It is a schematic flowchart of the second embodiment of the point cloud saliency object detection method for feature fusion of the present invention;
[0056] Figure 7 It is a schematic flowchart of the third embodiment of the point cloud saliency object detection method for feature fusion of the present invention;
[0057] Figure 8 It is a schematic flowchart of the fourth embodiment of the point cloud saliency object detection method for feature fusion of the present invention;
[0058] Figure 9 It is a schematic flowchart of the fifth embodiment of the point cloud saliency object detection method for feature fusion of the present invention;
[0059] Figure 10 It is a schematic flowchart of the sixth embodiment of the point cloud saliency object detection method for feature fusion of the present invention;
[0060] Figure 11 It is a schematic diagram of the functional modules of the first embodiment of the point cloud saliency object detection method for feature fusion of the present invention.
[0061] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0062] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0063] As Figure 1 shown, Figure 1It is a schematic diagram of the device structure of the hardware operating environment involved in the embodiment solution of the present invention.
[0064] The device in the embodiment of the present invention can be a mobile terminal or a server device.
[0065] As Figure 1 shown, the device may include: a processor 1001, such as a CPU, a network interface 1004, a user interface 1003, a memory 1005, and a communication bus 1002. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display) and an input unit such as a keyboard (Keyboard). Optionally, the user interface 1003 may further include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 1005 may be a high-speed RAM memory or a stable memory (non-volatile memory), such as a disk memory. The memory 1005 may optionally be a storage device independent of the aforementioned processor 1001.
[0066] Those skilled in the art can understand that Figure 1 the device structure shown in
[0067] does not constitute a limitation on the device, and may include more or fewer components than shown in the figure, or combine certain components, or have a different component layout. Figure 1 As
[0068] shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a feature fusion point cloud saliency target detection program.
[0069] Among them, the operating system is a program for managing and controlling the feature fusion point cloud saliency target detection device and software resources, and supports the operation of the network communication module, the user interface module, the feature fusion point cloud saliency target detection program, and other programs or software; the network communication module is used to manage and control the network interface 1002; the user interface module is used to manage and control the user interface 1003. Figure 1
[0069]
[0070] Based on the above hardware structure, an embodiment of the feature fusion point cloud saliency target detection method of the present invention is proposed.
[0071] Reference Figure 2 , Figure 2 is a schematic flowchart of the first embodiment of the point cloud saliency target detection method for feature fusion of the present invention. The point cloud saliency target detection method for feature fusion includes:
[0072] Step S10, obtain a point cloud map, and perform multi-level feature extraction on the point cloud map through a preset backbone network model to obtain multi-level features of different sampling scales corresponding to the point cloud map;
[0073] Step S20, based on the topmost feature of the multi-level features, perform preliminary fusion of features through the point cloud perception module to obtain semantic features corresponding to the topmost feature;
[0074] Step S30, based on the multi-level features, perform feature aggregation through the feature fusion module to obtain compact features corresponding to the multi-level features;
[0075] Step S40, based on the compact features, perform hierarchical fusion through the point cloud perception module to obtain multi-scale features corresponding to the compact features;
[0076] Step S50, based on the semantic features and the multi-scale features, perform prediction through the saliency perception module to generate a prediction result corresponding to the point cloud map.
[0077] In this embodiment, a point cloud map is obtained, and multi-level feature extraction is performed on the point cloud map through a preset backbone network to obtain multi-level features of different sampling scales; the topmost feature in the multi-level features is input into the point cloud perception module for preliminary fusion to obtain corresponding semantic features; the multi-level features are input into the feature fusion module for feature aggregation to obtain corresponding compact features; the compact features are input into the point cloud perception module for hierarchical fusion to obtain corresponding multi-scale features; the semantic features and the multi-scale features are input into the saliency perception module for prediction to obtain the prediction result of the point cloud map; thereby accurately determining the saliency target of the point cloud map.
[0078] The following will explain each step in detail:
[0079] Step S10, obtain a point cloud map, and perform multi-level feature extraction on the point cloud map through a preset backbone network model to obtain multi-level features of different sampling scales corresponding to the point cloud map.
[0080] In this embodiment, the point cloud saliency object detection method with feature fusion can be applied to an object detection model, and the object detection model includes a feature fusion module FAB, a point cloud perception module PPB, and a saliency perception module SPB. Point cloud maps are obtained from different channels. Among them, the point cloud map can be a point cloud map of any scene, which can be a point cloud map containing people, or a point cloud map without people, such as a point cloud map of plants or animals. This embodiment does not limit the channels for obtaining the point cloud map.
[0081] Among them, the preset backbone network model includes but is not limited to the PointNet network model, the PointNet++ network model, etc. In this embodiment, the PointNet++ network model is preferably used as the preset backbone network model. Refer to Figure 3 , Figure 3 which is a schematic structural diagram of the overall model of the point cloud saliency object detection method.
[0082] The PointNet++ network model outputs four different levels of features, namely the first-level feature F1, the second-level feature F2, the third-level feature F3, and the fourth-level feature F4. As the level gradually increases, the number of points retained in each layer is 1 / 4 of the previous layer, and the dimension of the feature doubles. For example, the number of points in the second-level feature F2 is 1 / 4 of the first-level feature F1, and the dimension of the second-level feature F2 is twice that of the first-level feature F1.
[0083] The obtained point cloud map is input into the PointNet++ network model, and the PointNet++ network model performs multi-level feature extraction on the point cloud map to obtain multi-level features of different sampling scales corresponding to the point cloud map.
[0084] Further, in one embodiment, refer to Figure 4 , step S10 includes:
[0085] Step S11, obtain a point cloud map and input the point cloud map into the preset backbone network model.
[0086] In this embodiment, the point cloud map is obtained from different channels, the obtained point cloud map is input into the PointNet++ network model, and the PointNet++ network model performs multi-level feature extraction on the point cloud map.
[0087] Step S12, perform multi-level feature extraction on the point cloud map through the preset backbone network model to obtain the first-level feature, the second-level feature, the third-level feature, and the fourth-level feature of different sampling scales corresponding to the point cloud map.
[0088] In this embodiment, the PointNet++ network model is used to perform multi-level feature extraction on the point cloud map, that is, downsampling by 4, 16, 64, and 128 times respectively is achieved, and multi-level features corresponding to different sampling scales of the point cloud map are obtained; among them, the multi-level features include the first-level feature F1, the second-level feature F2, the third-level feature F3, and the fourth-level feature F4, and the fourth-level feature F4 is also called the top-level feature of the multi-level features; among them, the setting of the number of sampling scales and the specific sampling scale values can be selected according to needs, and no specific limitation is made in this embodiment.
[0089] Step S20, based on the top-level feature of the multi-level features, perform preliminary fusion of features through the point cloud perception module to obtain the semantic feature corresponding to the top-level feature.
[0090] In this embodiment, among them, the multi-level features include the first-level feature F1, the second-level feature F2, the third-level feature F3, and the fourth-level feature F4, and the fourth-level feature F4 is also called the top-level feature of the multi-level features. That is, the fourth-level feature F4 is input into the point cloud perception module PPB in the target detection model, and the point cloud perception module PPB performs preliminary fusion of features on the fourth-level feature F4 to obtain the semantic feature Fs corresponding to the fourth-level feature F4. Among them, the semantic feature Fs contains rich semantic information, which is very beneficial for accurately locating the salient objects in the point cloud map.
[0091] Step S30, based on the multi-level features, perform feature aggregation through the feature fusion module to obtain the compact feature corresponding to the multi-level features.
[0092] In one embodiment, refer to Figure 5 , input the first-level feature F1, the second-level feature F2, the third-level feature F3, and the fourth-level feature F4 into the feature fusion module FAB in the target detection model, and perform feature aggregation on the first-level feature F1, the second-level feature F2, the third-level feature F3, and the fourth-level feature F4 through the feature fusion module FAB; then use a preset multi-layer perceptron to splice the aggregated result Fc to obtain the compact feature Fa corresponding to the first-level feature F1, the second-level feature F2, the third-level feature F3, and the fourth-level feature F4.
[0093] Among them, the feature fusion module FAB uses common trilinear interpolation as the upsampling operation, and the aggregation operation can be addition, multiplication, or concatenation, and then a preset multi-layer perceptron is used to splice the aggregated result.
[0094] Among them, it is preferable that the MLP (Multilayer Perceptron) is a preset multilayer perceptron. The MLP is also called an artificial neural network (ANN, Artificial Neural Network). Except for the input and output layers, it can have multiple hidden layers in the middle. The simplest MLP only contains one hidden layer, that is, a three-layer structure. The MLP is fully connected between layers (fully connected means that any neuron in the upper layer is connected to all neurons in the lower layer). The bottom layer of the MLP is the input layer, the middle is the hidden layer, and the last is the output layer.
[0095] Step S40: Based on the compact feature, perform hierarchical fusion through the point cloud perception module to obtain multi-scale features corresponding to the compact feature.
[0096] In this embodiment, referring to Figure 5 , the compact feature Fa is input into the point cloud perception module PPB in the target detection model. The point cloud perception module PPB performs hierarchical fusion on the compact feature Fa to obtain multi-scale features Fm corresponding to the compact feature Fa. The configuration of the point cloud perception module PPB in this embodiment is different from the configuration of the point cloud perception module PPB used in step S20.
[0097] Step S50: Based on the semantic feature and the multi-scale feature, perform prediction through the saliency perception module to generate a prediction result corresponding to the point cloud map.
[0098] In one embodiment, referring to Figure 5 , the semantic feature Fs and the multi-scale feature Fm are input into the saliency perception module SPB in the target detection model. The saliency perception module SPB performs prediction on the semantic feature Fs and the multi-scale feature Fm to obtain a prediction result p corresponding to the point cloud map.
[0099] In this embodiment, a point cloud map is obtained, and multi-level feature extraction is performed on the point cloud map through a preset backbone network to obtain multi-layer features corresponding to different sampling scales; the topmost feature in the multi-layer features is input into the point cloud perception module for preliminary fusion to obtain a corresponding semantic feature; the multi-layer features are input into the feature fusion module for feature aggregation to obtain a corresponding compact feature; the compact feature is input into the point cloud perception module for hierarchical fusion to obtain a corresponding multi-scale feature; the semantic feature and the multi-scale feature are input into the saliency perception module for prediction to obtain a prediction result of the point cloud map; thereby accurately determining the salient object of the point cloud map.
[0100] Furthermore, based on the first embodiment of the point cloud salient object detection method with feature fusion of the present invention, a second embodiment of the point cloud salient object detection method with feature fusion of the present invention is proposed.
[0101] The second embodiment of the point cloud saliency object detection method based on feature fusion is different from the first embodiment of the point cloud saliency object detection method based on feature fusion in that, before the step S10 of obtaining a point cloud map and performing multi-level feature extraction on the point cloud map through a preset backbone network model to obtain multi-level features of different sampling scales corresponding to the point cloud map, with reference to Figure 6 , the point cloud saliency object detection method based on feature fusion further includes:
[0102] Step A10: Obtain sample point cloud maps of multiple different scenes, and construct a training set for training the initial model with the sample point cloud maps;
[0103] Step A20: Obtain the saliency point cloud maps of the labeled significant objects corresponding to the respective sample point cloud maps;
[0104] Step A30: Use the sample point cloud maps in the training set as the input of the initial model, use the saliency point cloud maps as the output of the initial model, and perform iterative training on the initialized model to obtain the point cloud saliency object detection model.
[0105] In this embodiment, a large number of sample point cloud maps in multiple different scenes are obtained, and a training set for training the initial model is constructed with these sample point cloud maps; the saliency point cloud maps of the labeled significant objects corresponding to the respective sample point cloud maps are obtained; the respective sample point cloud maps in the training set are used as the input of the initial model, the saliency point cloud maps of the corresponding labeled significant objects are used as the output of the initial model, and the initial model is iteratively trained to train the target detection model; thereby improving the prediction accuracy of the trained target detection model.
[0106] The following will elaborate on each step in detail:
[0107] Step A10: Obtain sample point cloud maps of multiple different scenes, and construct a training set for training the initial model with the sample point cloud maps.
[0108] In this embodiment, sample point cloud maps of multiple different scenes are obtained, and based on these sample point cloud maps, a training set for the initial model is constructed. During the process of constructing the training set, these sample point cloud maps need to be pre-processed so that the size and format of each sample point cloud map are unified, which is for the convenience of batch processing.
[0109] It should be noted that to ensure the accuracy of the target detection model, the sample point cloud maps in the training set should be sufficient. This embodiment does not limit the number of sample point cloud maps in the training set. Generally, a sample point cloud map is used as a training set. In practical applications, the more sample point cloud maps in the training set, the more accurate the saliency target prediction result output by the target detection model.
[0110] Step A20: Obtain the saliency point cloud maps corresponding to the labeled salient objects in each sample point cloud map.
[0111] In this embodiment, the saliency point cloud maps of the salient objects labeled by the user for each sample point cloud map are obtained. The saliency point cloud map is the result of the user's prior annotation on the sample point cloud map. Among them, the saliency point cloud maps also need to be pre-processed so that the size and format of each saliency point cloud map are unified, which is for the convenience of batch processing.
[0112] Step A30: Use the sample point cloud maps in the training set as the input of the initial model, use the saliency point cloud maps as the output of the initial model, and perform iterative training on the initialized model to obtain the target detection model.
[0113] In one embodiment, each sample point cloud map in the training set is used as the input of the initial model, the saliency point cloud map corresponding to the labeled salient object is used as the output of the initial model, and the initial model is iteratively trained to obtain the target detection model; thereby improving the prediction accuracy of the trained target detection model.
[0114] In one embodiment, a large number of sample point cloud maps in multiple different scenarios are obtained, and these sample point cloud maps are used to construct a training set for training the initial model; the saliency point cloud maps corresponding to the labeled salient objects in each sample point cloud map are obtained; each sample point cloud map in the training set is used as the input of the initial model, the saliency point cloud map corresponding to the labeled salient object is used as the output of the initial model, and the initial model is iteratively trained to obtain the target detection model; thereby improving the prediction accuracy of the trained target detection model.
[0115] Furthermore, based on the first and second embodiments of the point cloud saliency object detection method with feature fusion of the present invention, the third embodiment of the point cloud saliency object detection method with feature fusion of the present invention is proposed.
[0116] The difference between the third embodiment of the point cloud saliency object detection method with feature fusion and the first and second embodiments of the point cloud saliency object detection method with feature fusion is that in this embodiment, for step S20, the topmost feature based on the multi-layer features is preliminarily fused through the point cloud perception module to obtain the refinement of the semantic feature corresponding to the topmost feature. Refer to Figure 7 , and this step specifically includes:
[0117] Step S21: Input the fourth-layer feature into the point cloud perception module and branch the fourth-layer feature to obtain the branched fourth-layer feature;
[0118] Step S22, find the first nearest neighbor points of each point of the fourth-layer features after branching according to a preset nearest neighbor algorithm;
[0119] Step S23, obtain the first coordinate information of each point of each fourth-layer feature and each first nearest neighbor point, and calculate the first relative position between each point of the fourth-layer feature and the first nearest neighbor point through the first coordinate information;
[0120] Step S24, fuse and splice the first nearest neighbor point and the first relative position according to a preset splicing formula to obtain the first local area corresponding to the fourth-layer feature;
[0121] Step S25, preliminarily fuse the fourth-layer feature and the first local area according to a preset fusion formula to obtain the semantic feature corresponding to the fourth-layer feature.
[0122] In this embodiment, the fourth-layer feature F4 is input into the point cloud perception module PPB in the target detection model, and the point cloud perception module PPB performs preliminary fusion features on the fourth-layer feature F4 to obtain the semantic feature Fs corresponding to the fourth-layer feature F4; thus, the significant target in the site cloud map can be accurately determined through the high-level semantic feature Fs.
[0123] The following will elaborate on each step in detail:
[0124] Step S21, input the fourth-layer feature into the point cloud perception module and branch the fourth-layer feature to obtain the fourth-layer feature after branching.
[0125] In this embodiment, the multi-layer features include the first-layer feature F1, the second-layer feature F2, the third-layer feature F3, and the fourth-layer feature F4. The fourth-layer feature F4 is also called the topmost feature of the multi-layer features. The fourth-layer feature F4 is input into the point cloud perception module PPB, and the point cloud perception module PPB branches the fourth-layer feature F4 to obtain the fourth-layer feature F4 with 4 different branches.
[0126] Step S22, find the first nearest neighbor points of each point of the fourth-layer features after branching according to a preset nearest neighbor algorithm.
[0127] In this embodiment, it is preferred that the k-nearest neighbor algorithm is the preset nearest neighbor algorithm. The k nearest neighbor points (also known as the first nearest neighbor points) around each point of the fourth-layer feature F4 of the four different branches are found through the k-nearest neighbor algorithm; among them, the number k of the first nearest neighbor points of the four different branches is 1 (the first branch), 4 (the second branch), 9 (the third branch), and 16 (the fourth branch) respectively. The configuration of the point cloud perception module PPB in this embodiment is different from the configuration of the point cloud perception module PPB used in step S40. Specifically, the value of k in the k-nearest neighbor algorithm is different. In this embodiment, the values of the number k of the surrounding points selected for the four branches are {1, 4, 9, 25}.
[0128] Step S23: Obtain the first coordinate information of each point of each fourth-layer feature and each first nearest neighbor point, and calculate the first relative position between each point of the fourth-layer feature and the first nearest neighbor point through the first coordinate information.
[0129] In this embodiment, the coordinate information (also known as the first coordinate information) of each point of the fourth-layer feature F4 of the four different branches and the first nearest neighbor points is obtained, and the relative position (also known as the first relative position) between each point of the fourth-layer feature F4 of the four different branches and the first nearest neighbor points is calculated through the first coordinate information. The calculation formula is as follows:
[0130]
[0131] Among them, taking the point represents the center point of the fourth-layer feature F4; p represents the spatial position coordinates of represents the jth nearest neighbor point of the center point ; represents the center point and the nearest neighbor point the difference in spatial position coordinates; represents the center point and the nearest neighbor point the Euclidean distance between; MLP represents using a multi-layer perceptron to splice the center point of the fourth-layer feature F4 and the nearest neighbor point ; represents calculating the first relative position between each point of the fourth-layer feature F4 of the four different branches and the first nearest neighbor points.
[0132] Step S24: According to the preset splicing formula, fuse and splice the first nearest neighbor point and the first relative position to obtain the first local area corresponding to the fourth-layer feature.
[0133] In this embodiment, the first nearest neighbor point and the first relative position are fused and spliced through a preset splicing formula to obtain the local area (also called the first local area) corresponding to the fourth-layer feature F4 of 4 different branches; the preset splicing formula is as follows:
[0134]
[0135] where L i represents the result of splicing the first relative position with the corresponding feature ; represents the feature of the center point of the fourth-layer feature F4; represents the center point ; the key features and global features of the local area are obtained by using max pooling and average pooling; MLP represents performing a splicing operation on the results of max pooling and average pooling by using a multi-layer perceptron.
[0136] Step S25, according to a preset fusion formula, preliminarily fuse the fourth-layer feature and the first local area to obtain the semantic feature corresponding to the fourth-layer feature.
[0137] In this embodiment, the fourth-layer feature F4 of 4 different branches and the first local area are preliminarily fused through a preset fusion formula; the preset fusion formula is as follows:
[0138]
[0139] where represents the feature of the first branch in the fourth-layer feature; represents the feature of the second branch in the fourth-layer feature F4; represents the feature of the third branch in the fourth-layer feature F4; represents the feature of the fourth branch in the fourth-layer feature F4; F P represents the fourth-layer feature F4 of 4 different branches; represents the semantic feature F s after preliminary fusion.
[0140] In this embodiment, the fourth-layer feature F4 is input into the point cloud perception module PPB in the target detection model, and the fourth-layer feature F4 is preliminarily fused by the point cloud perception module PPB to obtain the semantic feature Fs corresponding to the fourth-layer feature F4; thus, the significant target in the point cloud map can be accurately determined through the high-level semantic feature Fs.
[0141] Furthermore, based on the first, second, and third embodiments of the point cloud significant target detection method with feature fusion of the present invention, a fourth embodiment of the point cloud significant target detection method with feature fusion of the present invention is proposed.
[0142] The fourth embodiment of the point cloud saliency object detection method with feature fusion is different from the first, second, and third embodiments of the point cloud saliency object detection method with feature fusion in that in this embodiment, for step S30, the multi-layer features include the first-layer feature, the second-layer feature, the third-layer feature, and the fourth-layer feature. Based on the multi-layer features, feature aggregation is performed through the feature fusion module to obtain a refinement of the compact features corresponding to the multi-layer features. Refer to Figure 8 , and this step specifically includes:
[0143] Step S31, input the first-layer feature, the second-layer feature, the third-layer feature, and the fourth-layer feature into the feature fusion module;
[0144] Step S32, upsample the fourth-layer feature, and splice and aggregate the upsampled fourth-layer feature and the third-layer feature through a preset multi-layer perceptron to obtain a third-layer splicing result;
[0145] Step S33, upsample the third-layer splicing result, and splice and aggregate the upsampled third-layer splicing result and the second-layer feature through a preset multi-layer perceptron to obtain a second-layer splicing result;
[0146] Step S34, upsample the second-layer splicing result, and splice and aggregate the upsampled second-layer splicing result and the first-layer feature through a preset multi-layer perceptron to obtain the compact features corresponding to the first-layer feature, the second-layer feature, the third-layer feature, and the fourth-layer feature.
[0147] In this embodiment, the first-layer feature F1, the second-layer feature F2, the third-layer feature F3, and the fourth-layer feature F4 are input into the feature fusion module FAB. The feature fusion module FAB upsamples the fourth-layer feature, and the upsampled fourth-layer feature F4 and the third-layer feature F3 are spliced and aggregated through the MLP to obtain a third-layer splicing result; the third-layer splicing result is upsampled, and the upsampled third-layer splicing result and the second-layer feature F2 are spliced and aggregated through the MLP to obtain a second-layer splicing result; the second-layer splicing result is upsampled, and the upsampled second-layer splicing result and the first-layer feature F1 are spliced and aggregated through the MLP to obtain the corresponding compact feature Fa; the saliency object in the site cloud map can be accurately determined through the compact feature Fa.
[0148] The following will elaborate on each step in detail:
[0149] Step S31, input the first-layer feature, the second-layer feature, the third-layer feature, and the fourth-layer feature into the feature fusion module.
[0150] In this embodiment, the first-layer feature F1, the second-layer feature F2, the third-layer feature F3, and the fourth-layer feature F4 are input into the feature fusion module FAB, so as to splice and aggregate the first-layer feature F1, the second-layer feature F2, the third-layer feature F3, and the fourth-layer feature F4 through the feature fusion module FAB to obtain the corresponding compact feature Fa. Since the sampling scales (also known as dimensions) of the first-layer feature F1, the second-layer feature F2, the third-layer feature F3, and the fourth-layer feature F4 are different, these features need to be upsampled before they can be spliced and fused.
[0151] Step S32: Upsample the fourth-layer feature, and splice and aggregate the upsampled fourth-layer feature and the third-layer feature through an MLP to obtain a third-layer splicing result.
[0152] In this embodiment, the fourth-layer feature F4 is upsampled to obtain the upsampled fourth-layer feature F4; then the upsampled fourth-layer feature F4 and the third-layer feature F3 are spliced and aggregated through an MLP to obtain a third-layer splicing result.
[0153] Step S33: Upsample the third-layer splicing result, and splice and aggregate the upsampled third-layer splicing result and the second-layer feature through a preset multi-layer perceptron to obtain a second-layer splicing result.
[0154] In this embodiment, the third-layer splicing result is upsampled to obtain the upsampled third-layer splicing result; then the upsampled third-layer splicing result and the second-layer feature F2 are spliced and aggregated through an MLP to obtain a second-layer splicing result.
[0155] Step S34: Upsample the second-layer splicing result, and splice and aggregate the upsampled second-layer splicing result and the first-layer feature through a preset multi-layer perceptron to obtain the compact feature corresponding to the first-layer feature, the second-layer feature, the third-layer feature, and the fourth-layer feature.
[0156] In this embodiment, the second-layer splicing result is upsampled to obtain the upsampled second-layer splicing result; then the upsampled second-layer splicing result and the first-layer feature F1 are spliced and aggregated through an MLP to obtain the compact feature Fa.
[0157] In this embodiment, the first-layer feature F1, the second-layer feature F2, the third-layer feature F3, and the fourth-layer feature F4 are input into the feature fusion module FAB. The feature fusion module FAB performs upsampling on the fourth-layer feature, and the upsampled fourth-layer feature F4 and the third-layer feature F3 are concatenated and aggregated through an MLP to obtain a third-layer concatenation result. The third-layer concatenation result is upsampled, and the upsampled third-layer concatenation result and the second-layer feature F2 are concatenated and aggregated through an MLP to obtain a second-layer concatenation result. The second-layer concatenation result is upsampled, and the upsampled second-layer concatenation result and the first-layer feature F1 are concatenated and aggregated through an MLP to obtain the corresponding compact feature Fa. The significant target in the point cloud map can be accurately determined through the compact feature Fa.
[0158] Furthermore, based on the first, second, third, and fourth embodiments of the point cloud significant target detection method with feature fusion of the present invention, the fifth embodiment of the point cloud significant target detection method with feature fusion of the present invention is proposed.
[0159] The difference between the fifth embodiment of the point cloud significant target detection method with feature fusion and the first, second, third, and fourth embodiments of the point cloud significant target detection method with feature fusion is that in this embodiment, for step S40, based on the compact feature, hierarchical fusion is performed through the point cloud perception module to obtain the refinement of the multi-scale feature corresponding to the compact feature. Refer to Figure 9 , and this step specifically includes:
[0160] Step S41: Input the compact feature into the point cloud perception module, and branch the compact feature to obtain the branched compact feature;
[0161] Step S42: Find the second nearest neighbor points of each point of the branched compact feature according to the preset nearest neighbor algorithm;
[0162] Step S43: Obtain the second coordinate information of each point of each compact feature and each second nearest neighbor point, and calculate the second relative position between each point of the compact feature and the second nearest neighbor point through the second coordinate information;
[0163] Step S44: According to the preset splicing formula, fuse and splice the second nearest neighbor point and the second relative position to obtain the second local area corresponding to the compact feature;
[0164] Step S45: According to the preset fusion formula, perform preliminary fusion on the compact feature and the second local area to obtain the multi-scale feature corresponding to the compact feature.
[0165] In this embodiment, the compact feature Fa is input into the point cloud perception module PPB in the target detection model, and the compact feature Fa is fused and spliced by the point cloud perception module PPB to obtain the multi-scale feature Fm corresponding to the compact feature Fa; thus, the significant target in the site cloud map can be accurately determined through the multi-scale feature Fm.
[0166] The following will explain each step in detail:
[0167] Step S41: Input the compact feature into the point cloud perception module and branch the compact feature to obtain the branched compact feature.
[0168] In this embodiment, the compact feature Fa is input into the point cloud perception module PPB in the target detection model, and the compact feature Fa is processed by the point cloud perception module PPB to obtain the compact feature Fa with 4 different branches.
[0169] Step S42: Find the second nearest neighbor points of each point of the branched compact feature according to the preset nearest neighbor algorithm.
[0170] In this embodiment, the k-nearest neighbor algorithm is preferably the preset nearest neighbor algorithm, and the k nearest neighbor points (also called the second nearest neighbor points) around each point of the compact feature Fa with 4 different branches are found through the k-nearest neighbor algorithm; among them, the number k of the second nearest neighbor points of the 4 different branches is 1 (the first branch), 9 (the second branch), 25 (the third branch), and 49 (the fourth branch) respectively. The configuration of the point cloud perception module PPB in this embodiment is different from the configuration of the point cloud perception module PPB used in step S20. Specifically, the value of k in the k-nearest neighbor algorithm is different, and the value of the number k of the surrounding points selected for the four branches in this embodiment is {1, 9, 25, 49}.
[0171] Step S43: Obtain the second coordinate information of each point of each compact feature and each second nearest neighbor point, and calculate the second relative position between each point of the compact feature and the second nearest neighbor point through the second coordinate information.
[0172] In this embodiment, the coordinate information (also called the second coordinate information) of each point of the compact feature Fa with 4 different branches and the second nearest neighbor points is obtained, and the relative position (also called the second relative position) between each point of the compact feature Fa with 4 different branches and the second nearest neighbor points is calculated through the second coordinate information. The calculation formula is as follows:
[0173]
[0174] Among them, taking the point represents the center point of the compact feature Fa; p represents the spatial position coordinate of Represents the center point The j-th nearest neighbor point of; Represents the center point And the nearest neighbor point The spatial position coordinate difference of; Represents the center point And the nearest neighbor point The Euclidean distance between; MLP represents using a multi-layer perceptron to splice the center point of the compact feature Fa And the second nearest neighbor point Perform a splicing operation; Represents the second relative position of each point of the compact feature Fa with 4 different branches and the second nearest neighbor point.
[0175] Step S44, according to a preset splicing formula, fuse and splice the second nearest neighbor point and the second relative position to obtain the second local region corresponding to the compact feature.
[0176] In this embodiment, the second nearest neighbor point and the second relative position are fused and spliced through a preset splicing formula to obtain the local regions (also called the second local regions) corresponding to the compact features Fa with 4 different branches; the preset splicing formula is as follows:
[0177]
[0178] Among them, L i Represents the second relative position Spliced with the corresponding feature Result; Represents the feature of the center point of the compact feature Fa Of; Represents the center point Of the local region; the key features and global features of the local region are obtained by using max pooling and average pooling; MLP represents using a multi-layer perceptron to splice the results of max pooling and average pooling.
[0179] Step S45, according to a preset fusion formula, preliminarily fuse the compact feature and the second local region to obtain the multi-scale feature corresponding to the compact feature.
[0180] In this embodiment, the compact features Fa with 4 different branches and the second local region are preliminarily fused through a preset fusion formula; the preset fusion formula is as follows:
[0181]
[0182] Among them, Represents the feature of the first branch in the compact feature Fa; Represents the feature of the second branch in the compact feature Fa; Represents the feature of the third branch in the compact feature Fa; Represents the feature of the fourth branch in the compact feature Fa; F P Represents the compact feature Fa with 4 different branches; represents the multi-scale feature F after preliminary fusion m 。
[0183] In this embodiment, the compact feature Fa is input into the point cloud perception module PPB in the target detection model, and the compact feature Fa is fused and stitched through the point cloud perception module PPB to obtain the multi-scale feature Fm corresponding to the compact feature Fa; thus, the significant target in the site cloud map can be accurately determined through the multi-scale feature Fm.
[0184] Further, based on the first, second, third, fourth, and fifth embodiments of the point cloud significant target detection method with feature fusion of the present invention, the sixth embodiment of the point cloud significant target detection method with feature fusion of the present invention is proposed.
[0185] The difference between the sixth embodiment of the point cloud significant target detection method with feature fusion and the first, second, third, fourth, and fifth embodiments of the point cloud significant target detection method with feature fusion is that in this embodiment, step S50 is refined, that is, based on the semantic feature and the multi-scale feature, through the significant perception module, the prediction is made to generate the prediction result corresponding to the point cloud map. Refer to Figure 10 , and this step specifically includes:
[0186] Step S51, input the semantic feature and the multi-scale feature into the significant perception module;
[0187] Step S52, perform upsampling on the semantic feature, and process the upsampled semantic feature through a preset multi-layer perceptron to obtain the processed semantic feature;
[0188] Step S53, according to a preset stitching algorithm, fuse and stitch the processed semantic feature and the multi-scale feature to obtain the stitching result corresponding to the processed semantic feature and the multi-scale feature;
[0189] Step S54, perform prediction on the stitching result through a preset multi-layer perceptron to obtain the prediction result corresponding to the stitching result.
[0190] In this embodiment, by inputting the semantic feature Fs and the multi-scale feature Fm into the significant perception module SPB in the target detection model, the semantic feature Fs and the multi-scale feature Fm are stitched and fused through the significant perception module SPB, and the result after stitching and fusion is predicted to obtain the prediction result corresponding to the point cloud map; thus, the significant target of the point cloud map is accurately determined.
[0191] The following will explain each step in detail:
[0192] Step S51: Input the semantic feature and the multi-scale feature into the saliency perception module.
[0193] In this embodiment, the semantic feature Fs and the multi-scale feature Fm are input into the saliency perception module SPB in the object detection model, so as to splice and fuse the semantic feature Fs and the multi-scale feature Fm through the saliency perception module SPB, and predict the result after splicing and fusion.
[0194] Step S52: Upsample the semantic feature, and process the upsampled semantic feature through a preset multi-layer perceptron to obtain the processed semantic feature.
[0195] In this embodiment, the semantic feature Fs is upsampled through the saliency perception module SPB to obtain the upsampled semantic feature Fs; and the upsampled semantic feature Fs is spliced through the MLP to obtain the processed semantic feature Fs.
[0196] Step S53: According to a preset splicing algorithm, fuse and splice the processed semantic feature and the multi-scale feature to obtain the splicing result corresponding to the processed semantic feature and the multi-scale feature.
[0197] In this embodiment, the processed semantic feature Fs and the multi-scale feature Fm are fused and spliced through a preset splicing algorithm, and the fused and spliced semantic feature Fs and multi-scale feature Fm are spliced through the MLP to obtain the corresponding splicing result p'. The preset splicing formula is as follows:
[0198] p' = [MLP(∪(F s )), F m
[0199] Among them, MLP represents a preset multi-layer perceptron; ∪ represents the upsampling operation; p' represents the splicing result.
[0200] Step S54: Predict the splicing result through a preset multi-layer perceptron to obtain the prediction result corresponding to the splicing result.
[0201] In this embodiment, the splicing result p' is predicted through the MLP to obtain the prediction result p corresponding to the splicing result, where p = MLP([MLP(∪(F s )), F m ) and use four relatively common metrics {Mean Absolute Error (MAE), F - measure, E - measure, Intersection over Union (IoU)} to evaluate the final prediction results; among them, MAE represents the average of the absolute differences between the probability value of each point and its predicted value, and the smaller the MAE value, the better; F - measure represents the harmonic mean of precision and recall, and the larger the F - measure value, the better; E - measure represents the matching degree of each point and the overall average matching degree, and the larger the E - measure value, the better; IoU represents the intersection of the prediction result divided by the union, and the larger the IoU value, the better.
[0202] In this embodiment, the semantic feature Fs and the multi - scale feature Fm are input into the saliency perception module SPB in the object detection model. The saliency perception module SPB splices and fuses the semantic feature Fs and the multi - scale feature Fm, and predicts the result after splicing and fusing to obtain the prediction result corresponding to the point cloud map; thus, the saliency target of the point cloud map is accurately determined.
[0203] The present invention also provides a point cloud saliency target detection device for feature fusion. Referring to Figure 11 , the point cloud saliency target detection device for feature fusion of the present invention includes:
[0204] An acquisition module 10, configured to acquire a point cloud map, and perform multi - level feature extraction on the point cloud map through a preset backbone network model to obtain multi - layer features of different sampling scales corresponding to the point cloud map;
[0205] A fusion module 20, configured to perform preliminary fusion of features through the point cloud perception module based on the top - layer feature of the multi - layer features to obtain a semantic feature corresponding to the top - layer feature;
[0206] An aggregation module 30, configured to perform feature aggregation through the feature fusion module based on the multi - layer features to obtain a compact feature corresponding to the multi - layer features;
[0207] A hierarchical module 40, configured to perform hierarchical fusion through the point cloud perception module based on the compact feature to obtain a multi - scale feature corresponding to the compact feature;
[0208] A prediction module 50, configured to perform prediction through the saliency perception module based on the semantic feature and the multi - scale feature to generate a prediction result corresponding to the point cloud map.
[0209] In addition, the present invention further provides a medium, which is a computer-readable storage medium, and stores a program for point cloud saliency object detection with feature fusion. When the program for point cloud saliency object detection with feature fusion is executed by a processor, the steps of the method for point cloud saliency object detection with feature fusion as described above are implemented.
[0210] Among them, the method implemented when the program for point cloud saliency object detection with feature fusion running on the processor is executed can refer to each embodiment of the method for point cloud saliency object detection with feature fusion of the present invention, which will not be elaborated here.
[0211] It should be noted that in this article, the term "including", "comprising" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or system. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or system including that element.
[0212] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages and disadvantages of the embodiments.
[0213] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present invention.
[0214] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present invention.
Claims
1. A method for point cloud saliency object detection with feature fusion, characterized in that, The above-mentioned point cloud saliency object detection method based on feature fusion is applied to an object detection model, which includes a feature fusion module, a point cloud perception module, and a saliency perception module. The point cloud saliency object detection method based on feature fusion includes the following steps: Obtain a point cloud map, and perform multi-level feature extraction on the point cloud map through a preset backbone network model to obtain multi-layer features of different sampling scales corresponding to the point cloud map; Based on the topmost feature of the multi-layer features, perform preliminary fusion of features through the point cloud perception module to obtain semantic features corresponding to the topmost feature; Based on the multi-layer features, perform feature aggregation through the feature fusion module to obtain compact features corresponding to the multi-layer features; Based on the compact features, perform hierarchical fusion through the point cloud perception module to obtain multi-scale features corresponding to the compact features; Based on the semantic features and the multi-scale features, perform prediction through the saliency perception module to generate a prediction result corresponding to the point cloud map; Among them, the multi-layer features include a first-layer feature, a second-layer feature, a third-layer feature, and a fourth-layer feature. The fourth-layer feature is the topmost feature of the multi-layer features. The step of performing preliminary fusion of features through the point cloud perception module based on the topmost feature of the multi-layer features to obtain semantic features corresponding to the topmost feature includes: Input the fourth-layer feature into the point cloud perception module, and branch the fourth-layer feature to obtain the branched fourth-layer feature; Find the first nearest neighbor points of each point of the branched fourth-layer feature according to a preset nearest neighbor algorithm, where the preset nearest neighbor algorithm is the k-nearest neighbor algorithm; Obtain the first coordinate information of each point of each fourth-layer feature and each first nearest neighbor point, and calculate the first relative position between each point of the fourth-layer feature and the first nearest neighbor point through the first coordinate information; According to a preset splicing formula, fuse and splice the first nearest neighbor point and the first relative position to obtain a first local area corresponding to the fourth-layer feature; According to a preset fusion formula, perform preliminary fusion of the fourth-layer feature and the first local area to obtain semantic features corresponding to the fourth-layer feature; Among them, the step of performing hierarchical fusion through the point cloud perception module based on the compact features to obtain multi-scale features corresponding to the compact features includes: Input the compact feature into the point cloud perception module, and branch the compact feature to obtain the branched compact feature; Find the second nearest neighbor points of each point of the branched compact feature according to a preset nearest neighbor algorithm, where the preset nearest neighbor algorithm is the k-nearest neighbor algorithm; Obtain the second coordinate information of each point of each compact feature and each second nearest neighbor point, and calculate the second relative position between each point of the compact feature and the second nearest neighbor point through the second coordinate information; According to a preset splicing formula, fuse and splice the second nearest neighbor point and the second relative position to obtain a second local area corresponding to the compact feature; According to a preset fusion formula, preliminarily fuse the compact feature with the second local region to obtain a multi-scale feature corresponding to the compact feature; Among them, the value of k of the preset nearest neighbor algorithm in finding the first nearest neighbor points of each point of the fourth-layer feature after branching according to the preset nearest neighbor algorithm is different from the value of k of the preset nearest neighbor algorithm in finding the second nearest neighbor points of each point of the compact feature after branching according to the preset nearest neighbor algorithm.
2. The method for point cloud saliency object detection with feature fusion according to claim 1, characterized in that, Before the step of obtaining a point cloud map and performing multi-level feature extraction on the point cloud map through a preset backbone network model to obtain multi-level features of different sampling scales corresponding to the point cloud map, the point cloud saliency object detection method based on feature fusion further includes: Obtain sample point cloud maps of multiple different scenarios, and construct a training set for training an initial model with the sample point cloud maps; Obtain a saliency point cloud map annotating a significant object corresponding to each sample point cloud map; Use the sample point cloud maps in the training set as the input of the initial model, use the saliency point cloud map as the output of the initial model, and perform iterative training on the initial model to obtain the target detection model.
3. The method for point cloud saliency object detection with feature fusion according to claim 1, characterized in that, The multi-level features include a first-level feature, a second-level feature, a third-level feature, and a fourth-level feature. The step of obtaining a point cloud map and performing multi-level feature extraction on the point cloud map through a preset backbone network model to obtain multi-level features of different sampling scales corresponding to the point cloud map includes: Obtain a point cloud map and input the point cloud map into a preset backbone network model; Perform multi-level feature extraction on the point cloud map through the preset backbone network model to obtain a first-level feature, a second-level feature, a third-level feature, and a fourth-level feature of different sampling scales corresponding to the point cloud map.
4. The method for point cloud saliency object detection with feature fusion according to claim 1, characterized in that, The multi-level features include a first-level feature, a second-level feature, a third-level feature, and a fourth-level feature. The step of aggregating features through the feature fusion module based on the multi-level features to obtain a compact feature corresponding to the multi-level features includes: Input the first-level feature, the second-level feature, the third-level feature, and the fourth-level feature into the feature fusion module; Upsample the fourth-level feature, and splice and aggregate the upsampled fourth-level feature with the third-level feature through a preset multi-layer perceptron to obtain a third-level splicing result; Upsample the third-level splicing result, and splice and aggregate the upsampled third-level splicing result with the second-level feature through a preset multi-layer perceptron to obtain a second-level splicing result; Upsample the second-level splicing result, and splice and aggregate the upsampled second-level splicing result with the first-level feature through a preset multi-layer perceptron to obtain a compact feature corresponding to the first-level feature, the second-level feature, the third-level feature, and the fourth-level feature.
5. The method for point cloud saliency object detection with feature fusion according to claim 1, characterized in that, The step of generating a prediction result corresponding to the point cloud map through prediction by the saliency perception module based on the semantic feature and the multi-scale feature includes: Input the semantic feature and the multi-scale feature into the saliency perception module; Upsample the semantic features, and process the upsampled semantic features through a preset multi-layer perceptron to obtain processed semantic features; According to a preset splicing algorithm, fuse and splice the processed semantic features with the multi-scale features to obtain a splicing result corresponding to the processed semantic features and the multi-scale features; Predict the splicing result through a preset multi-layer perceptron to obtain a prediction result corresponding to the splicing result.
6. A point cloud saliency object detection device with feature fusion, characterized in that, The point cloud saliency object detection device for feature fusion includes: An acquisition module, configured to acquire a point cloud map, and perform multi-level feature extraction on the point cloud map through a preset backbone network model to obtain multi-layer features of different sampling scales corresponding to the point cloud map; A fusion module, configured to perform preliminary fusion of features through a point cloud perception module based on the topmost feature of the multi-layer features to obtain semantic features corresponding to the topmost feature; An aggregation module, configured to perform feature aggregation through the feature fusion module based on the multi-layer features to obtain compact features corresponding to the multi-layer features; A grading module, configured to perform hierarchical fusion through the point cloud perception module based on the compact features to obtain multi-scale features corresponding to the compact features; A prediction module, configured to perform prediction through a saliency perception module based on the semantic features and the multi-scale features to generate a prediction result corresponding to the point cloud map; Wherein, the multi-layer features include a first layer feature, a second layer feature, a third layer feature, and a fourth layer feature, the fourth layer feature is the topmost feature of the multi-layer features, and the fusion module is further configured to: Input the fourth layer feature into the point cloud perception module, and branch the fourth layer feature to obtain a branched fourth layer feature; Find the first nearest neighbor points of each point of the branched fourth layer feature according to a preset nearest neighbor algorithm, where the preset nearest neighbor algorithm is a k-nearest neighbor algorithm; Obtain the first coordinate information of each point of each fourth layer feature and each first nearest neighbor point, and calculate the first relative position between each point of the fourth layer feature and the first nearest neighbor point through the first coordinate information; According to a preset splicing formula, fuse and splice the first nearest neighbor point and the first relative position to obtain a first local area corresponding to the fourth layer feature; According to a preset fusion formula, perform preliminary fusion of the fourth layer feature and the first local area to obtain semantic features corresponding to the fourth layer feature; Wherein, the grading module is further configured to: Input the compact feature into the point cloud perception module, and branch the compact feature to obtain a branched compact feature; Find the second nearest neighbor points of each point of the branched compact feature according to a preset nearest neighbor algorithm, where the preset nearest neighbor algorithm is a k-nearest neighbor algorithm; Obtain the second coordinate information of each point of each compact feature and each second nearest neighbor point, and calculate the second relative position between each point of the compact feature and the second nearest neighbor point through the second coordinate information; According to a preset splicing formula, fuse and splice the second nearest neighbor point and the second relative position to obtain a second local area corresponding to the compact feature; According to a preset fusion formula, preliminarily fuse the compact feature and the second local area to obtain a multi-scale feature corresponding to the compact feature; Among them, the value of the preset nearest neighbor algorithm k in the fusion module is different from the value of the preset nearest neighbor algorithm k in the grading module.
7. A device, the device is a point cloud saliency object detection device with feature fusion, characterized in that, The point cloud saliency target detection device for feature fusion includes: a memory, a processor, and a point cloud saliency target detection program for feature fusion stored on the memory and executable on the processor. When the point cloud saliency target detection program for feature fusion is executed by the processor, the steps of the point cloud saliency target detection method for feature fusion according to any one of claims 1 to 5 are implemented.
8. A medium, the medium is a computer-readable storage medium, characterized in that, A point cloud saliency target detection program for feature fusion is stored on the computer-readable storage medium. When the point cloud saliency target detection program for feature fusion is executed by the processor, the steps of the point cloud saliency target detection method for feature fusion according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Three-dimensional point cloud semantic segmentation method and system based on fully-fusion network
CN111860138A
RGBD saliency detection method based on feature aggregation
CN111931787A