Air-ground point cloud scene identification method and system based on hierarchical adaptive feature fusion

By adopting a hierarchical adaptive feature fusion method in the recognition of empty-site point cloud scenes, the problem of degradation of recognition performance in heterogeneous point cloud environments is solved, and higher recognition accuracy and robustness are achieved.

CN120014394APending Publication Date: 2025-05-16SHANDONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411970371.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing point cloud-based empty-site point cloud scene recognition method performs poorly in heterogeneous point cloud environments and cannot effectively deal with differences in view angle, density and spatial coverage, resulting in a degradation of recognition performance.

Method used

Using a hierarchical adaptive feature fusion method, the features are divided into high-level features and underlying features according to the number of layers through the feature pyramid network, and three feature fusion modules are used to perform hierarchical adaptive fusion of the feature map, assigning weights to the features of each scale to enhance features that are more important to the scene recognition task and reduce the domain gap between heterogeneous point clouds.

Benefits of technology

It effectively enhances the features that are more important for scene recognition tasks, improves recognition performance, and reduces the domain gap between heterogeneous point clouds, significantly improving the accuracy and robustness of empty-site point cloud scene recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014394A_ABST
    Figure CN120014394A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of autonomous navigation and environmental perception, and provides an air-ground point cloud scene recognition method and system based on hierarchical adaptive feature fusion, and the technical scheme is as follows: obtaining a point cloud data set; performing feature extraction based on the point cloud data set and the trained feature coding network to obtain global feature descriptors, and constructing a global feature descriptor database; extracting a global feature descriptor of the to-be-queried point cloud based on the trained feature coding network; searching in a global feature descriptor database to obtain a point cloud global feature descriptor most similar to the global feature descriptor of the to-be-queried point cloud, and in the feature coding network, dividing features in the feature pyramid network architecture into high-level features and low-level features according to the number of layers, the three feature fusion modules are used for carrying out hierarchical adaptive fusion on feature maps, weights are distributed to features of all scales, features which are more important for scene recognition tasks are effectively enhanced, and domain gaps among heterogeneous point clouds are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of autonomous navigation and environmental perception, and in particular to an air-space point cloud scene recognition method and system based on hierarchical adaptive feature fusion. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Point cloud scene recognition is a key technology for air-to-ground robot systems to achieve precise positioning. Point cloud, as an expression of the three-dimensional spatial information of the environment, has the advantages of being insensitive to lighting and weather conditions and providing rich geometric information, making it particularly suitable for diverse and unstructured environments.

[0004] At present, most point cloud-based scene recognition methods are mainly targeted at ground robot systems. When applied to air-ground heterogeneous point clouds (such as data collected from air and ground platforms), their performance drops significantly. This is because air-ground point clouds have significant differences in viewpoint, density, and spatial coverage. Projection-based methods first project the original point cloud into a 2D image, and then convert it into a hand-crafted feature descriptor, or extract features using a deep learning network. Most of these methods are designed for ground mobile robots and specific types of LiDAR sensors, so they are susceptible to viewpoint changes and occlusion. Point-based methods directly use the original point cloud to extract features without additional preprocessing. However, the density difference between heterogeneous point clouds will affect the accuracy and robustness of such methods. Voxel-based methods first voxelize the point cloud and then extract features from the point cloud, which can reduce the density difference between heterogeneous point clouds to a certain extent. Among them, MinkFPN and other networks composed of 3D sparse convolutions have achieved good performance in capturing the features of homogeneous point clouds, but do not consider the domain gap of heterogeneous point clouds. This method directly adds the points after unifying the size in a top-down order, and does not select the features according to their importance. In addition, low-level features can only reach the output layer through layers of convolution, which has little impact on the final output features and is not fully utilized. Summary of the invention

[0005] In order to solve at least one technical problem existing in the above-mentioned background technology, the present invention provides an air-space point cloud scene recognition method and system based on hierarchical adaptive feature fusion, which divides the features in the feature pyramid network (FPN) architecture into high-level features and low-level features according to the number of layers, and uses three feature fusion modules to perform hierarchical adaptive fusion on the feature maps, assigns weights to features of each scale, effectively enhances features that are more important to the scene recognition task, and reduces the domain gap between heterogeneous point clouds.

[0006] In order to achieve the above object, the present invention adopts the following technical solution:

[0007] A first aspect of the present invention provides an air-space cloud scene recognition method based on hierarchical adaptive feature fusion, comprising the following steps:

[0008] Get point cloud dataset;

[0009] Based on the point cloud data set and the trained feature encoding network, feature extraction is performed to obtain a global feature descriptor, and a global feature descriptor database is constructed; wherein the construction of the feature encoding network includes:

[0010] The point cloud data is converted into a sparse tensor, and the sparse tensor is used as input to a feature pyramid network composed of multiple 3D sparse convolutions to extract a multi-scale feature map. The multi-scale feature map is divided into low-level features and high-level features according to the number of layers in the feature pyramid network. The multi-scale features corresponding to the low-level features and the high-level features are fused to obtain fused low-level features and high-level features, respectively. The fused low-level features and high-level features are fused to obtain local features; the local features are encoded to obtain encoded local features, and the global feature descriptors are obtained after aggregation;

[0011] Extract the global feature descriptor of the query point cloud based on the trained feature encoding network;

[0012] The global feature descriptor of the point cloud that is most similar to the global feature descriptor of the query point cloud is searched in the global feature descriptor database.

[0013] Furthermore, the point cloud training dataset is represented as:

[0014] D={S1,S2,...,S N},

[0015] S i = {P i ,T i ,L i},

[0016] The 3D point cloud map M constructed by the laser radar carried by the drone a , and the 3D point cloud map M constructed by the laser radar carried by the unmanned vehicle g Divide into N subgraphs P sub = {P1,P2,...,P N}, the overlapping area between each sub-image is used as the point cloud data P of each sample i , where P i ∈R N×3 , the coordinates of each point are expressed as (x, y, z), T iTo obtain the coordinates of the sub-graph in the global coordinate system, in TUM format, L i is the label of the sample, indicating that the sample is obtained by the laser radar carried by a drone or unmanned vehicle.

[0017] Furthermore, converting point cloud data into sparse tensors includes:

[0018] The sparse_quantize function in MinkowskiEngine is used to convert the point cloud into a sparse tensor according to the preset voxel size voxels using the SparseTensor function.

[0019] Furthermore, the feature pyramid network includes five bottom-up convolution blocks, the output features extracted by the previous convolution block are used as the input features of the next convolution block, and the output feature map corresponding to each convolution block is obtained, the features extracted by the first convolution block, the second convolution block and the third convolution block are used as low-level features, and the features extracted by the fourth convolution block and the fifth convolution block are used as high-level features.

[0020] Furthermore, when the multi-scale features corresponding to the low-level features and the high-level features are fused respectively, the following steps are included:

[0021] Adjust multi-scale features of different scales to the same scale in spatial resolution;

[0022] Project the resized feature map to the direction of the adaptive viewing angle to obtain the projection length;

[0023] Calculate the weight corresponding to each feature according to the projection length;

[0024] Combine each feature and the corresponding weight to get the fused feature.

[0025] Furthermore, the loss function of the feature encoding network is:

[0026] L=ω P · P L+ω O · O L+ T L,

[0027]

[0028] in, Represents a given sample pair {S x ,S y}, the point loss of the sample pair is, Represents a given heterogeneous sample pair {S g ,S a}, the overlap loss of the sample pair, T LB is the heterogeneous triplet loss of the entire batch during training, P L is the number of sample pairs that meet the conditions in the entire batch during training. The average value of O L is the number of sample pairs that meet the conditions in the entire batch during training. , M represents the set of triplet samples of constructed point loss, where x and y are the anchor samples and their positive sample index numbers, q represents the supervoxel in sample x and it has corresponding points and non-corresponding points in sample y, p represents any corresponding point of q in sample y, n represents the non-corresponding point of q in sample y with the smallest local feature distance, t represents any non-corresponding point of q in sample y, α is a pre-defined hyperparameter, Represents the index set of corresponding point pairs of sample x in sample y, The set of indices representing the supervoxels of non-corresponding point pairs of sample x in sample y, l F x (q) represents the local feature corresponding to q, l F y (p) represents the local feature corresponding to p, l F y (n) represents the local feature corresponding to n, l F y (t) represents the local feature corresponding to t, represents the set of supervoxels in sample g that have corresponding points in its heterogeneous positive sample a, A represents the set of supervoxels in sample g that do not have corresponding points in its heterogeneous positive sample a. g (s) represents the attention score of each supervoxel in sample g, m is a predefined margin parameter, i is the index number of the anchor sample, j is the index number of the positive sample with the largest global feature distance from anchor sample i in the positive sample set corresponding to anchor sample i, k is the index number of the negative sample with the smallest global feature distance from anchor sample i in the negative sample set corresponding to anchor sample i, a represents the index number of any positive sample corresponding to anchor sample i, b represents the index number of any negative sample corresponding to anchor sample i, B represents the constructed triplet sample set, g F i is the global feature corresponding to the anchor sample i, g F j is the global feature corresponding to sample j, g F k is the global feature corresponding to sample k, g F a is the global feature corresponding to sample a, g F b is the global feature corresponding to sample b, T i To obtain the pose of sample i sub-image, Ta To obtain the pose of sample a sub-image, T b To obtain the pose of sample b sub-image, τ TP is the positive sample coordinate distance threshold, τ TN is the negative sample coordinate distance threshold.

[0029] Furthermore, the GeM generalized average pooling layer is used to aggregate local features to obtain the global feature descriptor.

[0030] A second aspect of the present invention provides an air-space point cloud scene recognition system based on hierarchical adaptive feature fusion, comprising:

[0031] A data acquisition module, used to acquire point cloud data sets;

[0032] The feature encoding module extracts features based on the point cloud dataset and the trained feature encoding network to obtain global feature descriptors and build a global feature descriptor database. The construction of the feature encoding network includes:

[0033] The point cloud data is converted into a sparse tensor, and the sparse tensor is used as input to a feature pyramid network composed of multiple 3D sparse convolutions to extract a multi-scale feature map. The multi-scale feature map is divided into low-level features and high-level features according to the number of layers in the feature pyramid network. The multi-scale features corresponding to the low-level features and the high-level features are fused to obtain fused low-level features and high-level features, respectively. The fused low-level features and high-level features are fused to obtain local features; the local features are encoded to obtain encoded local features, and the global feature descriptors are obtained after aggregation;

[0034] The scene recognition module is used to extract the global feature descriptor of the point cloud to be queried based on the trained feature encoding network; and to search in the global feature descriptor database to obtain the point cloud global feature descriptor that is most similar to the global feature descriptor of the point cloud to be queried.

[0035] A third aspect of the present invention provides a computer-readable storage medium.

[0036] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps in the above-mentioned method for identifying an empty space cloud scene based on hierarchical adaptive feature fusion.

[0037] A fourth aspect of the present invention provides a computer device.

[0038] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the above-mentioned method for identifying an air-space point cloud scene based on hierarchical adaptive feature fusion are implemented.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] The present invention divides the features in the feature pyramid network architecture into high-level features and low-level features according to the number of layers, uses three feature fusion modules to perform hierarchical adaptive fusion of feature maps, assigns weights to features of each scale, effectively enhances features that are more important for scene recognition tasks, and reduces the domain gap between heterogeneous point clouds.

[0041] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0043] Figure 1 It is a flow chart of an air-space point cloud scene recognition method based on hierarchical adaptive feature fusion provided by an embodiment of the present invention;

[0044] Figure 2 is a schematic diagram of a feature coding network provided by an embodiment of the present invention;

[0045] Figure 3 Schematic diagram of the structure of Conv1-Conv5 provided in an embodiment of the present invention;

[0046] Figure 4 It is a schematic diagram of a feature fusion module provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0047] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0048] It should be noted that the following detailed descriptions are all illustrative and intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0049] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0050] In view of the problem that the existing aerial point cloud scene recognition methods do not consider the domain gap of heterogeneous point clouds and cannot make full use of the output features, the present invention proposes an aerial point cloud scene recognition method based on hierarchical adaptive feature fusion, which divides the features in the feature pyramid network architecture into high-level features and low-level features according to the number of layers, and uses three feature fusion modules to perform hierarchical adaptive fusion on the feature maps, assigns weights to features of each scale, effectively enhances the features that are more important to the scene recognition task, and reduces the domain gap between heterogeneous point clouds.

[0051] Embodiment 1

[0052] like Figure 1 As shown, this embodiment provides an air-space point cloud scene recognition method based on hierarchical adaptive feature fusion, comprising the following steps:

[0053] Step 1: Get point cloud dataset;

[0054] In this embodiment, the point cloud data set is represented by D = {S1, S2, ..., S N}, each sample in the data set is represented as S i = {P i ,T i ,L i};

[0055] When acquiring point cloud data, drones and unmanned vehicles equipped with laser radars are used to move along the same or repeated tracks and collect point cloud data of the environment to ensure that the point cloud data have sufficient overlapping areas. The unmanned vehicle drives along the road and the drone flies at high altitude to obtain the three-dimensional point cloud map M constructed by the laser radar carried by the drone. a , and the 3D point cloud map M constructed by the laser radar carried by the unmanned vehicle g ; Then the large 3D point cloud map M a and M g Divided into multiple subgraphs, a total of N, denoted as P sub = {P1,P2,...,P N}, the overlapping area between each sub-image is used as the point cloud data P of each sample i , where P i ∈R N×3 , the coordinates of each point are expressed as (x, y, z). i To obtain the coordinates of the sub-graph in the global coordinate system, in TUM format, L i is the label of the sample, indicating that the sample is obtained by the laser radar carried by a drone or unmanned vehicle.

[0056] Step 2: Sparsely quantize the point cloud data and convert it into a sparse tensor;

[0057] In this embodiment, the sparse_quantize function in MinkowskiEngine is used to convert the point cloud into a sparse tensor according to a preset voxel size using the SparseTensor function, and the dimension of the sparse tensor is N×3.

[0058] Step 3: Extract features based on the sparse tensor and the trained feature encoding network to obtain global feature descriptors and build a global feature descriptor database;

[0059] like Figure 2 As shown, the feature encoding network includes a multi-scale feature extraction module, a hierarchical adaptive feature fusion module, a Transformer encoder and a GeM pooling layer.

[0060] The specific steps include:

[0061] Step 301: input the sparse tensor into a multi-scale feature extraction module to extract a multi-scale feature map;

[0062] Among them, the multi-scale feature extraction module is an FPN architecture composed of 3D sparse convolutions, which is used to generate multi-scale feature maps.

[0063] The FPN architecture consists of five bottom-up convolutional blocks, namely Conv1-Conv5, with resolutions of 1, 1 / 2, 1 / 4, 1 / 8, and 1 / 16, and channels of 64, 64, 128, 64, and 32, respectively. It is used to gradually downsample the point cloud and increase the receptive field to extract features of different scales.

[0064] Conv1 contains a convolutional layer with a kernel of 5×5×5 to obtain larger neighborhood information. Conv2-Conv5 has a stride of 2 and a kernel of 2×2×2, which can reduce the spatial resolution of voxels to 1 / 2 of the previous layer for downsampling. In addition, Conv2-Conv5 is followed by a residual block consisting of two convolutional layers and a channel attention ECA layer. The convolutional layer has a 3×3×3 kernel and a stride of 1.

[0065] The output feature maps corresponding to the five convolution blocks are X1-X5, which are divided into low-level features and high-level features according to their number of layers in the FPN architecture. The low-level features are X1-X3, and the high-level features are X4 and X5. The schematic diagram of the Conv1-Conv5 structure is shown in the figure. Figure 3 As shown, Batch Normalization is the batch normalization layer and ReLU is the revised rectifier unit.

[0066] Step 302: Combining the multi-scale feature map and the hierarchical adaptive feature fusion module to perform feature fusion;

[0067] In order to adaptively enhance features that are more beneficial to the sky-high point cloud scene recognition task, this embodiment designs a hierarchical adaptive feature fusion module to fuse multi-scale features.

[0068] The hierarchical adaptive feature fusion module consists of three feature fusion modules, which are represented by feature fusion module 1 to feature fusion module 3. Feature fusion module 1 is used to fuse low-level features, i.e., X1-X3, and the output is f F1, feature fusion module 2 is used to fuse high-level features, namely X4 and X5, and the output is f F2. Feature fusion module 3 is used to fuse f F1 and f F2, the output is l F.

[0069] The basic fusion process of the feature fusion module consists of three parts, such as Figure 4 As shown, they are feature size adjustment, weight calculation, and feature fusion;

[0070] The specific steps include:

[0071] Step 3021: Adjust the multi-scale features of different scales to the same scale in terms of spatial resolution. The adjustment formula is:

[0072] Y i =ReLU(BatchNorm( u Conv i (X i ))) (1),

[0073] Where X i represents the i-th scale feature, Y i represents the i-th resized feature, BatchNorm is the batch normalization layer, u Conv i () is the size adjustment convolution layer corresponding to the i-th feature to be fused.

[0074] Step 3022: Project the resized feature map to the direction of the adaptive viewing angle to obtain the projection length, and calculate the weight corresponding to each feature according to the projection length. In this embodiment, the resized features Y1, ..., Y n , use formula (2) to project each feature map to the direction of the adaptive viewing angle to obtain the projection length d i :

[0075] d i =MLP i (GMP(Y i ))·l (2),

[0076] Among them, GMP() is the global maximum pooling, MLP i () is the multi-layer perceptron corresponding to each feature, and l is the direction of adaptive viewing angle.

[0077] The weight corresponding to each feature is calculated using the softmax formula:

[0078]

[0079] Among them, w i is the weight corresponding to the i-th feature to be fused.

[0080] Step 3023: Combine each feature and the corresponding feature weight to obtain a fused feature;

[0081] The feature fusion formula is:

[0082]

[0083] in, f F is the fused feature, Conv() is a convolutional layer with a kernel of 1×1×1 for adjusting the number of feature channels, and Concat() is a feature concatenation operation.

[0084] Feature fusion module 1 fuses low-level features X1-X3 and outputs f F1, where u Conv1 kernel is 4×4×4, stride is 4, u Conv2 kernel is 2×2×2, stride is 2, u The Conv3 kernel is 1×1×1, the stride is 1, the number of channels of the input feature is not changed, l is a 128-dimensional vector, the Conv kernel is 1×1×1, the stride is 1, the input is 256-dimensional, and the output is 256-dimensional.

[0085] Feature fusion module 2 fuses high-level features X4 and X5, and the output is f F2, where u Conv4 kernel is 2×2×2, stride is 2, u The Conv5 kernel is 1×1×1, and the stride is 1. The transposed convolution layer does not change the number of channels of the input feature. l is a 64-dimensional vector. The Conv kernel is 1×1×1, the stride is 1, the input is 96-dimensional, and the output is 256-dimensional. In addition, in the feature fusion module 2, in order to adjust the spatial resolution of the feature, the output of Conv is input into a transposed convolution layer with a kernel of 2×2×2 and a stride of 2.

[0086] Feature fusion module 3 is used to fuse f F1 and f F2, the output is l F, due to fF1 and f The F2 features have the same size and do not require resizing, l is a 256-dimensional vector, the Conv kernel is 1×1×1, the stride is 1, and the output is 256-dimensional.

[0087] After the hierarchical adaptive feature fusion module, the final local features are obtained l F and supervoxel l P, l N is the number of supervoxels, that is, the number of local features.

[0088] Step 3024: local features l F performs feature encoding and obtains the encoded local features e F, after aggregation, we get the global feature descriptor g F;

[0089] In this embodiment, the local features l F is input into the Transformer encoder for feature encoding, and the multi-head self-attention mechanism (MHSA), feedforward network and normalization layer are used for l After F, the encoded local features are obtained e F.

[0090] In this embodiment, the GeM generalized average pooling layer is used to e F is aggregated into a global feature descriptor with a dimension of 256 g F, as shown in formula (5):

[0091]

[0092] Here, p is a shared learnable parameter.

[0093] Step 3025: Based on the obtained global feature descriptor g F constructs a global feature descriptor database; in this embodiment, when the feature encoding network is trained, the loss function of the feature encoding network includes three parts: point loss, overlap loss and heterogeneous triplet loss.

[0094] In order to facilitate metric learning, this embodiment introduces the concept of positive samples and negative samples. If the coordinates of two samples in the global coordinate system T i and T j The distance between them is less than the threshold τ TP , then the two samples are considered to be positive samples of each other. On the contrary, if the coordinates of the two samples in the global coordinate system T i and T j The distance between them is greater than the threshold τ TN, then the two samples are considered to be negative samples of each other. If the labels of the two samples are the same, they are isomorphic samples. If the labels of the two point clouds are different, they are heterogeneous samples.

[0095] The mid-point loss is used to narrow the domain gap between aerial point clouds and ground point clouds.

[0096] Given a sample pair {S x ,S y}, where S x is the anchor sample, S y is the anchor sample S x Positive samples, using T x and T y Supervoxel l P x and l P y Transformed to global coordinates, we get l P' x and l P' y Then perform ICP registration and get l P x and l P y The corresponding points in The Euclidean distance between corresponding points should be less than the threshold τ pp The set of non-corresponding points is Accordingly, the Euclidean distance between non-corresponding points should be greater than the threshold τ pn .

[0097] The most difficult triplet determined according to the feature distance is shown in formula (6), and the expression of point loss is shown in formula (7).

[0098]

[0099]

[0100] Among them, M represents the set of triplet samples of constructed point loss, where x and y are the index numbers of anchor point samples and their positive samples, Represents the index set of corresponding point pairs of sample x in sample y, represents the index set of supervoxels of non-corresponding point pairs of sample x in sample y, q represents a supervoxel in sample x and it has corresponding points and non-corresponding points in sample y, p represents any corresponding point of q in sample y, n represents the non-corresponding point of q in sample y with the smallest local feature distance, t represents any non-corresponding point of q in sample y, α is a pre-defined hyperparameter, l F x (q) represents the local feature corresponding to q, l F y(p) represents the local feature corresponding to p, l F y (n) represents the local feature corresponding to n, l F y (t) represents the local feature corresponding to t.

[0101] The purpose of overlap loss is to maximize the heterogeneous positive sample pair {S g ,S a The attention scores of the points in the overlapping area are minimized, while the attention scores of other points are minimized. The expression is shown in formula (8):

[0102]

[0103] in Represents sample S g In its heterogeneous positive sample S a There is a supervoxel set with corresponding points in , Represents sample S g In its heterogeneous positive sample S a There is no corresponding point in the supervoxel set, A g (s) represents sample S g The attention score of each supervoxel in .

[0104] Introducing global descriptors into heterogeneous triplet loss g F, minimizes the feature distance between the positive sample and the anchor point, while maximizing the distance between the negative sample and the anchor point, thereby enhancing the distinguishability of the feature. When there are enough heterogeneous positive sample pairs, the heterogeneous triplet is constructed as shown in formula (9), and the heterogeneous triplet loss is shown in formula (10), where []+ represents the hinge loss function, m represents the predefined margin parameter, and B represents the constructed triplet sample set.

[0105]

[0106] Among them, i is the index number of the anchor sample, j is the index number of the positive sample with the largest global feature distance from i in the positive sample set corresponding to anchor sample i, k is the index number of the negative sample with the smallest global feature distance from i in the negative sample set corresponding to anchor sample i, a represents the index number of any positive sample corresponding to anchor sample i, b represents the index number of any negative sample corresponding to anchor sample i, g F i is the global feature corresponding to the anchor sample i, g F j is the global feature corresponding to sample j, g F k is the global feature corresponding to sample k, g F a is the global feature corresponding to sample a,g F b is the global feature corresponding to sample b, T i To obtain the pose of sample i sub-image, T a To obtain the pose of sample a sub-image, T b To obtain the pose of sample b sub-image, τ TP is the positive sample coordinate distance threshold, τ TN is the negative sample coordinate distance threshold.

[0107] Represents a given sample pair {S x ,S y}, the point loss of the sample pair is, Represents a given heterogeneous sample pair {S g ,S a}, the overlap loss of the sample pair, T L B is the heterogeneous triplet loss of the entire batch during training, P L is the number of sample pairs that meet the conditions in the entire batch during training. The average value of O L is the number of sample pairs that meet the conditions in the entire batch during training. The average value of .

[0108] The final loss expression is shown in formula (11), ω P and ω O is a predefined weight parameter.

[0109] L=ω P · P L+ω O · O L+ T L (11),

[0110] Step 4: Given a query point cloud P Q and its label L Q , and a global feature descriptor database, the query point cloud and its labels are input into the trained feature encoding network to obtain the global feature descriptor of the query point cloud, and the most similar point cloud global feature descriptor is searched in the global feature descriptor database according to the global feature descriptor of the query point cloud, and the corresponding pose information is obtained.

[0111] In order to verify the accuracy of the algorithm, the present invention was tested on the GAPR dataset, which is currently the only air-ground point cloud scene recognition dataset containing sample pose information. The experimental equipment platform is Intel i7-13700 CPU and Nvidia RTX 3080 GPU. On this dataset, the method was evaluated by TopN-Recall and Precision-Recall. Under the four query modes of ground-ground, ground-air, air-ground, and air-air, the average recall rate AR@1 of this method is 99.92%, 96.07%, 98.14%, and 99.68%, respectively, and AR@5 is 100.00%, 98.10%, 99.64%, and 99.84%, respectively. Regardless of the mode, a high average recall rate can be achieved, which can be effectively used for large-scale scene exploration, map construction, global positioning and other tasks of air-ground robot collaboration systems.

[0112] Embodiment 2

[0113] This embodiment provides an air-space cloud scene recognition method based on hierarchical adaptive feature fusion, comprising the following steps:

[0114] A data acquisition module, used to acquire point cloud data sets;

[0115] The feature encoding module extracts features based on the point cloud dataset and the trained feature encoding network to obtain global feature descriptors and build a global feature descriptor database. The construction of the feature encoding network includes:

[0116] The point cloud data is converted into a sparse tensor, and the sparse tensor is used as input to a feature pyramid network composed of multiple 3D sparse convolutions to extract a multi-scale feature map. The multi-scale feature map is divided into low-level features and high-level features according to the number of layers in the feature pyramid network. The multi-scale features corresponding to the low-level features and the high-level features are fused to obtain fused low-level features and high-level features, respectively. The fused low-level features and high-level features are fused to obtain local features; the local features are encoded to obtain encoded local features, and the global feature descriptors are obtained after aggregation;

[0117] The scene recognition module is used to extract the global feature descriptor of the point cloud to be queried based on the trained feature encoding network; and to search in the global feature descriptor database to obtain the point cloud global feature descriptor that is most similar to the global feature descriptor of the point cloud to be queried.

[0118] Embodiment 3

[0119] This embodiment provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the steps in the above-mentioned method for identifying an empty space cloud scene based on hierarchical adaptive feature fusion are implemented.

[0120] Embodiment 4

[0121] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the above-described method for identifying an air-space point cloud scene based on hierarchical adaptive feature fusion are implemented.

[0122] Embodiment 5

[0123] This embodiment provides a program product, which is a computer program product, including a computer program. When the computer program is executed by a processor, the steps in the above-mentioned method for identifying an air-space point cloud scene based on hierarchical adaptive feature fusion are implemented.

[0124] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) containing computer-usable program code.

[0125] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0126] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0127] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0128] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0129] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A method for identifying empty space cloud scenes based on hierarchical adaptive feature fusion, characterized in that: The steps include: Get point cloud dataset; Based on the point cloud data set and the trained feature encoding network, feature extraction is performed to obtain a global feature descriptor, and a global feature descriptor database is constructed; wherein the construction of the feature encoding network includes: The point cloud data is converted into a sparse tensor, and the sparse tensor is used as input to a feature pyramid network composed of multiple 3D sparse convolutions to extract a multi-scale feature map. The multi-scale feature map is divided into low-level features and high-level features according to the number of layers in the feature pyramid network. The multi-scale features corresponding to the low-level features and the high-level features are fused to obtain fused low-level features and high-level features, respectively. The fused low-level features and high-level features are fused to obtain local features; the local features are encoded to obtain encoded local features, and the global feature descriptors are obtained after aggregation; Extract the global feature descriptor of the query point cloud based on the trained feature encoding network; The global feature descriptor of the point cloud that is most similar to the global feature descriptor of the query point cloud is searched in the global feature descriptor database.

2. The method for identifying sky-space cloud scenes based on hierarchical adaptive feature fusion according to claim 1, characterized in that: The point cloud training dataset is represented as: D={S1,S2,...,S N }, S i ={P i ,T i ,L i }, The 3D point cloud map M constructed by the laser radar carried by the drone a , and the 3D point cloud map M constructed by the laser radar carried by the unmanned vehicle g Divide into N subgraphs P sub = {P1,P2,...,P N }, the overlapping area between each sub-image is used as the point cloud data P of each sample i , where P i ∈R N×3 , the coordinates of each point are expressed as (x, y, z), T i To obtain the coordinates of the sub-graph in the global coordinate system, in TUM format, L i is the label of the sample, indicating that the sample is obtained by the laser radar carried by a drone or unmanned vehicle.

3. The method for identifying sky-space cloud scenes based on hierarchical adaptive feature fusion according to claim 1, characterized in that: Converting point cloud data into sparse tensors includes: The sparse_quantize function in MinkowskiEngine is used to convert the point cloud into a sparse tensor according to the preset voxel size voxels using the SparseTensor function.

4. The method for identifying sky-space cloud scenes based on hierarchical adaptive feature fusion according to claim 1, characterized in that: The feature pyramid network includes five bottom-up convolution blocks, the output features extracted by the previous convolution block are used as the input features of the next convolution block, and the output feature map corresponding to each convolution block is obtained. The features extracted by the first convolution block, the second convolution block and the third convolution block are used as low-level features, and the features extracted by the fourth convolution block and the fifth convolution block are used as high-level features.

5. The method for identifying sky-space cloud scenes based on hierarchical adaptive feature fusion according to claim 1, characterized in that: When fusing the multi-scale features corresponding to the low-level features and the high-level features, the following steps are included: Adjust multi-scale features of different scales to the same scale in spatial resolution; Project the resized feature map to the direction of the adaptive viewing angle to obtain the projection length; Calculate the weight corresponding to each feature according to the projection length; Combine each feature and the corresponding weight to get the fused feature.

6. The method for identifying sky-space cloud scenes based on hierarchical adaptive feature fusion according to claim 1, characterized in that: The loss function of the feature encoding network is: L=ω P · P L+ω O · O L+ T L, in, Represents a given sample pair {S x ,S y }, the point loss of the sample pair is, Represents a given heterogeneous sample pair {S g ,S a }, the overlap loss of the sample pair, T L B is the heterogeneous triplet loss of the entire batch during training, P L is the number of sample pairs that meet the conditions in the entire batch during training. The average value of O L is the number of sample pairs that meet the conditions in the entire batch during training. , M represents the set of triplet samples of constructed point loss, where x and y are the anchor samples and their positive sample index numbers, q represents the supervoxel in sample x and it has corresponding points and non-corresponding points in sample y, p represents any corresponding point of q in sample y, n represents the non-corresponding point of q in sample y with the smallest local feature distance, t represents any non-corresponding point of q in sample y, α is a pre-defined hyperparameter, Represents the index set of corresponding point pairs of sample x in sample y, The set of indices representing the supervoxels of non-corresponding point pairs of sample x in sample y, l F x (q) represents the local feature corresponding to q, l F y (p) represents the local feature corresponding to p, l F y (n) represents the local feature corresponding to n, l F y (t) represents the local feature corresponding to t, represents the set of supervoxels in sample g that have corresponding points in its heterogeneous positive sample a, A represents the set of supervoxels in sample g that do not have corresponding points in its heterogeneous positive sample a. g (s) represents the attention score of each supervoxel in sample g, m is a predefined margin parameter, i is the index number of the anchor sample, j is the index number of the positive sample with the largest global feature distance from anchor sample i in the positive sample set corresponding to anchor sample i, k is the index number of the negative sample with the smallest global feature distance from anchor sample i in the negative sample set corresponding to anchor sample i, a represents the index number of any positive sample corresponding to anchor sample i, b represents the index number of any negative sample corresponding to anchor sample i, B represents the constructed triplet sample set, g F i is the global feature corresponding to the anchor sample i, g F j is the global feature corresponding to sample j, g F k is the global feature corresponding to sample k, g F a is the global feature corresponding to sample a, g F b is the global feature corresponding to sample b, T i To obtain the pose of sample i sub-image, T a To obtain the pose of sample a sub-image, T b To obtain the pose of sample b sub-image, τ TP is the positive sample coordinate distance threshold, τ TN is the negative sample coordinate distance threshold.

7. The method for identifying sky-space cloud scenes based on hierarchical adaptive feature fusion according to claim 1, characterized in that: The GeM generalized average pooling layer is used to aggregate local features to obtain the global feature descriptor.

8. A sky-point cloud scene recognition system based on hierarchical adaptive feature fusion, characterized in that: include: A data acquisition module, used to acquire point cloud data sets; The feature encoding module extracts features based on the point cloud dataset and the trained feature encoding network to obtain global feature descriptors and build a global feature descriptor database. The construction of the feature encoding network includes: The point cloud data is converted into a sparse tensor, and the sparse tensor is used as input to a feature pyramid network composed of multiple 3D sparse convolutions to extract a multi-scale feature map. The multi-scale feature map is divided into low-level features and high-level features according to the number of layers in the feature pyramid network. The multi-scale features corresponding to the low-level features and the high-level features are fused to obtain fused low-level features and high-level features, respectively. The fused low-level features and high-level features are fused to obtain local features; the local features are encoded to obtain encoded local features, and the global feature descriptors are obtained after aggregation; The scene recognition module is used to extract the global feature descriptor of the point cloud to be queried based on the trained feature encoding network; and to search in the global feature descriptor database to obtain the point cloud global feature descriptor that is most similar to the global feature descriptor of the point cloud to be queried.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps in the method for identifying an air-space point cloud scene based on hierarchical adaptive feature fusion as described in any one of claims 1 to 7 are implemented.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps in the method for identifying an air-space point cloud scene based on hierarchical adaptive feature fusion as described in any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Scientific and technological dispatcher scene practice teaching assisting method and system based on large model

    CN120894203A