Feature extraction method, network training method, electronic device and storage medium

By adopting a multi-scale sparse feature processing method in point cloud feature extraction, the problem of insufficient point cloud segmentation efficiency and performance in the existing technology is solved, and more efficient and real-time feature extraction is achieved.

CN114494807BActive Publication Date: 2025-06-27SHENZHEN DEEPROUTE AI CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202111580868.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2025-06-27
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

The prior art is difficult to improve efficiency, performance and memory consumption in large-scale outdoor point cloud segmentation, especially in terms of real-time and efficientness.

Method used

A feature extraction method is proposed. By obtaining the first sparse feature of the input point cloud and sampling it based on a predefined multiple different scales, the full connection layer is used to obtain the weights assigned by the second sparse feature corresponding to each scale, and the second sparse feature of the multiple different scales are fused to obtain the multi-scale geometrically enhanced fusion sparse feature.

Benefits of technology

This method can not only ensure geometric information and context information, but also significantly improve the real-time and efficient nature of point cloud feature extraction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494807B_ABST
    Figure CN114494807B_ABST
Patent Text Reader

Abstract

The present application discloses a feature extraction method, a network training method based on sparse feature processing, an electronic device, and a computer storage medium. The feature extraction method includes: obtaining a first sparse feature of an input point cloud; sampling the first sparse feature based on a plurality of predefined different scales to obtain a second sparse feature corresponding to each scale; using a fully connected layer to obtain weights assigned to the second sparse feature corresponding to each scale; and fusing the second sparse features corresponding to the plurality of different scales by using the assigned weights to obtain a fused sparse feature with multi-scale geometric enhancement. In this way, geometric information and context information can be guaranteed, and the real-time performance and efficiency of point cloud feature extraction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of point cloud learning, and in particular to a feature extraction method, a network training method based on sparse feature processing, an electronic device, and a computer storage medium. Background Art

[0002] Large-scale outdoor point cloud segmentation has become a key task in autonomous driving systems and the like. This task has strict requirements for efficiency, performance, and memory consumption. In recent years, the method of fusing multiple representation methods (such as DRINet, SPVCNN) has gradually become the mainstream. The method of fusing multiple representation methods can make up for the deficiencies of a single representation method. The general framework of multi-representation fusion is to use point-level operations to extract geometric information and use sparse convolution for context learning. This framework will bring additional computational complexity, and the real-time performance and efficiency are relatively low. Summary of the Invention

[0003] The main technical problem solved by this application is: to provide a feature extraction method, a network training method based on sparse feature processing, a terminal device, and a computer storage medium to ensure geometric information and context information, and improve the real-time performance and efficiency of point cloud feature extraction.

[0004] To solve the above technical problem, a technical solution provided by this application is: to provide a feature extraction method, which includes: obtaining a first sparse feature of the input point cloud; sampling the first sparse feature based on a plurality of predefined different scales to obtain a second sparse feature corresponding to each scale; using a fully connected layer to obtain the weights assigned to the second sparse feature corresponding to each scale; and fusing the second sparse features corresponding to multiple different scales by using the assigned weights to obtain a multi-scale geometric enhanced fused sparse feature.

[0005] To solve the above technical problem, a technical solution provided by this application is: to provide a network training method based on sparse feature processing. The network training method includes: obtaining a first training sparse feature of the training point cloud; obtaining a training fused sparse feature based on the first training sparse feature, and performing semantic supervision based on the training fused sparse feature to train the model of the fully connected layer, including repeating the following steps iteratively: sampling the first training sparse feature based on a plurality of predefined different scales to obtain a second training sparse feature corresponding to each scale; using a fully connected layer to obtain the weights assigned to the second training sparse feature corresponding to each scale; fusing the second training sparse features corresponding to multiple different scales by using the assigned weights to obtain a multi-scale geometric enhanced training fused sparse feature; inputting the fused sparse feature into a semantic prediction model for semantic prediction to obtain the predicted semantic classification of the fused sparse feature; and training the model of the fully connected layer based on the predicted semantic classification and the true semantic classification.

[0006] To solve the above technical problems, a technical solution provided by this application is: to provide an electronic device. The electronic device includes a memory and a processor coupled to the memory; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the above feature extraction method and / or network training method.

[0007] To solve the above technical problems, a technical solution provided by this application is: to provide a computer-readable storage medium. The computer storage medium is used to store program data, and when the program data is executed by a processor, it is used to implement the above feature extraction method and / or network training method.

[0008] The feature extraction method provided by this application first obtains the first sparse feature of the input point cloud, samples the first sparse feature based on a plurality of predefined different scales to obtain the second sparse feature corresponding to each scale; uses a fully connected layer to obtain the weights assigned to the second sparse feature corresponding to each scale; and fuses the second sparse features corresponding to multiple different scales by using the assigned weights to obtain a fused sparse feature with multi-scale geometric enhancement. This application uses multi-scale pooling and attention-based scale selection to enhance the geometric characteristics of the point cloud, which can not only ensure geometric information and context information, but also improve the real-time performance and efficiency of point cloud feature extraction. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:

[0010] Figure 1 is a schematic flowchart of an embodiment of the feature extraction method of this application;

[0011] Figure 2 is a specific flowchart of an embodiment of the feature extraction method of this application;

[0012] Figure 3 is Figure 2 a specific flowchart of multi-scale sparse pooling and attention scale selection in the embodiment;

[0013] Figure 4 is Figure 1 a specific flowchart of step S14 in the embodiment;

[0014] Figure 5 is a schematic flowchart of an embodiment of the feature extraction method of this application;

[0015] Figure 6 is Figure 5 Schematic diagram of the processing flow from the second sparse feature to the fused sparse feature in the embodiment;

[0016] Figure 7 Schematic diagram of the process of an embodiment of the feature extraction method of the present application;

[0017] Figure 8 Schematic diagram of the process of an embodiment of the feature extraction method of the present application;

[0018] Figure 9 Schematic diagram of the process of an embodiment of the network training method based on sparse feature processing of the present application;

[0019] Figure 10 is Figure 9 Schematic diagram of the specific process of step S95 in the embodiment;

[0020] Figure 11 Schematic diagram of the process of an embodiment of the network training method based on sparse feature processing of the present application;

[0021] Figure 12 Schematic diagram of the process of an embodiment of the network training method based on sparse feature processing of the present application;

[0022] Figure 13 Schematic diagram of the structure of an embodiment of the electronic device of the present application;

[0023] Figure 14 Schematic diagram of the structure of an embodiment of the computer-readable storage medium of the present application. Detailed implementation manners

[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0025] Next, the present application will be described in detail with reference to the drawings and embodiments.

[0026] Existing methods based on element - point representations directly operate on the original point cloud to maintain and utilize point - by - point geometry (accurate measurement information). However, due to insufficient memory consumption and runtime efficiency, these methods are difficult to apply to outdoor scenarios. Another type of method is based on rasterized representations and uses sparse convolution for feature extraction. In recent years, the method of fusing multiple representation methods has gradually become mainstream. The method of fusing multiple representation methods can make up for the deficiencies of a single representation method, but at the same time, it also brings additional computational complexity. The general framework of multi - representation fusion is to use point - level operations to extract geometric information and sparse convolution for context learning, but the real - time performance and efficiency are relatively low. To improve the real - time performance and efficiency of point - cloud feature extraction, this application proposes a feature extraction method applied to an electronic device. Here, the electronic device of this application can be a server or a system composed of a server and terminal devices cooperating with each other. Correspondingly, each part included in the electronic device, such as each unit, subunit, module, and sub - module, can be all set in the server or can be respectively set in the server and the terminal device.

[0027] Furthermore, the above - mentioned server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules used to provide a distributed server, or as a single software or software module, and no specific limitation is made here. In some possible implementation manners, the feature extraction method of the embodiments of this application can be implemented by a processor calling computer - readable instructions stored in a memory.

[0028] This application first proposes a feature extraction method, as Figures 1 to 3 shown, Figure 1 is a schematic flowchart of an embodiment of the feature extraction method of this application; Figure 2 is a specific flowchart of an embodiment of the feature extraction method of this application; Figure 3 is Figure 2 a specific flowchart of multi - scale sparse pooling and attention scale selection in the embodiment. The feature extraction method of this embodiment specifically includes the following steps:

[0029] Step S11: Obtain the first sparse feature of the input point cloud.

[0030] Point cloud, that is, point cloud data, refers to a set of vectors in a three-dimensional coordinate system. Scanning data is recorded in the form of points, and each point contains three-dimensional coordinates. Some may contain color information or reflection intensity information. In addition to geometric position information, some point clouds also have color information. The color information is usually obtained by a camera to capture a color image, and then the color information of the corresponding pixels at the corresponding positions is assigned to the corresponding points in the point cloud. The acquisition of intensity information is the echo intensity collected by the receiving device of the laser scanner. This intensity information is related to the surface material, roughness, incident angle direction of the target, as well as the emission energy and laser wavelength of the instrument.

[0031] The purpose of point cloud segmentation is to extract different objects in the point cloud, so as to achieve the purpose of dividing and conquering, highlighting the key points, and processing them separately.

[0032] Specifically, first perform multi-scale voxelization on the point cloud to obtain the voxel features of the point cloud; then input the voxel features into a sparse feature encoder to obtain sparse voxel features (sparse voxel representation), that is, the first sparse feature; the sparse feature encoder can quickly and efficiently extract the context features of the point cloud and expand the receptive field of the point cloud (referring to how large a range of the point cloud a feature encodes).

[0033] Among them, voxelization is to convert the geometric form representation of an object into the voxel representation form closest to the object, generating a voxel dataset, which not only contains the surface information of the model, but also can describe the internal attributes of the model. The spatial voxels representing the model are similar to the two-dimensional pixels representing an image, except that it extends from two-dimensional points to three-dimensional cube units, and the three-dimensional model based on voxels has many applications.

[0034] The sparse voxel features can easily apply standard convolution operations to extract local context information. Therefore, in this embodiment, using sparse convolution operations to extract sparse voxel features can maintain high running efficiency and explore more locality.

[0035] One of the great advantages of sparse convolution lies in its sparsity. The convolution operation only considers non-empty voxels. Based on this, in this embodiment, using sparse convolution to construct a sparse feature encoder can quickly expand the receptive field with less computational cost.

[0036] Furthermore, the sparse feature encoder adopts a residual network Bottleneck and replaces the ReLU activation with LeakyReLU activation. For example, in order to maintain high efficiency, all the number of channels of the sparse convolution is set to 64.

[0037] After obtaining the first sparse feature from the sparse feature encoder, further enhance the geometric features of the first sparse feature to enhance the first sparse feature with more geometric guidance and obtain the fused sparse feature.

[0038] Step S12: Sample the first sparse feature based on multiple predefined different scales to obtain the second sparse feature corresponding to each scale.

[0039] The multi-scale sparse pooling algorithm is used to sample the first sparse feature to obtain the second sparse feature corresponding to each scale.

[0040] Multi-scale context information helps enhance the feature extraction ability, especially for point clouds, which have inherent scale invariance and geometry. Multi-scale features bring more geometric enhancement to point cloud features because each voxel scale reflects a specific physical dimension attribute. However, due to the sparsity of point clouds, this embodiment proposes a multi-scale sparse pooling algorithm (shown as Algorithm 1 below) to utilize multi-scale features at the sparse voxel level with a lower memory cost.

[0041]

[0042] Multi-scale sparse pooling is described in Algorithm 1. Given the input first sparse feature F and the predefined pooling scale set S, different scales contain different geometric priors because the scale semantics in point clouds are proportional to the real dimensions, reflecting the physical metric space.

[0043] Specifically, this embodiment uses the depth residual algorithm in the multi-scale sparse pooling algorithm to extract the second sparse features of the first sparse feature at different scales, that is, the depth residual algorithm downsamples the same first sparse feature at different scales.

[0044] The depth residual algorithm provided in this embodiment uses skip connections (skipping unactivated regions) to make up for the loss of texture information during downsampling and maintain the maximum local response of each scale in each sub-region. The local pooled features are upsampled to the input sparse voxel feature scale through the nearest interpolation algorithm. For example, the first sparse feature is downsampled at the finest scale through the depth residual algorithm, and then the second sparse feature at the finest scale is upsampled for scale amplification through the depth residual algorithm. The second sparse feature at the finest scale and the upsampled sparse feature are laterally connected through a 1×1×1 convolutional channel to reduce parameters, and they are stacked step by step to stack the geometric features at different scales together.

[0045] Step S13: Use a fully connected layer to obtain the weights assigned to the second sparse feature corresponding to each scale.

[0046] The fully connected layer is used to assign weights to the second sparse features at multiple scales.

[0047] Put the second sparse features corresponding to different scales into the fully connected layer of the multi-scale attention selection algorithm. The multi-scale attention selection layer automatically assigns different attention weights to the second sparse features of different scales according to the response strength of the sparse features themselves, enabling the entire sparse geometric feature enhancement layer to have layer selection capabilities, thereby solving the problem of unfocused feature learning and improving the real-time performance and efficiency of point cloud feature extraction.

[0048] Step S14: Fuse the second sparse features corresponding to multiple different scales using the assigned weights to obtain a fused sparse feature with multi-scale geometric enhancement.

[0049] In this embodiment, the fully connected layer of the above multi-scale sparse pooling algorithm and multi-scale attention selection algorithm is used to complete the operations of multi-scale pooling and scale selection of the first sparse features output by the sparse feature encoder, obtaining a fused sparse feature with geometric enhancement.

[0050] Different from the prior art, this embodiment uses multi-scale sparse pooling and attention-based scale selection to enhance the geometric characteristics of the point cloud, which can not only ensure geometric information and context information, but also improve the real-time performance and efficiency of point cloud feature extraction.

[0051] Specifically, this embodiment can adopt the method as Figure 4 shown to implement step S14. This embodiment includes step S41 and step S42.

[0052] Step S41: Perform weighted processing on multiple second sparse features and the corresponding assigned weights.

[0053] Step S42: Sum up the multiple weighted second sparse features to obtain a fused sparse feature with multi-scale geometric enhancement.

[0054] In this embodiment, the weights of the second sparse features of multiple scales are obtained using a fully connected layer, and these weights are used to perform weighted averaging on the second sparse features of multiple scales, obtaining a fused sparse feature with multi-scale geometric enhancement that integrates multi-scale geometry.

[0055] This application further proposes another embodiment of the feature extraction method, as Figure 5 、 Figure 2 and Figure 3 shown. The network training method of this embodiment specifically includes the following steps:

[0056] Step S51: Obtain the first sparse features of the input point cloud.

[0057] Step S51 is similar to the above step S11 and will not be elaborated here.

[0058] Step S52: Sample the first sparse feature based on multiple predefined different scales to obtain the second sparse feature corresponding to each scale.

[0059] Step S52 is similar to the above-mentioned step S12 and will not be elaborated here.

[0060] Step S53: Upsample the second sparse features corresponding to multiple different scales so that the lengths of the upsampled second sparse features are equal.

[0061] Upsample the second sparse feature nearest neighbor to the finest scale.

[0062] Specifically, first obtain the minimum scale among the multiple predefined different scales, and then upsample the multiple second sparse features according to the minimum scale to unify the feature lengths of the multiple second sparse features to the feature length of the minimum scale.

[0063] Step S54: Fuse the multiple upsampled second sparse features to obtain the global sparse feature.

[0064] The multi-scale sparse pooling layer obtains the feature vector of each scale, the global pooling feature vector, by a global pooling method for the second sparse feature corresponding to each scale. For example, the second sparse features corresponding to four scales are respectively: Then, the feature vectors corresponding to the four scales are respectively obtained through global pooling as: Among them, the formula for global pooling is as follows:

[0065] V i =global pooling(X i )

[0066] where i ∈ {1, 2, 3, 4}.

[0067] Concatenate the feature vectors V1, V2, V3, V4 of different layers to obtain the global feature vector V g , that is, the global sparse feature. Among them, the formula for feature vector concatenation is as follows:

[0068] V g =concat(V1, V2, V3, V4)

[0069] Step S55: Input the global sparse feature into the fully connected layer to perform weight assignment for the second sparse feature corresponding to each scale.

[0070] Input the global feature vector V g into a fully connected layer in the multi-scale attention selection layer to perform weight assignment for the second sparse feature corresponding to each scale.

[0071] Step S56: Use a fully connected layer to obtain the weights assigned to the second sparse features corresponding to each scale.

[0072] Step S56 is similar to the above-mentioned Step S13.

[0073] Input the global feature vector V g into a fully connected layer in the multi-scale attention selection algorithm to predict the weights of the second sparse features at each scale, and normalize the predicted weights to obtain the attention weights ω of the second sparse features at each scale i (i ∈ {1, 2, 3, 4}). Among them, the output dimension of the fully connected layer is the number of preset scales, that is, the number of predicted attention weights is the same as the number of preset scales. The formula for weight normalization is as follows:

[0074] ω = Sigmoid(fc(V g ))

[0075] where ω = {ω1, ω2, ω3, ω4}, and fc() is the weight prediction function.

[0076] Through the above weight prediction and assignment process, the distribution of the second sparse features at all scales can be perceived, and a weight with global information can be customized for the second sparse features at each scale according to this overall information.

[0077] Step S57: Perform weighted processing on multiple second sparse features and their corresponding assigned weights.

[0078] Step S58: Sum multiple weighted second sparse features to obtain a fused sparse feature with multi-scale geometric enhancement.

[0079] Furthermore, after obtaining the weights corresponding to the second sparse features at each scale, broadcast the weights ω i of the second sparse features at each scale, and multiply them with the corresponding second sparse features X i to achieve attention weighting of the second sparse features at each scale, and obtain the final feature P with attention between multiple scales i . The formula for attention weighting is as follows:

[0080] P i = ω i X i

[0081] where i ∈ {1, 2, 3, 4}.

[0082] After obtaining the multi-scale second sparse features, a simple way to fuse the multi-scale second sparse features is to apply tensor concatenation or tensor summation. Therefore, all features from different scales share the same weights and are treated equally. A general consensus in point clouds is that features from different scales have different geometric priors and focus on scene understanding. From this perspective, driven by SENet and SKNet, a better way to fuse multi-scale features is to apply a re-weighting strategy to each feature channel along the scale dimension to redistribute the importance of each scale.

[0083] Figure 6 As shown, this embodiment first sums all input tensors (second sparse features) fused in the first stage to collect all information from different scales. A scale-by-scale MLP layer (fully connected layer) with sigmoid activation is applied to the result to obtain the attention embedding of each scale. Finally, the tensor multiplication between the attention weight tensor and the multi-scale features (second sparse features) is applied, and then the tensor sum of the scale dimension is obtained (fused sparse features enhanced by multi-scale geometry).

[0084] In another embodiment, if Figure 7 , Figure 2 and Figure 3 As shown, the network training method based on sparse feature processing in this embodiment specifically includes the following steps:

[0085] Step S71: Obtain the first sparse feature of the input point cloud.

[0086] Step S72: sampling the first sparse features based on a plurality of predefined different scales to obtain second sparse features corresponding to each scale.

[0087] Step S73: up-sample the second sparse features corresponding to multiple different scales so that the lengths of the up-sampled second sparse features are equal.

[0088] Step S74: fuse multiple upsampled second sparse features to obtain a global sparse feature.

[0089] Step S75: input the global sparse features into the fully connected layer to distribute the weights of the second sparse features corresponding to each scale.

[0090] Step S76: Use a fully connected layer to obtain the weight of the second sparse feature assignment corresponding to each scale.

[0091] Step S77: performing weighted processing on the plurality of second sparse features and the corresponding assigned weights.

[0092] Step S78: summing up the multiple weighted second sparse features to obtain a multi-scale geometrically enhanced fused sparse feature.

[0093] Steps S71 to S78 are similar to the above-mentioned steps S51 to S58 and will not be elaborated here.

[0094] Step S79: Input the fused sparse features into the semantic prediction model to perform semantic prediction and obtain the predicted semantic classification of the fused sparse features.

[0095] Step S791: Map the predicted semantic classification to the corresponding positions of the input point cloud according to the position mapping relationship between the fused sparse features and the input point cloud to obtain the overall predicted semantic classification of the input point cloud.

[0096] Input the fused sparse features enhanced by multi-scale geometry into the semantic prediction model, and perform semantic prediction classification on the fused sparse features through the semantic prediction model to obtain the overall predicted semantic classification of the input point cloud.

[0097] The semantic prediction model can be a well-trained and mature object detection model, etc., and its model parameters are fixed.

[0098] Among them, the mapping method for mapping the predicted semantic classification to the overall predicted semantic classification is the nearest neighbor interpolation method. Specifically, map the predicted semantic classification to the corresponding positions of the input point cloud according to the above position mapping relationship to obtain the semantic classification of the corresponding positions of the input point cloud; for other positions of the input point cloud, obtain the predicted semantic classification of its nearest neighbor position as the output after interpolation.

[0099] For the final semantic prediction, fuse the multi-stage features from the outputs of each sparse geometric feature enhancement layer by nearest upsampling to the finest voxel granularity scale. To map to the per-point result, apply the nearest interpolation strategy, through which each point is attached with the semantic information or features from its voxel. The entire algorithm is shown in Algorithm 2 below.

[0100]

[0101] Specifically, the semantic prediction model outputs the predicted semantic classification of the fused sparse features. The predicted semantic classification includes information such as confidence. Construct the loss function of the fully connected layer according to the predicted semantic classification and the true semantic classification. Among them, the loss function of the fully connected layer can be the cross-entropy loss function, etc.

[0102] Furthermore, the model of the fully connected layer can also be trained based on the predicted semantic classification and the true semantic classification.

[0103] In each feature extraction process, the model of the fully connected layer is trained using the predicted semantic classification of the fused sparse features and the true semantics of the fused sparse features, so that the fully connected layer adjusts the weights of the second sparse features corresponding to different scales in the next feature extraction, thereby improving the accuracy of the network based on sparse feature processing.

[0104] On the basis of the above-mentioned embodiments, this embodiment performs semantic supervision on the acquired fused sparse features to adjust the model parameters of the fully connected layer so that the weights of the second sparse features at multiple scales are adjusted during the next feature extraction, thereby making the predicted semantic classification obtained by the second sparse features at multiple scales closer to the true semantic classification, thereby improving the prediction accuracy.

[0105] This embodiment uses sparse supervision to handle the supervision of sparse features. Since the fused sparse features stack the second sparse features of different scales, the sparse supervision can be gradually applied to the output voxel features as auxiliary losses. All auxiliary branches can be disabled to maintain runtime efficiency.

[0106] In another embodiment, if Figure 8 , Figure 2 and Figure 3 As shown, the feature extraction method of this embodiment specifically includes the following steps:

[0107] Step S81: Obtain the first sparse feature of the input point cloud.

[0108] Step S82: sampling the first sparse features based on a plurality of predefined different scales to obtain second sparse features corresponding to each scale.

[0109] Step S83: up-sample the second sparse features corresponding to multiple different scales so that the lengths of the up-sampled second sparse features are equal.

[0110] Step S84: Fusing multiple upsampled second sparse features to obtain a global sparse feature.

[0111] Step S85: input the global sparse features into the fully connected layer to distribute the weights of the second sparse features corresponding to each scale.

[0112] Step S86: Use a fully connected layer to obtain the weight of the second sparse feature assignment corresponding to each scale.

[0113] Step S87: performing weighted processing on the plurality of second sparse features and the corresponding assigned weights.

[0114] Step S88: summing up the multiple weighted second sparse features to obtain a multi-scale geometrically enhanced fused sparse feature.

[0115] Step S89: Input the fused sparse features into the semantic prediction model for semantic prediction to obtain the predicted semantic classification of the fused sparse features.

[0116] Step S891: Train the perceptron model based on the predicted semantic classification and the true semantic classification.

[0117] Loop and execute steps S82 to S891 to complete the iteration for a preset number of times.

[0118] The fused sparse features obtained in each iteration are used as the first sparse features for the next iteration and the model parameters for training the fully connected layer in the next iteration.

[0119] For the specific implementation manners of steps S81 to S891, reference can be made to the above embodiments and will not be elaborated here.

[0120] After the iteration for a preset number of times is completed or the loss function converges, execute steps S892 to S895.

[0121] Step S892: Stack the fused sparse features obtained in multiple iterations respectively to obtain the stacked fused sparse features.

[0122] In each iteration process, through the above method, the second sparse features corresponding to each scale have been upsampled to the minimum scale, that is, the finest scale. Therefore, the fused sparse features obtained in each iteration are the fused features at the finest scale; stack the fused sparse features at the finest scale in multiple iteration processes.

[0123] Step S893: Input the stacked fused sparse features into the semantic prediction model.

[0124] Input the stacked fused sparse features into the semantic prediction model; the semantic prediction model can be a trained and mature object detection model, etc., and its model parameters are fixed.

[0125] Step S894: Train the model of the fully connected layer based on the predicted semantic classification and the true semantic classification of the stacked fused sparse features.

[0126] Specifically, obtain the predicted semantic classification of the stacked fused sparse features; map the predicted semantic classification to the corresponding positions of the input point cloud according to the position mapping relationship between the stacked fused sparse features and the input point cloud to obtain the overall predicted semantic classification of the input point cloud; train the model of the fully connected layer based on the overall predicted semantic classification and the true semantic classification.

[0127] Based on the above embodiments, the stacked fused sparse features of multiple fused sparse features obtained by the preset number of iterations in this embodiment are subjected to final semantic supervision to obtain the overall predicted semantic classification of the input point cloud, so as to further adjust the model parameters of the fully connected layer and improve the accuracy of the network.

[0128] This application proposes a network training method based on sparse feature processing, and its corresponding network structure is a geometric information learning framework (GASN) based on a grid representation method.

[0129] The network training method based on sparse feature processing and its network structure described in the embodiments of this application can be applied to the above-mentioned electronic device.

[0130] As Figure 9 shown, Figure 9 is a flowchart of an embodiment of the network training method based on sparse feature processing of this application. Figure 9 The network training method of this embodiment specifically includes the following steps:

[0131] Step S91: Obtain the first training sparse feature of the training point cloud.

[0132] Specifically, the network based on sparse feature processing first performs multi-scale voxelization on the training point cloud to obtain the voxel feature of the training point cloud; then inputs the voxel feature into the sparse feature encoder of the network to obtain the sparse voxel feature (sparse voxel representation), that is, the first training sparse feature; the sparse feature encoder can quickly and efficiently extract the context feature of the point cloud and expand the receptive field of the point cloud (referring to the range of point cloud encoded by a feature).

[0133] The sparse voxel feature can easily apply standard convolution operations to extract local context information. Therefore, this embodiment uses a sparse convolution layer to extract the sparse voxel feature, which can maintain high operating efficiency and explore more locality.

[0134] Furthermore, the sparse feature encoder adopts a residual network Bottleneck and replaces the ReLU activation with LeakyReLU activation. For example, in order to maintain high efficiency, the number of all channels of the sparse convolution is set to 64.

[0135] Step S92: Obtain the training fusion sparse feature based on the first training sparse feature, and perform semantic supervision based on the training fusion feature to train the model of the fully connected layer.

[0136] After obtaining the first training sparse feature from the sparse feature encoder, further input the first training sparse feature into the sparse geometric feature enhancement layer of the network based on sparse feature processing to enhance the first training sparse feature with more geometric guidance to obtain the training fusion sparse feature.

[0137] Specifically, this embodiment can implement step S92 by repeatedly executing step S93 to step S97 in an iterative manner.

[0138] Step S93: Sample the first training sparse feature based on multiple predefined different scales to obtain the second training sparse feature corresponding to each scale.

[0139] Use the multi-scale sparse pooling layer in the sparse geometric feature enhancement layer to sample the first training sparse feature to obtain the second training sparse feature corresponding to each scale.

[0140] This embodiment can use a multi-scale sparse pooling layer (refer to Algorithm 1 in the above embodiment) to utilize multi-scale features at the sparse voxel level with a lower memory cost.

[0141] Specifically, in this embodiment, the first training sparse feature is input into the multi-scale sparse pooling layer, and the depth residual network in the multi-scale sparse pooling layer extracts the second training sparse feature of the first training sparse feature at different scales, that is, the depth residual network performs downsampling of the same first training sparse feature at different scales.

[0142] The depth residual network provided in this embodiment uses skip connections (skipping unactivated regions) to make up for the loss of texture information during downsampling and maintain the maximum local response of each scale in each sub-region. The local pooled features are upsampled to the input sparse voxel feature scale through nearest interpolation. For example, the first training sparse feature is downsampled at the finest scale through the depth residual network, and then, the second training sparse feature at the finest scale is upsampled through the depth residual network for scale amplification. The second training sparse feature at the finest scale and the upsampled training sparse feature are laterally connected through a 1×1×1 convolutional channel to reduce parameters, and they are stacked step by step to stack the geometric features shown at different scales together.

[0143] Step S94: Use a fully connected layer to obtain the weights assigned to the second training sparse feature corresponding to each scale.

[0144] Input the second training sparse features corresponding to different scales into the fully connected layer of the multi-scale attention selection layer in the sparse geometric feature enhancement layer. The multi-scale attention selection layer automatically assigns different attention weights to the second training sparse features at different scales according to the response strength of the sparse features themselves, enabling the entire sparse geometric feature enhancement layer to have layer selection capabilities, thereby solving the problem of unfocused feature learning and improving the real-time performance and efficiency of point cloud learning.

[0145] Step S95: Use the assigned weights to fuse the second training sparse features corresponding to multiple different scales to obtain a multi-scale geometric enhanced training fused sparse feature.

[0146] Among them, the training fusion sparse features obtained in each iteration are used as the first training sparse features for the next iteration and the model parameters for training the fully connected layer in the next iteration, so that the fully connected layer adjusts the weights for the second training sparse features corresponding to different scales in the next iteration, thereby improving the accuracy of the network based on sparse feature processing.

[0147] In this embodiment, the sparse geometric feature enhancement layer uses the output of the sparse feature encoder to complete the operations of multi-scale pooling of sparse features and scale selection, and obtains the geometrically enhanced training fusion sparse features as the input of the sparse feature encoder in the next iteration.

[0148] Different from the prior art, this embodiment uses multi-scale sparse pooling and attention-based scale selection to enhance the geometric characteristics of the point cloud, which can not only ensure geometric information and context information, but also improve the real-time efficiency of point cloud learning.

[0149] Specifically, this embodiment can adopt the method as Figure 10 shown to implement step S95. This embodiment includes step S101 and step S102.

[0150] Step S101: Perform weighted processing on multiple second training sparse features and the corresponding assigned weights.

[0151] Step S102: Sum multiple weighted second training sparse features to obtain multi-scale geometrically enhanced training fusion sparse features.

[0152] In this embodiment, the fully connected layer learns the weights of the multi-scale second training sparse features obtained in the previous step, and uses these weights to perform weighted averaging on the multi-scale second training sparse features to obtain multi-scale geometrically enhanced training fusion sparse features that fuse multi-scale geometry.

[0153] Step S96: Input the training fusion sparse features into the semantic prediction model to perform semantic prediction and obtain the predicted semantic classification of the fusion sparse features.

[0154] In each iteration process, the multi-scale geometrically enhanced training fusion sparse features are input into the semantic prediction model, and the semantic prediction model performs semantic prediction classification on the training fusion sparse features obtained in each iteration to obtain the predicted semantic classification of the training fusion sparse features in each iteration.

[0155] The semantic prediction model can be a trained and mature object detection model, etc., and its model parameters are fixed.

[0156] Step S97: Train the model of the fully connected layer based on the predicted semantic classification and the true semantic classification.

[0157] Among them, the fully connected layer is used to assign weights to the second training sparse features at multiple scales.

[0158] In each iteration process, the model of the fully connected layer is trained by using the predicted semantic classification of the training fused sparse features and the true semantics of the training fused sparse features, so that the fully connected layer adjusts the weights of the second training sparse features corresponding to different scales in the next iteration, thereby improving the accuracy of the network based on sparse feature processing. Specifically, the semantic prediction model outputs the predicted semantic classification of the training fused sparse features, and the predicted semantic classification includes information such as confidence. The loss function of the fully connected layer is constructed according to the predicted semantic classification and the true semantic classification. Among them, the loss function of the fully connected layer can be a cross-entropy loss function, etc.

[0159] Steps S93 to S97 are repeatedly executed. After a preset number of iterations or when the loss function converges, the network training is completed.

[0160] Based on the above embodiments, in this embodiment, semantic supervision is performed on the training fused sparse features obtained in each iteration process to adjust the model parameters of the fully connected layer, so that the weights of the second training sparse features at multiple scales are adjusted in the next iteration, thereby making the predicted semantic classification obtained by the second training sparse features at multiple scales closer to the true semantic classification, that is, improving the prediction accuracy.

[0161] This embodiment uses sparse supervision to handle the supervision of sparse features. Since the training fused sparse features stack the second training sparse features at different scales, sparse supervision can be gradually applied to the output voxel features as an auxiliary loss. All auxiliary branches can be disabled to maintain the runtime efficiency.

[0162] The present application further proposes another embodiment of a network training method based on sparse feature processing, as Figure 11 、 Figure 2 and Figure 3 shown. The network training method of this embodiment specifically includes the following steps:

[0163] Step S111: Obtain the first sparse feature of the training point cloud.

[0164] Step S111 is similar to step S91 above and will not be elaborated here.

[0165] Step S112: Sample the first training sparse feature based on a predefined plurality of different scales to obtain the second training sparse feature corresponding to each scale.

[0166] Step S112 is similar to step S93 above and will not be elaborated here.

[0167] Step S113: Upsample the second training sparse features corresponding to multiple different scales so that the lengths of the upsampled second training sparse features are equal.

[0168] The multi-scale sparse pooling layer upsamples the second training sparse features of the neighbors to the finest scale.

[0169] Specifically, the multi-scale sparse pooling layer first obtains the smallest scale among the predefined multiple different scales, and then upsamples the multiple second training sparse features according to the smallest scale to unify the feature lengths of the multiple second training sparse features to the length of the features at the smallest scale.

[0170] Step S114: Fuse the multiple upsampled second training sparse features to obtain the global training sparse features.

[0171] The multi-scale sparse pooling layer obtains the feature vectors of each scale by a global pooling method for the second training sparse features corresponding to each scale, the global pooling feature vectors. For example, the second training sparse features corresponding to four scales are respectively: Then, the feature vectors corresponding to the four scales are respectively obtained through global pooling as: Among them, the formula for global pooling is as follows:

[0172] V i = global pooling(X i )

[0173] where i ∈ {1, 2, 3, 4}.

[0174] The multi-scale sparse pooling layer concatenates the feature vectors V1, V2, V3, V4 of different layers to obtain the global feature vector V g , that is, the global training sparse features. Among them, the formula for feature vector concatenation is as follows:

[0175] V g = concat(V1, V2, V3, V4)

[0176] Step S115: Input the global training sparse features into the fully connected layer to perform weight allocation for the second training sparse features corresponding to each scale.

[0177] The multi-scale sparse pooling layer inputs the global feature vector V g into a fully connected layer in the multi-scale attention selection layer to perform weight allocation for the second training sparse features corresponding to each scale.

[0178] Step S116: Use the fully connected layer to obtain the weights allocated to the second training sparse features corresponding to each scale.

[0179] Step S116 is similar to the above-mentioned step S94.

[0180] Input the global feature vector V g into a fully connected layer in the multi-scale attention selection layer to predict the weights of the second training sparse features at each scale, and normalize the predicted weights to obtain the attention weights ω of the second training sparse features at each scale i (i ∈ {1, 2, 3, 4}). Among them, the output dimension of the fully connected layer is the number of preset scales, that is, the number of predicted attention weights is the same as the number of preset scales. The formula for weight normalization is as follows:

[0181] ω = Sigmoid(fc(V g ))

[0182] where ω = {ω1, ω2, ω3, ω4}, and fc() is the weight prediction function.

[0183] Through the above weight prediction and assignment process, the multi-scale attention selection layer can perceive the distribution of the second training sparse features at all scales, and customize a weight with global information for the second training sparse features at each scale according to this overall information.

[0184] Step S117: Perform weighted processing on multiple second training sparse features and the corresponding assigned weights.

[0185] Step S118: Sum multiple weighted second training sparse features to obtain a multi-scale geometric enhanced training fusion sparse feature.

[0186] Among them, the training fusion sparse feature obtained in each iteration is used as the first sparse feature in the next iteration and the model parameters of the fully connected layer in the next iteration are trained, so that the fully connected layer adjusts the weights of the second training sparse features corresponding to different scales in the next iteration, thereby improving the accuracy of the network based on sparse feature processing.

[0187] Furthermore, after the multi-scale attention selection layer obtains the weights corresponding to the second training sparse features at each scale, it broadcasts the weights ω i of the second training sparse features at each scale, and multiplies them with the corresponding second training sparse features X i to implement attention weighting on the second training sparse features at each scale, and obtain the final feature P with attention among multiple scales i . The formula for attention weighting is as follows:

[0188] P i = ω i X i

[0189] where \(i\in\{1,2,3,4\}\).

[0190] After obtaining the multi-scale second training sparse features from the multi-scale sparse pooling layer, a simple method for fusing the multi-scale second training sparse features is to apply tensor concatenation or tensor summation. Therefore, all features from different scales share the same weights and are treated equally. A general consensus in point clouds is that features from different scales have different geometric priors and focus on scene understanding. From this perspective, driven by SENet and SKNet, a better method for fusing multi-scale features is to apply a reweighting strategy along the scale dimension for each feature channel, reallocating the importance of each scale.

[0191] Step S119: Input the training fused sparse features into the semantic prediction model for semantic prediction to obtain the predicted semantic classification of the training fused sparse features.

[0192] Step S119 is similar to the above-mentioned step S96 and will not be elaborated here.

[0193] Step S120: Train the model of the fully connected layer based on the predicted semantic classification and the true semantic classification.

[0194] Step S120 is similar to the above-mentioned step S97 and will not be elaborated here.

[0195] Loop through steps S112 to S120. After a preset number of iterations or when the loss function converges, complete the network training.

[0196] In another embodiment, as Figure 12 , Figure 2 and Figure 3 shown, the network training method based on sparse feature processing in this embodiment specifically includes the following steps:

[0197] Step S121: Obtain the first training sparse features of the training point cloud.

[0198] Step S122: Sample the first training sparse features based on a predefined number of different scales to obtain the second training sparse features corresponding to each scale.

[0199] Step S123: Upsample the second training sparse features corresponding to multiple different scales so that the lengths of the upsampled second training sparse features are equal.

[0200] Step S124: Fuse the multiple upsampled second training sparse features to obtain the global training sparse features.

[0201] Step S125: Input the global training sparse features into the fully connected layer for weight assignment of the second training sparse features corresponding to each scale.

[0202] Step S126: Use the fully connected layer to obtain the weights assigned to the second training sparse features corresponding to each scale.

[0203] Step S127: Perform weighted processing on multiple second training sparse features and their corresponding assigned weights.

[0204] Step S128: Sum up multiple weighted second training sparse features to obtain the training fusion sparse features with multi-scale geometric enhancement.

[0205] Step S129: Input the training fusion sparse features into the semantic prediction model to perform semantic prediction and obtain the predicted semantic classification of the training fusion sparse features.

[0206] Step S130: Train the model of the fully connected layer based on the predicted semantic classification and the true semantic classification.

[0207] Loop and execute steps S122 to S130 to complete the iteration for a preset number of times.

[0208] Steps S121 to S130 are similar to the above steps S111 to S120, and will not be elaborated here.

[0209] After the iteration for a preset number of times is completed, execute steps S131 to S135.

[0210] Step S131: Stack the training fusion sparse features respectively obtained from multiple iterations to obtain the stacked training fusion sparse features.

[0211] In each iteration process, the second training sparse features corresponding to each scale have been upsampled to the minimum scale, that is, the finest scale, through the above method. Therefore, the training fusion sparse features obtained in each iteration are the fusion features of the finest scale; stack the training fusion sparse features of the finest scale in multiple iteration processes.

[0212] Step S132: Input the stacked training fusion sparse features into the semantic prediction model.

[0213] Input the stacked training fusion sparse features into the semantic prediction model; the semantic prediction model can be a trained and mature object detection model, etc., and its model parameters are fixed.

[0214] Optionally, it further includes:

[0215] Step S133: Obtain the predicted semantic classification of the stacked training fusion sparse features.

[0216] Step S134: Map the predicted semantic classification to the corresponding positions of the training point cloud according to the superposition training fusion sparse feature and the position mapping relationship of the input point cloud, so as to obtain the overall predicted semantic classification of the training point cloud.

[0217] Among them, the mapping method for mapping the predicted semantic classification to the overall predicted semantic classification is the nearest neighbor interpolation method. Specifically, map the predicted semantic classification to the corresponding positions of the training point cloud according to the above position mapping relationship to obtain the semantic classification of the corresponding positions of the training point cloud; for other positions of the input point cloud, obtain the predicted semantic classification of its nearest neighbor position as the output after interpolation.

[0218] For the final semantic prediction, fuse the multi-stage features of the outputs from each sparse geometric feature enhancement layer by nearest upsampling to the finest voxel granularity scale. To map to the per-point results, apply the nearest interpolation strategy, through which each point is attached with the semantic information or features from its voxel. The entire algorithm refers to Algorithm 2 in the above embodiment.

[0219] Step S135: Train the model of the fully connected layer based on the overall predicted semantic classification and the true semantic classification.

[0220] Specifically, obtain the predicted semantic classification of the superposition training fusion sparse feature; map the predicted semantic classification to the corresponding positions of the training point cloud according to the superposition training fusion sparse feature and the position mapping relationship of the input point cloud, so as to obtain the overall predicted semantic classification of the training point cloud; train the model of the fully connected layer based on the overall predicted semantic classification and the true semantic classification.

[0221] Based on the above embodiment, in this embodiment, the superposition training fusion sparse features obtained by presetting the number of iterations are used for final semantic supervision to obtain the overall predicted semantic classification of the input point cloud, so as to further adjust the model parameters of the fully connected layer and improve the accuracy of the network.

[0222] The above embodiments are only one common case of the present application, and do not limit the technical scope of the present application. Therefore, any minor modifications, equivalent changes or decorations made to the above content based on the essence of the present application scheme still fall within the scope of the technical scheme of the present application.

[0223] Please refer to Figure 13 , Figure 13 which is a schematic structural diagram of an embodiment of the electronic device provided by the present application. The electronic device includes a memory 52 and a processor 51 that are connected to each other.

[0224] The memory 52 is used to store program data, and the processor 51 is used to execute the program data to implement: obtaining the first sparse feature of the input point cloud; sampling the first sparse feature based on a plurality of predefined different scales to obtain the second sparse feature corresponding to each scale; using a fully connected layer to obtain the weights assigned to the second sparse feature corresponding to each scale; and fusing the second sparse features corresponding to a plurality of different scales by using the assigned weights to obtain a fused sparse feature with multi-scale geometric enhancement.

[0225] The processor 51 further executes the program data to implement the feature extraction method and / or network training method of the above embodiments.

[0226] Among them, the processor 51 can also be called a CPU (Central Processing Unit). The processor 51 may be an integrated circuit chip with the ability to process signaling. The processor 51 may also be a general-purpose processor, a digital signaling processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0227] The memory 52 can be a memory stick, a TF card, etc., and can store all the information in the target detection device. The input raw data, computer programs, intermediate operation results, and final operation results are all stored in the memory. It stores and retrieves information according to the positions specified by the controller. With the memory, the string matching prediction device has a memory function and can ensure normal operation. The memory of the string matching prediction device can be classified into a main memory (RAM) and an auxiliary memory (external memory) according to its use, or there is also a classification method of dividing it into an external memory and an internal memory. The external memory is usually a magnetic medium or an optical disc, etc., which can store information for a long time. The RAM refers to the storage component on the motherboard, which is used to store the data and programs being currently executed, but only temporarily stores the programs and data. When the power is turned off or interrupted, the data will be lost.

[0228] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces. The indirect coupling or communication connection of the device or unit may be in an electrical, mechanical or other form.

[0229] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0230] In addition, in each embodiment of the present application, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0231] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a system server, or a network device, etc.) or a processor to execute all or part of the steps of the methods of each embodiment of the present application.

[0232] Please refer to Figure 14 , which is a schematic structural diagram of the computer-readable storage medium of the present application. The storage medium of the present application stores program data 61, and when the program data 61 is executed by a processor, it can achieve: obtaining the first sparse feature of the input point cloud; sampling the first sparse feature based on a plurality of predefined different scales to obtain the second sparse feature corresponding to each scale; using a fully connected layer to obtain the weight assigned to the second sparse feature corresponding to each scale; and fusing the second sparse features corresponding to a plurality of different scales by using the assigned weights to obtain a multi-scale geometric enhanced fused sparse feature.

[0233] The program data 61 can further implement the feature extraction method and / or network training method of the above embodiments when executed by a processor. The program data 61 can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods according to the various embodiments of the present application. The aforementioned storage device includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, or a target detection device such as a computer, a server, a mobile phone, or a tablet.

[0234] The feature extraction method provided by the present application first obtains the first sparse feature of the input point cloud, samples the first sparse feature based on a plurality of predefined different scales to obtain the second sparse feature corresponding to each scale; uses a fully connected layer to obtain the weights assigned to the second sparse feature corresponding to each scale; and fuses the second sparse features corresponding to multiple different scales by using the assigned weights to obtain a fused sparse feature with multi-scale geometric enhancement. The present application uses multi-scale pooling and attention-based scale selection to enhance the geometric characteristics of the point cloud, which can not only ensure geometric information and context information, but also improve the real-time performance and efficiency of point cloud feature extraction.

[0235] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A feature extraction method, characterized in that, The feature extraction method includes: Obtain the first sparse feature of the input point cloud; Sample the first sparse feature based on a plurality of predefined different scales to obtain the second sparse feature corresponding to each scale; Use a fully connected layer to obtain the weights assigned to the second sparse feature corresponding to each scale; Fuse the second sparse features corresponding to multiple different scales by using the assigned weights to obtain a fused sparse feature with multi-scale geometric enhancement; Wherein, before using the fully connected layer to obtain the weights assigned to the second sparse feature corresponding to each scale, the feature extraction method further includes: Upsample the second sparse features corresponding to multiple different scales so that the lengths of the upsampled second sparse features are equal; Fuse the multiple upsampled second sparse features to obtain a global sparse feature; Input the global sparse feature into the fully connected layer for weight assignment of the second sparse feature corresponding to each scale; Wherein, the point cloud refers to a set of vectors in a three-dimensional coordinate system; the scanning data is recorded in the form of points, and each of the points only includes three-dimensional coordinates, or each of the points includes three-dimensional coordinates and color information or reflection intensity information.

2. The feature extraction method according to claim 1, wherein After obtaining the fused sparse feature with multi-scale geometric enhancement, The feature extraction method further includes: Perform nearest neighbor upsampling on the fused sparse feature according to a predefined minimum scale to obtain the final fused sparse feature.

3. The feature extraction method according to claim 1, wherein The upsampling of the second sparse features corresponding to multiple different scales includes: Obtain the minimum scale among the multiple different scales; Upsample the multiple second sparse features according to the minimum scale to unify the feature lengths of the multiple second sparse features to the feature length of the minimum scale.

4. The feature extraction method according to claim 1, wherein The fusion of the second sparse features corresponding to multiple different scales by using the assigned weights includes: Perform weighted processing on the multiple second sparse features and the corresponding assigned weights; Sum the multiple weighted second sparse features to obtain the fused sparse feature with multi-scale geometric enhancement.

5. The feature extraction method according to claim 1, characterized in that The feature extraction method further includes: Input the fused sparse feature into a semantic prediction model to perform semantic prediction and obtain the predicted semantic classification of the fused sparse feature; Map the predicted semantic classification to the corresponding position of the input point cloud according to the position mapping relationship between the fused sparse feature and the input point cloud to obtain the overall predicted semantic classification of the input point cloud.

6. The feature extraction method according to claim 5, wherein The mapping method for mapping the predicted semantic classification to the overall predicted semantic classification is the nearest neighbor interpolation method.

7. The feature extraction method according to claim 5, wherein The feature extraction method further includes: Train the model of the fully connected layer based on the predicted semantic classification and the true semantic classification.

8. The feature extraction method according to claim 7, wherein The feature extraction method further includes: Repeat the following steps in an iterative manner: Sample the first sparse feature based on a plurality of predefined different scales to obtain the second sparse feature corresponding to each scale; Use a fully connected layer to obtain the weights assigned to the second sparse feature corresponding to each scale; Fuse the second sparse features corresponding to multiple different scales by using the assigned weights to obtain a fused sparse feature with multi-scale geometric enhancement; Among them, the fused sparse feature obtained in each iteration serves as the first sparse feature for the next iteration and the model parameters of the fully connected layer in training the next iteration.

9. The feature extraction method according to claim 8, wherein, The feature extraction method further includes: Superposing the fused sparse features respectively obtained in multiple iterations to obtain a superposed fused sparse feature; Inputting the superposed fused sparse feature into the semantic prediction model; Training the model of the fully connected layer based on the predicted semantic classification and the true semantic classification of the superposed fused sparse feature.

10. A network training method based on sparse feature processing, characterized in that, The network training method includes: Obtaining a first training sparse feature of the training point cloud; Obtaining a training fused sparse feature based on the first training sparse feature, and performing semantic supervision based on the training fused sparse feature to train the model of the fully connected layer, including repeatedly performing the following steps in an iterative manner: Sampling the first training sparse feature based on a plurality of predefined different scales to obtain a second training sparse feature corresponding to each scale; Using the fully connected layer to obtain the weights assigned to the second training sparse features corresponding to each scale; Fusing the second training sparse features corresponding to multiple different scales by using the assigned weights to obtain a multi-scale geometric enhanced training fused sparse feature; Inputting the training fused sparse feature into the semantic prediction model for semantic prediction to obtain the predicted semantic classification of the training fused sparse feature; Training the model of the fully connected layer based on the predicted semantic classification and the true semantic classification; Among them, the point cloud refers to a set of vectors in a three-dimensional coordinate system; the scanning data is recorded in the form of points, and each point only contains three-dimensional coordinates, or each point contains three-dimensional coordinates and color information or reflection intensity information.

11. An electronic device, characterized in that, The electronic device includes a memory and a processor coupled to the memory; Among them, the memory is used to store program data, and the processor is used to execute the program data to implement the feature extraction method according to any one of claims 1 to 9 and / or the network training method according to claim 10.

12. A computer storage medium, characterized in that, The computer storage medium is used to store program data, and when the program data is executed by the processor, it is used to implement the feature extraction method according to any one of claims 1 to 9 and / or the network training method according to claim 10.