Rapid three-dimensional construction and identification method for complex scene
Voxelized three-dimensional scenes are generated through multi-sensor data fusion and multi-level feature fusion, and semantic recognition is combined with PV-RCNN network, which solves the problem of low three-dimensional construction and recognition efficiency in complex scenarios, and achieves efficient and accurate semantic understanding, and supports high-speed autonomous passage of intelligent mobile robots or unmanned vehicles in complex scenarios.
Patent Information
- Application Number
- CN202510563263.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-08
AI Technical Summary
The existing technology has low efficiency and insufficient accuracy in three-dimensional construction and recognition in complex scenarios, which cannot meet the needs of intelligent mobile robots or unmanned vehicles for high-speed autonomous passage in complex scenarios.
A multi-sensor data fusion method is adopted, combining lidar, millimeter-wave radar and RGB-D cameras, and voxelized three-dimensional scenes are generated through multi-scale feature extraction and multi-level feature fusion, and semantic recognition is performed using PV-RCNN network.
It significantly improves the three-dimensional construction efficiency and semantic recognition accuracy of complex scenarios, meets the real-time requirements, improves feature expression and semantic understanding capabilities, and supports high-speed autonomous passage of intelligent mobile robots or unmanned vehicles in complex scenarios.
Smart Images

Figure CN120451403A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent mobile robots or unmanned vehicles, and in particular to a method for rapid three-dimensional construction and recognition of complex scenes. Background Art
[0002] In both the wild and urban environments, there are diverse and complex scenes, such as city streets, hills, and mountains. Enabling intelligent mobile robots or unmanned vehicles to navigate autonomously and at high speed in these complex and diverse scenes requires rapid and accurate construction and recognition of these scenes. Several existing approaches have attempted to address this problem, but they have drawbacks.
[0003] For example, some existing methods rely on a single sensor, such as lidar or a camera, to acquire scene information. However, a single sensor cannot provide sufficiently comprehensive and accurate environmental information. While lidar can generate highly accurate 3D point cloud data, it is sensitive to ambient light and is relatively expensive. While cameras can provide rich texture information, they lack depth perception. The limitations of a single sensor lead to insufficient accuracy in scene construction and recognition.
[0004] Some methods attempt to improve the accuracy of scene construction and recognition by fusing multi-sensor data. For example, they combine lidar and camera data, or millimeter-wave radar and camera data. However, existing multi-sensor fusion methods often suffer from the following problems:
[0005] (1) Data fusion and processing issues: Existing data fusion methods are usually complex and slow to process, and cannot meet real-time requirements. Complex feature extraction and matching of point cloud data and image data are required, which is computationally intensive and difficult to complete efficiently within a limited time.
[0006] (2) Feature extraction and expression issues: Existing feature extraction methods can usually only extract features from a single modality and cannot fully utilize the complementarity of multi-sensor data. Extracting only geometric features from point cloud data or only texture features from image data cannot effectively fuse the two features, resulting in insufficient feature expression capabilities;
[0007] (3) 3D construction and recognition accuracy issues: Existing 3D construction methods can usually only generate relatively sparse 3D point clouds and cannot accurately reflect the details of the scene. At the same time, existing semantic recognition methods can usually only perform simple classification of target objects in the scene and cannot provide detailed semantic understanding of complex scenes.
[0008] In summary, existing methods have many shortcomings in the three-dimensional construction and recognition of complex scenes, and cannot meet the needs of intelligent mobile robots or unmanned vehicles for high-speed autonomous passage in complex scenes. Summary of the Invention
[0009] In view of the above analysis, an embodiment of the present invention aims to provide a method for rapid three-dimensional construction and recognition of complex scenes, so as to solve the technical problems of low efficiency of three-dimensional construction of complex scenes and low accuracy of semantic recognition in existing methods.
[0010] The purpose of the present invention is mainly achieved through the following technical solutions:
[0011] The present invention provides a method for rapid three-dimensional construction and recognition of complex scenes, comprising the following steps:
[0012] Obtain point cloud registration maps and RGB-D images of complex scenes;
[0013] Perform multi-scale feature extraction based on the RGB-D image to obtain multi-scale image features; obtain multi-scale scene point cloud features of the point cloud registration image;
[0014] Performing multi-level feature fusion on the multi-scale image features and the multi-scale scene point cloud features to obtain a multi-scale fused feature vector, and performing voxel representation based on the multi-scale fused feature vector to obtain a voxelized three-dimensional scene;
[0015] Reasoning is performed on the voxelized three-dimensional scene to obtain a semantic recognition result corresponding to the complex scene.
[0016] Furthermore, the RGB-D image is subjected to multi-scale feature extraction based on the ResNet network to obtain multi-scale image features at the P2, P3, P4 and P5 levels of the feature pyramid;
[0017] The point cloud registration map is subjected to feature extraction through an improved PoinNet++ model to obtain scene point cloud features S1, S2, S3 and S4 multi-scale scene point cloud features;
[0018] The multi-scale image features at the P2, P3, P4 and P5 levels of the feature pyramid are respectively fused with the corresponding S1, S2, S3 and S4 multi-scale scene point cloud features.
[0019] Furthermore, the preprocessed RGB-D image is subjected to multi-scale feature extraction based on the ResNet network to obtain the corresponding multi-scale image features, including:
[0020] The preprocessed RGB-D image is subjected to feature extraction at different scales by first, second, and third image feature extraction units; wherein the first, second, and third image feature extraction units each include first, second, and third convolutional layers with a convolution kernel of 3×3, a 2×2 maximum pooling layer, and a residual connection from the output of the first convolutional layer to the maximum pooling layer;
[0021] The convolution is used to increase the number of image channels and extract local features of RGB-D images;
[0022] The maximum pooling layer is used to downsample the image and reduce the spatial resolution of the image;
[0023] The residual connection adds the image features input to the first convolutional layer directly to the output of the maximum pooling layer to alleviate the gradient disappearance;
[0024] The preprocessed RGB image features and the output feature maps of the first, second and third image feature extraction units are respectively the P2, P3, P4 and P5 level image features of the feature pyramid.
[0025] Furthermore, the point cloud registration map is subjected to feature extraction based on the improved PoinNet++ model to obtain scene point cloud features, including:
[0026] The point cloud registration graph is sequentially subjected to the first, second, third and fourth cycles to obtain S1, S2, S3 and S4 scene point cloud features respectively;
[0027] Each cycle includes downsampling, point-wise MLP, local feature extraction, and point feature extraction;
[0028] The downsampling of the second and third loops adopts FPS downsampling and voxel grid downsampling respectively; the third and fourth loops adopt adaptive sampling based on point cloud density to obtain the downsampled point cloud;
[0029] The point-by-point MLP extracts the normal and curvature of the point cloud after the sampling point to obtain the basic features of the point cloud;
[0030] A point is randomly selected from the point cloud as the first centroid point. With the centroid point as the center, a radius neighborhood is dynamically divided until the number of points in the neighborhood reaches a predetermined threshold. A neighborhood is obtained, and shared MLP + maximum pooling is applied to the points in the neighborhood to generate local features of the point cloud. The point farthest from the selected centroid point is selected as the new centroid point. The distance calculation and selection of the farthest point are repeated until a neighborhood formed by enough centroid points covers the geometric structure of the entire point cloud, and multiple centroid points are obtained.
[0031] The local features of the point cloud corresponding to each centroid and the corresponding basic features of the point cloud are spliced together, and weighted fusion is performed using the attention mechanism to obtain multi-scale scene point cloud features.
[0032] Furthermore, multi-level feature fusion of the multi-scale image features and the multi-scale scene point cloud features is performed, including:
[0033] The first, second, third and fourth features of P2 and S1, P3 and S2, P4 and S3, and P5 and S4 are respectively fused to obtain fused feature vectors of corresponding scales to form the multi-scale fused feature vector; wherein obtaining the fused feature vector of corresponding scale includes:
[0034] Mapping image pixels in the corresponding scale image features to point cloud coordinates;
[0035] In the point cloud feature space, the features of the 32 point cloud points with the closest Euclidean distance in the neighborhood centered on each centroid are weighted and fused into the cloud feature vector F;
[0036] Calculate the weighted fusion of the feature vectors of the three pixels closest to the centroid point in Euclidean distance to obtain the image feature vector f;
[0037] Concatenate the image feature vector f and the point cloud feature vector F along the channel dimension to form a fusion feature F′=[F,f];
[0038] Perform full connection processing on F′ to obtain the fused feature vector of the corresponding scale.
[0039] Furthermore, performing voxel representation based on the multi-scale fusion feature vector to obtain a voxelized three-dimensional scene includes:
[0040] The multi-scale fusion feature vectors are segmented with different voxel side lengths to obtain the corresponding multi-scale voxel grids;
[0041] The internal features of the multi-scale voxel grid are fused to obtain the fused voxel grid V fusion ;
[0042] The multi-scale image features are back-projected to the fused voxel grid V fusion , get the projected image feature I proj ;
[0043] Calculate V fusion to I proj The guidance weight is calculated based on the guidance weight to obtain the image feature I updated ;
[0044] Based on image features I updated By enhancing voxels through cross attention, the enhanced voxel features are obtained as the global semantic V global ;
[0045] The global semantic V global The current voxel grid is compared with the multi-scale image features acquired in real time to obtain feature differences. When the feature difference between two consecutive iterations is less than a predetermined threshold or reaches the maximum number of iterations, the global semantic features are iteratively updated to complete the voxelized 3D scene construction of the complex scene.
[0046] Furthermore, the PV-RCNN network is used to reason about the voxelized 3D scene to obtain semantic recognition results of the complex scene, including:
[0047] Querying non-empty voxels in the voxelized three-dimensional scene, and convolving each non-empty voxel with its neighboring voxels to obtain a global voxel feature;
[0048] Performing three-dimensional sparse convolution on the global voxel features to generate K candidate regions; performing a pooling operation on each candidate region to extract features of a fixed size to obtain all candidate regions and corresponding candidate region features;
[0049] A fixed number of key points are sampled in each candidate area to generate a local point cloud, and the local point cloud is passed through the first, second, third, and fourth cycles to obtain the corresponding local point cloud feature S4; 1x1 convolution is used to adjust the number of channels of the candidate area to align it with the local point cloud feature; the candidate area feature is spliced with the local point cloud feature to form a multi-scale joint feature F fuse ;
[0050] F fuse After three-layer full connection processing, a three-layer multi-layer perceptron dimensionality reduction is performed to obtain the corresponding classification probability and the semantic recognition results corresponding to complex scenes.
[0051] Furthermore, the multi-scale joint feature F fuse Gradually fuse the features F through three-layer full connection processing final ,as follows:
[0052] F final =FC3(ReLU(FC2(ReLU(FC1(F fuse )))))
[0053] Among them, FC1, FC2, and FC3 are the first, second, and third fully connected layers in the three-layer fully connected processing;
[0054] The fusion feature F final Classification is performed through three layers of perceptrons; the first layer is used for feature dimensionality reduction, and the dimensionality reduction feature F1 is obtained as follows:
[0055] F1=ReLU(FC1(F final ))
[0056] The second layer is used to further fuse the F1 features as follows:
[0057] F2=ReLU(FC2(F1))
[0058] The third layer is used to output semantic recognition probability, as follows:
[0059] p=Softmax(FC3(F2))
[0060] Where p is the probability distribution of each category.
[0061] Furthermore, point cloud data is obtained using a laser radar; speed information is obtained using a millimeter wave radar; and a point cloud registration map is obtained based on the point cloud data and the speed information.
[0062] Furthermore, a binocular visible light camera is used to obtain a depth map and an RGB image;
[0063] The depth map and the RGB image are preprocessed to obtain an RGB-D image.
[0064] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects:
[0065] 1. This invention integrates LiDAR (to obtain high-precision point clouds), millimeter-wave radar (to obtain velocity information), and an RGB-D camera (to obtain depth and texture). It combines the point cloud registration map generated by LiDAR and millimeter-wave radar with the RGB-D image acquired by a binocular visible light camera to address the problem that a single sensor cannot provide sufficiently comprehensive and accurate environmental information. It also employs multi-level feature fusion to significantly reduce computational complexity and meet real-time requirements in complex scenarios.
[0066] 2. This invention performs multi-scale feature extraction on RGB-D images and point cloud registration maps, respectively, to capture detailed information at different levels of the scene. For images, features ranging from fine local to global structure can be obtained; for point clouds, features ranging from local geometry to global shape can be extracted. Compared to existing technologies that can only extract single-modal features, this technical solution provides richer and more complete feature expression.
[0067] 3. This invention uses multi-level feature fusion to enhance feature expression, innovatively fusing multi-scale image features with multi-scale scene point cloud features. This overcomes the limitations of existing technologies in fully utilizing the complementarity of multi-sensor data. Through operations such as pixel and point cloud coordinate mapping and weighted fusion of features within a neighborhood, a fused feature vector is constructed, fully exploiting and integrating the advantages of the two modal features, significantly improving feature expression capabilities.
[0068] 4. This invention uses voxelized representation to process the fused feature vectors, segmenting them into multi-scale voxel grids with varying voxel edge lengths. It then enhances the voxel global semantics through operations such as feature backprojection and cross-attention weight calculation. This is then combined with an iterative optimization loop to dynamically update the global semantic features based on feature differences until convergence conditions are met, generating a high-precision voxelized 3D scene. This addresses the shortcomings of existing 3D construction methods, which often generate sparse point clouds and lack detail.
[0069] 5. This invention leverages the PV-RCNN network to reason about voxelized 3D complex scenes, implementing a complete process from non-empty voxel queries, global feature generation through convolution, candidate region generation, keypoint sampling, and multi-scale joint feature construction. Finally, through fully connected multi-layer perceptron processing, accurate classification probabilities are output. This process overcomes the shortcomings of existing semantic recognition methods, which can only simply classify target objects and have difficulty understanding complex scenes. It provides reliable semantic understanding support for intelligent mobile robots or unmanned vehicles to navigate high-speed autonomously in complex scenes.
[0070] In the present invention, the above-mentioned technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of the present invention will be described in the following description, and some advantages will become apparent from the description or be learned through practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the contents particularly pointed out in the description and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0071] The accompanying drawings are only for the purpose of illustrating particular embodiments and are not to be considered limiting of the present invention. Like reference symbols denote like parts throughout the drawings.
[0072] Figure 1 This is a flow chart of a method for rapid three-dimensional construction and recognition of complex scenes in an embodiment of the present invention;
[0073] Figure 2 A schematic diagram of rapid 3D construction and semantic recognition of complex scenes in an embodiment of the present invention;
[0074] Figure 3a Schematic diagram of a multi-scale image feature extraction mechanism in an embodiment of the present invention;
[0075] Figure 3b Schematic diagram of a multi-scale image feature pyramid for an RGB-D image in an embodiment of the present invention;
[0076] Figure 4a This is a flowchart of multi-scale scene point cloud feature extraction of a point cloud registration graph in an embodiment of the present invention;
[0077] Figure 4b Schematic diagram of the specific process of the first, second, third and fourth cycles of multi-scale scene point cloud feature extraction of the point cloud registration graph in an embodiment of the present invention;
[0078] Figure 5 Schematic diagram of multi-level feature fusion of multi-scale image features and multi-scale scene point cloud features in an embodiment of the present invention;
[0079] Figure 6Schematic diagram of the adaptive conversion mechanism between image pixel coordinates and point cloud space in an embodiment of the present invention;
[0080] Figure 7 Schematic diagram of the fusion relationship between image pixels and point cloud points in an embodiment of the present invention. DETAILED DESCRIPTION
[0081] The preferred embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings, wherein the accompanying drawings constitute a part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, and are not used to limit the scope of the present invention.
[0082] The purpose of this invention is to design a method for rapid three-dimensional construction and recognition of complex scenes, which can perform rapid three-dimensional construction and accurate semantic recognition of complex scenes such as cities, mountains, and hills.
[0083] First, laser radar / millimeter-wave radar and a binocular visible light camera are used to acquire multi-source scene information, and feature extraction and feature fusion of radar point cloud information and image information are performed. On this basis, the fused features are represented in three-dimensional voxels to construct feature voxels, completing the rapid construction of complex three-dimensional scenes. To ensure rapidity, feature extraction such as three-dimensional convolution is applied in parallel, and the PV-RCNN network (Point-Voxel Region-based Convolutional Neural Network) is used to recognize three-dimensional point cloud scenes, completing the rapid semantic recognition of complex scenes such as cities, hills, mountains, forests, deserts, grasslands, waters, and vehicles.
[0084] A specific embodiment of the present invention discloses a method for rapid three-dimensional construction and recognition of complex scenes, such as Figure 1 and Figure 2 As shown, the following steps are included:
[0085] Step 1: Obtain the point cloud registration map and RGB-D image of the complex scene;
[0086] Step 2: Perform multi-scale feature extraction based on the RGB-D image to obtain multi-scale image features; obtain multi-scale scene point cloud features of the point cloud registration image;
[0087] Step 3: Perform multi-level feature fusion on the multi-scale image features and the multi-scale scene point cloud features to obtain a multi-scale fusion feature vector, and perform voxel representation based on the multi-scale fusion feature vector to obtain a voxelized three-dimensional scene;
[0088] Step 4: Inferring the voxelized three-dimensional scene to obtain a semantic recognition result corresponding to the complex scene.
[0089] Step S1, specifically.
[0090] Deploy three types of sensors directly in front of an intelligent mobile robot or unmanned vehicle: a binocular visible light camera, a lidar sensor above the camera, and a millimeter-wave radar sensor below. Multi-sensor data acquisition includes:
[0091] Point cloud data is obtained using a laser radar; speed information is obtained using a millimeter wave radar; and a point cloud registration map is obtained based on the point cloud data and the speed information.
[0092] The LiDAR acquires a high-precision point cloud of the scene, and the millimeter-wave radar acquires the velocity information of the object. The high-precision point cloud and velocity information acquired by the LiDAR and millimeter-wave radar are aligned to obtain a point cloud registration map, providing high-precision geometric information (such as the shape and distance of the object).
[0093] Use binocular visible light cameras to obtain depth maps and RGB images;
[0094] The depth map and the RGB image are preprocessed to obtain an RGB-D image.
[0095] Preprocessing the depth map and the RGB image to obtain an RGB-D image includes:
[0096] Performing Z-Scroe normalization on the RGB channels of the RGB image to obtain a normalized RGB image;
[0097] The depth map is spatially aligned with the normalized RGB image; the channels of the RGB image are concatenated with the channels of the depth map, and the number of channels is adjusted by convolution to obtain an RGB-D image.
[0098] The image features of RGB-D images are 540×200×3, with a resolution of 540×200 and three channels of R, G, and B.
[0099] The first step is to perform RGB standardization and Z-Score normalization on the RGB channels (mean is 0 and variance is 1);
[0100] The second step is to align the depth map with the RGB image space and normalize it (range [0,1]);
[0101] The third step is to concatenate the RGB (3 channels) and depth map (1 channel) channels into 540×200×4, and then adjust the number of channels to 64 through 1×1 convolution. The preprocessed RGB-D image features are 540×200×64.
[0102] Step S1 acquires and preprocesses RGB-D image data by deploying multiple sensors to generate RGB-D images and point cloud registration maps, providing a high-precision geometry and texture information foundation for subsequent voxelized 3D scene construction and semantic recognition of complex scenes.
[0103] Step S2 includes steps S21-S22.
[0104] Feature extraction is performed on the point cloud registration map and RGB-D image information.
[0105] The RGB-D image is subjected to multi-scale feature extraction based on the ResNet network to obtain multi-scale image features at the P2, P3, P4 and P5 levels of the feature pyramid;
[0106] The point cloud registration map is subjected to feature extraction through an improved PoinNet++ model to obtain scene point cloud features S1, S2, S3 and S4 multi-scale scene point cloud features;
[0107] The multi-scale image features at the P2, P3, P4 and P5 levels of the feature pyramid are respectively fused with the corresponding S1, S2, S3 and S4 multi-scale scene point cloud features.
[0108] Step S21: Perform multi-scale feature extraction based on the RGB-D image to obtain multi-scale image features.
[0109] The RGB-D image is acquired by a binocular visible light camera, and then the ResNet residual network is used to extract multi-level features from the RGB-D image through a multi-scale convolutional neural network to output multi-scale image features.
[0110] The preprocessed RGB-D image is subjected to multi-scale feature extraction based on the ResNet network to obtain the corresponding multi-scale image features, including:
[0111] The preprocessed RGB-D image is subjected to feature extraction at different scales by first, second, and third image feature extraction units; wherein the first, second, and third image feature extraction units each include first, second, and third convolutional layers with a convolution kernel of 3×3, a 2×2 maximum pooling layer, and a residual connection from the output of the first convolutional layer to the maximum pooling layer;
[0112] The convolution is used to increase the number of image channels and extract local features of RGB-D images;
[0113] The maximum pooling layer is used to downsample the image and reduce the spatial resolution of the image;
[0114] The residual connection adds the image features input to the first convolutional layer directly to the output of the maximum pooling layer to alleviate the gradient disappearance;
[0115] The preprocessed RGB image features and the output feature maps of the first, second and third image feature extraction units are respectively the P2, P3, P4 and P5 level image features of the feature pyramid.
[0116] like Figure 3a As shown in the figure, multi-scale image features are extracted from RGB-D images, and a hierarchical feature extraction mechanism in convolutional neural networks is used to expand the convolution receptive field layer by layer while reducing the resolution of the feature map and increasing the dimension of the feature vector.
[0117] First, the preprocessed RGB-D image is input, and then high-level features are gradually extracted through the first, second, and third image extraction units to extract multi-scale image features.
[0118] This process gradually reduces the spatial resolution from 540×200 to 68×25, compresses the spatial information but retains the position information; the number of feature channels gradually increases from 64 to 512, the receptive field gradually expands, the color complexity increases, and the semantic information is enhanced.
[0119] The final result is a three-dimensional tensor of 68×25×512, where the spatial dimension (resolution 68×25) retains position information and the high number of channels (512) encodes rich semantics.
[0120] The first image feature extraction unit:
[0121] Feature map size change: 540×200×64 (input) → 270×100×128 (output).
[0122] a) First, second and third convolutional layers:
[0123] Based on the preprocessed RGB-D image feature map size of 540×200×64; three consecutive layers of Conv3×3 convolution (each convolution layer is followed by ReLU activation), the number of channels increases from 64 to 128; the output feature map size is 540×200×128, the resolution remains unchanged, and the number of channels is doubled;
[0124] The convolution kernel size of the three convolution layers is 3×3, which is used to extract local features of RGB-D images.
[0125] b) Maximum pooling operation MaxPool2×2:
[0126] Downsampling is performed with a 2×2 window stride of 2, and the output feature map size is 270×100×128, the resolution is halved, and the number of channels remains unchanged.
[0127] Reduce spatial resolution and expand receptive field.
[0128] Second image feature extraction unit:
[0129] Feature map size change: 270×100×128 (input) → 135×50×256 (output).
[0130] a) First, second, and third convolutional layers: Based on a feature map of size 270 × 100 × 128, three consecutive layers of Conv3 × 3 convolution are performed (resolution remains unchanged, but the number of channels is doubled from 128 to 256); the resulting feature map size is 270 × 100 × 256;
[0131] b)MaxPool2×2:
[0132] Downsampling is performed, the resolution is halved, the number of channels remains unchanged, and the output feature map size is 135×50×256.
[0133] The third image feature extraction unit:
[0134] Feature map size change: 135×50×256 (input) → 68×25×512 (output).
[0135] a) The first, second, and third convolutional layers are based on a feature map of size 135 × 50 × 256. Three consecutive layers of Conv3 × 3 (number of channels 256 → 512) are repeated with the same resolution and doubled number of channels. The resulting feature map size is 135 × 50 × 512.
[0136] b)MaxPool2×2:
[0137] Downsampling is performed, the resolution is halved, the number of channels remains unchanged, and the feature map size is 68×25×512.
[0138] Finally, the multi-scale image features are obtained with a size of 68×25×512.
[0139] like Figure 3b As shown, using the feature pyramid, the feature levels are defined as follows:
[0140] P2 level: RGB-D image resolution after preprocessing, feature map size 540×200×64, retaining rich detail information, for high-resolution feature maps;
[0141] P3 level: medium resolution, feature map size 270×100×128, balance between semantics and details, medium resolution feature map;
[0142] P4 level: low resolution, strong semantics, feature map size 135×50×256, captures high-level semantics, low-resolution feature map;
[0143] P5 level: lowest resolution, feature map size 68×25×512, abstract semantics, lowest resolution.
[0144] The function of step S21 is to extract multi-scale features from the RGB-D image and generate a feature pyramid containing semantic information at different levels, which provides a basis for subsequent feature fusion.
[0145] Step S22: Obtain multi-scale scene point cloud features of the point cloud registration image.
[0146] The point cloud registration map is obtained through LiDAR / millimeter wave radar, and then PointNet++ is used to extract features from the point cloud registration map data to capture local and global scene geometric features, thereby obtaining multi-scale scene point cloud features.
[0147] The point cloud registration map is subjected to feature extraction based on the improved PoinNet++ model to obtain scene point cloud features, including:
[0148] The point cloud registration graph is sequentially subjected to the first, second, third and fourth cycles to obtain S1, S2, S3 and S4 scene point cloud features respectively;
[0149] Each cycle includes downsampling, point-wise MLP, local feature extraction, and point feature extraction;
[0150] The downsampling of the third and second loops adopts FPS downsampling and voxel grid downsampling respectively; the third and fourth loops adopt adaptive sampling based on point cloud density to obtain the downsampled point cloud;
[0151] The point-by-point MLP extracts the normal and curvature of the point cloud after the sampling point to obtain the basic features of the point cloud;
[0152] A point is randomly selected from the point cloud as the first centroid point. With the centroid point as the center, a radius neighborhood is dynamically divided until the number of points in the neighborhood reaches a predetermined threshold. A neighborhood is obtained, and shared MLP + maximum pooling is applied to the points in the neighborhood to generate local features of the point cloud. The point farthest from the selected centroid point is selected as the new centroid point. The distance calculation and selection of the farthest point are repeated until a neighborhood formed by enough centroid points covers the geometric structure of the entire point cloud, and multiple centroid points are obtained.
[0153] The local features of the point cloud corresponding to each centroid and the corresponding basic features of the point cloud are spliced together, and weighted fusion is performed using the attention mechanism to obtain multi-scale scene point cloud features.
[0154] The point cloud registration map is extracted through PoinNet++. The specific process is through the first, second, third and fourth loop iterations, such as Figure 4aAs shown in Figure 1, the first, second, third, and fourth loops are connected in series to extract and fuse multi-scale features layer by layer, ultimately generating multi-scale scene point cloud features S1, S2, S3, and S4 containing global semantics and local details.
[0155] The first, second, third and fourth loops all include four steps: downsampling, point-by-point MLP (Multilayer Perceptron), local feature extraction, and point feature generation. Figure 4b As shown, the details are as follows:
[0156] 1) Downsampling: Use FPS (Farthest Point Sampling) or voxel grid downsampling to gradually reduce the point cloud density and retain key geometric structures.
[0157] First loop: downsample using FPS to preserve global structure (sampling rate is 1 / 4);
[0158] Second loop: downsampling using voxel grid (voxel size 0.1m) to preserve local details;
[0159] The third and fourth loops: dynamically adjust the sampling rate for adaptive sampling (e.g. based on point density adaptation).
[0160] Downsampling output: The point cloud after downsampling is N / 4 points after the first cycle sampling, N / 16 points after the second cycle sampling, and N / 2 points after the third cycle sampling. 6 , after the fourth cycle sampling is N / 2 8 .
[0161] 2) Point-wise MLP: Apply ML with shared weights to each point to extract point-wise basic features (such as the normal and curvature of the point).
[0162] The point-by-point MLP part outputs the MLP features of each point, which is lightweight and preserves local details.
[0163] 3) Local feature extraction: First, group the points and divide them into radius neighborhoods with the centroid as the center. Then apply shared MLP+maximum pooling to the neighborhood points to generate local descriptors to capture local geometric details (such as edges and surfaces).
[0164] Grouping radius: 0.05m for the first cycle, 0.1m for the second cycle, and gradually doubled for the third and fourth cycles.
[0165] The number of neighborhood points, for example, each group of neighborhood points is 32-64 points.
[0166] Local features extract local features of some output points.
[0167] 4) Point feature generation: The local features of the point are concatenated with the MLP features and weighted fused through the attention mechanism to generate robust point-by-point features (fusing local geometry and global context).
[0168] The point feature generation part outputs enhanced point features.
[0169] For feature transfer between the first, second, third, and fourth loops: feature upsampling is used to transfer high-level features back to the original resolution through interpolation or feature propagation, retaining detailed information and providing richer input for the next layer of loop.
[0170] After the first, second, third and fourth loop iterations, each point contains multi-scale and multi-level features, and the obtained multi-scale scene point cloud features continue to perform subsequent feature fusion and segmentation tasks.
[0171] Through four iterations, multi-scale features are extracted and fused layer by layer, ultimately generating a multi-scale scene point cloud feature that combines global semantics and local details. The first, second, third, and fourth feature outputs are S1, S2, S3, and S4, respectively. This step combines the layered approach and attention mechanism of PointNet++, resulting in outstanding performance in complex scenes.
[0172] The layer-by-layer feature extraction mechanism of the point cloud registration map has three advantages:
[0173] The first is multi-scale feature fusion, which covers features from local details to global structures through four cycles of the first, second, third and fourth cycles;
[0174] Then there is adaptive sampling, which dynamically adjusts the sampling rate and grouping radius to adapt to scenarios with different point cloud densities;
[0175] Finally, end-to-end optimization is performed to jointly optimize all parameters through back-propagation.
[0176] The purpose of step S22 is to perform multi-level feature extraction on the point cloud registration map through the improved PointNet++ model to generate multi-scale scene point cloud features S1, S2, S3 and S4 containing local details and global semantic information, providing a basis for subsequent feature fusion and scene understanding.
[0177] Step S2 extracts multi-scale features, obtains multi-scale image features based on RGB-D images, and obtains multi-scale scene point cloud features based on the point cloud registration map, constructing multi-level features with strong complementarity and rich semantics, providing high-precision and robust representation for subsequent voxelized three-dimensional scene construction and semantic recognition of complex scenes.
[0178] Step S3 includes steps S31-S32.
[0179] Step S31: Perform multi-level feature fusion on the multi-scale image features and the multi-scale scene point cloud features to obtain a multi-scale fusion feature vector.
[0180] Performing multi-level feature fusion on the multi-scale image features and the multi-scale scene point cloud features, including:
[0181] The first, second, third and fourth features of P2 and S1, P3 and S2, P4 and S3, and P5 and S4 are respectively fused to obtain fused feature vectors of corresponding scales to form the multi-scale fused feature vector; wherein obtaining the fused feature vector of corresponding scale includes:
[0182] Mapping image pixels in the corresponding scale image features to point cloud coordinates;
[0183] In the point cloud feature space, the features of the 32 point cloud points with the closest Euclidean distance in the neighborhood centered on each centroid are weighted and fused into the cloud feature vector F;
[0184] Calculate the weighted fusion of the feature vectors of the three pixels closest to the centroid point in Euclidean distance to obtain the image feature vector f;
[0185] Concatenate the image feature vector f and the point cloud feature vector F along the channel dimension to form a fusion feature F′=[F,f];
[0186] Perform full connection processing on F′ to obtain the fused feature vector of the corresponding scale.
[0187] The fused feature vectors of four corresponding scales form a multi-scale fused feature vector.
[0188] Adopting a multi-level fusion structure, such as Figure 5 As shown in Figure 1, the feature vectors P2, P3, P4, and P5 of the image pixels are fused with the feature vectors S1, S2, S3, and S4 of the point cloud points to obtain new point cloud features. This fusion structure ensures that the multi-scale image feature extraction in step S21 and the multi-scale scene point cloud feature extraction in step S22 are independent of the multi-level feature fusion process, so that they will not be interfered with by the fused features.
[0189] Multi-level feature fusion, such as Figure 5 shown.
[0190] After obtaining the RGB-D image and point cloud registration map, the point cloud feature space that already has accurate spatial structure information is enhanced by migrating and fusing the multi-scale image features with richer texture information into the multi-scale scene point cloud feature space.
[0191] Figure 5 There are four fusion processes in this paper. The details of each fusion are as follows:
[0192] 1) A layer-by-layer adaptive conversion mechanism for image pixel coordinates, mapping image pixel coordinates to the corresponding point cloud coordinate space: In the image data input to the network model, the coordinate information in the image-point cloud registration map is used to establish the correspondence between the image pixels to be fused and the point cloud points. However, as the network layer deepens, after the hierarchical feature extraction of the RGB-D image, the dimension of its feature image increases layer by layer, but the resolution decreases layer by layer. In order to fully utilize the advantages of the hierarchical feature extraction structure, a layer-by-layer feature fusion mechanism is adopted, and the image features output by each layer of the ResNet model are fused into the point cloud feature space output by the PontNet++ model. The first, second, third, and fourth features of P2 are fused with S1, P3 with S2, P4 with S3, and P5 with S4, respectively.
[0193] like Figure 6 As shown in Figure 1, the adaptive conversion mechanism from image pixel coordinates to point cloud space realizes the mapping from 2D image pixels to 3D point cloud coordinates through dynamic parameter matrix and geometric transformation.
[0194] The left matrix of the input terminal represents the pixel coordinates (including depth information), and each grid represents the 3D coordinates of a pixel point (X i ,Y i ,Z i )), where X i ,Y i is the pixel coordinate, Z i The matrix on the right represents the adaptive transformation parameters a1, a2, a3, and a4, which are used to dynamically adjust the projection weights. The "×" between the two matrices represents the weighted fusion operation of the two. The output represents the conversion to point cloud space coordinates. Through matrix operations, the image pixel coordinates are mapped to the point cloud space, and the corresponding 3D point cloud coordinates (X5, Y5, Z5) are obtained.
[0195] The image pixel coordinates in P2 should set the point cloud coordinates in the point cloud space in S1;
[0196] The image pixel coordinates in P3 should set the point cloud coordinates in the point cloud space in S2;
[0197] The image pixel coordinates in P4 should set the point cloud coordinates in the point cloud space in S3;
[0198] The image pixel coordinates in P5 should set the point cloud coordinates in the point cloud space in S4.
[0199] 2) Construction of the fusion relationship between image pixels and point cloud points: In the image and point cloud feature extraction structure, each layer of feature space is obtained by downsampling and local feature fusion based on the feature space of the previous layer.
[0200] Therefore, the resolution of image pixels and the number and distribution of point cloud points change layer by layer. Therefore, before the layer-by-layer feature fusion mechanism can achieve feature fusion, it is necessary to reconstruct the correspondence between the image pixels to be fused and the point cloud points.
[0201] In the point cloud feature space, the three pixel points with the closest Euclidean distance to the centroid point cloud point are selected to form a fusion relationship.
[0202] For each target point obtained by downsampling the point cloud layer, the overall features of the 32 point cloud points with the closest Euclidean distance within the spherical neighborhood formed by the point as the center and the features of the 3 image pixels with the closest Euclidean distance to the point will be fused.
[0203] The feature vectors of the above points and pixels will be fused into a new feature vector and assigned to the center point.
[0204] The relationship between image pixels and point cloud fusion is as follows Figure 7 As shown in the figure, 1x1xh(m) represents the size parameters of the point cloud block. The length and width of the point cloud block are both 1 meter, indicating the horizontal coverage of the point cloud block. h is the height, which can change dynamically.
[0205] The orange dot represents the centroid point selected from the original point cloud, the orange circle represents the spherical neighborhood centered on the centroid point, the blue dot represents the initial point cloud, the green dot represents the image pixel point, and the red dot represents the three pixels closest to the Euclidean distance of the point cloud center point.
[0206] Extract the initial point cloud features and downsample to get the centroid point (see Figure 7 (a) Point cloud feature extraction downsampling), converting image pixels into point cloud coordinates, the distribution of image pixels in the point cloud coordinate system (see Figure 7 (b) Image pixels are distributed in the point cloud coordinate system), the overall features of the 32 nearest point cloud points in the spherical neighborhood formed by the centroid point as the center, and the features of the 3 nearest image pixels to the point are fused into a new feature vector and assigned to the centroid point (see Figure 7 (c) Corresponding fusion relationship between image pixels and point cloud).
[0207] 3) Fusion Mechanism for Pixel and Point Features: Once the correspondence between the image pixels and point cloud points at each layer is known, a feature fusion mechanism is needed to achieve layer-by-layer fusion of the two data feature vectors. By leveraging the rich texture information of image features to enhance the rich spatial structure of point cloud features, this mechanism can highlight the feature differences between different object categories and help optimize the accuracy of point cloud semantic segmentation.
[0208] This network model adopts a layer-by-layer fusion mechanism. After achieving layer-by-layer feature extraction of the image and point cloud, it is necessary to fuse the corresponding image pixel feature vector with the point feature vector point by point. The steps for achieving image feature fusion at any point cloud point are as follows:
[0209] a. Each point cloud target point to be fused with image features has obtained a new feature vector F through the point cloud feature extraction mechanism of this layer;
[0210] b. Based on the established correspondence between the image pixels to be fused and the point cloud points, use the inverse weighted average of the distance to extract the overall features of the three pixel feature vectors corresponding to the point, which is used as the image feature vector f to be fused to the centroid point cloud point. The specific steps are as follows:
[0211] First, calculate the distance between the three pixels and use the Euclidean distance to calculate the distance between the three pixels. Assuming that the positions of the three pixels are P1(x1,y1), P2(x2,y2), and P3(x3,y3), the distance calculation formula is:
[0212]
[0213] The weight of each pixel is the sum of the reciprocals of its distances to the other two pixels, and the calculation formula is:
[0214]
[0215] Normalize the weights so that their sum is 1, and the calculation formula is:
[0216]
[0217] Use the normalized weights to perform weighted averaging on the feature vectors f1, f2, and f3 of the three pixels to obtain the image feature vector f. The calculation formula is:
[0218] f=w1′·f1+w2′·f2+w3′·f3 Formula (6)
[0219] c. Use the concatenation operation to concatenate the weighted image feature vector f and the point cloud feature vector F along the channel dimension to form a fusion feature F′=[F,f];
[0220] d. In order to improve the fusion degree of the two features, the fully connected layer is used to process F′, and the resulting fusion feature vector F fusion The function of the fully connected layer is to enhance the representation ability of features through multi-layer nonlinear transformation and fuse information of different dimensions. The calculation process can be expressed as:
[0221] F fusion=FC(F′)=σ(W2·ReLU(W1F′+b1)+b2) Formula (7)
[0222] Among them, W1 and b1 are the weight and bias of the first layer, W2 and b2 are the weight and bias of the second layer, and σ is the activation function Sigmoid function of the output layer.
[0223] In the feature vector output by the fully connected layer, each value of the feature vector is the weighted average sum of all values in the input vector based on different weights, so F fusion Each value in contains the feature information of the image and point cloud. The final output point cloud feature vector contains both the spatial structure information of the point cloud and the texture information of the image.
[0224] Step S32: performing voxel representation based on the multi-scale fusion feature vector to obtain a voxelized three-dimensional scene.
[0225] Through the multi-scale image feature extraction of the obtained point cloud information and image information, the multi-scale scene point cloud feature extraction, and the fusion with multi-level features, the fused features are represented by three-dimensional voxelization, feature voxels are constructed, and the rapid voxelization of the scene three-dimensional scene construction is completed.
[0226] Performing voxel representation based on the fused feature vector to obtain a voxelized three-dimensional scene includes:
[0227] The multi-scale fusion feature vectors are segmented with different voxel side lengths to obtain the corresponding multi-scale voxel grids;
[0228] The internal features of the multi-scale voxel grid are fused to obtain the fused voxel grid V fusion ;
[0229] The multi-scale image features are back-projected to the fused voxel grid V fusion , get the projected image feature I proj ;
[0230] Calculate V fusion to I proj的 Guide weight, based on the guide weight calculation to obtain the image feature I updated ;
[0231] Based on image features I updated By enhancing voxels through cross attention, the enhanced voxel features are obtained as the global semantic V global ;
[0232] The global semantic V globalThe current voxel grid is compared with the multi-scale image features acquired in real time to obtain feature differences. When the feature difference between two consecutive iterations is less than a predetermined threshold or reaches the maximum number of iterations, the global semantic features are iteratively updated to complete the voxelized 3D scene construction of the complex scene.
[0233] The point aggregation and feature set extraction operations in conventional point cloud operations are very time-consuming, which will lead to the inability to guarantee the real-time performance of the overall algorithm; conventional simple voxelization can convert irregular three-dimensional point clouds into regular voxel grid representations and use normalization operations such as three-dimensional convolution for scene understanding, but it will result in excessive occupation of video memory space and a certain amount of information loss.
[0234] This method uses a voxel representation method that combines the two. In order to reduce the problem of large memory consumption caused by directly encoding the huge voxel scene, the multi-scale scene point cloud features S1, S2, S3 and S4 after each layer of fusion in step S22 are voxelized and represented respectively. The voxel side length defined in each layer is different. The fused feature vectors of the first, second, third and fourth features are r1, r2, r3 and r4 as the voxel side length respectively, and a multi-scale voxel grid is obtained;
[0235] For example, r1 = 0.1m; r2 = 0.2m; r3 = 0.3m; r4 = 0.5m, and finally the multi-scale voxel grids corresponding to the multi-scale scene point cloud features S1, S2, S3, and S4 are obtained. The point cloud voxel representation method for each layer of the multi-scale voxel grids corresponding to S1, S2, S3, and S4 is the same, and the steps for one layer are described below:
[0236] First, the farthest point sampling method is used to select N points and then operate the features of the scene voxels based on these N points to summarize the entire scene information. The specific steps are to uniformly select N representative points from the point cloud as voxelized seed points, calculate the distances from the remaining points to the selected points, and each time select the point farthest from the current selected point set until all N points are selected.
[0237] Then, the N sampled point clouds are converted to voxel grids, and the voxel index of each key point obtained by sampling the farthest point is calculated according to the voxel size. The calculation formula is:
[0238]
[0239] Among them, x, y, and z represent the point cloud coordinates, x min 、y min 、z min is the voxel origin coordinate, and r is the voxel side length.
[0240] Geometric features (point density, normal vector mean) and semantic features (color mean) are calculated for each voxel, and a hash table is used to store the voxel index, geometric features, and semantic features of non-empty voxels. The voxel hash table key is (i, j, k), and the value is the voxel feature vector.
[0241] The original point cloud is reduced to a small voxel grid and the point cloud is converted into a multi-scale voxel grid, ensuring real-time and fast scene understanding.
[0242] On this basis, within the multi-scale voxel grid obtained above, a three-dimensional convolution operation is first performed to achieve multi-scale voxel internal feature fusion, and the fused voxel grid V is obtained. fusion .
[0243] At the same time, multi-scale cross-modal fusion is performed, and the multi-scale image features obtained by image extraction in step S21 are fused into the multi-scale voxel grid to achieve multi-scale image and point cloud feature fusion from a global perspective. The interactive fusion is performed with the feature grid from different scales in the form of iterative updates, and the voxel three-dimensional feature representation of the entire scene is iteratively updated to complete the rapid three-dimensional construction of the three-dimensional scene.
[0244] The specific steps are:
[0245] First, the multi-scale image features are back-projected into the voxel space to obtain the projected image features I proj ;
[0246] Then perform bidirectional cross-modality and calculate the fused voxel grid V fusion To the projected image feature I proj The guidance weight of the new image feature I updated , calculation formula:
[0247] α=Softmax(W q V fusion +W k I proj ) Formula (9)
[0248] I updated =α·I proj +(1-α)·I raw Formula (10)
[0249] Among them, W q 、W k is the weight matrix. And the global semantics of the voxel is enhanced by cross attention V global , calculation formula:
[0250] V global =CrossAttn(V fusion ,I updated ) Formula (11)
[0251] Finally, an iterative optimization cycle is performed to enhance the voxel features V after global semantics global As the initial features, the historical features are updated by the reconstruction loss of the multi-scale image features corresponding to the current voxel grid and the RGB-D image acquired in real time;
[0252] When the feature difference between two consecutive iterations is less than a predetermined threshold or reaches the maximum number of iterations, the process stops (up to 5 iterations), and the voxel 3D feature representation of the entire scene is iteratively updated to complete the rapid 3D construction of the 3D scene. The feature difference between two iterations is calculated as follows:
[0253]
[0254] Among them, V t 、V t+1 are the iterative features at time t and t+1 respectively, is a predetermined threshold; illustratively, is 10 -4 .
[0255] The function of step S3 is to fuse multi-scale image features with multi-scale scene point cloud features at multiple levels, and generate voxelized three-dimensional scenes through voxel representation, providing fused features and three-dimensional scenes for subsequent semantic recognition of complex scenes.
[0256] Step S4, specifically.
[0257] Reasoning is performed on voxelized 3D scenes to obtain semantic recognition results corresponding to complex scenes.
[0258] The PV-RCNN network is used to reason about the voxelized 3D scene to obtain semantic recognition results of the complex scene, including:
[0259] Querying non-empty voxels in the voxelized three-dimensional scene, and convolving each non-empty voxel with its neighboring voxels to obtain a global voxel feature;
[0260] Performing three-dimensional sparse convolution on the global voxel features to generate K candidate regions; performing a pooling operation on each candidate region to extract features of a fixed size to obtain all candidate regions and corresponding candidate region features;
[0261] A fixed number of key points are sampled in each candidate area to generate a local point cloud, and the local point cloud is passed through the first, second, third, and fourth cycles to obtain the corresponding local point cloud feature S4; 1x1 convolution is used to adjust the number of channels of the candidate area to align it with the local point cloud feature; the candidate area feature is spliced with the local point cloud feature to form a multi-scale joint feature F fuse ;
[0262] F fuse After three-layer full connection processing, a three-layer multi-layer perceptron dimensionality reduction is performed to obtain the corresponding classification probability and the semantic recognition results corresponding to complex scenes.
[0263] Based on the constructed voxelized scene representation, the PV-RCNN network is used to infer the scene to obtain the classification category of each local point cloud area, realizing semantic recognition of the terrain scene.
[0264] The PV-RCNN network structure implements semantic recognition of terrain scenes based on the voxelized 3D scene constructed in step 4. The main operations are divided into four steps as follows:
[0265] The first step is voxel feature extraction to obtain global voxel features.
[0266] Cascaded three-dimensional convolution is introduced to perform local feature fusion on the voxel area, and taking into full consideration that the voxels come from uneven point clouds, point feature fusion is also performed on the point information inside each voxel.
[0267] First, a neighborhood query operation is performed. For each non-empty voxel, all non-empty voxels within the defined receptive field (such as 3×3×3) around it are queried. Then, local feature aggregation is performed. Each non-empty voxel is convolved with its neighboring voxels (3×3×3 window) to fuse local context information and obtain the global voxel feature.
[0268] The second step is to obtain candidate regions and candidate region features based on global voxel features.
[0269] A 3D sparse convolution operation is performed on the global voxel features to generate candidate regions (Rols, number K). Each region contains the coordinates of the center point, size (length, width, height), and orientation. A pooling operation is then performed on each candidate region (Rols, Region of Interest Feature) to extract fixed-size features (e.g., 7×7×C, where C is the number of channels), resulting in Rol features (K×7×7×C). Finally, the candidate regions (Rols) and Rol features (K×7×7×C) are output.
[0270] Step 3: Get the multi-scale joint feature F fuse .
[0271] The different point feature information within each voxel area is used to extract the point cloud level features within the voxel. First, a fixed number of key points (such as N points) are sampled in each RoI to generate a local point cloud (K×N×3, where N is the number of key points), and the features of the local point cloud (K×N×C′, where C′ is the number of channels) are extracted through PointNet++ (see the point cloud extraction operation in step 2 for this step); then, a 1x1 convolution is used to adjust the number of channels of the RoI feature to align it with the local point cloud feature, and the adjusted RoI feature is concatenated with the local point cloud feature to form a multi-scale joint feature F. fuse ;
[0272] Step 4: Combine the multi-scale joint features F fuse Gradually fuse the features F through three-layer full connection processing final , enhance the nonlinear expression ability, and integrate the fusion feature F final After classification by a multi-layer perceptron, the semantic recognition results corresponding to complex scenes are obtained.
[0273] Output fusion feature F final (K×N×C″, C″ is the number of channels after fusion).
[0274] The multi-scale joint feature F fuse Gradually fuse the features F through three-layer full connection processing final ,as follows:
[0275] F final =FC3(ReLU(FC2(ReLU(FC1(F fuse ))))) Formula (13)
[0276] Among them, FC1, FC2, and FC3 are the first, second, and third fully connected layers in the three-layer fully connected processing;
[0277] The fusion feature F final Classification is performed through three layers of perceptrons; the first layer is used for feature dimensionality reduction, and the dimensionality reduction feature F1 is obtained as follows:
[0278] F1=ReLU(FC1(F final )) Formula (14)
[0279] The second layer is used to further fuse the F1 features as follows:
[0280] F2=ReLU(FC2(F1)) Formula (15)
[0281] The third layer is used to output semantic recognition probability, as follows:
[0282] p=Softmax(FC3(F2)) Formula (16)
[0283] Where p is the probability distribution of each category.
[0284] The third layer is mainly used to output classification probability, where p is the probability distribution of each category (such as city, hill, mountain, etc.), FC is the fully connected layer, ReLU activation function is used for the hidden layer, and Softmax activation function is used for the final classification.
[0285] Finally, the category with the highest probability is selected as the prediction result. The combination of voxel feature extraction and point cloud feature fusion modules effectively compensates for the problem of ignoring features between adjacent local areas in the point cloud processing model. The PV-RCNN network generates refined fusion features and local features through the above steps, and then fuses these two parts of features through a multi-layer perceptron to give point-by-point classification results. This can effectively solve the problem of not being able to capture local structural information between points and feature information between neighborhoods, and improve the accuracy and robustness of semantic recognition classification. The output result is a perceptual understanding of complex scenes, which is assigned specific semantic categories, completes the rapid semantic recognition of scenes such as cities, hills, and mountains, and realizes the rapid acquisition of rapid three-dimensional construction and semantic recognition of complex scenes such as cities, hills, and mountains.
[0286] The function of step S4 is to use the PV-RCNN network to infer the voxelized 3D scene and output the semantic recognition results of the complex scene.
[0287] In summary, the method for rapid 3D construction and recognition of complex scenes according to the embodiment of the present invention has the following beneficial effects:
[0288] 1. This invention integrates LiDAR (to obtain high-precision point clouds), millimeter-wave radar (to obtain velocity information), and an RGB-D camera (to obtain depth and texture). It combines the point cloud registration map generated by LiDAR and millimeter-wave radar with the RGB-D image acquired by a binocular visible light camera to address the problem that a single sensor cannot provide sufficiently comprehensive and accurate environmental information. It also employs multi-level feature fusion to significantly reduce computational complexity and meet real-time requirements in complex scenarios.
[0289] 2. This invention performs multi-scale feature extraction on RGB-D images and point cloud registration maps, respectively, to capture detailed information at different levels of the scene. For images, features ranging from fine local to global structure can be obtained; for point clouds, features ranging from local geometry to global shape can be extracted. Compared to existing technologies that can only extract single-modal features, this technical solution provides richer and more complete feature expression.
[0290] 3. This invention uses multi-level feature fusion to enhance feature expression, innovatively fusing multi-scale image features with multi-scale scene point cloud features. This overcomes the limitations of existing technologies in fully utilizing the complementarity of multi-sensor data. Through operations such as pixel and point cloud coordinate mapping and weighted fusion of features within a neighborhood, a fused feature vector is constructed, fully exploiting and integrating the advantages of the two modal features, significantly improving feature expression capabilities.
[0291] 4. This invention uses voxelized representation to process the fused feature vectors, segmenting them into multi-scale voxel grids with varying voxel edge lengths. It then enhances the voxel global semantics through operations such as feature backprojection and cross-attention weight calculation. This is then combined with an iterative optimization loop to dynamically update the global semantic features based on feature differences until convergence conditions are met, generating a high-precision voxelized 3D scene. This addresses the shortcomings of existing 3D construction methods, which often generate sparse point clouds and lack detail.
[0292] 5. This invention leverages the PV-RCNN network to reason about voxelized 3D complex scenes, implementing a complete process from non-empty voxel queries, global feature generation through convolution, candidate region generation, keypoint sampling, and multi-scale joint feature construction. Finally, through fully connected multi-layer perceptron processing, accurate classification probabilities are output. This process overcomes the shortcomings of existing semantic recognition methods, which can only simply classify target objects and have difficulty understanding complex scenes. It provides reliable semantic understanding support for intelligent mobile robots or unmanned vehicles to navigate high-speed autonomously in complex scenes.
[0293] Those skilled in the art will appreciate that all or part of the process steps of the above-described embodiments can be implemented by instructing related hardware through a computer program, and the program can be stored in a computer-readable storage medium, such as a magnetic disk, an optical disk, a read-only memory, or a random access memory.
[0294] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by any technician familiar with this technical field within the technical scope disclosed by the present invention should be covered by the scope of protection of the present invention.
Claims
1. A method for rapid 3D construction and recognition of complex scenes, characterized by: include: Obtain point cloud registration maps and RGB-D images of complex scenes; Performing multi-scale feature extraction based on the RGB-D image to obtain multi-scale image features; Obtaining multi-scale scene point cloud features of the point cloud registration image; Performing multi-level feature fusion on the multi-scale image features and the multi-scale scene point cloud features to obtain a multi-scale fused feature vector, and performing voxel representation based on the multi-scale fused feature vector to obtain a voxelized three-dimensional scene; Reasoning is performed on the voxelized three-dimensional scene to obtain a semantic recognition result corresponding to the complex scene.
2. The method according to claim 1, characterized in that The RGB-D image is subjected to multi-scale feature extraction based on the ResNet network to obtain multi-scale image features at the P2, P3, P4 and P5 levels of the feature pyramid; The point cloud registration map is subjected to feature extraction through an improved PoinNet++ model to obtain scene point cloud features S1, S2, S3 and S4 multi-scale scene point cloud features; The multi-scale image features at the P2, P3, P4 and P5 levels of the feature pyramid are respectively fused with the corresponding S1, S2, S3 and S4 multi-scale scene point cloud features.
3. The method according to claim 2, characterized in that The preprocessed RGB-D image is subjected to multi-scale feature extraction based on the ResNet network to obtain the corresponding multi-scale image features, including: The preprocessed RGB-D image is subjected to feature extraction at different scales by first, second, and third image feature extraction units; wherein the first, second, and third image feature extraction units each include first, second, and third convolutional layers with a convolution kernel of 3×3, a 2×2 maximum pooling layer, and a residual connection from the output of the first convolutional layer to the maximum pooling layer; The convolution is used to increase the number of image channels and extract local features of RGB-D images; The maximum pooling layer is used to downsample the image and reduce the spatial resolution of the image; The residual connection adds the image features input to the first convolutional layer directly to the output of the maximum pooling layer to alleviate the gradient disappearance; The preprocessed RGB image features and the output feature maps of the first, second and third image feature extraction units are respectively the P2, P3, P4 and P5 level image features of the feature pyramid.
4. The method according to claim 2, characterized in that The point cloud registration map is subjected to feature extraction based on the improved PoinNet++ model to obtain scene point cloud features, including: The point cloud registration graph is sequentially subjected to the first, second, third and fourth cycles to obtain S1, S2, S3 and S4 scene point cloud features respectively; Each cycle includes downsampling, point-wise MLP, local feature extraction, and point feature extraction; The first and second loops use FPS downsampling and voxel grid downsampling respectively; the third and fourth loops use adaptive sampling based on point cloud density to obtain the downsampled point cloud; The point-by-point MLP extracts the normal and curvature of the point cloud after the sampling point to obtain the basic features of the point cloud; A point is randomly selected from the point cloud as the first centroid point. With the centroid point as the center, a radius neighborhood is dynamically divided until the number of points in the neighborhood reaches a predetermined threshold. A neighborhood is obtained, and shared MLP + maximum pooling is applied to the points in the neighborhood to generate local features of the point cloud. The point farthest from the selected centroid point is selected as the new centroid point. The distance calculation and selection of the farthest point are repeated until a neighborhood formed by enough centroid points covers the geometric structure of the entire point cloud, and multiple centroid points are obtained. The local features of the point cloud corresponding to each centroid and the corresponding basic features of the point cloud are spliced together, and weighted fusion is performed using the attention mechanism to obtain multi-scale scene point cloud features.
5. The method according to claim 3, characterized in that: Performing multi-level feature fusion on the multi-scale image features and the multi-scale scene point cloud features, including: The first, second, third and fourth features of P2 and S1, P3 and S2, P4 and S3, and P5 and S4 are respectively fused to obtain fused feature vectors of corresponding scales to form the multi-scale fused feature vector; wherein obtaining the fused feature vector of corresponding scale includes: Mapping image pixels in the corresponding scale image features to point cloud coordinates; In the point cloud feature space, the features of the 32 point cloud points with the closest Euclidean distance in the neighborhood centered on each centroid are weighted and fused into the cloud feature vector F; Calculate the weighted fusion of the feature vectors of the three pixels closest to the centroid point in Euclidean distance to obtain the image feature vector f; Concatenate the image feature vector f and the point cloud feature vector F along the channel dimension to form a fusion feature F′=[F,f]; Perform full connection processing on F′ to obtain the fused feature vector of the corresponding scale.
6. The method according to claim 5, characterized in that Performing voxel representation based on the multi-scale fusion feature vector to obtain a voxelized three-dimensional scene includes: The multi-scale fusion feature vectors are segmented with different voxel side lengths to obtain the corresponding multi-scale voxel grids; The internal features of the multi-scale voxel grid are fused to obtain the fused voxel grid V fusion ; The multi-scale image features are back-projected to the fused voxel grid V fusion , get the projected image feature I proj ; Calculate V fusion to I proj The guidance weight is calculated based on the guidance weight to obtain the image feature I updated ; Based on image features I updated By enhancing voxels through cross attention, the enhanced voxel features are obtained as the global semantic V global ; The global semantic V global The current voxel grid is compared with the multi-scale image features acquired in real time to obtain feature differences. When the feature difference between two consecutive iterations is less than a predetermined threshold or reaches the maximum number of iterations, the global semantic features are iteratively updated to complete the voxelized 3D scene construction of the complex scene.
7. The method according to claim 6, characterized in that The PV-RCNN network is used to reason about the voxelized 3D scene to obtain semantic recognition results for the complex scene, including: Querying non-empty voxels in the voxelized three-dimensional scene, and convolving each non-empty voxel with its neighboring voxels to obtain a global voxel feature; Performing three-dimensional sparse convolution on the global voxel features to generate K candidate regions; performing a pooling operation on each candidate region to extract features of a fixed size to obtain all candidate regions and corresponding candidate region features; A fixed number of key points are sampled in each candidate area to generate a local point cloud, and the local point cloud is passed through the first, second, third, and fourth cycles to obtain the corresponding local point cloud feature S4; 1x1 convolution is used to adjust the number of channels of the candidate area to align it with the local point cloud feature; the candidate area feature is spliced with the local point cloud feature to form a multi-scale joint feature F fuse ; F fuse After three-layer full connection processing, a three-layer multi-layer perceptron dimensionality reduction is performed to obtain the corresponding classification probability and the semantic recognition results corresponding to complex scenes.
8. The method according to claim 7, characterized in that: The multi-scale joint feature F fuse Gradually fuse the features F through three-layer full connection processing final ,as follows: F final =FC3(ReLU(FC2(ReLU(FC1(F fuse ))))) Among them, FC1, FC2, and FC3 are the first, second, and third fully connected layers in the three-layer fully connected processing; The fusion feature F final Classification is performed through three layers of perceptrons; the first layer is used for feature dimensionality reduction, and the dimensionality reduction feature F1 is obtained as follows: F1=ReLU(FC1(F final )) The second layer is used to further fuse the F1 features as follows: F2=ReLU(FC2(F1)) The third layer is used to output semantic recognition probability, as follows: p=Softmax(FC3(F2)) Where p is the probability distribution of each category.
9. The method according to any one of claims 1 to 8, characterized in that Point cloud data is obtained using a laser radar; speed information is obtained using a millimeter wave radar; and a point cloud registration map is obtained based on the point cloud data and the speed information.
10. The method according to claim 9, characterized in that: Use binocular visible light cameras to obtain depth maps and RGB images; The depth map and the RGB image are preprocessed to obtain an RGB-D image.
Citation Information
Cited By
Three-dimensional semantic scene data construction method and device and medium
CN121120997A
Power distribution network wire identification method and system based on unmanned aerial vehicle multi-mode perception
CN121582917A
A power distribution network wire identification method and system based on unmanned aerial vehicle multi-modal perception
CN121582917B