Multi-modal coupled perception method for target recognition and region segmentation in confined space
By adopting a multimodal coupled perception method in confined space, the deep learning network processes multiple sensor data, solving the problems of low illumination, dust and noise in traditional single-modal perception technology, and achieving efficient target recognition and area segmentation.
Patent Information
- Application Number
- PCT/CN2023/138695
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-07
- Filing Date
- 2023-12-14
- Publication Date
- 2025-06-12
AI Technical Summary
In confined spaces, traditional single-mode environment perception technology has problems such as difficulty in collecting image data under low illumination, degradation of lidar performance under the influence of dust, and high noise from millimeter-wave radar and low data utilization.
The multimodal coupled perception method is adopted to deeply mine the original data of multiple sensors (lidar, millimeter wave radar, camera) through deep learning networks, and a point cloud data processing network of three-dimensional lidar and two-dimensional millimeter wave radar, as well as camera image data processing networks, realize the feature fusion and multi-scale feature integration of multi-sensor data, and output the results of target recognition and region segmentation.
It effectively improves the detection accuracy of target recognition, reduces perception errors, improves sensor data utilization, and realizes efficient target recognition and area segmentation in complex confined spaces.
Smart Images

Figure CN2023138695_12062025_PF_FP_ABST
Abstract
Description
Multimodal coupled perception method for target recognition and region segmentation in confined space Technical Field
[0001] The present invention relates to the technical field of environment perception in confined spaces, and in particular to a multimodal coupled perception method for target recognition and region segmentation in confined spaces. Background Art
[0002] Environmental perception within confined spaces is a key area of research in the field of intelligent coal mining. Within the well-visible environment of a mine, sensors such as network cameras and lidar are effective means of scene monitoring. However, complex confined spaces such as tunneling faces, coal mining faces, and mining tunnels are often characterized by high dust levels, low illumination, and a narrow field of view. This results in network cameras being unable to effectively capture image data in low illumination conditions, and lidar performance being degraded by dust. Furthermore, millimeter-wave radars, commonly used for environmental perception, often suffer from high noise levels and low data utilization.
[0003] Fusion of multimodal data from various sensors and giving full play to the performance advantages of each sensor is an effective means to overcome the above problems. However, current multimodal data fusion methods are often based on post-fusion of multimodal data, that is, after analyzing the perception data of each sensor to obtain the perception results, the final perception results are obtained by fusing them based on the perception results of each sensor. This has the problems of many perception calculation results, large perception errors, and low utilization of each sensor data.
[0004] Summary of the Invention
[0005] In order to overcome the shortcomings of the above-mentioned prior art, the present invention discloses a multimodal coupled perception method for target recognition and region segmentation in confined space. It can utilize multimodal raw data perceived by multiple sensors, deeply mine the multimodal data through a deep learning network, and realize target recognition and region segmentation in confined space.
[0006] The multimodal coupled perception method for target recognition and region segmentation in a confined space proposed according to the purpose of the present invention includes the following steps:
[0007] Step 1: Build a 3D laser radar 3D point cloud data processing network to deeply mine the features of laser point cloud data and output the bird's-eye view feature data of the laser point cloud;
[0008] Step 2: Build a millimeter-wave radar 2D point cloud data processing network to deeply mine the features of millimeter-wave point cloud data and output the bird's-eye view feature data of the millimeter-wave point cloud;
[0009] Step 3: Build a camera image data processing network to deeply mine image data features and output bird's-eye view feature data of camera data;
[0010] Step 4: Construct a feature fusion engine to output a bird's-eye view coupled with multi-sensor information;
[0011] Step 5: Construct a multi-scale feature fusion extraction network based on the feature pyramid network and output multi-scale integrated feature data;
[0012] Step 6: Build a network output head to output target recognition results and region segmentation results. At this point, the multimodal coupled perception network is built.
[0013] Step 7: Design a loss function based on the network output head to train the network weights;
[0014] Step 8: Input the lidar 3D point cloud data, millimeter wave radar 2D point cloud data, and camera image data into the multimodal coupled perception network, infer the target recognition and region segmentation prediction results, and visualize the prediction results.
[0015] Preferably, in step one and step two, the 3D point cloud data processing network and the 2D point cloud data processing network are mainly composed of a voxelization layer, a voxelization coding layer, and a residual coding layer; wherein the voxelization layer performs a voxelization operation in a loop for each point cloud sample, and outputs voxelized data, voxel coordinates, and the number of points of each voxel; the voxelization coding layer performs in-depth processing on the original voxel features, mines the offset features of the point cloud in the voxelized block and the voxel block itself, and outputs feature data based on a bird's-eye view; the residual coding layer further mines features after feature compression of the voxelized encoded data, and then scales the compressed learned features back to feature data based on a bird's-eye view.
[0016] Preferably, in step three, the image data processing network is mainly composed of a basic feature extraction layer, a depth feature extraction layer, a geometric transformation layer, and a voxel pooling layer; wherein the basic feature extraction layer is constructed based on the Resnet101 framework, extracts and outputs high-dimensional basic feature data of the image; the depth feature extraction layer performs feature mapping on the basic feature data, extracts and outputs data containing depth and other features; the geometric transformation layer converts the cone features in the image coordinate system to the lidar coordinate system through affine transformation and coordinate system conversion methods; the voxel pooling layer processes the feature data into feature data based on a bird's-eye view perspective.
[0017] Preferably, in step four, the feature fusion device is mainly composed of a convolution layer and a batch normalization layer; the 3D point cloud bird's-eye view feature data output by the 3D point cloud data processing network, the 2D point cloud bird's-eye view feature data output by the 2D point cloud data processing network, and the image bird's-eye view feature data output by the image data processing network are input into the feature fusion device, the data of multiple sensors are fused, and a BEV feature map coupled with multi-sensor information is output.
[0018] Preferably, in step five, the multi-scale feature fusion extraction network is mainly composed of a convolutional layer sequence of two downsampling layers and two upsampling layers; wherein the downsampling layer is used to reduce the spatial size of the feature map and increase the number of channels of the feature map; the upsampling layer is used to increase the spatial size of the feature map and integrate features at different levels; and finally the network outputs multi-scale integrated feature data.
[0019] Preferably, in step six, in terms of target recognition, the output target recognition result mainly includes target category information and target 3D frame information; the target 3D frame information includes target 3D frame center point information, target 3D frame center point deviation information, target 3D frame length and width size information, target 3D frame height size information and target 3D frame yaw angle information; wherein, the target category information and target 3D frame center point information are designed to be obtained through CenterHead; the target 3D frame center point deviation information is designed to be obtained through DeltaHead; the length and width size information of the target 3D frame is obtained through DimsHead; the height size information of the target 3D frame is obtained through HeightZHead; the yaw angle information of the target 3D frame is obtained through RotationHead; in terms of region segmentation, the output region segmentation result is obtained through SegmentationHead, and the coordinate points of the drivable area from the BEV perspective are output, and the region segmentation map can be drawn according to the obtained coordinate points.
[0020] Preferably, in step 7, CenterHead is used to predict the target category and target center point. Since the area with targets in the driving scene is far less than the area without targets, based on the imbalanced characteristics of sample categories in this scene, Focal loss is selected as the loss function for this output;
[0021] DeltaHead, DimsHead, HeightZHead, and RotationHead are used to predict continuous values describing the 3D target box, and the MSE function is used as the loss function of these output heads;
[0022] SegmentationHead is used to predict the coordinate point values describing the region segmentation, and Dice loss is selected as the loss function for image region segmentation.
[0023] Compared with the existing technology, the advantages of the multimodal coupled perception method for target recognition and region segmentation in confined space disclosed by the present invention are:
[0024] (1) In view of the characteristics of millimeter-wave radar itself, such as cluttered point cloud data, many detection noise points, and difficult data filtering, the present invention constructs a multimodal coupled perception network, which can deeply mine the effective information in the massive point cloud data of millimeter-wave radar and effectively improve the detection accuracy of target recognition.
[0025] (2) The present invention addresses the performance disadvantages of various sensors under complex conditions, such as the limited detection distance of laser radar under dusty conditions, the large noise data of millimeter-wave radar under complex environments, and the unclear field of view of cameras under low-light conditions. By deeply mining the multimodal raw data of each sensor through a deep learning network and fusing the available information of different modal data, the present invention overcomes the above problems, realizes target recognition and area segmentation in confined spaces, reduces perception errors, and improves sensor data utilization. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the embodiments of the present invention or the technical solutions of the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0027] FIG1 is a schematic diagram of a network model of the present invention.
[0028] FIG2 is a diagram showing the prediction effect from the main perspective in the embodiment.
[0029] FIG3 is a BEV perspective diagram of the prediction effect in the embodiment.
[0030] FIG4 is a pixel coordinate diagram of the prediction effect in the embodiment. DETAILED DESCRIPTION
[0031] The following is a brief description of the specific embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts are also within the scope of protection of the present invention.
[0032] 1 to 4 show preferred embodiments of the present invention and provide a detailed analysis thereof.
[0033] This embodiment uses a laser radar that samples 24,000 point clouds per frame as the laser point cloud data source.
[0034] The multimodal coupled perception method for target recognition and region segmentation in a confined space disclosed in the present invention includes the following steps:
[0035] Step 1: Construct a 3D laser radar 3D point cloud data processing network LidarNet, which is used to deeply mine the laser point cloud data features Lidar Backbone and output the bird's-eye view BEV feature data of the laser point cloud. The 3D point cloud data processing network LidarNet is mainly composed of the voxelization layer PillarLayer, the voxelization encoding layer PillarEncoderLayer, and the residual encoding layer BottleneckLayer. The voxelization layer accepts point cloud samples of dimension N×24000×3, where N is the number of input point cloud frames. After the voxelization operation, the voxelized data is output with a shape of N p ×N pmc ×3, where N p is the number of voxel blocks, N pmc The maximum number of point clouds in a voxel block; the coordinates of the output voxels, the shape is N p ×3, the number of point clouds in each voxel N p ×1. The voxel encoding layer receives the three types of data output by the voxelization layer mentioned above, mines the offset features between the point cloud and the voxel block itself, and outputs BEV-based feature data. The feature data has a shape of N × 256 × 64 × 256, representing the number of inputs, the number of features, the height of the BEV map, and the width of the BEV map, respectively. The residual coding layer receives the output of the voxel encoding as its input. After compression learning, the output data has a shape of N × 256 × 64 × 256, which is the BEV feature map output of the 3D lidar 3D point cloud data processing network.
[0036] Step 2. Construct the millimeter-wave radar 2D point cloud data processing network RadarNet, which is used to deeply mine the millimeter-wave point cloud data feature Radar Backbone and output the BEV feature data of the millimeter-wave point cloud; the 2D point cloud data processing network RadarNet is mainly composed of the voxelization layer PillarLayer, the voxelization encoding layer PillarEncoderLayer, and the residual encoding layer BottleneckLayer. Before the data is input into this network, the 2D data containing (x, y) coordinate information needs to be upgraded to 3D data. The third dimension z is determined according to the installation height of the millimeter-wave radar. The data dimensionality can be upgraded by adding the z value to each sample point. The voxelization layer accepts millimeter-wave point cloud samples with a dimension of N×3 after dimensionality upgrade, where N is the number of input millimeter-wave point clouds. After the voxelization operation, the voxelized data is output with a shape of N p ×N pmc ×3, where N p is the number of voxel blocks, N pmc The maximum number of millimeter wave point clouds in a voxel block; the coordinates of the output voxels are in the shape of N p ×3, the number of millimeter wave point clouds in each voxel Np ×1. The voxel encoding layer receives the three types of data output by the voxelization layer mentioned above, mines the offset features between the millimeter-wave point cloud and the voxel block itself within the voxelized block, and outputs BEV-based feature data. The feature data has a shape of N × 256 × 64 × 256, representing the number of inputs, the number of features, the height of the BEV map, and the width of the BEV map, respectively. The residual coding layer receives the output of the voxel encoding as its input. After compression learning, the output data has a shape of N × 256 × 64 × 256, which is the BEV feature map output of the millimeter-wave 2D point cloud data processing network.
[0037] Step 3: Build the camera image data processing network, CameraNet, to deeply mine the image data features (Camera Backbone) and output the camera data's BEV feature data. The CameraNet image data processing network primarily consists of a BaseNet layer, a DepthNet layer, a Geometry Transformation layer, and a Voxel Pooling layer. The BaseNet layer accepts image data of dimensions N × C × 368 × 640, where N is the number of images, C is the number of image channels, and 368 × 640 represents the height and width of the image input to the model. After BaseNet processing, the output is high-dimensional basic feature data of the image, with a shape of N × 1024 × 23 × 40, where 1024 represents the number of high-dimensional basic features and 23 × 40 represents the image size after BaseNet processing. The depth feature extraction layer accepts the output data of BaseNet. This layer performs feature mapping on the basic feature data, extracts and outputs data containing depth and other features. The data shape is N×376×23×40, where 376 is the sum of the number of depth features and the number of other features. The geometric transformation layer converts the cone features in the image coordinate system to the lidar coordinate system through affine transformation and coordinate system conversion methods. The input of this layer is the coordinate transformation matrix required to transform the image coordinate system to the lidar coordinate system. The output cone features in the lidar coordinate system have a shape of N×120×23×40×3, where 120 is the number of cone depth features and 3 is the number of image channels. The transformation formula is shown in Equation 1: L =P L ×P IN ×X C (Formula 1)
[0038] Among them, P IN is the image internal parameter matrix, P L is the external parameter matrix from camera to radar coordinate system, X C is the feature point of the viewing cone in the camera pixel coordinate system, XL is the feature point of the viewing cone in the radar coordinate system.
[0039] The input of the voxel pooling layer is the cone features obtained by the geometric transformation layer and the image depth features obtained by the depth feature extraction layer. The output data shape is N×256×64×256. At this time, the BEV feature map output of the camera image data processing network is obtained.
[0040] Step 4: Construct a feature fuser Fuser to output a BEV feature map that couples multi-sensor information. The feature fuser is mainly composed of a convolutional layer Conv Layer and a batch normalization layer BatchNormalLayer. The 3D point cloud BEV feature data output by the 3D point cloud data processing network, the 2D point cloud BEV feature data output by the 2D point cloud data processing network, and the image BEV feature data output by the image data processing network are input into the feature fuser. Specifically, the three BEV feature maps of shape N×256×64×256 provided in steps 1, 2, and 3 are fused into a shape of N×(256+256+256)×64×256 as input, where 256+256+256 represents the feature coupling of the three types of data into a type of high-dimensional feature. The output is a BEV feature map of shape N×256×64×256 after the convolutional neural network coupling features.
[0041] Step 5. Based on the feature pyramid network (FPN), a multi-scale feature fusion extraction network (Multi ScaleNeck) is constructed to output multi-scale integrated feature data. The multi-scale feature fusion extraction network mainly consists of a sequence of convolutional layers, including two downsampling layers and two upsampling layers. The downsampling layer uses the output of step 4 as its input to reduce the spatial size of the feature map and increase the number of channels. The output is N×512×64×256, where 512 is the number of channels in the increased feature map. The upsampling layer uses the output of the downsampling layer as its input to increase the spatial size of the feature map and integrate features at different levels. The network finally outputs multi-scale integrated feature data with a shape of N×256×64×256.
[0042] Step 6: Build a network output head to output target recognition and segmentation results. For target recognition, the output target recognition results primarily include target category information and target 3D bounding box information. The target 3D bounding box information includes the target 3D bounding box center point information, target 3D bounding box center point deviation information, target 3D bounding box length, width, height, and yaw angle information. The target category information and target 3D bounding box center point information are obtained through CenterHead; the target 3D bounding box center point deviation information is obtained through DeltaHead; the target 3D bounding box length and width information is obtained through DimsHead; the target 3D bounding box height information is obtained through HeightZHead; and the target 3D bounding box yaw angle information is obtained through RotationHead. For segmentation, the output segmentation results are obtained through SegmentationHead, which outputs the coordinates of the drivable area from the BEV's perspective. The obtained coordinates can then be used to draw a segmentation map. The model is now complete, and the model network structure is shown in Figure 1.
[0043] Step 7: Design the loss function based on the network output head to train the network weights; CenterHead is used to predict the target category and target center point. Since the area with targets in the driving scene is far less than the area without targets, based on the imbalanced characteristics of the sample categories in this scene, Focal loss is selected as the loss function of this output, as shown in Equations 2, 3, and 4. The corresponding loss L is composed of the positive sample loss L pos and negative sample loss L neg Composition, calculated as follows:
[0044] Among them, for each position i, if gt i =1 indicates a positive sample, otherwise it is a negative sample; pred i is the predicted value of the model at position i; Npos is the number of positive samples; Positives is the set of all positive samples, and Negatives is the set of all negative samples.
[0045] DeltaHead, DimsHead, HeightZHead, and RotationHead are used to predict continuous values describing the 3D target box, so the MSE function is used as the loss function of these output heads; as shown in Equation 5:
[0046] Among them, i is the sample number, N is the number of samples, predict is the predicted value, and target is the true value.
[0047] SegmentationHead is used to predict the coordinate point values describing the region segmentation. For this type of segmentation problem, Dice loss is selected as the loss function for image region segmentation, as shown in Equation 6:
[0048] Among them, I flat is the model prediction value vector; T flat is a real value vector.
[0049] Step 8: Input the lidar 3D point cloud data, millimeter wave radar 2D point cloud data, and camera image data into the multimodal coupled perception network, infer the target recognition and region segmentation prediction results, and visualize the prediction results. The prediction results are shown in Figures 2, 3, and 4.
[0050] The above description of the disclosed embodiments will enable one skilled in the art to implement and use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit and scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.
Claims
1. A multi-modal coupled perception method for target recognition and area segmentation in a confined space, characterized in that, it includes the following steps: Step 1: Construct a 3D lidar point cloud data processing network to deeply mine the characteristics of lidar point cloud data and output the bird's-eye view feature data of the lidar point cloud; Step 2: Construct a millimeter-wave radar 2D point cloud data processing network to deeply mine the characteristics of millimeter-wave point cloud data and output the bird's-eye view feature data of the millimeter-wave point cloud; Step 3: Construct a camera image data processing network to deeply mine the characteristics of image data and output the bird's-eye view feature data of the camera data; Step 4: Construct a feature fuser to output a bird's-eye view that couples multi-sensor information; Step 5: Construct a multi-scale feature fusion extraction network based on a feature pyramid network to output multi-scale integrated feature data; Step 6: Construct a network output head to output the target recognition result and the area segmentation result; thus, the multi-modal coupled perception network is constructed; Step 7: Design a loss function according to the network output head to train the network weights; Step 8: Input the lidar 3D point cloud data, millimeter-wave radar 2D point cloud data, and camera image data into the multi-modal coupled perception network constructed in Steps 1 to 7, infer the target recognition and area segmentation prediction results, and perform visual rendering on the prediction results.
2. The multi-modal coupled perception method for target recognition and area segmentation in a confined space according to claim 1, characterized in that, in Steps 1 and 2, the 3D point cloud data processing network and the 2D point cloud data processing network are mainly composed of a voxelization layer, a voxelization encoding layer, and a residual encoding layer; among them, the voxelization layer performs voxelization operations on each point cloud sample in a loop, outputs voxelized data, the coordinates of the voxels, and the number of points in each voxel; the voxelization encoding layer deeply processes the original voxel features, mines the offset features between the point cloud in the voxel block and the voxel block itself, and outputs feature data based on the bird's-eye view; the residual encoding layer further mines the features after compressing the features of the voxelized encoded data, and then scales the features learned by compression back to the feature data based on the bird's-eye view.
3. The multi-modal coupled perception method for target recognition and area segmentation in a confined space according to claim 1, characterized in that, in Step 3, the image data processing network is mainly composed of a basic feature extraction layer, a depth feature extraction layer, a geometric transformation layer, and a voxel pooling layer; among them, the basic feature extraction layer is constructed based on the Resnet101 framework, extracts and outputs high-dimensional basic feature data of the image; the depth feature extraction layer performs feature mapping on the basic feature data, extracts and outputs data containing depth and other features; the geometric transformation layer transforms the frustum features in the image coordinate system to the lidar coordinate system through affine transformation and coordinate system conversion methods; the voxel pooling layer processes the feature data into feature data based on the bird's-eye view perspective.
4. The multi-modal coupled perception method for target recognition and area segmentation in a confined space according to claim 1, characterized in that, In Step 4, the feature fusion unit mainly consists of a convolutional layer and a batch normalization layer. The bird's-eye view feature data of the 3D point cloud output by the 3D point cloud data processing network, the bird's-eye view feature data of the 2D point cloud output by the 2D point cloud data processing network, and the bird's-eye view feature data of the image output by the image data processing network are input into the feature fusion unit to fuse the data of multiple sensors and output a BEV feature map that couples multi-sensor information.
5. The multi-modal coupled perception method for target recognition and area segmentation in a restricted space according to claim 1, wherein, In Step 5, the multi-scale feature fusion extraction network mainly consists of a convolutional layer sequence of two downsampling layers and two upsampling layers. Among them, the downsampling layer is used to reduce the spatial size of the feature map and increase the number of channels of the feature map; the upsampling layer is used to increase the spatial size of the feature map and integrate features at different levels. Finally, the network outputs multi-scale integrated feature data.
6. The multi-modal coupled perception method for target recognition and area segmentation in a restricted space according to claim 1, wherein, In Step 6, in terms of target recognition, the output target recognition results mainly include target category information and target 3D box information. The target 3D box information includes the center point information of the target 3D box, the deviation information of the center point of the target 3D box, the length and width dimension information of the target 3D box, the height dimension information of the target 3D box, and the yaw angle information of the target 3D box. Among them, the target category information and the center point information of the target 3D box are designed to be obtained through CenterHead; the deviation information of the center point of the target 3D box is designed to be obtained through DeltaHead; the length and width dimension information of the target 3D box is obtained through DimsHead; the height dimension information of the target 3D box is obtained through HeightZHead; the yaw angle information of the target 3D box is obtained through RotationHead. In terms of area segmentation, the output area segmentation result is obtained through SegmentationHead, and the coordinate points of the drivable area in the BEV view are output. The area segmentation map can be drawn according to the obtained coordinate points.
7. The multi-modal coupled perception method for target recognition and area segmentation in a restricted space according to claim 6, wherein, CenterHead is used to predict the target category and the target center point. Since the area with targets in the driving scenario is much less than the area without targets, based on the characteristic of unbalanced sample categories in this scenario, the Focal loss is selected as the loss function for this output. DeltaHead, DimsHead, HeightZHead, and RotationHead are used to predict the continuous values describing the 3D target box, and the MSE function is selected as the loss function for these output heads. SegmentationHead is used to predict the coordinate point values describing the area segmentation, and the Dice loss is selected as the loss function for image area segmentation.
Citation Information
Patent Citations
Target detection network system and method applied to multi-sensor data fusion in rainy and snowy weather scene
CN114140672A
Multi-source heterogeneous perception information multi-level fusion representation and target identification method
CN115100618A
Three-dimensional target detection method and system based on multi-modal feature fusion under cross view angle
CN115965847A
Multi-task environment sensing method based on point cloud fusion of camera and laser radar
CN117036895A
Learning method and learning device for integrating image acquired by camera and point-cloud map acquired by radar or LiDAR corresponding to image at each of convolution stages in neural network and testing method and testing device using the same
US10408939B1
Cited By
3D target detection system and method based on neural architecture search
CN120808003A
A 3D object detection system and method based on neural architecture search
CN120808003B
Semantic environment perception method and system based on multi-modal dynamic fusion and storage medium
CN121121746A
Robust adaptive aerial view fusion algorithm and application thereof in three-dimensional target detection
CN121747077A
Robust adaptive bird's eye fusion algorithm and its application in 3d object detection
CN121747077B