A Multi-Objective Visual Detection Method for Autonomous Driving Scenarios Based on Improved CenterNet

The modified CenterNet method addresses precision issues in complex driving scenarios by employing a feature pyramid network and average boundary extraction, improving detection accuracy and robustness for autonomous vehicles.

CN114581866BActive Publication Date: 2025-07-15JIANGSU UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210077170.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-07-15
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

The existing CenterNet algorithms do not have enough multi-object detection accuracy in complex driving scenarios, making it difficult to meet the real-time and robustness requirements of autonomous driving.

Method used

The improved CenterNet method is adopted to generate a feature pyramid structure through the feature extraction module, and combined with the average boundary extraction module, the detection accuracy and robustness are improved, and the detection speed is optimized by the introduction of RepVGG feature extractor and structural reparameterization technology.

Benefits of technology

It improves the detection accuracy of small targets and detection accuracy in occlusion scenarios, enhances the detection robustness of autonomous vehicles, and meets the requirements of real-time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114581866B_ABST
    Figure CN114581866B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligent vehicle driving technology, and in particular to a multi-object visual detection method for an autonomous driving scenario based on an improved CenterNet, including: extracting features from pictures around an autonomous vehicle captured by an on-vehicle camera to obtain feature maps of different scales, sampling the generated feature maps to generate a feature pyramid composed of feature maps of different scales, using different feature maps as inputs to a detection head module, performing convolutional operations on different feature maps to generate a final prediction result. The multi-object visual detection method for an autonomous driving scenario based on an improved CenterNet of the present invention, by using a feature pyramid structure composed of feature maps of different scales generated by a feature extraction module, improves the detection accuracy of small targets by an autonomous vehicle in a driving environment; improves the robustness of the detection of an autonomous vehicle; and meets the real-time requirements of autonomous driving detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent vehicle driving, and in particular to a multi-object visual detection method for an autonomous driving scenario based on an improved CenterNet. Background Art

[0002] In recent years, with the rapid development of deep learning, the continuous improvement of the computing power of computing platform hardware, and the continuous reduction of the costs of in-vehicle sensors such as cameras, radars, and lidar, the progress of autonomous driving perception technology has been promoted. A reliable perception system is a prerequisite for the normal operation of autonomous vehicles under complex traffic conditions. Vision-based object detection algorithms have advantages such as high perception accuracy and low cost, and are widely used in academia and industry.

[0003] Computer vision algorithms based on deep learning far exceed traditional detection algorithms based on handcrafted features in terms of both speed and accuracy. The current mainstream object detection algorithms mainly include: anchor box-based detection algorithms and anchor-free detection algorithms. Anchor-free detection algorithms have advantages such as a simple network structure and fast detection speed compared with anchor box-based detection algorithms. CenterNet is a classic anchor-free general detection algorithm, but its detection accuracy is insufficient in complex driving scenarios. Summary of the Invention

[0004] The technical problem to be solved by the present invention is: to provide an improved multi-object visual detection method for an autonomous driving scenario based on an improved CenterNet to solve the problems existing in the above background art.

[0005] The present invention proposes a multi-object visual detection method for an autonomous driving scenario based on an improved CenterNet, including the following steps:

[0006] S1, extracting features from the pictures around the autonomous vehicle captured by the in-vehicle camera to obtain feature maps of different scales;

[0007] S2, performing sampling processing on the feature maps generated in step S1 to generate a feature pyramid composed of feature maps of different scales;

[0008] S3, extracting different feature maps in step S2 as the input of the detection head module, and performing convolution operations on the different feature maps to generate the final prediction results.

[0009] Preferably, S1 specifically includes:

[0010] The feature extraction network consists of 5 feature extraction stages. Each feature extraction stage is composed of several feature extraction blocks. Each feature extraction block consists of a 3×3 convolution, a 1×1 convolution, a ReLu activation function, an identity mapping branch, and a batch normalization layer. The numbers of feature extraction blocks in the 5 feature extraction stages are 1, 4, 6, 16, and 1 respectively.

[0011] In the training stage, except that the first feature extraction block in each feature extraction stage consists of a 3×3 convolution, a 1×1 convolution, a ReLu activation function, an identity mapping branch, and a batch normalization layer, the remaining feature extraction blocks are all composed of a 3×3 convolution, a 1×1 convolution, an identity mapping branch, and a batch normalization layer.

[0012] In the inference stage, each feature extraction block is transformed into a 3×3 convolution through the structural reparameterization technique. The picture is sent into the feature extraction network, and finally, a feature map C1 with a size of [64, 256, 256], a feature map C2 with a size of [128, 128, 128], a feature map C3 with a size of [256, 64, 64], a feature map C4 with a size of [512, 32, 32], and a feature map C5 with a size of [2048, 16, 16] are generated.

[0013] Preferably, S2 specifically includes:

[0014] Perform an upsampling operation on the feature map C5 generated in step S1, and then use the deformable convolution operation to set the number of channels from 2048 to 512, and finally generate a feature map P5 with a size of [512, 16, 16]. Add the feature map C4 and the feature map P5 element by element, and then perform an upsampling operation and use the deformable convolution on the added feature map. And so on, to form a feature pyramid. Among them, the output of the last layer and the output of the penultimate layer of the feature pyramid will be used as the input of step S3.

[0015] Preferably, S3 specifically includes:

[0016] Use the last layer of the feature map to regress the heat map of the object, predict the offset between the center point and the actual center point, and the preliminary predicted size box of the object. Perform a convolution operation on the feature map A1 of the penultimate layer of the feature pyramid, with a convolution kernel of 1×1, a stride of 1, and a padding of 1.

[0017] Perform different convolution operations on the generated feature maps respectively to generate a feature map with a size of [H, W, num_classes] and two feature maps with a size of [H, W, 2]. The two feature maps with a size of [H, W, 2] respectively regress and predict the offset between the center point and the actual center point and the preliminary predicted bounding box of the object.

[0018] The feature map of the penultimate layer of the feature pyramid is convolved three times to generate a high-semantic feature map with a size of [H, W, 5C]. Then, the generated feature map with a size of [H, W, 5C] and the regression parameters of the rough object size box are fed into the average boundary extraction module to generate accurate object size box information. The average boundary extraction module is an object size box regression module that directly uses boundary features to enhance center point features.

[0019] Preferably, the structural process of the average boundary extraction module includes:

[0020] First, the feature map with a size of [H, W, 5C] generated by the convolution operation and the regression parameters of the preliminarily predicted object size box are used as inputs, and then the generated rough size box is projected onto the feature map with 5C channels.

[0021] Secondly, each boundary is divided into N points, where N represents the convolution kernel size of the average pooling operation in the next step, and then an average pooling operation is performed channel by channel to generate an average boundary.

[0022] Then, the average boundary extraction module can use the average points of the boundary to represent the boundary features. For the feature map with 5C channels, an average pooling operation is performed channel by channel, that is, a pooling operation is performed on each boundary separately.

[0023] Finally, the feature map with clear boundary information generated by the average boundary extraction module is convolved twice to finally predict the size and position of the size box.

[0024] Preferably, the num_classes represents the category of each pixel on the feature map, and the 5C can be expressed as (4 + 1)C, where C represents the category, 4C represents the 4 boundaries of each category, and C represents the center point.

[0025] Preferably, the network structure conversion process of the feature extraction network includes:

[0026] First, the convolution and its corresponding batch normalization layer are converted into a convolution with bias. Then, the 1×1 convolution branch and the identity mapping branch are converted into 3×3 convolution branches by padding with zeros, and then they are respectively converted with their corresponding batch normalization layers. Finally, the convolution kernels and biases from the conversions of the 3 branches are added together to obtain the final convolution kernel and bias.

[0027] A multi - object visual detection method for autonomous driving scenarios based on improved CenterNet of the present invention uses a feature pyramid structure composed of feature maps of different scales generated by a feature extraction module, which improves the detection accuracy of small objects by the driverless vehicle in the driving environment; the present invention uses an average boundary extraction module to assist in the regression of the boundary size of the center point, improves the regression accuracy of the target boundary, and at the same time improves the detection accuracy in the occlusion scenario, enhancing the robustness of the driverless vehicle detection; the present invention introduces a RepVGG feature extractor as the feature extraction module of the detection algorithm, and reduces the scale of the detection algorithm through the structural re - parameterization technology, greatly improving the detection speed and meeting the real - time requirements of driverless detection. Brief Description of the Drawings

[0028] The present invention will be further described below in conjunction with the drawings and embodiments.

[0029] Figure 1 is the flowchart of the multi - object visual detection algorithm of the present invention.

[0030] Figure 2 is the structure diagram of the feature extraction module of the present invention.

[0031] Figure 3 is the network structure conversion diagram of the feature extraction network in the training stage and the inference stage of the present invention.

[0032] Figure 4 is the parameter conversion diagram of the feature extraction network in the training stage and the inference stage of the present invention.

[0033] Figure 5 is the average boundary extraction module of the present invention. Detailed Embodiment

[0034] To make the technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described more clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0035] Figure 1 and Figure 2 The present invention shown as such proposes a multi - object visual detection method for autonomous driving scenarios based on improved CenterNet. Its overall implementation process is as Figure 1 shown, mainly including the following steps:

[0036] Step S1: Extract features from the pictures around the autonomous vehicle captured by the in-vehicle camera to obtain feature maps of different scales. The specific steps are as follows;

[0037] As Figure 3 and Figure 4 shown, the feature extraction network consists of 5 feature extraction stages, and each feature extraction stage consists of several feature extraction blocks. Each feature extraction block consists of a 3×3 convolution, a 1×1 convolution, a ReLu activation function, an identity mapping branch, and a batch normalization layer.

[0038] The number of feature extraction blocks in the 5 feature extraction stages is [1, 4, 6, 16, 1]. The feature extraction stage stage1 consists of only 1 feature extraction block block because the resolution of the obtained picture is very large and the processing time is long. To improve the speed, one feature extraction block is selected to extract the picture features. In order to obtain a high-resolution feature map and a faster inference time, the number of channels in the last 1 feature extraction stage stage is very large, so only 1 feature extraction block block is used to save the parameters.

[0039] In the training stage, except that the first feature extraction block in each feature extraction stage consists of a 3×3 convolution, a 1×1 convolution, a ReLu activation function, an identity mapping branch, and a batch normalization layer, the remaining feature extraction blocks all consist of a 3×3 convolution, a 1×1 convolution, an identity mapping branch, and a batch normalization layer. The picture is sent into the feature extraction network, and finally feature maps C1, C2, C3, C4, C5 with sizes of [64, 256, 256], [128, 128, 128], [256, 64, 64], [512, 32, 32], [2048, 16, 16] are generated. The structure diagram is shown in Figure 2 。

[0040] In the inference stage, each feature extraction block is transformed into a 3×3 convolution through the structure reparameterization technique. The transformation schematic diagram is shown in Figure 3 。

[0041] Among them, the specific steps of the network structure conversion of the feature extraction network in the training stage and the inference stage are as follows: Assume that the number of channels of the output feature and the output feature channels are both 2, then the convolution kernel size of the 3×3 convolution is W (3) =R 2×2×3×3 ,and the convolution kernel size of the 1×1 convolution is W (1) =R 2×2×1×1, μn, σn, γn, and βn represent the cumulative average, standard deviation, scale factor, and bias of the batch normalization layer after convolution. n ∈ {3, 1, 0} represents the 3×3 convolution operation, 1×1 convolution operation, and identity mapping branch. First, the convolution and its corresponding batch normalization layer are transformed into a convolution with bias. Then, the 1×1 convolution branch and the identity mapping branch are transformed into 3×3 convolution branches by padding, and then transformed with their corresponding batch normalization layers respectively. Finally, we add the convolution kernels and biases transformed from the three branches to obtain the final convolution kernel and bias. The schematic diagram of parameter transformation is shown in Figure 4 .

[0042] Step S2: Sample the feature map generated in Step S1 to generate a feature pyramid composed of feature maps of different scales. The specific steps are as follows;

[0043] First, upsample the feature map C5 generated in Step S1, and then use the deformable convolution operation to change the number of channels from 2048 to 512, finally generating a feature map P5 with a size of [512, 16, 16]. Add the feature map C4 and the feature map P5 element by element, and then perform upsample operation and use deformable convolution on the added feature map, and so on, to form a feature pyramid. The output of the last layer and the penultimate layer of the feature pyramid will be used as the input of Step 3.

[0044] Step S3: Extract different feature maps in Step 2 as the input of the detection head module, perform convolution operations on different feature maps to generate the final prediction results. The specific steps are as follows;

[0045] Use the last layer of feature maps to regress the heatmap of the object, predict the offset between the center point and the actual center point, and the size bounding box of the preliminarily predicted object. Perform a convolution operation on the feature map A1 of the penultimate layer of the feature pyramid with a convolution kernel of 1×1, a stride of 1, and a padding of 1 to eliminate the feature overlap effect brought about during the upsampling process. Then, perform different convolution operations on the generated feature maps to generate a feature map of size [H, W, num_classes] and two feature maps of size [H, W, 2] respectively. Here, num_classes represents the category of each pixel on the feature map, and the two feature maps of size [H, W, 2] are used to regress and predict the offset between the center point and the actual center point and the bounding box of the preliminarily predicted object. The reason for using the last layer of feature maps to regress the parameters of the rough size bounding box is that the resolution of the last layer of feature maps is high, which is beneficial for regressing the size bounding box of small objects. Use the feature map of the second-to-last layer of the feature pyramid to perform three convolution operations to generate a high-semantic feature map of size [H, W, 5C], where 5C can be expressed as (4 + 1)C, where C represents the category, 4C represents the 4 boundaries (top, bottom, left, and right) of each category, and C represents the center point. Then, send the feature map of size [H, W, 5C] generated in the previous step and the regression parameters of the rough object size bounding box into the average boundary extraction module to generate accurate object size bounding box information.

[0046] As Figure 5 shown, the average boundary extraction module is a newly designed object size bounding box regression module that directly uses boundary features to strengthen center point features. Its schematic diagram is shown in Figure 5 . The specific structural process of the average boundary extraction module is as follows: First, use the feature map of size [H, W, 5C] generated by the convolution operation and the regression parameters of the preliminarily predicted object size bounding box as inputs. Then, project the generated rough size bounding box onto the feature map with 5C channels. Next, divide each boundary into N points, where N represents the convolution kernel size of the average pooling operation in the next step, and then use the per-channel average pooling operation to generate the average boundary. The reason for dividing the boundary into several points for operation is that we believe that extracting boundary features point by point on the boundary is time-consuming and more memory-consuming. The subsequent average boundary extraction module can use the average points of the boundary to represent the boundary features. We use the per-channel average pooling operation on the feature map with 5C channels, that is, perform the pooling operation on each boundary separately, which can better represent the features of the boundary. Finally, perform two convolution operations on the feature map with clear boundary information generated by the average boundary extraction module to finally predict the size and position of the size bounding box.

[0047] A multi-object visual detection method for autonomous driving scenarios based on improved CenterNet of the present invention uses a feature pyramid structure composed of feature maps of different scales generated by a feature extraction module, improving the detection accuracy of small targets by unmanned vehicles in the driving environment; the present invention uses an average boundary extraction module to assist in the regression of the boundary size of the center point, improving the regression accuracy of the target boundary, while improving the detection accuracy in occlusion scenarios and the robustness of unmanned vehicle detection; the present invention introduces a RepVGG feature extractor as the feature extraction module of the detection algorithm, reducing the scale of the detection algorithm through the structural reparameterization technology, greatly improving the detection speed and meeting the real-time requirements of unmanned detection.

[0048] Inspired by the ideal embodiments of the present invention described above, through the above description, relevant staff can make various changes and modifications completely within the scope not deviating from the technical idea of the present invention. The technical scope of the present invention is not limited to the content in the specification, and its technical scope must be determined according to the scope of the claims.

Claims

1. A multi-object visual detection method for autonomous driving scenarios based on improved CenterNet, characterized in that Including the following steps: S1. Extract features from the pictures around the autonomous vehicle captured by the in-vehicle camera to obtain feature maps of different scales; S2. Perform sampling processing on the feature maps generated in step S1 to generate a feature pyramid composed of feature maps of different scales; S3. Extract different feature maps in step S2 as the input of the detection head module, perform convolution operations on different feature maps, and generate the final prediction results; S1 specifically includes: The feature extraction network consists of 5 feature extraction stages. Each feature extraction stage consists of several feature extraction blocks. Each feature extraction block consists of a 3×3 convolution, a 1×1 convolution, a ReLu activation function, an identity mapping branch, and a batch normalization layer. The numbers of feature extraction blocks in the 5 feature extraction stages are 1, 4, 6, 16, and 1 respectively; In the training stage, except that the first feature extraction block in each feature extraction stage consists of a 3×3 convolution, a 1×1 convolution, a ReLu activation function, an identity mapping branch, and a batch normalization layer, the remaining feature extraction blocks all consist of a 3×3 convolution, a 1×1 convolution, an identity mapping branch, and a batch normalization layer; In the inference stage, each feature extraction block is transformed into a 3×3 convolution through the structural reparameterization technique. The picture is sent into the feature extraction network, and finally a feature map C1 with a size of [64, 256, 256], a feature map C2 with a size of [128, 128, 128], a feature map C3 with a size of [256, 64, 64], a feature map C4 with a size of [512, 32, 32], and a feature map C5 with a size of [2048, 16, 16] are generated; S2 specifically includes: Perform an upsampling operation on the feature map C5 generated in step S1, and then use the deformable convolution operation to set the number of channels from 2048 to 512, and finally generate a feature map P5 with a size of [512, 16, 16]. Add the feature map C4 and the feature map P5 element by element, and then perform an upsampling operation and use the deformable convolution on the added feature map; and so on, to form a feature pyramid. Among them, the output of the last layer and the output of the penultimate layer of the feature pyramid will be used as the input of step S3; S3 specifically includes: Use the last layer of the feature map to regress the heat map of the object, predict the offset between the center point and the actual center point, and the size box of the preliminarily predicted object. Perform a convolution operation on the feature map A1 of the penultimate layer of the feature pyramid, with a convolution kernel of 1×1, a stride of 1, and a padding of 1; Perform different convolution operations on the generated feature maps respectively to generate a feature map with a size of [H, W, num_classes] and two feature maps with a size of [H, W, 2]. The two feature maps with a size of [H, W, 2] respectively regress and predict the offset between the center point and the actual center point and the bounding box of the preliminarily predicted object; The second-to-last layer feature map of the feature pyramid is subjected to three convolutions to generate a high-semantic feature map with a size of [H, W, 5C]. Then, the generated feature map with a size of [H, W, 5C] and the rough object size box regression parameters are fed into the average boundary extraction module together to generate accurate object size box information. The average boundary extraction module is an object size box regression module that directly uses boundary features to enhance center point features.

2. The multi-object visual detection method for an autonomous driving scenario based on the improved CenterNet according to claim 1, wherein, The structural process of the average boundary extraction module includes: First, the feature map with a size of [H, W, 5C] generated by the convolution operation and the preliminary predicted object size box regression parameters are used as inputs, and then the generated rough size box is projected onto the feature map with 5C channels. Secondly, each boundary is divided into N points, where N represents the convolution kernel size of the average pooling operation in the next step, and then an average pooling operation is performed channel by channel to generate an average boundary. Then, the average boundary extraction module can use the average points of the boundary to represent the boundary features. For the feature map with 5C channels, an average pooling operation is performed channel by channel, that is, a pooling operation is performed on each boundary separately. Finally, the feature map with clear boundary information generated by the average boundary extraction module is subjected to two convolution operations to finally predict the size and position of the size box.

3. A multi-object visual detection method for an autonomous driving scenario based on improved CenterNet according to claim 1, characterized in that, The num_classes represents the category of each pixel on the feature map, and the 5C can be expressed as (4 + 1)C, where C represents the category, 4C represents the 4 boundaries of each category, and C represents the center point.

4. A multi-object visual detection method for an autonomous driving scenario based on an improved CenterNet, characterized in that, The network structure conversion process of the feature extraction network includes: First, the convolution and its corresponding batch normalization layer are converted into a convolution with bias, then the 1×1 convolution branch and the identity mapping branch are converted into 3×3 convolution branches by padding with zeros, and then they are respectively converted with their corresponding batch normalization layers. Finally, the convolution kernels and biases from the conversions of the 3 branches are added together to obtain the final convolution kernel and bias.