Target detection method and device, terminal equipment and storage medium

Through the extraction of pillar feature and multi-scale feature fusion in the single-stage object detection model, combined with the augmentation of data of the region of interest and the aggregation of neighborhood information, the problem of insufficient small object detection capabilities of lidar in sparse point cloud environments is solved, and the detection accuracy and accuracy are improved.

CN120279240AInactive Publication Date: 2025-07-08VANJEE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311845563.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-28
Publication Date
2025-07-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing lidars have sparsity problems when detecting small objects, especially in scenarios where pedestrians and cyclists have limited detection capabilities.

Method used

A single-stage object detection model with a pillar feature extraction layer, a multi-scale feature fusion network and an object detection layer is adopted, and combined with the data enhancement of the region of interest and the neighborhood information aggregation layer, the three-dimensional point cloud images are directly processed, feature extraction and data enhancement are enhanced, and small object detection capabilities are improved.

Benefits of technology

有效减少数据损失,提高了稀疏点云小目标的检测精度和准确性,特别是在盲区内的行人检测效果显著提升。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279240A_ABST
    Figure CN120279240A_ABST
Patent Text Reader

Abstract

The invention is applicable to the field of target detection, and provides a target detection method and device, a terminal device and a storage medium, N frames of point cloud images of a traffic road section to be detected are acquired firstly, then the N frames of point cloud images are input to a preset single-stage target detection model, and the single-stage target detection model outputs at least one target detection result. According to the roadside-oriented sparse point cloud small target detection method provided by the invention, by using the target detection model comprising the pillar extraction layer, compared with a conventional model in which three-dimensional point cloud data cannot be input, the three-dimensional point cloud image can be directly processed, the data loss is reduced, and the feature is enhanced by the pillar extraction layer, so that the detection accuracy is improved. And small target detection of sparse point clouds can be dealt with.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of object detection, and particularly relates to an object detection method, device, terminal device, and storage medium. Background Art

[0002] To achieve L5 autonomous driving, accurate and reliable 3D object detection results are crucial. Currently, the idea of cooperative vehicles and roadside infrastructure is more feasible. In roadside perception systems, LiDAR is used as the main high-precision sensor to provide supplementary information outside the vehicle's field of view to improve the safety of autonomous vehicles. In this process, the roadside perspective is very important. However, due to the sparsity of LiDAR, its detection ability for small objects is limited, making this technology extremely limited in supplementing vehicle information in scenarios with high pedestrian and cyclist densities. Summary of the Invention

[0003] The purpose of this application is to provide an object detection method, device, terminal device, and storage medium, aiming to solve the problem of limited detection ability for small objects due to the sparsity of LiDAR.

[0004] The first aspect embodiment of this application provides an object detection method, and the object detection method includes:

[0005] Obtain point cloud images of N frames of the traffic section to be detected, where the point cloud images are obtained by scanning with a LiDAR installed on the roadside of the traffic section to be detected;

[0006] Input the N frames of the point cloud images into a preset single-stage object detection model, and the single-stage object detection model outputs at least one object detection result;

[0007] Among them, the single-stage object detection model includes a pillar feature extraction layer, a multi-scale feature fusion network, and an object detection layer. The pillar feature extraction layer is used to input the N frames of the point cloud images and output a three-dimensional feature map. The multi-scale feature fusion network is used to input the three-dimensional feature map and output a three-dimensional aggregated feature map. The object detection layer is used to input the three-dimensional aggregated feature map and output at least one object detection result. The three-dimensional features include the number of feature channels, the height of the feature map, and the width of the feature map.

[0008] In an optional embodiment, the multi-scale feature fusion network includes:

[0009] Multiple convolutional layers. The input of the first convolutional layer is the three-dimensional feature map, and the output is a feature image of a set scale. The input of the next adjacent convolutional layer is the output of the previous adjacent convolutional layer, and at least two convolutional layers correspond to different set scales;

[0010] A plurality of pooling layers, corresponding one-to-one with the plurality of convolutional layers. The input of the first pooling layer is the output of the first convolutional layer, and the input of the next adjacent pooling layer is the output of the corresponding convolutional layer and the output of the previous adjacent pooling layer. The output of each pooling layer is a feature image of the same scale;

[0011] A connection layer that combines the outputs of each pooling layer by increasing the number of channels to generate the three-dimensional aggregated feature map;

[0012] Inputting the point cloud image into a preset single-stage object detection model, and the single-stage object detection model outputs at least one object detection result, including:

[0013] Inputting the point cloud image into the first convolutional layer to obtain the three-dimensional aggregated feature map.

[0014] In an alternative embodiment, the single-stage object detection model further includes:

[0015] A region of interest data augmentation layer. The input of the region of interest data augmentation layer is N frames of the point cloud image, and the output is a point cloud image including pseudo-samples, and the output is used as the input of the pillar feature extraction layer;

[0016] Before inputting the point cloud image into the first convolutional layer, inputting the point cloud image into a preset single-stage object detection model further includes:

[0017] Inputting N frames of the point cloud image into the region of interest data augmentation layer to obtain a point cloud image including pseudo-samples.

[0018] In an alternative embodiment, inputting N frames of the point cloud image into the region of interest data augmentation layer to obtain a point cloud image including pseudo-samples includes:

[0019] Taking N frames of point cloud images as training samples, randomly selecting one frame of historical point cloud image for training, and setting a random number with a range including a first preset value and a second preset value; wherein the first preset value is greater than a preset K value;

[0020] If the random number corresponding to the currently trained point cloud image is greater than the preset K value, randomly add the object point clouds in other frames of historical point cloud images to the currently trained point cloud image to generate a point cloud image including pseudo-samples.

[0021] In an alternative embodiment, before inputting the point cloud image into the first convolutional layer to obtain the three-dimensional aggregated feature map, inputting the point cloud image into a preset single-stage object detection model, and the single-stage object detection model outputs at least one object detection result further includes:

[0022] Perform global data augmentation operations on the N-frame point cloud images to obtain multiple preprocessed point cloud images, and replace the N-frame point cloud images with the set of multiple preprocessed point cloud images and the N-frame point cloud images as the input to the first convolutional layer.

[0023] In an alternative embodiment, the single-stage object detection model further includes a neighborhood information aggregation layer, and the neighborhood information aggregation layer includes multiple computing nodes. The input of each computing node is the output of all corresponding adjacent computing nodes, and the output is used as one of the inputs of all corresponding adjacent computing nodes.

[0024] Before inputting the point cloud image into the first convolutional layer to obtain the three-dimensional aggregation feature map, when inputting the point cloud image into a preset single-stage object detection model, and the single-stage object detection model outputs at least one object detection result, it further includes:

[0025] Input the N-frame point cloud images into the neighborhood information aggregation layer, and the neighborhood information aggregation layer outputs N-frame three-dimensional feature maps containing neighborhood information and inputs them into the first convolutional layer.

[0026] In an alternative embodiment, inputting the N-frame point cloud images into the region of interest data augmentation layer to obtain a point cloud image including pseudo-samples includes:

[0027] Use the N-frame point cloud images as training samples, and randomly select one frame of historical point cloud image for training, and set a random number with a range including a first preset value and a second preset value; where the first preset value is greater than the preset K value.

[0028] If the random number corresponding to the currently trained point cloud image is greater than the preset K value, use the GT-sample algorithm to randomly add the current frame of object point cloud to other frame point cloud images to generate a point cloud image including pseudo-samples.

[0029] Another embodiment of the present application provides an object detection device, and the object detection device includes:

[0030] An acquisition module that acquires N-frame point cloud images of a traffic section to be detected, and the point cloud images are obtained by scanning with a lidar installed on the roadside of the traffic section to be detected.

[0031] A detection module that inputs the N-frame point cloud images into a preset single-stage object detection model, and the single-stage object detection model outputs at least one object detection result.

[0032] Among them, the single-stage object detection model includes a backbone feature extraction layer, a multi-scale feature fusion network, and an object detection layer. The backbone feature extraction layer is used to input N frames of the point cloud images and output a three-dimensional feature map. The multi-scale feature fusion network is used to input the three-dimensional feature map and output a three-dimensional aggregated feature map. The object detection layer is used to input the three-dimensional aggregated feature map and output at least one object detection result. The three-dimensional features include the number of feature channels, the height of the feature map, and the width of the feature map.

[0033] A third aspect of the embodiments of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method described in the first aspect above is implemented.

[0034] A fourth aspect of the embodiments of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.

[0035] A fifth aspect of the embodiments of the present application provides a computer program product, which when running on a terminal device causes the terminal device to execute the method described in the first aspect above.

[0036] Beneficial effects of this application

[0037] A method, device, terminal device, and storage medium for object detection provided by the present application first obtain N frames of point cloud images of a traffic section to be detected, and then input the N frames of point cloud images into a preset single-stage object detection model. The single-stage object detection model outputs at least one object detection result. The present application proposes a method for detecting small objects in sparse point clouds facing the roadside. By using an object detection model including a backbone extraction layer, compared with a conventional model that cannot input three-dimensional point cloud data, it can directly process three-dimensional point cloud images, reduce data loss, and because the backbone extraction layer strengthens the features, it can handle the detection of small objects in sparse point clouds. Description of the Drawings

[0038] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 It is a schematic diagram of the overall architecture of the single-stage object detection model SPP provided by the embodiments of the present application;

[0040] Figure 2 Schematic diagram of the enhanced effect of the region of interest provided by the embodiment of the present application;

[0041] Figure 3 Enlarged schematic diagram of the enhanced effect of the region of interest provided by the embodiment of the present application;

[0042] Figure 4 One of the schematic diagrams of the experimental demonstration data provided by the embodiment of the present application;

[0043] Figure 5 Another schematic diagram of the experimental demonstration data provided by the embodiment of the present application;

[0044] Figure 6 Schematic diagram of the experimental demonstration data provided by the embodiment of the present application Figure 3 ;

[0045] Figure 7 Schematic diagram of the strut feature extraction model provided by the embodiment of the present application;

[0046] Figure 8 Schematic flow diagram of a target detection method provided by the embodiment of the present application;

[0047] Figure 9 Schematic structural diagram of a target detection device provided by the embodiment of the present application;

[0048] Figure 10 Schematic structural diagram of a terminal device provided by the embodiment of the present application. Detailed implementation manners

[0049] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. In the following description, specific details such as specific system structures and technologies are proposed for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, the detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0050] It should be understood that when used in the specification and the appended claims of the present application, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0051] It should also be understood that the term "and / or" as used in the specification and appended claims of this application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0052] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrase "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.

[0053] In addition, in the description of the specification and appended claims of this application, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0054] Reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a specific feature, structure, or characteristic described in connection with that embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all of the embodiments", unless otherwise specifically emphasized in other ways. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0055] It should be understood that the magnitudes of the sequence numbers of the steps in this embodiment do not mean the order of execution is prior or subsequent, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of this application.

[0056] Currently, the performance of LiDAR 3D object detection is still limited by its low detection accuracy for small object categories such as pedestrians. One of the common solutions to improve the detection accuracy of small objects is to introduce multi-sensor data, such as BEVFusion and RaLiBEV. Yang et al. proposed RaLiBEV, which uses transformations to perform feature-level fusion of radar and LiDAR data, and also proposed a new label assignment strategy. Liu et al. proposed BEVFusion, an efficient and general multi-task multi-sensor fusion framework. It unifies multi-modal features in a shared bird's-eye view (BEV) representation space, well preserving geometric and semantic information. However, the detection accuracy of multi-modal methods highly depends on the alignment of different sensors. The fusion of misaligned data from different sensors introduces noise, thus reducing the final perception performance. Therefore, some methods have been proposed on how to align multi-modal data.

[0057] Another way to improve the ability to detect small objects in the perception system is achieved through some traditional machine learning clustering methods, such as Density-Based Spatial Clustering of Applications with Noise (DBSCAN). DBSCAN relies on the density-based cluster concept and aims to discover clusters of arbitrary shapes. Liu et al. used the DBSCAN method to identify nearby traffic objects such as vehicles and road users.

[0058] However, traditional machine learning methods are restricted by manual parameter settings, and the performance of the algorithm will significantly decline with the change of scenarios.

[0059] To solve the above problems, this application provides an object detection method. By using an object detection model including a pillar extraction layer, compared with a conventional model that cannot input 3D point cloud data, it can directly process 3D point cloud images, reducing data loss. And because the pillar extraction layer strengthens the features, it can handle small object detection in sparse point clouds.

[0060] In the embodiments of this application, the overall architecture of the object detection model is improved, including the following three points:

[0061] 1. Configure a pillar feature extraction layer in the single-stage object detection model to strengthen the features, which can handle small object detection in sparse point clouds.

[0062] 2. Design a region of interest enhancement layer, which can significantly improve the accuracy of pedestrian detection.

[0063] 3. A neighborhood information aggregation layer is configured. By increasing the stride of the convolutional kernel while keeping the number of channels unchanged and expanding the receptive field, it aggregates more extensive neighborhood information into the current pixel. In cooperation with the region of interest enhancement layer, it compensates for the unreasonable features brought by pseudo-samples in training by aggregating neighborhood information, enabling the network to better detect small objects in blind spots and forming a synergistic effect.

[0064] The above three core inventive points of this application will be described separately below.

[0065] 1. Configure a backbone feature extraction layer in a single-stage object detection model

[0066] In the embodiment of this application, a lidar is a radar system that detects the position, speed and other characteristic quantities of a target by emitting laser beams. Its working principle is to emit a detection signal (laser beam) to the target, and then compare the received signal (target echo) reflected from the target with the emitted signal. After appropriate processing, relevant information about the target can be obtained, such as parameters such as target distance, azimuth, altitude, speed, attitude, and even shape, so as to realize the detection, tracking and recognition of targets such as vehicles.

[0067] Among them, the current traffic section can be the section within the scanning range of the lidar at the current moment.

[0068] Among them, the point cloud image is three-dimensional point cloud image data.

[0069] Exemplarily, the lidar is installed at a preset position on the current traffic section, and the lidar collects the three-dimensional point cloud image data of this traffic section and stores and transmits this data respectively.

[0070] Figure 1 Shows the overall architecture of the single-stage object detection model in the embodiment of this application, as Figure 1 shown, the single-stage object detection model at least includes: a backbone feature extraction layer, a multi-scale fusion layer and an object detection layer. Among them, the backbone feature extraction layer is used to input N frames of the point cloud image and output a three-dimensional feature map, the multi-scale feature fusion network is used to input the three-dimensional feature map and output a three-dimensional aggregated feature map, the object detection layer is used to input the three-dimensional aggregated feature map and output at least one object detection result, and the three-dimensional feature includes the number of feature channels, the height of the feature map and the width of the feature map.

[0071] Combined with Figure 1 , this application provides an object detection method, as Figure 2 shown, including:

[0072] S1: Obtain N frames of point cloud images of the traffic section to be detected, and the point cloud images are obtained by scanning with a lidar installed on the roadside of the traffic section to be detected;

[0073] S2: Input the N frames of the point cloud images into a preset single-stage object detection model, and the single-stage object detection model outputs at least one object detection result.

[0074] A target detection method provided by this application first obtains N frames of point cloud images of the traffic section to be detected, and then inputs the N frames of the point cloud images into a preset single-stage object detection model. The single-stage object detection model outputs at least one object detection result. This application proposes a small object detection method for sparse point clouds facing the roadside. By using an object detection model including a pillar extraction layer, compared with a conventional model that cannot input three-dimensional point cloud data, it can directly process three-dimensional point cloud images, reduce data loss, and since the pillar extraction layer strengthens the features, it can handle the small object detection of sparse point clouds.

[0075] First, the pillar extraction layer, multi-scale fusion network, and object detection layer will be described in detail below.

[0076] The three-dimensional features of the pillar extraction layer include the number of feature channels, the height of the feature map, and the width of the feature map. The pillar extraction layer, multi-scale fusion network, and object detection layer cooperate to form a point-pillar network structure, which converts the point cloud into pillar features for feature extraction and calculation.

[0077] Exemplarily, the point-pillar network structure composed of the cooperation of the pillar extraction layer, multi-scale fusion network, and object detection layer can include the following steps when specifically used:

[0078] 1) Divide the point cloud space corresponding to each frame into fixed-size cylinders, which are the Pillars here. Each Pillar contains several points.

[0079] 2) Each point is represented by 10-dimensional features, which are: <x, y, z, r / v, x c , y c , z c , x p , y p , z p .

[0080] In kitti, the fourth dimension is intensity r, and in millimeter waves, it can be velocity v. x c , y c , z c is the deviation of the current point from the average value of all points within the current pillar, and is calculated by the following formula:

[0081] x c =(x - x mean )

[0082] yc =(y - y mean )

[0083] z c =(z - z mean )

[0084]

[0085] In the formula, N point represents the number of points inside the current Pillar. x p , y p , z p respectively represent the distances in the x, y, and z dimensions of the current point relative to the center point of the current Pillar. Data shape: [batch, P, N, 10], where N represents the number of points and P represents the number of Pillars into which one frame of data is divided.

[0086] 3) Increase the dimension of the points to 64 using a fully connected layer: [batch, P, N, 9] -> [batch, P, N, 64]. The deviation of the current point from the average value of all points inside the current pillar is calculated using the following formula:

[0087] 4) After MaxPooling, compress the number of points in each pillar to 1, and only take the representative points through max pooling, [batch, P, N, 64] -> [batch, P, N, 64].

[0088] 5) Map the data obtained above back to the original space and convert it into a 2D pseudo-image, that is, convert the data into the format of [batch, C, H, W]. C represents the number of channels. The specific process is as follows:

[0089] a. According to the size of each pillar, the H and W in the original space can be obtained, where W represents the number of points inside the current Pillar. Since the pillar occupies the entire height, there is no height dimension here. That is, finally GridSize = [H, W, 1].

[0090] b. According to the coordinates of each pillar in the original data, the coordinates in the original space can be obtained, that is, the position of the current pillar in the H*W space can be obtained.

[0091] c. After the above process, the data of a certain feature dimension of all pillars in one frame is stored in the space of size H*W for each channel. There are a total of 64 feature dimensions, so the data shape is: [batch, 64, H, W].

[0092] In short, the point cloud data is divided into grids according to the X and Y axes (ignoring the Z axis) where the point cloud data is located. All the point cloud data falling into a grid is regarded as being in a pillar (Pillar), or it can be understood that they form a Pillar.

[0093] After that, each point cloud is represented by a 9-dimensional vector: (x, y, z, r, x c , y c , z c , x p , y p ), where x, y, z, and r are the real coordinate information (three-dimensional) and reflection intensity of the point cloud, and x c , y c , z c are the geometric centers of all points in the Pillar where the point cloud is located, and x p , y p are x - x c , y - y c , reflecting the relative position of the point to the geometric center.

[0094] When performing feature extraction, the dimensions of the point cloud can be processed. The original dimension of the point cloud is D = 9, and the processed dimension is C. Therefore, a tensor of (C, P, N) can be obtained at this time; then, the present application performs MaxPooling operation according to the dimension where the Pillar is located, that is, a feature map of (C, P) dimension is obtained. After that, in order to obtain pseudo-image features, the P features are transformed into H×W, so as to obtain a three-dimensional feature map of (C, H, W).

[0095] The input of the pseudo-image 2D CNN is used to further extract image features. As can be seen from the figure, the 2D CNN uses two networks. One of the networks continuously reduces the resolution of the feature map while increasing the dimension of the feature map, so three feature maps with different resolutions are obtained.

[0096] After that, through the pooling layer, the three feature maps with different resolutions are transformed into the same scale, and then the channel increasing operation is performed. The three feature maps of the same scale are combined together and sent to the target detection layer.

[0097] In this embodiment, the multi-scale feature fusion network includes:

[0098] Multiple convolutional layers. The input of the first convolutional layer is the three-dimensional feature map, and the output is a feature image of a set scale. The input of the next adjacent convolutional layer is the output of the previous adjacent convolutional layer, and at least two convolutional layers have different set scales;

[0099] Multiple pooling layers, corresponding one by one to the multiple convolutional layers. The input of the first pooling layer is the output of the first convolutional layer, and the input of the next adjacent pooling layer is the output of the corresponding convolutional layer and the output of the adjacent previous pooling layer. The output of each pooling layer is a feature image of the same scale;

[0100] A connection layer that combines the outputs of each pooling layer by increasing the number of channels to generate the three-dimensional aggregated feature map;

[0101] Inputting the point cloud image into a preset single-stage object detection model, and the single-stage object detection model outputs at least one object detection result, including:

[0102] Inputting the point cloud image into the first convolutional layer to obtain the three-dimensional aggregated feature map.

[0103] In a preferred embodiment, anchor points of a fixed size determined according to the sizes and central positions of all ground truths, because the objects to be detected have approximately fixed sizes. For the regression target, the present application uses the following box encoding function:

[0104]

[0105]

[0106] 0 = 0 - 0(3)

[0107] where t, a, and g represent the encoded value, the anchor point, and the ground truth respectively; and,

[0108] is the diagonal of the bottom of the anchor box.

[0109] Another network upsamples the three feature maps to the same size and then performs concatenation.

[0110] In the embodiments of the present application, feature maps of different resolutions are responsible for detecting objects of different sizes. For example, feature maps with a large resolution often have a small receptive field and are suitable for capturing small objects (pedestrians in KITTI), so they are especially suitable for extracting small targets under sparse point clouds.

[0111] A loss function similar to SECOND, where each 3D bbox is represented by a 7-dimensional vector (x, y, z, w, h, l, θ) respectively, and is trained using the smooth L1 loss function:

[0112]

[0113] L θ = SmoothL1(sin(θ p - θ t ) (5)

[0114] Where θp represents the predicted angle.

[0115] For the selection of the classification loss, this application uses focal loss, which can solve the problem of class imbalance:

[0116] L cls = -α t (1 - pt) γ log(pt) (6)

[0117] Where pt is the estimated probability of the model, and α and γ are the parameters of focal loss.

[0118] By combining the losses discussed above, this application can obtain the final form of the multi-task loss as follows:

[0119] L total = β1L cls + β2L loc + β3L dir (7)

[0120] Where Lcls is the classification loss, Lloc is the regression loss of location and dimension, and Ldir is the direction classification loss.

[0121] In a preferred embodiment, global data augmentation operations such as rotation, flipping, and scaling can be performed before input convolution. In this embodiment, before inputting the point cloud image into the first convolutional layer to obtain the three-dimensional aggregated feature map, when inputting the point cloud image into a preset single-stage object detection model, the single-stage object detection model outputs at least one object detection result, and further includes:

[0122] Performing global data augmentation operations on N frames of the point cloud image to obtain multiple preprocessed point cloud images, and using the multiple preprocessed point cloud images and the set of N frames of the point cloud image to replace N frames of the point cloud image as the input of the first convolutional layer.

[0123] Global augmentation can increase the data sample capacity, thereby making object detection more accurate.

[0124] Through the detailed description of the above single-stage object detection model, it can be seen that this application can directly process three-dimensional point cloud images by using an object detection model including a pillar extraction layer, reducing data loss. And because the pillar extraction layer strengthens the features, it can handle small object detection in sparse point clouds.

[0125] 2. Region of Interest Data Augmentation

[0126] Region of Interest Data Augmentation is a technique for enhancing data of regions of interest in computer vision tasks, mainly used in tasks such as object detection, semantic segmentation, and instance segmentation, aiming to improve the robustness and generalization ability of the model.

[0127] In the embodiments of this application, the region of interest enhancement is further used for a single-stage object detection model. In this embodiment, the single-stage object detection model further includes:

[0128] A region of interest data augmentation layer, the input of the region of interest data augmentation layer is N frames of the point cloud image, the output is a point cloud image including pseudo-samples, and the output is used as the input of the pillar feature extraction layer;

[0129] Before inputting the point cloud image into the first convolutional layer, inputting the point cloud image into a preset single-stage object detection model further includes:

[0130] Input N frames of the point cloud image into the region of interest data augmentation layer to obtain a point cloud image including pseudo-samples.

[0131] Specifically, the sample position distributions in the training set and the validation set may be different in the roadside scenario. When there are insufficient training samples in the region of interest in our training set, when we validate the model, it will result in the lack of objects in that region. This situation is more obvious for small objects such as pedestrians, because compared with vehicles, the features of pedestrians are more indistinct and variable.

[0132] Based on this, we propose Region of Interest Data Augmentation (RoIAug), which is a strategy for generating pseudo point clouds by copying target point clouds from other places in the training samples to a given region of interest. In this study, GT-sample and RoIAug are applied for data augmentation, and an artificial parameter k is used to determine the weights of the two strategies. The algorithm implementation steps are shown in Algorithm 1 below.

[0133] Algorithm 1 RoIAug: Region of Interest Data Augmentation

[0134] Initialization: S = {δ0, δ1,..., δN}, δt = ; S is the sample set of all training data, in this data set we have N frames of point clouds δt representing the object point cloud set in the current frame; and m in stm represents the object index in the current frame. k is an artificial parameter, and the value of k is inversely proportional to the weight of RoIAug in data augmentation, and ε represents the number of pseudo-objects created by RoIAug

[0135] Repeat n ← n + 1;

[0136] Randomly select Train, and set a random number p ∈ [0, 1];

[0137] If p > k: Use the GT sample strategy;

[0138] Otherwise, randomly place the sample s ∈ {S - δt} into the region of interest in the current frame;

[0139] Until n = ε, Another one Update to the current frame image.

[0140] Figure 2 (a) depicts the data frame after data augmentation through GT samples and RoIAug. Compared with the original data, the number of pedestrian samples has increased significantly. The blue boxes represent the samples added by GT samples, while the red boxes represent the artificial samples added by RoIAug. In addition, we magnify the samples added by the two strategies respectively to show their details, as Figure 2 (b) shows. It can be seen that due to the addition of pseudo-samples by RoIAug, there are inconsistencies with real samples in terms of details. Compared with the pseudo-samples, the point cloud distribution of the real samples replicated by GT-sample is on the side of the object close to the lidar. For the pseudo-samples, part of the point cloud is distributed on the other side, obviously there are significant differences between the pseudo-samples and the real samples. For the pseudo-samples added by RoIAug in this application, the distance between the point cloud distribution and the point cloud distribution of the real samples is relatively close, and the effect is better than that of GT-sample.

[0141] According to the above algorithm, it can be understood that in the embodiment of this application, inputting the N-frame point cloud images into the region of interest data augmentation layer to obtain a point cloud image including pseudo-samples includes:

[0142] Taking the N-frame point cloud images as training samples, randomly selecting one of the historical point cloud images for training, and setting a random number with a range including a first preset value and a second preset value; where the first preset value is greater than the preset K value;

[0143] If the random number corresponding to the currently trained point cloud image is greater than the preset K value, randomly add the object point clouds in other historical point cloud images to the current frame point cloud image to generate a point cloud image including pseudo-samples.

[0144] Through the above region of interest enhancement method, the accuracy of pedestrian detection can be significantly improved.

[0145] 3. Neighborhood Feature Aggregation

[0146] Since there are differences in the point cloud features of objects between pseudo-samples and real samples, such as point cloud density and direction. We propose a Neighborhood Information Aggregation (NIA) layer, which aggregates more extensive neighborhood information into the current pixel by increasing the stride of the convolutional kernel while keeping the number of channels unchanged and expanding the receptive field.

[0147] Based on this, the single-stage object detection model described in the embodiments of the present application further includes a neighborhood information aggregation layer, and the neighborhood information aggregation layer includes multiple computing nodes. The input of each computing node is the output of all corresponding adjacent computing nodes, and the output is used as one of the inputs of all corresponding adjacent computing nodes.

[0148] Before inputting the point cloud image into the first convolutional layer to obtain the three-dimensional aggregated feature map, when inputting the point cloud image into a preset single-stage object detection model, and the single-stage object detection model outputs at least one object detection result, it further includes:

[0149] Input N frames of the point cloud image into the neighborhood information aggregation layer, and the neighborhood information aggregation layer outputs N frames of three-dimensional feature maps containing neighborhood information and inputs them into the first convolutional layer.

[0150] Figure 3 Illustrates the difference between the baseline receptive field and the receptive field of the NIA layer. Obviously, compared with the baseline, adding the NIA layer can expand the receptive field.

[0151] When we adopt a non-linear activation function in the neural network, the backpropagation is as follows:

[0152]

[0153] Where g(i, j, p - 1) is the gradient of the pixel (i, j) at the (p - 1)th layer, a and b are the offsets of the positions, and g(i + a, j + b, p) is used to represent the gradient of the neighborhood information.

[0154] Use (σpi + a, j + b) to represent the gradient of the non-linear activation function of the pixel (i, j) at the pth layer. (ωPa, b) is the weight of the convolutional kernel in the Pth layer.

[0155] By using the NIA layer, k - 1 can be increased to 2k - 1, and it can be found from Equation 8 that more neighborhood information is used to calculate the gradient weight at (i, j). And by aggregating the neighborhood information to compensate for the unreasonable features brought by pseudo-samples in training, the network can better detect small objects in blind spots. The experimental results show that the neighborhood information aggregation effect is good, and when combined with the region of interest enhancement layer, by aggregating the neighborhood information to compensate for the unreasonable features brought by pseudo-samples in training, the network can better detect small objects in blind spots, forming a synergistic effect.

[0156] Experimental data verification

[0157] This application uses the Vanjee Gate No.8 dataset, including 4800 frames (2700 pedestrian samples) in the training set and 1062 frames (743 pedestrian samples) in the validation set. This application also calculates the position distribution of pedestrians in the training and validation sets using an 8m grid, as Figure 4 shown. By binarizing the heatmaps of the sample distributions in the training and validation sets, this application can obtain Figure 5 the results given in the regions, indicating that the sample distributions in the training and validation sets are different in the two regions.

[0158] For pedestrian object detection, this application starts with PointPillars, which can achieve 53.44 mAP at 0.25 IoU. The results of the pedestrian detection accuracy counts for 8m and 16m radii are then shown in Table 1 for the Figure 5 distribution region in the right figure in. It has been found that the pedestrian detection accuracy in the blind area drops significantly.

[0159] In this paper, the entire process is customized by the Openpcdet toolbox. The batch size is 8, the maximum number of epochs is 200, and the learning rate policy is single-cycle Adm. This application sets the IoU matching threshold to 0.3 to generate proposals in the NMS step. During the training process, this application sets the PFE layer from N = 9 to N = 64, and the shape of the feature map is (400, 400, 64). Finally, this application performs regression and classification on the boxes in the (200, 200, 328) feature map. For Equation 6Lcl, this application uses α = 0.25 and γ = 2, and β2 = 2.0 and β3 = 0.2 are the constant coefficients of Loss Equation 7. In Algorithm 1, it is recommended to use k = 0.15, and this application sets ε = 15 in this paper.

[0160] 8m 16m Location gt / dt 3d mAP@0.5 gt / dt 3d mAP@0.5 (10,10) 112 / 4 3.57% 200 / 68 13.8% (17.5,15) 81 / 12 9.88% 81 / 12 9.88%

[0161] Table 1. Results at 8m and 16m in different distribution regions

[0162] As Figure 6 shown, where the red boxes represent the ground truth and the green boxes represent the detection boxes, this application compares the pedestrian detection results of PointPillars and SSP, paying particular attention to the detection results in the blind area. Obviously, SSP can identify pedestrian objects in the blind area more accurately than PointPillars.

[0163] In Table 2, the present application provides metrics for evaluating the detection accuracy of blind spots. Based on the IoU accuracy values of 0.25 and 0.5, the present application calculates the 3d mAP for pedestrian detection. It can be found that using the RoIAug method, the mAP at 0.25 IoU is increased by 9.45%, and NIA can achieve a better improvement of about 7.36% at 0.5 IoU. By using the RoIAug and NIA layers, the mAP can achieve the best results at 0.5 and 0.25 IoU. This is consistent with the prediction because due to the pseudo-samples provided by RoIAug, the model will not be able to recognize small items with high IoU, but the recall rate can be greatly improved.

[0164]

[0165] It can be seen that the present application proposes a roadside lidar small target detection SSP, which can more effectively solve the blind spot problem. Specifically, the present application proposes an NIA layer, an architecture for aggregating neighborhood information in a neural network, and RoIAug, a data augmentation method based on regions of interest, which can add pseudo-samples to the blind spot. Experiments on the Vanjee GateNo.8 dataset show that the method of the present application performs better in small object detection. Using SSP, in a vehicle-infrastructure collaborative system, the reliability of roadside lidar small object perception information can be effectively improved. In addition, it can avoid harm to vulnerable road users by autonomous vehicles.

[0166] Figure 3 The figure is a schematic structural diagram of an object detection device provided by an embodiment of the present application. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown.

[0167] The object detection device may specifically include the following modules:

[0168] An acquisition module 101, which acquires point cloud images of N traffic sections to be detected, and the point cloud images are obtained by scanning with a lidar disposed on the roadside of the traffic section to be detected;

[0169] A detection module 102, which inputs the N point cloud images into a preset single-stage object detection model, and the single-stage object detection model outputs at least one object detection result;

[0170] Among them, the single-stage object detection model includes a backbone feature extraction layer, a multi-scale feature fusion network, and an object detection layer. The backbone feature extraction layer is used to input the N point cloud images and output a three-dimensional feature map. The multi-scale feature fusion network is used to input the three-dimensional feature map and output a three-dimensional aggregated feature map. The object detection layer is used to input the three-dimensional aggregated feature map and output at least one object detection result. The three-dimensional features include the number of feature channels, the height of the feature map, and the width of the feature map.

[0171] In a specific implementation, the target detection device described in the embodiments of the present application includes, but is not limited to, other portable devices such as mobile phones, laptop computers, or tablet computers having a touch-sensitive surface (e.g., a touch screen display and / or a touchpad). It should also be understood that in some embodiments, the above-mentioned target detection device is not a portable communication device, but a computer having a touch-sensitive surface (e.g., a touch screen display).

[0172] It should be understood that the target detection device may include one or more other physical user interface devices such as a physical keyboard, a mouse, and / or a joystick. Various application programs that can be executed on the target detection device may use at least one common physical user interface device such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and the corresponding information displayed on the terminal can be adjusted and / or changed between application programs and / or within the corresponding application programs. In this way, the common physical architecture of the terminal (e.g., the touch-sensitive surface) can support various application programs with a user interface that is intuitive and transparent to the user.

[0173] Figure 10 It is a schematic structural diagram of a terminal device provided by an embodiment of the present application. The terminal device 400 includes: at least one processor 401 ( Figure 10 only one is shown in the figure), a processor, a memory 402, and a computer program 403 stored in the memory 402 and executable on the at least one processor 401. When the processor 401 executes the computer program 403, the steps in the above-mentioned target detection method embodiment are implemented.

[0174] The terminal device 400 may be a computing device such as a desktop computer, a notebook, a palm computer, or a cloud server. The terminal device may include, but is not limited to, a processor 401 and a memory 402. Those skilled in the art can understand that Figure 10 merely examples of the terminal device 400, and do not constitute a limitation on the terminal device 400. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0175] The so-called processor 401 may be a Central Processing Unit (CPU), and this processor 401 may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0176] In some embodiments, the memory 402 may be an internal storage unit of the terminal device 400, such as the hard disk or memory of the terminal device 400. In other embodiments, the memory 402 may also be an external storage device of the terminal device 400, such as a plug-in hard disk equipped on the terminal device 400, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. Further, the memory 402 may also include both the internal storage unit of the terminal device 400 and the external storage device. The memory 402 is used to store an operating system, application programs, a BootLoader, data, and other programs, such as the program code of the computer program, etc. The memory 402 may also be used to temporarily store data that has been output or will be output.

[0177] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above-mentioned functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0178] In the above embodiments, the descriptions of the various embodiments each have their own emphases. For the parts that are not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0179] An embodiment of the present application provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0180] An embodiment of the present application provides a computer program product. When the computer program product runs on a mobile terminal, the mobile terminal can implement the steps in the above-mentioned various method embodiments when executed.

[0181] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0182] In the embodiments provided by the present application, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0183] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0184] In addition, the functional units in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above units can be implemented in the form of hardware or in the form of software.

[0185] When the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of this application, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0186] The above-described embodiments are only used to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A target detection method, characterized in that, The target detection method includes: Obtaining point cloud images of N frames of traffic sections to be detected, where the point cloud images are obtained by scanning with a lidar installed on the roadside of the traffic section to be detected; Inputting the N frames of point cloud images into a preset single-stage target detection model, and the single-stage target detection model outputs at least one target detection result; Wherein, the single-stage target detection model includes a backbone feature extraction layer, a multi-scale feature fusion network, and a target detection layer. The backbone feature extraction layer is used to input the N frames of point cloud images and output a three-dimensional feature map. The multi-scale feature fusion network is used to input the three-dimensional feature map and output a three-dimensional aggregated feature map. The target detection layer is used to input the three-dimensional aggregated feature map and output at least one target detection result. The three-dimensional features include the number of feature channels, the height of the feature map, and the width of the feature map.

2. The method according to claim 1, characterized in that, The multi-scale feature fusion network includes: Multiple convolutional layers. The input of the first convolutional layer is the three-dimensional feature map, and the output is a feature image of a set scale. The input of the next adjacent convolutional layer is the output of the previous adjacent convolutional layer, and at least two convolutional layers correspond to different set scales; Multiple pooling layers, which correspond one-to-one with the multiple convolutional layers. The input of the first pooling layer is the output of the first convolutional layer. The input of the next adjacent pooling layer is the output of the corresponding convolutional layer and the output of the previous adjacent pooling layer. The output of each pooling layer is a feature image of the same scale; A connection layer that combines the outputs of each pooling layer by increasing the number of channels to generate the three-dimensional aggregated feature map; The step of inputting the point cloud image into the preset single-stage target detection model, and the single-stage target detection model outputs at least one target detection result, includes: Inputting the point cloud image into the first convolutional layer to obtain the three-dimensional aggregated feature map.

3. The method according to claim 1, wherein The single-stage target detection model further includes: A region of interest data augmentation layer. The input of the region of interest data augmentation layer is the N frames of point cloud images, and the output is a point cloud image including pseudo samples, and the output is used as the input of the backbone feature extraction layer; Before inputting the point cloud image into the first convolutional layer, the step of inputting the point cloud image into the preset single-stage target detection model further includes: Inputting the N frames of point cloud images into the region of interest data augmentation layer to obtain a point cloud image including pseudo samples.

4. The method according to claim 3, characterized in that, Inputting the N frames of point cloud images into the region of interest data augmentation layer to obtain a point cloud image including pseudo samples, including: Taking the N frames of point cloud images as training samples, randomly selecting one frame of historical point cloud image for training, and setting a random number with a range including a first preset value and a second preset value; where the first preset value is greater than a preset K value; If the random number corresponding to the currently trained point cloud image is greater than the preset K value, randomly add the object point clouds in other frames of historical point cloud images to the currently trained point cloud image to generate a point cloud image including pseudo samples.

5. The method according to claim 2, wherein Before inputting the point cloud image into the first convolutional layer to obtain the three-dimensional aggregated feature map, inputting the point cloud image into a preset single-stage object detection model, where the single-stage object detection model outputs at least one object detection result, further comprising: Performing a global data augmentation operation on N frames of the point cloud image to obtain multiple preprocessed point cloud images, and replacing the N frames of the point cloud image with the set of the multiple preprocessed point cloud images and the N frames of the point cloud image as the input to the first convolutional layer.

6. The method according to claim 3, wherein The single-stage object detection model further includes a neighborhood information aggregation layer, and the neighborhood information aggregation layer includes multiple computing nodes. The input of each computing node is the output of all corresponding adjacent computing nodes, and the output is used as one of the inputs of all corresponding adjacent computing nodes; Before inputting the point cloud image into the first convolutional layer to obtain the three-dimensional aggregated feature map, inputting the point cloud image into a preset single-stage object detection model, where the single-stage object detection model outputs at least one object detection result, further comprising: Inputting N frames of the point cloud image into the neighborhood information aggregation layer, where the neighborhood information aggregation layer outputs N frames of three-dimensional feature maps containing neighborhood information and inputs them into the first convolutional layer.

7. The method according to claim 3, wherein Inputting N frames of the point cloud image into the region of interest data augmentation layer to obtain a point cloud image including pseudo-samples, including: Using N frames of point cloud images as training samples, randomly selecting one frame of historical point cloud image for training, and setting a random number with a range including a first preset value and a second preset value; where the first preset value is greater than a preset K value; If the random number corresponding to the currently trained point cloud image is greater than the preset K value, use the GT-sample algorithm to randomly add the current frame of object point cloud to other frames of point cloud images to generate a point cloud image including pseudo-samples.

8. A target detection device, characterized in that, The object detection device includes: An acquisition module that acquires N frames of point cloud images of a traffic section to be detected, where the point cloud images are obtained by scanning with a lidar installed on the roadside of the traffic section to be detected; A detection module that inputs the N frames of the point cloud image into a preset single-stage object detection model, and the single-stage object detection model outputs at least one object detection result; Wherein, the single-stage object detection model includes a pillar feature extraction layer, a multi-scale feature fusion network, and an object detection layer. The pillar feature extraction layer is used to input the N frames of the point cloud image and output a three-dimensional feature map. The multi-scale feature fusion network is used to input the three-dimensional feature map and output a three-dimensional aggregated feature map. The object detection layer is used to input the three-dimensional aggregated feature map and output at least one object detection result. The three-dimensional features include the number of feature channels, the height of the feature map, and the width of the feature map.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 7.