A traffic target detection method, device and electronic device

Through the combination of FMCW lidar and video camera and data fusion technology, a high-resolution traffic target detection method is generated, which solves the problem of insufficient detection accuracy of pedestrians and non-motor vehicles in the existing technology, and improves the small-target recognition capability of the autonomous driving system.

CN116129371BActive Publication Date: 2025-07-29ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310090468.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-09
Publication Date
2025-07-29
Estimated Expiration
2043-02-09

AI Technical Summary

Technical Problem

In the prior art, the traffic target detection method based on TOF lidar and video camera has low detection accuracy for pedestrians and non-motor vehicles, and is prone to missed and missed detection, affecting the safety of autonomous driving.

Method used

The combination of FMCW lidar and video camera is used to acquire and filter data with scattered echo intensity within a specific threshold range, 3D point cloud voxel feature vectors, SAR images and Doppler images are generated, and fuse them with RGB images. Convolutional neural networks are used for pedestrian and non-motor vehicle detection.

Benefits of technology

It improves the accuracy and recall of pedestrian and non-motor vehicle detection, reduces noise interference, and enhances the ability of the autonomous driving system to identify small targets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116129371B_ABST
    Figure CN116129371B_ABST
Patent Text Reader

Abstract

The present invention discloses a traffic target detection method, device and electronic device. The method includes: obtaining first perception data collected by an FMCW lidar and second perception data collected by a video camera; filtering data in the first perception data whose scattered echo intensity does not belong to the intensity threshold range, where the intensity threshold range represents the range to which the scattered echo intensity of pedestrians and / or non-motor vehicles belongs; performing structured processing on the filtered first perception data to obtain image processing structure data, where the image processing structure data includes 3D point cloud voxel feature vectors, SAR images and Doppler images; performing video preprocessing on the second perception data to obtain an RGB image synchronized with the first perception data; performing image fusion on the 3D point cloud voxel feature vectors, SAR images, Doppler images and RGB images, and performing pedestrian and / or non-motor vehicle detection based on the image fusion result, thereby improving the accuracy of pedestrian and / or non-motor vehicle detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent transportation, and particularly relates to a traffic target detection method, device and electronic device. Background Art

[0002] In intelligent transportation, multi-sensor fusion is commonly used for target detection, including fusing the data collected by lidar and video cameras and then performing target detection. In the prior art, whether at the vehicle end or the roadside, lidar uses the data collected by TOF (Time of Flight) lidar and video cameras as the input of the deep learning model to learn and identify targets.

[0003] The target detection method based on TOF lidar and video cameras has a relatively good detection effect on large target objects such as vehicles, but has poor detection effects on pedestrians, non-motor vehicles, etc., and often has problems such as missed detection and false detection. This is one of the important reasons restricting the entry of autonomous vehicles onto the road. In actual roads, pedestrians and non-motor vehicles are important participants, and their behaviors are sudden and variable. To improve the safety of autonomous driving, it is urgent to improve the accuracy of detecting road pedestrians and / or non-motor vehicles. Summary of the Invention

[0004] The present invention provides a traffic target detection method, device and electronic device, which are used to solve the technical problem of low accuracy in detecting road pedestrians and / or non-motor vehicles in the prior art.

[0005] In a first aspect, the present invention provides a traffic target detection method, which is applied to a roadside perception system. The roadside perception system includes at least one group of FMCW lidar and video cameras. The FMCW lidar and the video camera are arranged on the roadside and have overlapping sensing areas. The method includes:

[0006] Obtain first sensing data collected by the FMCW lidar and second sensing data collected by the video camera;

[0007] Filter data in the first sensing data whose scattered echo intensity does not belong to the intensity threshold range, where the intensity threshold range represents the range to which the scattered echo intensity of pedestrians and / or non-motor vehicles belongs;

[0008] Perform structured processing on the filtered first sensing data to obtain image processing structure data, where the image processing structure data includes 3D point cloud voxel feature vectors, SAR images, and Doppler images;

[0009] Perform video preprocessing on the second sensing data to obtain an RGB image synchronized with the first sensing data;

[0010] Perform image fusion on the 3D point cloud voxel feature vector, the SAR image, the Doppler image, and the RGB image, and perform pedestrian and / or non-motor vehicle detection based on the image fusion result.

[0011] Optionally, after performing video preprocessing on the second perception data to obtain an RGB image synchronized with the first perception data, the method further includes:

[0012] Obtain a target perception area where pedestrians and / or non-motor vehicles are located in the perception area of the video camera;

[0013] Extract target RGB data from the RGB image based on the target perception area, and update the RGB image based on the target RGB data.

[0014] Optionally, the Doppler image includes a time-Doppler spectrogram and a range-Doppler map.

[0015] Optionally, performing image fusion on the 3D point cloud voxel feature vector, the SAR image, the Doppler image, and the RGB image, and performing pedestrian and / or non-motor vehicle detection based on the image fusion result includes:

[0016] Perform image fusion on the 3D point cloud voxel feature vector, the SAR image, the Doppler image, and the RGB image to obtain a multi-source feature fusion feature map;

[0017] Perform secondary fusion on the multi-source feature fusion feature map and the original point cloud data collected by the FMCW lidar, and perform pedestrian and / or non-motor vehicle detection based on the data after secondary fusion.

[0018] Optionally, performing pedestrian and / or non-motor vehicle detection based on the image fusion result includes:

[0019] Input the multi-source feature fusion feature map, or the multi-source feature fusion feature map and the original point cloud data into a trained convolutional neural network for pedestrian and / or non-motor vehicle detection;

[0020] Wherein, the convolutional neural network includes three cascaded feature extraction layers, each feature extraction layer is composed of two convolutional layers and one pooling layer, and the convolutional kernel size of each convolutional layer is 3×3.

[0021] In a second aspect, the present invention provides a traffic target detection device, which is applied to a roadside perception system. The roadside perception system includes at least one group of FMCW lidars and video cameras. The FMCW lidars and the video cameras are arranged on the roadside and the perception areas overlap. The device includes:

[0022] An acquisition unit for acquiring first perception data collected by the FMCW lidar and second perception data collected by the video camera;

[0023] A filtering unit for filtering data in the first perception data whose scattered echo intensity does not belong to the intensity threshold range, where the intensity threshold range represents the range of scattered echo intensity of pedestrians and / or non-motor vehicles;

[0024] A structuring unit for performing structuring processing on the filtered first perception data to obtain image processing structure data, where the image processing structure data includes 3D point cloud voxel feature vectors, SAR images, and Doppler images;

[0025] A video processing unit for performing video preprocessing on the second perception data to obtain an RGB image synchronized with the first perception data;

[0026] A fusion detection unit for performing image fusion on the 3D point cloud voxel feature vectors, the SAR images, the Doppler images, and the RGB image, and performing pedestrian and / or non-motor vehicle detection based on the image fusion result.

[0027] Optionally, the device further includes: an extraction unit for, after performing video preprocessing on the second perception data to obtain an RGB image synchronized with the first perception data, acquiring a target perception area where pedestrians and / or non-motor vehicles are located in the perception area of the video camera; extracting target RGB data from the RGB image based on the target perception area, and updating the RGB image based on the target RGB data.

[0028] Optionally, the fusion detection unit is further configured to:

[0029] Perform image fusion on the 3D point cloud voxel feature vectors, the SAR images, the Doppler images, and the RGB image to obtain a multi-source feature fusion feature map; perform secondary fusion on the multi-source feature fusion feature map and the original point cloud data collected by the FMCW lidar, and perform pedestrian and / or non-motor vehicle detection based on the data after the secondary fusion.

[0030] In a third aspect, the present invention provides an electronic device, including a memory, and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors to implement any method described in the first aspect.

[0031] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements any method described in the first aspect.

[0032] One or more of the above technical solutions in the embodiments of the present application have at least the following technical effects:

[0033] A traffic target detection method provided by the present invention uses an FMCW lidar and a video camera as roadside perception devices. For pedestrians / or non-motor vehicles, 3D point cloud voxel feature vectors, SAR images, Doppler images, and RGB images are obtained based on the perception data of the FMCW lidar and the video camera. By the SAR image and the Doppler image, the resolution and micro-motion features of pedestrians / or non-motor vehicles are increased, so that the accuracy of data fusion and pedestrian and / or non-motor vehicle detection based on these four dimensions is greatly improved, and the technical problem of low accuracy of pedestrian and / or non-motor vehicle detection in the prior art is solved. At the same time, this method also filters the radar perception data through the scattered wave echo intensity, filtering out most of the noise other than pedestrians and / or non-motor vehicles, and further improving the accuracy of pedestrian and / or non-motor vehicle detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 It is a schematic diagram of a roadside perception system provided by an embodiment of the present application;

[0035] Figure 2 It is a flowchart of a traffic target detection provided by an embodiment of the present application;

[0036] Figure 3 It is a flowchart of image fusion processing provided by an embodiment of the present application;

[0037] Figure 4 It is a schematic diagram of the network structure of a convolutional neural network provided by an embodiment of the present application;

[0038] Figure 5 It is a schematic diagram of a traffic target detection device provided by an embodiment of the present application;

[0039] Figure 6 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] Before introducing the embodiments of the present disclosure, it should be noted that:

[0041] Some embodiments of the present disclosure are described as processing flows. Although the various operation steps of the flow may be labeled with sequential step numbers, the operation steps therein can be implemented in parallel, concurrently, or simultaneously.

[0042] The term "and / or" may be used in the embodiments of the present disclosure, and "and / or" includes any and all combinations of one or more of the listed related features.

[0043] It should be understood that when describing the connection relationship or communication relationship between two components, unless it is clearly specified that the two components are directly connected or directly communicate, otherwise, the connection or communication between the two components can be understood as direct connection or communication, or can be understood as indirect connection or communication through an intermediate component.

[0044] In order to make the technical solutions and advantages in the embodiments of the present disclosure clearer and more understandable, the exemplary embodiments of the present disclosure will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than an exhaustive list of all embodiments. It should be noted that, without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0045] Embodiment 1

[0046] Please refer to Figure 1 , this Embodiment 1 provides a roadside perception system, including: an FMCW (Frequency Modulated Continuous Wave) lidar, a video camera, a roadside computing module, and a perception result data communication module. The roadside computing module can be an edge computing device, and the perception result data communication module can be a device with data communication such as a 5G base station or an Internet of Things base station. The FMCW lidar and the video camera are arranged on a bracket on the roadside. There can be n groups of FMCW lidars and video cameras, n≥1. Each group of FMCW lidars and video cameras includes at least one FMCW lidar and at least one video camera, and their sensing areas completely overlap or partially overlap, that is, the data collected by each group of FMCW lidars and video cameras can be fused for detection.

[0047] For the FMCW lidar, without considering the target Doppler frequency shift, the beat frequency of the transmitted signal and the echo signal for a stationary target at the same radial distance is the same, and the beat frequencies are different for different radial distances. Moreover, if the target has a Doppler frequency shift, the beat signal will also be different. Therefore, the FMCW lidar is sensitive to the distance and speed parameters of the target. The inventor of the present invention found that small targets with dimensions smaller than a set threshold in the road, such as pedestrians and non-motor vehicles, have obvious micro-Doppler characteristics. However, compared with vehicles, pedestrians and / or non-motor vehicles on the road not only have variable positions but also have much lower resolution than vehicles, and the accuracy of their target detection is poor. Therefore, in this embodiment, the FMCW lidar is used to collect data in the sensing area and generate a SAR image to obtain small target information at high resolution, and then combined with the RGB information collected by the video camera to improve the accuracy and recall rate of pedestrian and / or non-motor vehicle detection.

[0048] Embodiment 2:

[0049] Based on the above roadside perception system, Embodiment 2 provides a traffic target detection method. Please refer to Figure 2 , the method includes:

[0050] S210. Obtain first perception data collected by an FMCW lidar and second perception data collected by a video camera;

[0051] S220. Perform structured processing on the first perception data to obtain image processing structure data, where the image processing structure data includes a 3D point cloud voxel feature vector, an SAR image, and a Doppler image;

[0052] S230. Perform video preprocessing on the second perception data to obtain an RGB image synchronized with the first perception data;

[0053] S240. Perform image fusion on the 3D point cloud voxel feature vector, the SAR image, the Doppler image, and the RGB image, and perform pedestrian and / or non-motor vehicle detection based on the image fusion result.

[0054] In the specific implementation process, the SAR imaging technology can use the relative motion between the radar and the target to form a large synthetic aperture, break through the limitation of the true aperture of the antenna, and achieve high-resolution imaging. The principle of the SAR imaging technology is as follows: Use a small antenna on the radar as a single radiation unit, continuously move this unit along a straight line, receive the echo signals of the same target object at different positions and process them, and then a higher-resolution image of the target object can be obtained. The above small antenna can synthesize an equivalent "large antenna" by moving. Therefore, the implementation condition of the SAR imaging technology is the relative motion between the radar and the target. In the application scenario of the roadside radar, the roadside radar is in a stationary state, and the pedestrians and vehicles on the road are in a moving state. Therefore, the SAR imaging technology can be used to obtain high-resolution images of vehicles and / or pedestrians.

[0055] When S220 processes the first perception data to obtain the SAR image, a series of data optimization processes such as Doppler parameter estimation, motion parameter estimation, motion compensation, and range compression can be performed on the first perception data first, and then filtering processing can be performed after the data optimization processing. Of course, filtering processing or synthetic aperture imaging can also be directly performed on the first perception data. Preferably, data accuracy and pertinence can be effectively improved through data optimization and filtering processing.

[0056] It should be understood that in the application scenario of roadside radar, the echo data obtained by the roadside radar includes the echo data of pedestrians / non-motor vehicles such as pedestrians and / or non-motor vehicles, and the echo data of guardrails, vehicles, etc. That is, the above first perception data not only includes the data of pedestrians and / or non-motor vehicles to be detected, but also mixes other target data other than the target object data. Different ground object targets, positions, ground object structures, surface morphologies, and dielectric properties have different scattered echo intensities for the radar beam. For example, bridges, transmission lines, houses, road edges, green belts, etc., their echo signals are very strong, and the echo signal intensities of the ground and road surface markings are relatively weak. Each type of target object has its own scattered echo intensity characteristics.

[0057] In order to form a SAR image that is more relevant to pedestrians and / or non-motor vehicles and has higher accuracy, the data after optimized processing is filtered. The data with scattered echo intensities not belonging to the intensity threshold range is deleted from the optimized data. The intensity threshold range represents the range to which the scattered echo intensity of pedestrians and / or non-motor vehicles belongs, that is. Through the filtering process, only the data with scattered echo intensities matching pedestrians and / or non-motor vehicles is retained. The synthetic aperture preprocessing algorithm is executed on the filtered data to obtain a SAR image containing pedestrians and / or non-motor vehicles. The above processing process is simple and easy to implement, effectively reducing the computing pressure of the data processing device and improving the efficiency of target object detection at the same time.

[0058] The synthetic aperture preprocessing algorithm decomposes the two-dimensional signal received by the radar receiver into one-dimensional signals cascaded in the range direction and the azimuth direction, and then performs signal processing on the signals in the two directions respectively. Among them, in the range direction, compression is performed through de-chirp processing. In the azimuth direction, due to the influence of range migration, the signals in the range and azimuth directions are coupled, and compression cannot be directly performed on the azimuth direction. It should be achieved by using the interpolation method before azimuth compression. Finally, the azimuth signal is focused and imaged through matched filtering, and the relevant SAR image is output.

[0059] When S220 processes the first perception data to obtain a Doppler image, two Doppler features are extracted for identifying pedestrians and / or non-motor vehicles. The Doppler features include the range-Doppler map and the time-Doppler spectrogram.

[0060] Range-Doppler map, which is used to describe the range and velocity of a target object and a radar in a radar data frame. The range-Doppler map can be obtained by successively performing windowing processing in the range dimension, FFT, windowing processing in the Doppler dimension, and FFT on the radar data frame. First, in order to reduce spectral leakage, windowing processing is performed on the range dimension of the FMCW radar data frame, and then FFT processing is performed on the range dimension, so that the range information of the target object can be obtained. Then, windowing processing is performed in the Doppler dimension, and FFT processing is performed in the Doppler dimension, so that the velocity information of the target object can be obtained. Finally, the range-Doppler map can be obtained.

[0061] Assume that the radar data frame is represented by the following formula:

[0062]

[0063] where, L represents the number of chirps in a radar data frame, K represents the number of sampling points of each chirp, A m represents the signal amplitude reflected by the target, f r represents the range frequency received from the target, f D represents the Doppler frequency caused by the radial velocity of the target, and j represents the imaginary unit.

[0064] The final range-Doppler map obtained through windowing processing in the range dimension, FFT, windowing processing in the Doppler dimension, and FFT can be expressed as:

[0065]

[0066] Time-Doppler spectrogram, which performs time-frequency analysis on the time-domain radar signal to obtain an image of the Doppler frequency of the echo signal changing with time. The short-time Fourier transform STFT is used:

[0067]

[0068] where, x[n] is the discrete-time signal, ω[n] is the window function, m is the sliding position of the window function, and ω is the angular frequency. The result of STFT is a distribution on the two-dimensional plane of time and frequency. Taking the square of the modulus of the result of STFT represents the power distribution of the input signal x[n] on the time and frequency plane, which is represented by a spectrogram.

[0069] Combined with experimental data, it is found that the time-Doppler spectral line pattern, that is, the spectrogram characteristics of the echo signal of pedestrians / non-motor vehicles, are different from those of motor vehicles. Therefore, if the time-Doppler spectrogram characteristics differences between pedestrians / non-motor vehicles and motor vehicles can be extracted from the Doppler images of the echo signals collected by the radar, then the type recognition of pedestrians / non-motor vehicles can be carried out based on this characteristics difference.

[0070] When S220 processes the first perception data to obtain the 3D point cloud voxel feature vector, the radar data can be grouped into different voxels and then the points within each voxel can be processed. Specifically, a network such as VoxelNet (Voxel Network), Voxel-FPN (Voxel-Feature Pyramid Network), etc. can be trained so that the network can encode the points within the voxel and finally form a feature vector within the voxel. The voxel-based method not only has better performance but also has a considerable calculation speed, which is beneficial for subsequent multi-source data fusion.

[0071] Converting the first perception data to obtain the voxel feature vector, SAR image, and Doppler image greatly increases the data dimension of small targets and also introduces a lot of noise. To suppress noise and improve data effectiveness, in this embodiment, before S220, the data with scattered echo intensity not belonging to the intensity threshold range in the first perception data can be filtered first, and then the filtered first perception data is structurally processed to obtain the voxel feature vector, SAR image, and Doppler image, so that the data dimension increases, the data volume decreases, and the data noise decreases, thereby greatly improving the efficiency and accuracy of traffic target detection.

[0072] Before, after, or at the same time as executing S220, execute S230 to perform video preprocessing on the second perception data to obtain an RGB image synchronized with the first perception data. The preprocessing includes image processing such as image extraction, filtering, image enhancement, and image difference, as well as time and space synchronization processing to ensure the data quality of the RGB image and its synchronization with the first perception data.

[0073] After performing video preprocessing on the second perception data to obtain an RGB image synchronized with the first perception data, the target perception area where pedestrians and / or non-motor vehicles are located in the perception area of the video camera can be further obtained; based on the target perception area, the target RGB data in the RGB image is extracted, and the RGB image is updated based on the target RGB data to extract the region of interest data and improve the subsequent image fusion efficiency. Among them, the target area can be delimited manually or constructed through traffic signs, such as taking the road area where lane lines, non-motor vehicles, and zebra crossings are marked as the target area. The target area can also be obtained through image recognition. The acquisition of the target RGB data can be based on making a target perception area mask for the target perception area, multiplying the target perception area mask by the RGB image to obtain a new RGB image. In this image, the image values within the target perception area remain unchanged, that is, the target RGB data is extracted, and the image values outside the area are all 0. The original RGB image is updated to the new RGB image.

[0074] Please refer to Figure 3, for the accuracy of subsequent data fusion, the following operations can also be performed after S230 in this embodiment:

[0075] S301. Determine whether there is data anomaly in the 3D point cloud voxel feature vector, RGB image, SAR image, and Doppler image. If so, perform frame-by-frame screening to obtain abnormal data frames for repair, discard, or resampling; if not, proceed to S302.

[0076] S302. Determine whether there is data imbalance in the above data. This step is used in the model training stage and does not need to be executed in the model detection stage. Data imbalance includes two types. One is the data imbalance of different targets, and the other is the data imbalance of the three-dimensional features of the same target. For the data imbalance of different targets, further measures are taken to balance the target types. For the data imbalance of feature data, data augmentation can be performed and then S303 is executed. For the case where the judgment result of S302 is yes, S303 can be directly executed.

[0077] S303. Data normalization. Through a series of transformations (i.e., using the invariant moments of the image to find a set of parameters to eliminate the influence of other transformation functions on the image transformation), the original image to be processed is converted into a corresponding unique standard form (the standard form image has invariant characteristics for affine transformations such as translation, rotation, and scaling).

[0078] After S303 or after S230, S240 is executed for image fusion and target recognition. Image fusion can output a multi-source feature fusion feature map. Based on the multi-source feature fusion feature map and the trained convolutional neural network, target recognition is performed to obtain the target object. Or, the multi-source feature fusion feature map is secondarily fused with the original radar point cloud data, incorporating corresponding distance, speed, and other information to generate more complete target object information as features, which is input into the subsequent convolutional neural network for target recognition to obtain the target object.

[0079] In this embodiment, a cascaded convolutional neural network is built based on the convolutional neural network using a stacked convolutional and pooling layer structure, and each convolutional layer uses a small-size convolutional kernel such as 3×3. The reason for using all small-size filters is that small-size filters are easier to extract the detailed features of the target image. As the forward calculation of the convolutional neural network progresses, the features extracted by the network are more detailed and more abstract. The essence of the convolutional layer is to map the features of the target image with a linear relationship. To fit more complex features, an activation function is introduced to non-linearly transform the features.

[0080] Please refer to Figure 4, the convolutional neural network for target recognition provided in this embodiment includes a convolutional layer, a pooling layer, and a fully connected layer. Among them, two convolutional layers and one pooling layer form a feature extraction layer. After three feature extraction layers are cascaded, they are propagated to the final layer of the network - the fully connected layer. The size of each convolutional layer is 32*3*3. The fully connected layer expands the features composed of each neuron in the hidden layer into a feature vector with a wide distribution for subsequent classification output. The fully connected layer can also be implemented by convolution. For a fully connected layer whose input layer is a fully connected layer, a filter can be used to perform linear convolution on the previous fully connected layer to map all target features. Figure 1 A link unfolds into a feature vector with a wide distribution for subsequent classification output. The fully connected layer can also be implemented by convolution. For a fully connected layer whose input layer is a fully connected layer, a filter can be used to perform linear convolution on the previous fully connected layer to map all target features.

[0081] For the training of the convolutional neural network with the above structure, 3D point cloud voxel feature vectors, SAR images, and Doppler images are used as model inputs, and pedestrian and / or non-motor vehicle object identifications are used as labels to construct training samples. Or, 3D point cloud voxel feature vectors, SAR images, Doppler images, and original radar point clouds are used as model inputs, and pedestrian and / or non-motor vehicle object identifications are used as labels to construct training samples. The network is trained based on a large number of samples until the network converges.

[0082] In the above embodiment, an FMCW lidar and a video camera are set on the roadside as roadside perception devices. For pedestrians / or non-motor vehicles, 3D point cloud voxel feature vectors, SAR images, Doppler images, and RGB images are obtained based on the perception data of the FMCW lidar and the video camera. By using SAR images and Doppler images, the resolution and micro-action features of pedestrians / or non-motor vehicles are increased, greatly improving the accuracy of data fusion and pedestrian and / or non-motor vehicle detection based on these four dimensions, and solving the technical problem of low accuracy of pedestrian and / or non-motor vehicle detection in the prior art on roads. At the same time, this method also filters the radar perception data through the scattered wave echo intensity, filtering out most of the noise other than pedestrians and / or non-motor vehicles, further improving the accuracy of pedestrian and / or non-motor vehicle detection.

[0083] Embodiment 3:

[0084] Based on Figure 2 A traffic target detection method provided, this embodiment 3 also correspondingly provides a traffic target detection device, which is applied to a roadside perception system. The roadside perception system includes at least one set of FMCW lidar and a video camera. The FMCW lidar and the video camera are set on the roadside and their perception areas overlap. Please refer to Figure 5 , the device includes:

[0085] An acquisition unit 51, configured to acquire first perception data acquired by the FMCW lidar and second perception data acquired by the video camera;

[0086] A filtering unit 52, configured to filter data in the first sensed data whose scattered echo intensity does not belong to the intensity threshold range, where the intensity threshold range represents the range to which the scattered echo intensity of pedestrians and / or non-motor vehicles belongs;

[0087] A structuring unit 53, configured to perform structuring processing on the filtered first sensed data to obtain image processing structure data, where the image processing structure data includes 3D point cloud voxel feature vectors, SAR images, and Doppler images;

[0088] A video processing unit 54, configured to perform video preprocessing on the second sensed data to obtain an RGB image synchronized with the first sensed data;

[0089] A fusion detection unit 55, configured to perform image fusion on the 3D point cloud voxel feature vectors, the SAR images, the Doppler images, and the RGB image, and perform pedestrian and / or non-motor vehicle detection based on the image fusion result.

[0090] As an optional implementation manner, the apparatus further includes:

[0091] An extraction unit 56, configured to, after performing video preprocessing on the second sensed data to obtain an RGB image synchronized with the first sensed data, acquire a target sensing area where pedestrians and / or non-motor vehicles are located in the sensing area of the video camera;

[0092] Extract target RGB data from the RGB image based on the target sensing area, and update the RGB image based on the target RGB data.

[0093] As an optional implementation manner, the fusion detection unit 55 is further configured to:

[0094] Perform image fusion on the 3D point cloud voxel feature vectors, the SAR images, the Doppler images, and the RGB image to obtain a multi-source feature fusion feature map;

[0095] Perform secondary fusion on the multi-source feature fusion feature map and the original point cloud data collected by the FMCW lidar, and perform pedestrian and / or non-motor vehicle detection based on the data after the secondary fusion.

[0096] As an optional implementation manner, the fusion detection unit is further configured to:

[0097] Input the multi-source feature fusion feature map, or the multi-source feature fusion feature map and the original point cloud data, into a trained convolutional neural network for pedestrian and / or non-motor vehicle detection;

[0098] Among them, the convolutional neural network includes three cascaded feature extraction layers, each feature extraction layer is composed of two convolutional layers and a pooling layer, and the convolutional kernel size of each convolutional layer is 3×3.

[0099] Regarding the device in the above embodiment, the specific manner in which each unit performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.

[0100] Embodiment 4:

[0101] Figure 6 It is a block diagram of an electronic device 600 for implementing a traffic target detection method shown according to an exemplary embodiment. For example, the electronic device 600 can be an industrial control computer, a computer, an edge server, an edge computing device, etc.

[0102] Referring to Figure 6 , the electronic device 600 may include one or more of the following components: a processing component 602, a memory 604, a power supply component 606, an input / presentation (I / O) interface 608, and a communication component 610.

[0103] The processing component 602 generally controls the overall operation of the electronic device 600, such as operations associated with data calculation, control, instruction issuance, and camera triggering. The processing element 602 may include one or more processors 620 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 602 may include one or more modules to facilitate the interaction between the processing component 602 and other components.

[0104] The memory 604 is configured to store various types of data to support the operation of the device 600. Examples of these data include instructions for any application or method operating on the electronic device 600, image data, associated data, configuration data, etc. The memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0105] The power supply component 606 provides power for various components of the electronic device 600. The power supply component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 600.

[0106] The communication component 610 is configured to facilitate communication between the electronic device 600 and other devices in a wired or wireless manner. The electronic device 600 can access a communication standard-based wireless network, such as WiFi, 4G, or 5G, or a combination thereof. In an exemplary embodiment, the communication component 410 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 410 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0107] In an exemplary embodiment, the electronic device 600 can be implemented by one or more Application Specific Integrated Circuits

[0108] (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0109] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, and the above instructions can be executed by a processor 620 of the electronic device 600 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc. When the instructions in the non-transitory computer-readable storage medium are executed by the processor 620 of the electronic device 600, the point cloud data processing method in the above embodiments can be implemented.

[0110] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention, which follow the general principles of the present invention and include known common knowledge or conventional technical means in the technical field not disclosed in this embodiment. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.

[0111] It should be understood that the present invention is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims. The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A traffic target detection method, characterized in that, The method is applied to a roadside perception system, which includes at least one set of FMCW lidar and a video camera. The FMCW lidar and the video camera are set on the roadside and their sensing areas overlap. The method includes: Obtain the first sensing data collected by the FMCW lidar and the second sensing data collected by the video camera; Filter the data in the first sensing data whose scattered echo intensity does not belong to the intensity threshold range, where the intensity threshold range represents the range of scattered echo intensity of pedestrians and / or non-motor vehicles; Perform structured processing on the filtered first sensing data to obtain image processing structure data. The image processing structure data includes 3D point cloud voxel feature vectors, SAR images, and Doppler images; Perform video preprocessing on the second sensing data to obtain an RGB image synchronized with the first sensing data; Perform image fusion on the 3D point cloud voxel feature vectors, SAR images, Doppler images, and RGB images, and perform pedestrian and / or non-motor vehicle detection based on the image fusion result; After performing video preprocessing on the second sensing data to obtain an RGB image synchronized with the first sensing data, the method further includes: Obtain the target sensing area where pedestrians and / or non-motor vehicles are located in the sensing area of the video camera; Extract the target RGB data in the RGB image based on the target sensing area, and update the RGB image based on the target RGB data; The Doppler image includes a time-Doppler spectrogram and a range-Doppler map; Performing image fusion on the 3D point cloud voxel feature vectors, SAR images, Doppler images, and RGB images, and performing pedestrian and / or non-motor vehicle detection based on the image fusion result includes: Perform image fusion on the 3D point cloud voxel feature vectors, the SAR image, the Doppler image, and the RGB image to obtain a multi-source feature fusion feature map; Perform secondary fusion of the multi-source feature fusion feature map with the original point cloud data collected by the FMCW lidar, and perform pedestrian and / or non-motor vehicle detection based on the data after the secondary fusion; The pedestrian and / or non-motor vehicle detection based on the image fusion result includes: Input the multi-source feature fusion feature map, or the data after the secondary fusion of the multi-source feature fusion feature map and the original point cloud data, into a trained convolutional neural network for pedestrian and / or non-motor vehicle detection; The convolutional neural network includes three cascaded feature extraction layers, each of which is composed of two convolutional layers and one pooling layer, and the convolutional kernel size of each convolutional layer is 3×3.

2. A traffic target detection device, characterized in that, The device is applied to a roadside perception system, which includes at least one set of FMCW lidar and a video camera. The FMCW lidar and the video camera are set on the roadside and their sensing areas overlap. The device includes: An acquisition unit for obtaining the first sensing data collected by the FMCW lidar and the second sensing data collected by the video camera; A filtering unit for filtering data in the first sensing data whose scattered echo intensity does not belong to the intensity threshold range, where the intensity threshold range represents the range of scattered echo intensity of pedestrians and / or non-motor vehicles; A structuring unit for performing structuring processing on the filtered first sensing data to obtain image processing structure data, where the image processing structure data includes 3D point cloud voxel feature vectors, SAR images, and Doppler images; A video processing unit for performing video preprocessing on the second sensing data to obtain an RGB image synchronized with the first sensing data; A fusion detection unit for performing image fusion on the 3D point cloud voxel feature vectors, the SAR images, the Doppler images, and the RGB image, and performing pedestrian and / or non-motor vehicle detection based on the image fusion result.

3. The traffic target detection device according to claim 2, characterized in that, The apparatus further includes: An extraction unit for obtaining a target sensing area where pedestrians and / or non-motor vehicles are located in the sensing area of the video camera after performing video preprocessing on the second sensing data to obtain an RGB image synchronized with the first sensing data; Extracting target RGB data from the RGB image based on the target sensing area, and updating the RGB image based on the target RGB data.

4. A traffic target detection device according to any one of claims 3, characterized in that, The fusion detection unit is further configured to: Perform image fusion on the 3D point cloud voxel feature vectors, the SAR images, the Doppler images, and the RGB image to obtain a multi-source feature fusion feature map; Perform secondary fusion on the multi-source feature fusion feature map and the original point cloud data collected by the FMCW lidar, and perform pedestrian and / or non-motor vehicle detection based on the data after the secondary fusion.

5. An electronic device, characterized in that, It includes a memory and one or more programs, where one or more programs are stored in the memory and are configured to be executed by one or more processors to implement the method according to claim 1.

6. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the program is executed by a processor, the steps of the method according to claim 1 are implemented.

Citation Information

Patent Citations

  • Vision-laser radar fusion method and system based on depth canonical correlation analysis

    CN113111974A

  • Data processing method, device and equipment and readable storage medium

    CN114529490A