Target detection method and device

By fusing feature vectors from radar point cloud data and image data in target detection, a comprehensive feature vector is generated and input into the target detection model, thus solving the problem of fusion between radar point cloud data and image data and improving the accuracy and robustness of target detection.

CN121010749APending Publication Date: 2025-11-25BEIJING VOYAGER TECH CO LTD

Patent Information

Application Number
CN202410658361.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-24
Publication Date
2025-11-25

AI Technical Summary

Technical Problem

In existing target detection methods based on multimodal data, effectively fusing radar point cloud data and image data to improve the accuracy of target detection is a challenge.

Method used

By acquiring radar point cloud data and image data of the target area, a priori information feature vectors are generated using a radar detection model. Based on the mapping relationship between radar point cloud data and image data, image feature vectors are determined. After feature fusion, the data are input into the target detection model to optimize the detection results.

Benefits of technology

It improves the accuracy and robustness of target detection, enhances the alignment effect of data from different modalities, and further improves the accuracy of target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121010749A_ABST
    Figure CN121010749A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a target detection method and device. According to the embodiment of the invention, after the radar point cloud data and the image data of the target area are obtained, the radar point cloud data are input into the radar detection model to obtain the radar detection result, then the priori information feature vector is determined according to the radar detection result, and the image feature vector is determined according to the mapping relation between the radar point cloud data and the image data. And fusing the two vectors to obtain a comprehensive feature vector, and finally inputting the comprehensive feature vector into a target detection model to obtain a target detection result. According to the embodiment of the invention, the radar detection result is used as the prior information, the prior information and the image data are subjected to feature extraction and fusion, and the fused feature information is input into the target detection model, so that the target area is subjected to secondary detection, and the detection accuracy is improved. In addition, the image feature vector is determined according to the mapping relation between the radar point cloud data and the image data, so that the alignment effect of the data of the two different modalities is higher.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of target detection technology, and more specifically, to a target detection method and apparatus. Background Technology

[0002] Object detection plays a crucial role in popular fields such as autonomous driving, augmented reality, and image analysis. Current object detection methods can be mainly categorized into those based on radar point cloud data, those based on image data, and those based on multimodal data. Among these, object detection methods based on multimodal data combine the advantages of radar point cloud data and image data, and offer higher accuracy compared to object detection methods based on single-modal data.

[0003] Current target detection methods based on multimodal data generate features from radar point cloud data and image data separately, and then fuse them for target detection. However, as radar point cloud data and image data are two heterogeneous types of data, how to fuse and utilize them to produce better results is a key focus driving the current development of target detection. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide a target detection method and apparatus to improve the target detection accuracy and optimize the target detection results.

[0005] Firstly, a target detection method is provided, the method comprising:

[0006] Acquire radar point cloud data and image data of the target area;

[0007] The radar point cloud data is input into the radar detection model to obtain the corresponding radar detection results;

[0008] The corresponding prior information feature vector is determined based on the radar detection results;

[0009] The corresponding image feature vector is determined based on the mapping relationship between the radar point cloud data and the image data;

[0010] The prior information feature vector and the image feature vector are fused to obtain the corresponding comprehensive feature vector;

[0011] The comprehensive feature vector is input into the target detection model to obtain the target detection result.

[0012] Secondly, a target detection device is provided, the device comprising:

[0013] The acquisition module is used to acquire radar point cloud data and image data of the target area;

[0014] The radar detection module is used to input the radar point cloud data into the radar detection model to obtain the corresponding radar detection results.

[0015] The prior information module is used to determine the corresponding prior information feature vector based on the radar detection result;

[0016] The determination module is used to determine the corresponding image feature vector based on the mapping relationship between the radar point cloud data and the image data;

[0017] The fusion module is used to fuse the prior information feature vector and the image feature vector to obtain the corresponding comprehensive feature vector;

[0018] The target detection module is used to input the comprehensive feature vector into the target detection model to obtain the target detection result.

[0019] Thirdly, an electronic device is provided, including a memory and a processor, the memory being used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in the first aspect above.

[0020] Fourthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed by a processor, the computer program implements the method described in the first aspect.

[0021] Fifthly, a computer program product is provided, comprising a computer program / instructions that, when executed by a processor, implement the method described in the first aspect.

[0022] This invention discloses a target detection method and apparatus. After acquiring radar point cloud data and image data of a target area, the radar point cloud data is input into a radar detection model to obtain the corresponding radar detection result. Then, based on the radar detection result, a corresponding prior information feature vector is determined. Based on the mapping relationship between the radar point cloud data and the image data, a corresponding image feature vector is determined. The prior information feature vector and the image feature vector are then fused to obtain a corresponding comprehensive feature vector. Finally, the comprehensive feature vector is input into the target detection model to obtain the target detection result. This invention uses the radar detection result output by the radar detection model as prior information, extracts and fuses features from the prior information and image data, and inputs the fused feature information into the target detection model, thereby performing secondary target detection on the corresponding data of the target area, thus improving the target detection accuracy and optimizing the detection result. Furthermore, determining the image feature vector based on the mapping relationship between the radar point cloud data and the image data can improve the alignment effect of the two different modalities of data, further improving the target detection accuracy. Attached Figure Description

[0023] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:

[0024] Figure 1 This is a flowchart of the target detection method according to an embodiment of the present invention;

[0025] Figure 2 This is a schematic diagram of the detection box corresponding to the target area in an embodiment of the present invention;

[0026] Figure 3 This is a flowchart of the image feature vector determination method according to an embodiment of the present invention;

[0027] Figure 4 This is a flowchart of a method for determining a sub-image corresponding to a radar-detected target according to an embodiment of the present invention;

[0028] Figure 5 This is a flowchart of a method for determining a sub-image corresponding to a radar-detected target according to an embodiment of the present invention;

[0029] Figure 6 This is a data flow diagram of the target detection method according to an embodiment of the present invention;

[0030] Figure 7 This is a schematic diagram of the target detection device according to an embodiment of the present invention;

[0031] Figure 8 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0032] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.

[0033] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.

[0034] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".

[0035] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0036] The solutions described in this specification and embodiments, if involving the processing of personal information, will be processed only under the premise of having a legal basis (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be processed within the scope stipulated or agreed upon. A user's refusal to process personal information beyond what is necessary for basic functions will not affect the user's use of basic functions.

[0037] Figure 1 This is a flowchart of a target detection method according to an embodiment of the present invention. Figure 1 As shown, the target detection method includes the following steps:

[0038] Step S101: Acquire radar point cloud data and image data of the target area.

[0039] The target area can be any region, and the radar point cloud data and image data of the target area refer to the radar point cloud data and image data corresponding to the same region. Radar (Light Detection and Ranging, LiDAR) is an active sensor that acquires information from the outside world. It can quickly acquire three-dimensional point cloud information of the surrounding environment and has advantages such as wide detection range, independence from external lighting, and high-precision ranging. Using radar point cloud data for target detection can accurately obtain the target's distance, contour, and other feature attributes. However, radar point cloud data contains relatively little information, which can easily lead to false detection or misclassification. Compared to radar point cloud data, image data contains more information, such as the shape or color of the target. However, when using image data for target detection, high-quality, high-resolution original images are usually required. Furthermore, the inherent limitations of image data devices mean that factors such as changes in lighting, target occlusion, and shadows can severely affect the quality of the acquired images. For example, severe weather conditions such as strong winds, rain, and snow, or changes in lighting, occlusion, and shadows can greatly reduce the reliability of target detection results. Therefore, combining radar point cloud data and image data can be used for target detection to obtain richer and clearer detection data, thereby enhancing the accuracy and robustness of target detection.

[0040] In one possible implementation, radar point cloud data and image data can be acquired through corresponding sensors. Specifically, radar point cloud data is acquired by scanning with a radar sensor, and image data is acquired using a visual sensor (i.e., a photosensitive element).

[0041] In one possible implementation, radar point cloud data and image data can also be acquired via network requests.

[0042] It is worth noting that the above acquisition methods can also be used in combination. For example, after obtaining image data of a certain area through a network request, a radar sensor can be used to scan that area to obtain the corresponding radar point cloud data.

[0043] Step S102: Input the radar point cloud data into the radar detection model to obtain the corresponding radar detection results.

[0044] The radar detection model is used to determine the corresponding radar detection results based on the input radar point cloud data. The radar detection results include multiple radar detection targets and corresponding detection information. The detection information may include one or more of the following: the detection category of the radar detection target, the point cloud information within the corresponding detection box of the radar detection target, and the position, orientation, and shape information of the corresponding detection box of the radar detection target. The radar detection target is also known as 3D labeled data, typically displayed in the form of a 3D labeled box (i.e., a detection box). The radar coordinates of the 3D labeled box represent its position in the radar coordinate system. For example, the radar coordinates can be (x, y, z, w, h, l), where x, y, and z represent the position of the center point of the 3D labeled box in the radar coordinate system, and w, h, and l represent the length, width, and height of the 3D labeled box, respectively. The radar point cloud information contained within the 3D labeled box may include one or more of the following: radar point cloud echo count, intensity, RGB, GPS time, scanning angle, and scanning direction. The detection category can be one or more of the preset categories. The preset categories are set according to the actual application scenario. For example, if the actual application scenario is autonomous driving, the corresponding preset categories may include vehicles, electric vehicles, bicycles, pedestrians and road obstacles.

[0045] In one possible implementation, the radar detection model processes radar point cloud data based on a grid-based structure, and then learns features through a 2D convolutional backbone network. The radar coordinates of the detection box corresponding to the radar detection target, the point cloud within the detection box, the detection category, and the confidence level corresponding to the detection category are determined through a neck and a multitaskhead.

[0046] Optionally, radar point cloud data can be extracted using grid-based feature extraction to obtain single-frame coarse features from a bird's-eye view (BEV). The coarse features are a set of features with low feature dimension, such as the acquisition time and reflection intensity of the radar point cloud. The feature extraction method can be selected as needed, such as Pointpillars, PointNet++, VoxelNet, Sparse Convolutional Networks, or RandLA-Net.

[0047] Figure 2 This is a schematic diagram of the detection box corresponding to the target area in an embodiment of the present invention. Figure 2 As shown, the output of the radar detection model includes 3D bounding boxes (i.e., detection boxes) used to characterize the presence of multiple radar-detected targets in the target area.

[0048] Step S103: Determine the corresponding prior information feature vector based on the radar detection results.

[0049] Specifically, the corresponding feature vector can be determined based on the information contained in the radar detection results, and this vector can be used as the prior information feature vector. One or more pieces of information can be selected from the various types of information contained in the radar detection results as prior information, thereby determining one or more prior information feature vectors.

[0050] Optionally, the detection category of each radar-detected target in the radar detection result can be used as prior information. In this case, step S103 can specifically be: determining the category feature vector corresponding to each radar-detected target based on the detection category of each radar-detected target.

[0051] Specifically, a category feature vector is determined based on the preset categories of the radar detection model and the detection categories of the radar-detected targets. For example, the predicted categories that the radar detection model can output are preset to include cars, pedestrians, bicycles, green belts, and roadblocks. If the detection category of the radar-detected target is a car, and is specified as 0.88, then a one-dimensional vector [0.88, 0, 0, 0, 0] can be generated. If the detection category of the radar-detected target is a roadblock, and is specified as 0.7, then a one-dimensional vector [0, 0, 0, 0, 0.7] can be generated. The length of the one-dimensional vector is determined according to the number of detection categories. Then, the one-dimensional vector is determined as the category feature vector. It should be understood that the category feature vector in this embodiment can use any discrete data encoding method to encode the detection categories of each radar-detected target to obtain the category feature vector, such as one-hot encoding, label encoding, ordinal encoding, dictionary encoding, mapping encoding, etc. This embodiment does not limit this.

[0052] Optionally, the point cloud information within the detection box corresponding to each radar detection target in the radar detection result can also be used as prior information. In this case, step S103 can specifically be: determining the point cloud feature vector corresponding to the radar detection target based on the point cloud information within the detection box corresponding to each radar detection target.

[0053] Specifically, one or more of the point cloud information within the detection frame are used to determine the point cloud feature vector. For example, the point cloud information may include the number of point clouds and the point cloud intensity. Based on the number of point clouds and the point cloud intensity, the point cloud feature vector corresponding to the radar-detected target can be determined. Furthermore, the corresponding vector can be determined by combining information such as the number of echoes, intensity, RGB values, GPS time, scanning angle, and / or scanning direction of the point cloud. Optionally, this embodiment can use a corresponding encoding method to encode various types of point cloud information and then concatenate them to obtain the corresponding point cloud feature vector. It should be understood that this embodiment can use any of the above-mentioned suitable encoding methods for the corresponding information, which will not be elaborated further here.

[0054] Optionally, the orientation information of each radar-detected target in the radar detection results can also be used as prior information. In this case, step S103 can specifically be: determining the orientation feature vector corresponding to each radar-detected target based on the orientation information of each radar-detected target.

[0055] Optionally, the shape information of each radar-detected target in the radar detection results can also be used as prior information. In this case, step S103 can specifically be: determining the shape feature vector corresponding to each radar-detected target based on the shape information of each radar-detected target.

[0056] In addition, prior information may also include topological relationships or texture information between radar-detected targets. The determination of the corresponding feature vectors is similar to the methods described above, and will not be repeated here.

[0057] Step S104: Determine the corresponding image feature vector based on the mapping relationship between the radar point cloud data and the image data.

[0058] Specifically, first, appropriate image data is selected, and then the image feature vector is determined based on the selected image data.

[0059] In one possible implementation, the entire radar point cloud data can be matched with the corresponding entire image data to determine the corresponding image feature vector. For example, if a set of radar point cloud data can be matched with multiple sets of image data, then the radar point cloud clustering results of the radar point cloud data and the image pixel clustering results of each set of image data can be matched, and then the image data with the highest matching degree can be selected to determine the image feature vector.

[0060] In one possible implementation, the image data can first be cropped into multiple sub-images. Then, a corresponding sub-image is selected for each radar-detected target, and the corresponding image feature vector is determined based on the sub-image. Specifically, from a BEV perspective, outputting the detection results of all targets in the target region based on the image of the entire target region is not as effective as detecting a single target based on its image. Therefore, the radar point cloud data can be combined to crop the image corresponding to the target region to obtain a sub-image corresponding to a single target. Then, image feature extraction can be performed based on the sub-image corresponding to the single target to improve target detection accuracy.

[0061] Figure 3 This is a flowchart illustrating the image feature vector determination method according to an embodiment of the present invention. Figure 3 As shown, the image feature vector determination method includes the following steps:

[0062] Step S301: Based on the mapping relationship between the radar point cloud data and the image data, and the detection box position information of each radar-detected target, determine the sub-image of each radar-detected target from the image data.

[0063] Figure 4 This is a flowchart illustrating a method for determining a sub-image corresponding to a radar-detected target according to an embodiment of the present invention. Figure 4 As shown, the method for determining the sub-image corresponding to the radar-detected target includes the following steps:

[0064] Step S401: Determine the coordinate transformation relationship between the radar point cloud data and the image data.

[0065] The coordinate transformation relationship is used to characterize the transformation relationship between the radar coordinate system and the image coordinate system.

[0066] In one possible implementation, the coordinate transformation relationship between the radar coordinate system and the world coordinate system (or Cartesian coordinate system) and the coordinate transformation relationship between the image coordinate system and the world coordinate system (or Cartesian coordinate system) can be determined separately. Then, the coordinates of the radar point cloud are first converted to coordinates in the world coordinate system or Cartesian coordinate system, and then the coordinates in the world coordinate system or Cartesian coordinate system are converted to coordinates in the image coordinate system.

[0067] In one possible implementation, the coordinate transformation matrix between the radar coordinate system and the world coordinate system (or the Cartesian coordinate system) can also be determined directly.

[0068] Step S402: Based on the coordinate transformation relationship, the radar coordinates of the detection box corresponding to the radar detected target are converted into image coordinates.

[0069] Step S403: Based on the image coordinates of the detection box corresponding to the radar-detected target, cut out the corresponding sub-image from the image data.

[0070] In the above method, the radar coordinates of the detection box are converted into image coordinates through the coordinate transformation relationship between the radar coordinate system and the image coordinate system. Then, the entire image of the target area is cropped according to the image coordinates to obtain sub-images of each radar-detected target in the target area.

[0071] In addition to mapping the detection boxes output by the radar detection model to the image based on the transformation relationship between coordinate systems, the image data can be semantically segmented first. Then, the semantic segmentation results of the image can be matched with the detection category of the radar target to determine the sub-image corresponding to each radar target. To improve the matching accuracy, the position information of the radar target (i.e., the coordinates of the detection box corresponding to the radar target) and the image position of each semantic segmentation image can be analyzed during the matching process.

[0072] Figure 5 This is a flowchart illustrating a method for determining a sub-image corresponding to a radar-detected target according to an embodiment of the present invention. Figure 5 As shown, the method for determining the sub-image corresponding to the radar-detected target includes the following steps:

[0073] Step S501: Perform semantic segmentation on the image data.

[0074] Step S502: Based on the semantic segmentation result, the image data is cut into multiple sub-images.

[0075] Specifically, semantic segmentation is performed on the image data corresponding to the target region. Semantic segmentation methods can include pixel-level thresholding methods, pixel-based clustering segmentation methods, graph partitioning segmentation methods, and deep learning. Taking the K-Means clustering-based image segmentation method as an example: K cluster centers are initialized in the image pixels, where K is any pre-defined positive integer. Then, for each pixel in the pixel dataset corresponding to the image, the distance between it and each cluster center is calculated, and the set corresponding to the nearest cluster center is selected as the set of pixels for that pixel. Then, for the updated set, the average value of each pixel in the set is recalculated as the new cluster center. This process is repeated until convergence, meaning the cluster centers no longer change. Finally, the clustering results are mapped onto the image, thereby achieving image segmentation and determining the semantic category of the segmented sub-images.

[0076] Step S503: Match the detection category of each radar-detected target with the semantic segmentation category of each sub-image to determine the sub-image corresponding to each radar-detected target.

[0077] Specifically, radar point cloud data and image data have different modal characteristics, leading to certain differences in target classification. Because radar point cloud data contains less information than image data, target classification based on radar point cloud data is broader, such as detecting cars, buildings, and pedestrians. Target classification based on image data is more refined, further specifying "car" as sedans, vans, trucks, buses, etc. Therefore, when matching categories obtained from these two types of data, attention should be paid to the semantic compatibility of the corresponding categories.

[0078] Furthermore, when the target area is large, there are many radar-detected targets in the target area, and it is difficult to accurately match based solely on category semantics, location information can be combined for filtering and matching. For example, if there are multiple vehicles in the current image, it is difficult to match them directly based solely on category. In this case, further matching can be performed based on the coordinates of the detection box corresponding to the radar-detected target and the coordinates of the sub-image to obtain the matching result.

[0079] Through the above Figure 4 or Figure 5 Once the sub-image is determined using the method shown, the image feature vector corresponding to the radar-detected target can be determined based on the sub-image.

[0080] Step S302: Determine the image feature vector of each radar-detected target based on the sub-image of each radar-detected target.

[0081] Methods for determining image feature vectors can be categorized into manual feature extraction and deep learning feature extraction. Manual feature extraction can utilize Scale-invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), Histogram of Oriented Gradients (HOG), or Oriented Fast and Rotated BRIEF, among others. Deep learning feature extraction involves using pre-trained deep convolutional neural networks (such as VGG, ResNet, and Inception) to extract image features; the output of the last fully connected layer is considered the image's feature vector.

[0082] Step S105: The prior information feature vector and the image feature vector are fused to obtain the corresponding comprehensive feature vector.

[0083] Before feature fusion, the prior information feature vectors and image feature vectors should meet certain dimensionality constraints to enable feature vector fusion. It should be understood that this embodiment can convert various prior information feature vectors and image feature vectors into a format suitable for fusion. The fusion method can be concatenation, matrix addition, matrix multiplication, etc. This embodiment does not limit the feature vector conversion process or the feature fusion method.

[0084] Based on the above, it can be seen that the prior information feature vector can be one or more. The following uses image feature vector and category feature vector as examples of prior information feature vectors for illustration. It is easy to understand that the prior information feature vector can be any one or more of the prior information feature vector examples mentioned above.

[0085] In one possible implementation, the category feature vector and point cloud feature vector are adjusted based on the dimension corresponding to the image feature vector so that they can be fused with the image feature vector.

[0086] For example, the method for adjusting the category feature vector can be as follows: determine a transformation matrix based on the image feature vector and the category feature vector, and then transform the category feature vector into a category feature vector to be fused based on the transformation matrix. For example, the image feature vector is: The category feature vector is [0, 0, 0, 0, 0.7], meaning the image feature vector has a dimension of (2, 3), and the category feature vector has a dimension of (1, 5). If the fusion method is matrix multiplication, to make the category feature vector multiplied by the image feature vector matrix, the transformation matrix can be set to a dimension of (3, 1). The category feature vector to be fused after multiplying the transformation matrix and the category feature vector matrix has a dimension of (3, 5). In this case, the category feature vector can be fused with the image feature vector through matrix multiplication. If the fusion method is matrix addition, to make the category feature vector addable to the image feature vector matrix, the transformation matrix can include multiple matrices. For example, the first transformation matrix has a dimension of (5, 3), and the second transformation matrix has a dimension of (2, 1). Multiplying the category feature vector by the first transformation matrix yields a matrix with a dimension of (1, 3). Then, multiplying the second transformation matrix by the matrix with a dimension of (1, 3) yields a matrix with a dimension of (2, 3) of the category feature vector to be fused. In this case, the category feature vector can be fused with the image feature vector through matrix addition. If the fusion method is vertical stitching, the dimension of the transformation matrix can be determined as (5, 3), and the dimension of the category feature vector to be fused obtained by multiplying the transformation matrix and the category feature vector matrix is ​​(1, 3). In this case, the category feature vector can be fused with the image feature vector in a vertical stitching manner. Alternatively, the fusion method can be horizontal stitching, in which case the dimension of the transformation matrix can be determined as (2, 1), and the dimension of the category feature vector to be fused obtained by multiplying the transformation matrix and the category feature vector matrix is ​​(2, 5). In this case, the category feature vector can be fused with the image feature vector in a horizontal stitching manner. The method for adjusting the point cloud feature vector can be the same as the method for adjusting the category feature vector described above. Alternatively, the output dimension of the point cloud feature vector can be set when generating the point cloud feature vector so that the dimension of the point cloud feature vector is consistent with that of the image feature vector along axis=1.

[0087] Optionally, the corresponding dimensions of the point cloud feature vectors can be used as a basis to fuse the category feature vectors and image feature vectors with the image feature vectors. The vector transformation method is similar to the method described above and will not be repeated here.

[0088] The above feature fusion steps use prior information feature vectors and image feature vectors as prior information, which can make the fused comprehensive feature vector have more accurate and richer information.

[0089] Step S106: Input the comprehensive feature vector into the target detection model to obtain the target detection result.

[0090] Specifically, the comprehensive feature vectors corresponding to each radar-detected target are input into the target detection model. The target detection model then re-outputs the corresponding verification category of the detected target. In addition to the verification category, the output results may also include information such as the target's position and orientation. The data type of the output results can be set according to actual needs.

[0091] The training method for the object detection model is as follows: obtain a sample dataset, which includes multiple comprehensive feature vectors and corresponding validation categories; build an end-to-end object detection model based on the sample dataset and train it; calculate the loss function value of the object detection model; update the object detection model according to the loss function value until the loss function meets the requirements.

[0092] The method of this invention, after acquiring radar point cloud data and image data of the target area, inputs the radar point cloud data into a radar detection model to obtain the corresponding radar detection result. Then, based on the radar detection result, it determines the corresponding prior information feature vector, and based on the mapping relationship between the radar point cloud data and the image data, it determines the corresponding image feature vector. The prior information feature vector and the image feature vector are then fused to obtain the corresponding comprehensive feature vector. Finally, the comprehensive feature vector is input into the target detection model to obtain the target detection result. This invention uses the radar detection result output by the radar detection model as prior information, extracts and fuses features from the prior information and image data, and inputs the fused feature information into the target detection model, thereby performing secondary target detection on the corresponding data of the target area, thus improving the target detection accuracy and optimizing the detection result. Furthermore, determining the image feature vector based on the mapping relationship between the radar point cloud data and the image data can improve the alignment effect of the two different modalities of data, further improving the target detection accuracy.

[0093] Figure 6 This is a data flow diagram of the target detection method according to an embodiment of the present invention. Figure 6 As shown, the data flow of the target detection method is as follows:

[0094] The following example uses prior information including predicted category, point cloud information within the detection box, and detection box location information. It is easy to understand that prior information can also be added to or replaced with information such as topological relationship, color, orientation, and shape.

[0095] Step S601: Acquire radar point cloud data.

[0096] Step S602: Acquire image data.

[0097] The data in steps S601 and S602 are related data corresponding to the same region.

[0098] Step S603: Input the radar point cloud data into the radar detection model.

[0099] Step S604: The radar detection model outputs the radar detection results, that is, the relevant information of multiple radar-detected targets, and uses the relevant information as prior information.

[0100] The relevant information includes at least the radar coordinates of the detection frame corresponding to the radar-detected target, the point cloud within the detection frame, and the detection category. The point cloud within the detection frame may include information such as the number of radar point cloud echoes, intensity, RGB values, GPS time, scanning angle, and scanning direction.

[0101] Step S605: Determine the category feature vector corresponding to each radar-detected target based on the detection category of each radar-detected target.

[0102] Specifically, based on the detection category and the confidence level corresponding to the detection category, a one-dimensional vector of a predetermined length is generated, and then the one-dimensional vector is determined as the category feature vector.

[0103] Step S606: Determine the point cloud feature vector corresponding to the radar-detected target based on the point cloud information within the detection frame corresponding to each radar-detected target.

[0104] Specifically, the point cloud feature vector corresponding to the radar-detected target is determined based on the number and intensity of the point cloud. Furthermore, the point cloud feature vector can also be determined by combining information such as the number of echoes, intensity, RGB values, GPS time, scanning angle, and scanning direction.

[0105] Step S607: Based on the radar coordinates of the detection box corresponding to the radar-detected target, cut out the sub-image corresponding to the radar-detected target from the image corresponding to the target area.

[0106] Specifically, the radar coordinates of the detection box corresponding to the radar-detected target are mapped to the corresponding image, and then the image is cropped to obtain the sub-image corresponding to the radar-detected target. Alternatively, the image corresponding to the target region can be semantically segmented to obtain multiple sub-images, and then the corresponding sub-images can be matched by combining the radar coordinates of the detection box corresponding to the radar-detected target.

[0107] Step S608: Input the sub-image corresponding to the radar-detected target into the deep learning model.

[0108] Step S609: The deep learning model outputs the image feature vector corresponding to the radar-detected target.

[0109] Step S610: The image feature vector, the category feature vector, and the point cloud feature vector are fused to obtain the corresponding comprehensive feature vector.

[0110] The fusion method can be point-by-point addition or vector concatenation.

[0111] Step S611: Input the comprehensive feature vector into the target detection model.

[0112] Step S612: The target detection model outputs the target detection results.

[0113] In this embodiment of the invention, after acquiring radar point cloud data and image data of the target area, the radar point cloud data is input into a radar detection model to obtain the corresponding radar detection result. Then, based on the radar detection result, the corresponding prior information feature vector is determined, and based on the mapping relationship between the radar point cloud data and the image data, the corresponding image feature vector is determined. The prior information feature vector and the image feature vector are then fused to obtain the corresponding comprehensive feature vector. Finally, the comprehensive feature vector is input into the target detection model to obtain the target detection result. This embodiment of the invention uses the radar detection result output by the radar detection model as prior information, and performs feature extraction and fusion on the prior information and image data. The fused feature information is then input into the target detection model, thereby performing secondary target detection on the corresponding data of the target area, thus improving the target detection accuracy and optimizing the detection result. Furthermore, determining the image feature vector based on the mapping relationship between the radar point cloud data and the image data can improve the alignment effect of the two different modalities of data, further improving the target detection accuracy.

[0114] Figure 7 This is a schematic diagram of a target detection device according to an embodiment of the present invention. Figure 7 As shown, the target detection device includes:

[0115] The acquisition module 701 is used to acquire radar point cloud data and image data of the target area;

[0116] The radar detection module 702 is used to input the radar point cloud data into the radar detection model to obtain the corresponding radar detection results.

[0117] The prior information module 703 is used to determine the corresponding prior information feature vector based on the radar detection result;

[0118] The determining module 704 is used to determine the corresponding image feature vector based on the mapping relationship between the radar point cloud data and the image data;

[0119] The fusion module 705 is used to fuse the prior information feature vector and the image feature vector to obtain a corresponding comprehensive feature vector;

[0120] The target detection module 706 is used to input the comprehensive feature vector into the target detection model to obtain the target detection result.

[0121] The apparatus of this invention, after acquiring radar point cloud data and image data of the target area, inputs the radar point cloud data into a radar detection model to obtain the corresponding radar detection result. Then, based on the radar detection result, it determines the corresponding prior information feature vector, and based on the mapping relationship between the radar point cloud data and the image data, it determines the corresponding image feature vector. The prior information feature vector and the image feature vector are then fused to obtain the corresponding comprehensive feature vector. Finally, the comprehensive feature vector is input into the target detection model to obtain the target detection result. This embodiment of the invention uses the radar detection result output by the radar detection model as prior information, extracts and fuses features from the prior information and image data, and inputs the fused feature information into the target detection model, thereby performing secondary target detection on the corresponding data of the target area, thus improving the target detection accuracy and optimizing the detection result. Furthermore, determining the image feature vector based on the mapping relationship between the radar point cloud data and the image data can improve the alignment effect of the two different modalities of data, further improving the target detection accuracy.

[0122] Figure 8 This is a schematic diagram of an electronic device according to an embodiment of the present invention. In this embodiment, the electronic device 800 includes a server, a terminal, etc. Figure 8 As shown, the electronic device 800 includes at least one processor 801; a memory 802 communicatively connected to at least one processor 801; and a communication component 803 communicatively connected to a scanning device, wherein the communication component 803 receives and transmits data under the control of the processor 801; wherein the memory 802 stores instructions executable by at least one processor 801, which are executed by at least one processor 801 to achieve the aforementioned target detection.

[0123] Specifically, the electronic device includes: one or more processors 801 and a memory 802. Figure 8 Taking a processor 801 as an example, the processor 801 and the memory 802 can be connected via a bus or other means. Figure 8 Taking a bus connection as an example, memory 802, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Processor 801 executes various functional applications and data processing of the device by running the non-volatile software programs, instructions, and modules stored in memory 802, thereby achieving the aforementioned target detection.

[0124] Memory 802 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store an option list, etc. Furthermore, memory 802 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 802 may optionally include memory remotely located relative to processor 801, and these remote memories can be connected to external devices via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0125] One or more modules are stored in memory 802 and, when executed by one or more processors 801, perform target detection in any of the above method embodiments.

[0126] The above-mentioned products can perform the methods provided in the embodiments of this application, and have the corresponding functional modules and beneficial effects of performing the methods. For technical details not described in detail in this embodiment, please refer to the methods provided in the embodiments of this application.

[0127] In this embodiment of the invention, after acquiring radar point cloud data and image data of the target area, the radar point cloud data is input into a radar detection model to obtain the corresponding radar detection result. Then, based on the radar detection result, the corresponding prior information feature vector is determined, and based on the mapping relationship between the radar point cloud data and the image data, the corresponding image feature vector is determined. The prior information feature vector and the image feature vector are then fused to obtain the corresponding comprehensive feature vector. Finally, the comprehensive feature vector is input into the target detection model to obtain the target detection result. This embodiment of the invention uses the radar detection result output by the radar detection model as prior information, and performs feature extraction and fusion on the prior information and image data. The fused feature information is then input into the target detection model, thereby performing secondary target detection on the corresponding data of the target area, thus improving the target detection accuracy and optimizing the detection result. Furthermore, determining the image feature vector based on the mapping relationship between the radar point cloud data and the image data can improve the alignment effect of the two different modalities of data, further improving the target detection accuracy.

[0128] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program for use by a computer to execute some or all of the above-described method embodiments.

[0129] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0130] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A target detection method, characterized in that, The method includes: Acquire radar point cloud data and image data of the target area; The radar point cloud data is input into the radar detection model to obtain the corresponding radar detection results; The corresponding prior information feature vector is determined based on the radar detection results; The corresponding image feature vector is determined based on the mapping relationship between the radar point cloud data and the image data; The prior information feature vector and the image feature vector are fused to obtain the corresponding comprehensive feature vector; The comprehensive feature vector is input into the target detection model to obtain the target detection result.

2. The method according to claim 1, characterized in that, The radar detection results include multiple radar-detected targets and corresponding detection information, wherein the detection information includes the detection category of the radar-detected targets; The step of determining the corresponding prior information feature vector based on the radar detection result includes: Based on the detection category of each radar-detected target, the category feature vector corresponding to the radar-detected target is determined.

3. The method according to claim 1, characterized in that, The radar detection results include multiple radar-detected targets and corresponding detection information, wherein the detection information includes point cloud information within the detection box corresponding to the radar-detected target; The step of determining the corresponding prior information feature vector based on the radar detection result includes: Based on the point cloud information within the detection frame corresponding to each radar-detected target, the point cloud feature vector corresponding to the radar-detected target is determined.

4. The method according to any one of claims 1-3, characterized in that, The radar detection results include multiple radar-detected targets and corresponding detection information, wherein the detection information includes the detection frame position information of the radar-detected targets; The step of determining the corresponding image feature vector based on the mapping relationship between the radar point cloud data and the image data includes: Based on the mapping relationship between the radar point cloud data and the image data, and the detection box position information of each radar-detected target, a sub-image of each radar-detected target is determined from the image data; Based on the sub-images of each radar-detected target, the image feature vector of each radar-detected target is determined.

5. The method according to claim 4, characterized in that, The step of determining the sub-image of each radar-detected target from the image data based on the mapping relationship between the radar point cloud data and the image data, and the detection box position information of each radar-detected target, includes: Determine the coordinate transformation relationship between the radar point cloud data and the image data; Based on the coordinate transformation relationship, the radar coordinates of the detection box corresponding to the radar-detected target are converted into image coordinates; Based on the image coordinates of the detection box corresponding to the radar-detected target, the corresponding sub-image is cut out from the image data.

6. The method according to claim 4, characterized in that, The detection information also includes the detection category of the target detected by the radar; The step of determining sub-images of each radar-detected target from the image data based on the mapping relationship between the radar point cloud data and the image data, and the detection bounding box position information of each radar-detected target, includes: Perform semantic segmentation on the image data; The image data is cut into multiple sub-images based on the semantic segmentation results; The detection category of each radar-detected target is matched with the semantic segmentation category of each sub-image to determine the sub-image corresponding to each radar-detected target.

7. The method according to claim 4, characterized in that, The step of determining the image feature vector of each radar-detected target based on the sub-image of each radar-detected target includes: The sub-image corresponding to the radar-detected target is input into a deep learning model to obtain the image feature vector corresponding to the radar-detected target.

8. The method according to claim 1, characterized in that, The step of fusing the prior information feature vector and the image feature vector to obtain the corresponding comprehensive feature vector includes: The transformation matrix is ​​determined based on the prior information feature vector and the image feature vector; Based on the transformation matrix, the prior information feature vector is transformed into a category feature vector to be fused; The image feature vector and the category feature vector to be fused are fused to obtain the corresponding comprehensive feature vector.

9. A target detection device, characterized in that, The device includes: The acquisition module is used to acquire radar point cloud data and image data of the target area; The radar detection module is used to input the radar point cloud data into the radar detection model to obtain the corresponding radar detection results. The prior information module is used to determine the corresponding prior information feature vector based on the radar detection result; The determination module is used to determine the corresponding image feature vector based on the mapping relationship between the radar point cloud data and the image data; The fusion module is used to fuse the prior information feature vector and the image feature vector to obtain the corresponding comprehensive feature vector; The target detection module is used to input the comprehensive feature vector into the target detection model to obtain the target detection result.

10. An electronic device comprising a memory and a processor, characterized in that, The memory is used to store one or more computer program instructions, wherein the one or more computer program instructions are executed by the processor to implement the method as described in any one of claims 1-8.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method as described in any one of claims 1-8.

12. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Vehicle target detection method and system based on Leiyu semantic segmentation adaptive fusion

    CN114724120A

  • Target detection method and device, electronic equipment and storage medium

    CN117456164A

Cited By

  • Patrol point location image multi-modal feature construction method and system and storage medium

    CN121415091A