Target detection method and device, electronic equipment and storage medium

By associating, segmenting, and performing depth search on the target bounding boxes of point clouds and images, the problems of noise interference and occlusion when point clouds are projected onto images in existing technologies are solved, thereby improving the robustness and accuracy of target detection.

CN115861746BActive Publication Date: 2026-02-17UISEE SHANGHAI AUTOMOTIVE TECH LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211424729.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-14
Publication Date
2026-02-17
Estimated Expiration
2042-11-14

AI Technical Summary

Technical Problem

Existing post-fusion methods are easily affected by noise when projecting point clouds onto images to obtain target depth information, and they are not good at detecting targets with occlusion problems, resulting in poor robustness.

Method used

By determining the association between the initial 3D target bounding box and the initial 2D target bounding box, the 2D target bounding box to be processed that failed to be associated is identified, and it is then processed in blocks and depth searched to recover the 3D target bounding box to be processed. Finally, the initial 3D target bounding box and the 3D target bounding box to be processed are fused together.

Benefits of technology

It improves the robustness of target detection, achieves accurate fusion of point cloud and image recognition results, and enhances the accuracy and reliability of target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115861746B_ABST
    Figure CN115861746B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a target detection method and device, electronic equipment and storage medium, the method comprising: determining at least one initial three-dimensional target frame according to initial point cloud data, and determining at least one initial two-dimensional target frame according to initial image data synchronized in time with the initial point cloud data; determining a converted two-dimensional target frame of the initial three-dimensional target frame in an image coordinate system; associating the initial two-dimensional target frame with the converted two-dimensional target frame, and determining an initial two-dimensional target frame that fails in association as a two-dimensional target frame to be processed; performing block processing on the two-dimensional target frame to be processed to obtain a plurality of sub-blocks; performing a depth search on each sub-block respectively to restore a three-dimensional target frame to be processed associated with the two-dimensional target frame to be processed; and determining a fused three-dimensional target frame according to the initial three-dimensional target frame and the three-dimensional target frame to be processed. The present disclosure realizes accurate fusion of point cloud recognition results and image recognition results, and improves the robustness of target detection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, and particularly relates to a target detection method and device, electronic equipment and storage medium. BACKGROUND

[0002] Current perception algorithms gradually evolve towards multi-sensor fusion. Since different sensors have different advantages, fusing data of different sensors can increase the reliability and stability of perception.

[0003] Point clouds can be obtained through radar sensors, and images can be obtained through image sensors. A method of fusing images and point clouds includes post-fusion.

[0004] However, the existing post-fusion method directly projects point clouds onto images to obtain depth information of a target. This method is easily disturbed by point cloud noise, and is not friendly to target detection of occlusion problems, and has poor robustness. SUMMARY

[0005] To solve the above technical problems or at least partially solve the above technical problems, the embodiments of the present disclosure provide a target detection method, device, electronic equipment and storage medium to realize accurate fusion of point cloud recognition results and image recognition results, and improve the robustness of target detection.

[0006] In a first aspect, the embodiments of the present disclosure provide a target detection method, which comprises:

[0007] According to initial point cloud data, at least one initial three-dimensional target box is determined, and according to initial image data that is time-synchronized with the initial point cloud data, at least one initial two-dimensional target box is determined.

[0008] A converted two-dimensional target box corresponding to the initial three-dimensional target box in an image coordinate system is determined.

[0009] The initial two-dimensional target box and the converted two-dimensional target box are associated, and an initial two-dimensional target box that fails in association is determined as a to-be-processed two-dimensional target box.

[0010] The to-be-processed two-dimensional target box is processed in blocks to obtain a plurality of sub-blocks.

[0011] A depth search is performed on each sub-block, and a to-be-processed three-dimensional target box associated with the to-be-processed two-dimensional target box is recovered.

[0012] According to the initial three-dimensional target box and the to-be-processed three-dimensional target box, a fused three-dimensional target box corresponding to the initial point cloud data is determined, wherein the fused three-dimensional target box is used to indicate three-dimensional information of a detected target.

[0013] In a second aspect, the present disclosure also provides a target detection device, the device comprising:

[0014] An initial target box determination module is configured to determine at least one initial three-dimensional target box according to initial point cloud data, and determine at least one initial two-dimensional target box according to initial image data that is time-synchronized with the initial point cloud data.

[0015] A converted two-dimensional target box determination module is configured to determine a converted two-dimensional target box corresponding to the initial three-dimensional target box in an image coordinate system.

[0016] A to-be-processed two-dimensional target box determination module is configured to associate the initial two-dimensional target box with the converted two-dimensional target box, and determine an initial two-dimensional target box that fails in association as a to-be-processed two-dimensional target box.

[0017] A block processing module is configured to perform block processing on the to-be-processed two-dimensional target box to obtain a plurality of sub-blocks.

[0018] A to-be-processed three-dimensional target box recovery module is configured to perform depth search on each sub-block respectively to recover a to-be-processed three-dimensional target box associated with the to-be-processed two-dimensional target box.

[0019] A fused three-dimensional target box determination module is configured to determine a fused three-dimensional target box corresponding to the initial point cloud data according to the initial three-dimensional target box and the to-be-processed three-dimensional target box, wherein the fused three-dimensional target box is used to indicate three-dimensional information of a detected target.

[0020] In a third aspect, the present disclosure also provides an electronic device, the electronic device comprising: one or more processors; a storage device configured to store one or more programs; and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the target detection method as described above.

[0021] In a fourth aspect, the present disclosure also provides a computer-readable storage medium having a computer program stored thereon, and the program is executed by a processor to implement the target detection method as described above.

[0022] The target detection method provided in the embodiments of the present disclosure comprises: performing target recognition on initial point cloud data and initial image data respectively to obtain an initial three-dimensional target frame and an initial two-dimensional target frame; converting the initial three-dimensional target frame to an image coordinate system to obtain a converted two-dimensional target frame; associating the initial two-dimensional target frame with the converted two-dimensional target frame; regarding an initial two-dimensional target frame that fails in association as a two-dimensional target frame to be processed, so as to perform subsequent three-dimensional recovery; further performing block processing on the two-dimensional target frame to be processed to obtain a plurality of sub-blocks, and performing depth search on each sub-block to recover a three-dimensional target frame to be processed associated with the two-dimensional target frame to be processed; and further fusing the initial three-dimensional target frame with the three-dimensional target frame to be processed to obtain a fused three-dimensional target frame, so as to fuse the image recognition result and the point cloud recognition result, solve the problem that the point cloud recognition result and the image recognition result are prone to interference and have poor robustness when being fused, and realize accurate fusion of the point cloud target frame and the image target frame, and improve the robustness of target detection. BRIEF DESCRIPTION OF DRAWINGS

[0023] The above and other features, advantages, and aspects of the present disclosure will become more apparent by referring to the following detailed description in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements throughout. It should be understood that the drawings are schematic and elements and features are not necessarily to scale.

[0024] Figure 1 A flowchart of a target detection method in the embodiments of the present disclosure;

[0025] Figure 2 A flowchart of another target detection method in the embodiments of the present disclosure;

[0026] Figure 3 A flowchart of another target detection method in the embodiments of the present disclosure;

[0027] Figure 4 A schematic diagram of a post-fusion process in the embodiments of the present disclosure;

[0028] Figure 5 A schematic diagram of block depth search of a two-dimensional target frame to be processed in the embodiments of the present disclosure;

[0029] Figure 6 A schematic diagram of a center point-based search method in the embodiments of the present disclosure;

[0030] Figure 7 A schematic diagram of an L-shaped structure in the embodiments of the present disclosure;

[0031] Figure 8 A structural schematic diagram of a target detection device in the embodiments of the present disclosure;

[0032] Figure 9 FIG. 1 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0033] Embodiments of the present disclosure will be described in more detail with reference to the drawings. Although certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be implemented in various forms and should not be interpreted as being limited to the embodiments set forth herein, but rather the embodiments are provided to more thoroughly and completely understand the present disclosure. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.

[0034] It should be noted that the concepts of "first", "second", and the like mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0035] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0036] Commonly used fusion methods of image targets and point cloud targets include early fusion, mid-fusion, and late fusion. Among them, early fusion mainly directly superimposes point cloud data and corresponding point image data at the input end of the network, or fuses point cloud data and abstract deep image data; mid-fusion mainly inputs point cloud and image into respective networks, extracts corresponding deep features, and then fuses the corresponding deep features. Late fusion mainly fuses the perception results of the point cloud network and the perception results of the image network.

[0037] Since early fusion and mid-fusion need to fuse corresponding feature points, they have very high requirements for actual radar sensor and image sensor time synchronization, and sensor time synchronization deviation will bring feature-level fusion misalignment, affecting the final perception result. Moreover, feature-level fusion needs more image and point cloud synchronization data driving. Late fusion can ignore the interference of point cloud and image feature level, maintain its own perception result, and is more flexible and has strong scalability. However, the existing late fusion method often directly projects the point cloud to the image to obtain the depth information of the detected target, which is easy to be disturbed by point cloud noise and is not friendly to the detection of occluded targets, and the obtained depth position information is not robust.

[0038] In view of the above problems, the embodiments of the present disclosure provide a target detection method, which reasonably and effectively fuses the target frame in the point cloud and the target frame in the image, and improves the robustness of target detection.

[0039] Figure 1 FIG. 1 is a flowchart of a target detection method according to an embodiment of the present disclosure. The method can be performed by a target detection device, which can be implemented in software and / or hardware, and can be configured in an electronic device. As shown in FIG. 1, the method can specifically include the following steps: Figure 1

[0040] S110, determining at least one initial three-dimensional target frame according to initial point cloud data, and determining at least one initial two-dimensional target frame according to initial image data that is time-synchronized with the initial point cloud data.

[0041] The initial point cloud data can be point cloud data collected based on a point cloud collection device such as a laser radar, or can be pre-stored point cloud data. The initial three-dimensional target frame can be a three-dimensional target frame obtained by a point cloud target recognition method, such as detection and recognition of a point cloud target by a pre-trained PointPillar network. The initial image data can be image data that is time-synchronized with the initial point cloud data and is collected based on an image collection device such as a camera, or can be pre-stored image data that is time-synchronized with the initial point cloud data. The initial two-dimensional target frame can be a two-dimensional target frame obtained by an image target recognition method, such as detection and recognition of an image target by a pre-trained Centernet network.

[0042] Specifically, the initial point cloud data and the initial image data that are time-synchronized can be collected based on a point cloud collection device and an image collection device, and the pre-stored initial point cloud data and the pre-stored initial image data that are time-synchronized can also be obtained. Then, the initial point cloud data is subjected to point cloud target recognition to extract at least one initial three-dimensional target frame in the initial point cloud data, and the initial image data is subjected to image target extraction to extract at least one initial two-dimensional target frame in the initial image data.

[0043] For example, the initial point cloud data can be P = {P1, P2, …, P N}, P j = {x, y, z, i}, P j represents a point cloud point in the initial point cloud data P, and x, y, z, i respectively represent three-dimensional coordinates and reflectivity of the point. The initial three-dimensional target frame can include a target class and a three-dimensional position, such as a target class of class_name: car and a three-dimensional position of location: x, y, z, w, l, h, yaw. Here, car represents that the target class of the detected target is a car; x, y, z represent three-dimensional coordinates of the center of the detected target; w, l, h represent dimensions of the detected target, i.e., width, length, and height; and yaw represents a heading angle of the detected target. The initial image data can be I = {I1, I2, …, In}, I​N}, I j = {r, g, b, x, y}, I j represents a pixel point in the initial image data I, r, g, b, x, y respectively represent color information and two-dimensional coordinates of the pixel point. The initial two-dimensional target frame can include a target category and a two-dimensional position, for example, the target category is class_name: car, and the two-dimensional position is location: x, y, w, h. Wherein, car represents that the target category of the detected target is a car; x, y represent the two-dimensional coordinates of the center of the detected target; w and h represent the size of the target, that is, the width and the length.

[0044] S120, determine the conversion two-dimensional target frame corresponding to the initial three-dimensional target frame in the image coordinate system.

[0045] Wherein, the conversion two-dimensional target frame can be a two-dimensional target frame converted from the point cloud coordinate system to the image coordinate system by the initial three-dimensional target frame.

[0046] Specifically, according to the conversion relationship between the point cloud coordinate system and the image coordinate system determined in advance, the three-dimensional coordinates in the point cloud coordinate system in the initial three-dimensional target frame can be converted into two-dimensional coordinates in the image coordinate system, and the conversion two-dimensional target frame corresponding to the initial three-dimensional target frame can be determined according to the two-dimensional coordinates.

[0047] Exemplarily, the conversion relationship between the point cloud coordinate system and the image coordinate system can be determined according to camera calibration:

[0048]

[0049] Wherein, (x, y) represents the two-dimensional coordinates of the pixel point in the image coordinate system, (X, Y, Z) represents the three-dimensional coordinates of the point cloud point in the point cloud coordinate system, represents the mapping table of the point cloud projection to the image determined according to the camera calibration.

[0050] S130, associate the initial two-dimensional target frame with the conversion two-dimensional target frame, and determine the initial two-dimensional target frame with association failure as the two-dimensional target frame to be processed.

[0051] Wherein, the two-dimensional target frame to be processed represents the initial two-dimensional target frame that cannot be associated with any conversion two-dimensional target frame. It can be understood as the initial two-dimensional target frame detected by image target detection, but not detected by point cloud target detection.

[0052] Specifically, the initial two-dimensional target frame and the converted two-dimensional target frame can be detected by overlap calculation, similarity calculation, or the like. The initial two-dimensional target frame and the converted two-dimensional target frame that meet a preset standard are associated. Then, the initial two-dimensional target frame that fails to be associated is determined, and the initial two-dimensional target frame is taken as a two-dimensional target frame to be processed, for subsequent reconstruction of a three-dimensional target frame to be processed.

[0053] If the detection target is detected by both the point cloud target detection and the image target detection, and the overlap corresponding to the detection target is greater than the overlap threshold, the initial two-dimensional target frame and the converted two-dimensional target frame are successfully associated. If the detection target is detected by both the point cloud target detection and the image target detection, but the overlap corresponding to the detection target is not greater than the overlap threshold, the initial two-dimensional target frame and the converted two-dimensional target frame fail to be associated, and the initial two-dimensional target frame needs to be three-dimensionally recovered. If the detection target is detected by the point cloud target detection but not by the image target detection, there is no initial two-dimensional target frame associated with the converted two-dimensional target frame. In this case, the point cloud target detection is more effective, and no processing is needed. If the detection target is detected by the image target detection but not by the point cloud target detection, there is no converted two-dimensional target frame associated with the initial two-dimensional target frame. In this case, the image target detection is more effective, and the detection target detected by the image target detection but not by the point cloud target detection can be three-dimensionally reconstructed to expand the three-dimensional target frame and improve the accuracy and robustness of the three-dimensional target detection.

[0054] S140, the two-dimensional target frame to be processed is divided into blocks to obtain a plurality of sub-blocks.

[0055] The sub-blocks can be obtained by dividing the two-dimensional target frame to be processed according to a preset block division rule. The preset block division rule can be 5x5, 7x7, 9x9, or the like. The specific rule can be set according to actual needs, and is not specifically limited in this embodiment.

[0056] Specifically, each part after the division of the two-dimensional target frame to be processed according to the preset block division rule is taken as a sub-block.

[0057] S150, a depth search is performed on each sub-block to recover a three-dimensional target frame to be processed associated with the two-dimensional target frame to be processed.

[0058] The depth search can be a search on the depth information of the sub-block. The three-dimensional target frame to be processed can be a three-dimensional target frame obtained by reconstructing the two-dimensional target frame to be processed.

[0059] Specifically, a depth search can be performed for each sub-block to determine depth information corresponding to the sub-block. For example, the depth information can be depth information matched by converting a preset pixel point in the sub-block to a point cloud coordinate system and matching the initial point cloud data, or can be an average of multiple depth information obtained by converting multiple preset pixel points in the sub-block to the point cloud coordinate system and matching the initial point cloud data, respectively. If there is no corresponding depth information after the depth search of the sub-block, it indicates that there is no point cloud point corresponding to the sub-block in the initial point cloud data. After the depth search of each sub-block, the obtained point cloud points can be integrated to obtain a three-dimensional box corresponding to the point cloud points, that is, a three-dimensional target box to be processed associated with the two-dimensional target box to be processed.

[0060] S160, determining a fusion three-dimensional target box corresponding to the initial point cloud data according to the initial three-dimensional target box and the three-dimensional target box to be processed.

[0061] The fusion three-dimensional target box is the sum of the initial three-dimensional target box and the three-dimensional target box to be processed. The fusion three-dimensional target box is used to indicate three-dimensional information of the detection target. The three-dimensional information can include three-dimensional coordinates and target categories.

[0062] Specifically, the initial three-dimensional target box is taken as a first part of the fusion three-dimensional target box, and the three-dimensional target box to be processed is taken as a second part of the fusion three-dimensional target box. The obtained fusion three-dimensional target box can fuse the point cloud target detection result and the image target detection result, thereby improving the accuracy and robustness of target detection.

[0063] For example, in the process of fusing the initial three-dimensional target box and the three-dimensional target box to be processed, an NMS (Non-maximum suppression) method can be used to fuse the overlapping parts of the initial three-dimensional target box and the three-dimensional target box to be processed, so that the initial three-dimensional target box and the three-dimensional target box to be processed that overlap with each other can be fused into one three-dimensional target box.

[0064] The target detection method provided in the embodiment performs target recognition on the initial point cloud data and the initial image data respectively, obtains an initial three-dimensional target frame and an initial two-dimensional target frame, then converts the initial three-dimensional target frame to the image coordinate system to obtain a converted two-dimensional target frame, associates the initial two-dimensional target frame with the converted two-dimensional target frame, takes the initial two-dimensional target frame that fails in association as a two-dimensional target frame to be processed, and performs subsequent three-dimensional recovery, further performs block processing on the two-dimensional target frame to be processed to obtain a plurality of sub-blocks, performs depth search on each sub-block, restores a three-dimensional target frame to be processed associated with the two-dimensional target frame to be processed, and then fuses the initial three-dimensional target frame with the three-dimensional target frame to be processed to obtain a fused three-dimensional target frame, so as to fuse the image recognition result and the point cloud recognition result, solve the problem that the point cloud recognition result and the image recognition result are prone to interference and have poor robustness when being fused, and realize accurate fusion of the point cloud target frame and the image target frame and improve the robustness of target detection.

[0065] Figure 2 A flowchart of another target detection method in the embodiment of the present disclosure is shown in FIG. 13. On the basis of the above-mentioned embodiments, the specific implementation of determining the converted two-dimensional target frame, the two-dimensional target frame to be processed, and the three-dimensional target frame to be processed can be referred to the detailed description of the present technical solution. The same or corresponding terms as those in the above-mentioned embodiments are not described herein again. As shown in FIG. 13, the method specifically can include the following steps: Figure 2

[0066] S210, determining at least one initial three-dimensional target frame according to the initial point cloud data, and determining at least one initial two-dimensional target frame according to the initial image data that is time-synchronized with the initial point cloud data.

[0067] S220, determining eight corner points of the initial three-dimensional target frame for each initial three-dimensional target frame.

[0068] Specifically, since the initial three-dimensional target frame is a hexahedral structure frame, there are eight corner points. Therefore, for each initial three-dimensional target frame, eight corner points can be determined.

[0069] S230, determining, for each corner point, corner point two-dimensional data of the corner point in the image coordinate system, and determining a converted two-dimensional target frame corresponding to the initial three-dimensional target frame according to the corner point two-dimensional data.

[0070] The corner point two-dimensional data can be two-dimensional data of the corner point converted from the point cloud coordinate system to the image coordinate system.

[0071] ​Specifically, after the eight corner points are determined, the three-dimensional coordinates of each corner point can be determined. Then, for each corner point, the three-dimensional coordinates of the corner point are converted to the image coordinate system, and the two-dimensional data of the corner point corresponding to the corner point can be obtained. After the eight corner point two-dimensional data are determined, the region surrounded by the eight corner point two-dimensional data can be processed to obtain the conversion two-dimensional target frame corresponding to the initial three-dimensional target frame.

[0072] For example, the circumscribed quadrilateral frame of the region surrounded by the eight corner point two-dimensional data can be taken as the conversion two-dimensional target frame.

[0073] Based on the above examples, the conversion two-dimensional target frame corresponding to the initial three-dimensional target frame can be determined according to the corner point two-dimensional data in the following manner:

[0074] According to the corner point two-dimensional data, the maximum value of the horizontal axis, the minimum value of the horizontal axis, the maximum value of the vertical axis, and the minimum value of the vertical axis are determined. According to the maximum value of the horizontal axis, the minimum value of the horizontal axis, the maximum value of the vertical axis, and the minimum value of the vertical axis, the conversion two-dimensional target frame corresponding to the initial three-dimensional target frame is determined.

[0075] The maximum value of the horizontal axis is the maximum value of the horizontal coordinate in the eight corner point two-dimensional data. The minimum value of the horizontal axis is the minimum value of the horizontal coordinate in the eight corner point two-dimensional data. The maximum value of the vertical axis is the maximum value of the vertical coordinate in the eight corner point two-dimensional data. The minimum value of the vertical axis is the minimum value of the vertical coordinate in the eight corner point two-dimensional data.

[0076] Specifically, according to the eight corner point two-dimensional data, eight horizontal coordinates are determined, and then the maximum value of the horizontal axis and the minimum value of the horizontal axis are determined therefrom. Furthermore, according to the eight corner point two-dimensional data, eight vertical coordinates are determined, and then the maximum value of the vertical axis and the minimum value of the vertical axis are determined therefrom. According to the maximum value of the horizontal axis, the minimum value of the horizontal axis, the maximum value of the vertical axis, and the minimum value of the vertical axis, four two-dimensional coordinates can be formed, and the frame formed by the four two-dimensional coordinates is taken as the conversion two-dimensional target frame corresponding to the initial three-dimensional target frame.

[0077] For example, the maximum value of the horizontal axis is xmax, the minimum value of the horizontal axis is xmin, the maximum value of the vertical axis is ymax, and the minimum value of the vertical axis is ymin. Then, the four two-dimensional coordinates formed are (xmax, ymax), (xmax, ymin), (xmin, ymax), and (xmin, ymin), and the quadrilateral frame formed by the four two-dimensional coordinates is the conversion two-dimensional target frame.

[0078] S240, for each initial two-dimensional target frame, the degree of overlap between the initial two-dimensional target frame and each conversion two-dimensional target frame is determined.

[0079] The overlap degree can be a result value of an intersection area of the initial two-dimensional target frame and the converted two-dimensional target frame divided by a union set area of the initial two-dimensional target frame and the converted two-dimensional target frame.

[0080] Specifically, for each initial two-dimensional target frame, the intersection area and the union set area of the initial two-dimensional target frame and any one of the converted two-dimensional target frames can be determined, and then the ratio of the intersection area to the union set area is taken as the overlap degree between the initial two-dimensional target frame and the converted two-dimensional target frame.

[0081] In S250, if the overlap degree is greater than a preset overlap threshold, it is determined that the initial two-dimensional target frame is associated successfully; if the overlap degrees are all not greater than the preset overlap threshold, it is determined that the initial two-dimensional target frame is associated unsuccessfully, and the initial two-dimensional target frame is determined as a two-dimensional target frame to be processed.

[0082] The preset overlap threshold can be a threshold value set for the overlap degree to determine whether the initial two-dimensional target frame is associated with the converted two-dimensional target frame. The specific value of the preset overlap threshold can be set according to actual needs, which is not limited in the embodiment.

[0083] Specifically, for each initial two-dimensional target frame, the overlap degree corresponding to each converted two-dimensional target frame can be calculated, and therefore the number of overlap degrees is the same as the number of converted two-dimensional target frames. If the overlap degree is greater than the preset overlap threshold, it indicates that the initial two-dimensional target frame corresponding to the overlap degree is associated successfully with the converted two-dimensional target frame, that is, the initial two-dimensional target frame is associated successfully. In this case, it indicates that the detection target is recognized by both the point cloud target recognition and the image target recognition. If the overlap degrees are all not greater than the preset overlap threshold, it indicates that there is no converted two-dimensional target frame corresponding to the initial two-dimensional target frame, that is, the initial two-dimensional target frame is associated unsuccessfully. In this case, the initial two-dimensional target frame can be determined as a two-dimensional target frame to be processed for subsequent three-dimensional recovery processing.

[0084] For example, the overlap degree iou between each converted two-dimensional target frame B 2d′ and the initial two-dimensional target frame B 2d is calculated. If the iou is greater than a preset overlap threshold, for example, 0.35, it is considered that the detection target is matched, that is, associated successfully, and the detection target is perceived in both the image target recognition and the point cloud face recognition. The initial two-dimensional target frame B 2d that is not matched (associated unsuccessfully) is taken as a two-dimensional target frame to be processed.

[0085] For example, if there are three converted two-dimensional target frames, B 2d′ 1, B 2d′ 2 and B 2s′ 3. For the initial two-dimensional target frame B 2sThe degree of overlap between B 2d The degree of overlap between B 2d′ 1, the degree of overlap between B 2d The degree of overlap between B 2d′ 2, and the degree of overlap between B 2d The degree of overlap between B 2d′ 3. If one or more of the above three degrees of overlap is greater than the overlap threshold, it can be considered that there is a converted two-dimensional target frame matched with the initial two-dimensional target frame B 2d , i.e., the initial two-dimensional target frame B 2d is successfully associated. If none of the above three degrees of overlap is greater than the overlap threshold, it means that there is no converted two-dimensional target frame matched with the initial two-dimensional target frame B 2d , i.e., the initial two-dimensional target frame B 2d fails to be associated.

[0086] S260, performing block processing on the two-dimensional target frame to be processed to obtain a plurality of sub-blocks.

[0087] S270, for each sub-block, determining a grid center point corresponding to the sub-block; and determining a target conversion point corresponding to the sub-block according to the grid center point and a preset search parameter.

[0088] The grid center point can be a point in the sub-block that is at the center position of the image. The preset search parameter can be a parameter for matching search movement, which can include, for example, the direction of search movement and the step length of search movement. The target conversion point can be a pixel point corresponding to the sub-block that has depth information.

[0089] Specifically, for each sub-block, it can be determined whether there is depth information of the grid center point. If there is, the pixel point corresponding to the depth information is determined as the target conversion point. Otherwise, the target conversion point can be searched one by one according to the preset search parameter until the target conversion point is determined or the last pixel point is searched. If the last pixel point is searched and the target conversion point corresponding to the sub-block is still not determined, the target conversion point corresponding to the sub-block is empty.

[0090] Based on the above example, the target conversion point corresponding to the sub-block can be determined according to the grid center point and the preset search parameter by the following steps:

[0091] Step one, for each point cloud point in the initial point cloud data, determining a point cloud conversion point of the point cloud point in the image coordinate system.

[0092] The point cloud conversion point can be a pixel point obtained by converting the point cloud point in the initial point cloud data to the image coordinate system.

[0093] Specifically, for each point cloud point in the initial point cloud data, the point cloud point can be converted from the point cloud coordinate system to the image coordinate system to obtain a pixel point corresponding to the point cloud point, i.e., a point cloud conversion point. For example, the conversion of the point cloud point to the pixel point can be performed according to a predetermined conversion relationship between the point cloud coordinate system and the image coordinate system.

[0094] Step two, taking the grid center point as a to-be-matched point.

[0095] The to-be-matched point can be a pixel point in the sub-block that is to be subsequently matched.

[0096] Step three, determining a matching result according to the to-be-matched point and each point cloud conversion point.

[0097] The matching result is used to indicate a matching situation between the to-be-matched point and the point cloud conversion point, and can include matching success and matching failure.

[0098] Specifically, the to-be-matched point and each point cloud conversion point can be matched respectively, i.e., it is determined whether the to-be-matched point has a corresponding point cloud point in the initial point cloud data, i.e., whether there is depth information. If there is a corresponding case between the to-be-matched point and the point cloud conversion point, it is determined that the matching result is matching success; if none of the point cloud conversion points corresponds to the to-be-matched point, it is determined that the matching result is matching failure.

[0099] Step four, if the matching result is matching success, the matched point cloud conversion point is taken as a target conversion point corresponding to the sub-block; if the matching result is matching failure, the to-be-matched point is updated according to a preset search parameter, and the operation of determining a matching result according to the to-be-matched point and each point cloud conversion point is performed again.

[0100] Specifically, if the matching result is matching success, it indicates that there is a point cloud conversion point corresponding to the to-be-matched point, at this time, the matched point cloud conversion point can be taken as a target conversion point corresponding to the sub-block, which is used to represent the sub-block. If the matching result is matching failure, it indicates that the depth information cannot be searched according to the current to-be-matched point, therefore, the next to-be-matched point can be determined according to the to-be-matched point and the preset search parameter to replace the current to-be-matched point, and then the operation of determining a matching result according to the to-be-matched point and each point cloud conversion point is performed again to further search the depth.

[0101] S280, determining a to-be-processed three-dimensional target frame corresponding to the to-be-processed two-dimensional target frame according to each target conversion point.

[0102] Specifically, after the target conversion points corresponding to each sub-block are determined, the target conversion points can be sorted and integrated, for example, clustering processing, and then a three-dimensional frame formed by the target conversion points is taken as a to-be-processed three-dimensional target frame corresponding to the to-be-processed two-dimensional target frame.

[0103] On the basis of the above examples, the to-be-processed three-dimensional target frame corresponding to the to-be-processed two-dimensional target frame can be determined according to the target conversion points by the following steps:

[0104] Step one, for each target conversion point, the to-be-compared slope of the target conversion point is determined according to the first coordinate information and the second coordinate information of the target conversion point.

[0105] Wherein, if the first coordinate information is the horizontal coordinate information, the second coordinate information can be the vertical coordinate information. If the first coordinate information is the vertical coordinate information, the second coordinate information can be the horizontal coordinate information. The to-be-compared slope can be the slope of the line connecting the target conversion point and the origin.

[0106] Specifically, taking the first coordinate information as the horizontal coordinate information and the second coordinate information as the vertical coordinate information as an example, after determining the first coordinate information and the second coordinate information of the target conversion point, the ratio of the second coordinate information to the first coordinate information can be taken as the to-be-compared slope of the target conversion point.

[0107] It can be understood that if the first coordinate information is the vertical coordinate information and the second coordinate information is the horizontal coordinate information, the ratio of the first coordinate information to the second coordinate information can be taken as the to-be-compared slope of the target conversion point.

[0108] Step two, the maximum slope and the minimum slope are determined according to the to-be-compared slopes, and the target conversion point corresponding to the maximum slope is taken as the first conversion point and the target conversion point corresponding to the minimum slope is taken as the second conversion point.

[0109] Wherein, the first conversion point is the point with the maximum to-be-compared slope among the target conversion points. The second conversion point is the point with the minimum to-be-compared slope among the target conversion points.

[0110] Specifically, after determining the to-be-compared slopes of each target conversion point, the maximum slope and the minimum slope can be determined, and then the target conversion point corresponding to the maximum slope can be taken as the first conversion point and the target conversion point corresponding to the minimum slope can be taken as the second conversion point.

[0111] Step three, the third conversion point is determined according to the target conversion points, the first conversion point and the second conversion point.

[0112] Wherein, the third conversion point can be the target conversion point farthest from the straight line formed by the first conversion point and the second conversion point among the target conversion points.

[0113] Specifically, a straight line can be formed according to the first conversion point and the second conversion point, and then the distances of the target conversion points to the straight line can be determined, and the target conversion point farthest from the straight line is determined as the third conversion point.

[0114] Optionally, the target conversion points can be screened first, for example, a straight line formed by the first conversion point and the second conversion point divides the image coordinate system into two parts, and the target conversion points located in the same part as the origin of the image coordinate system are selected to calculate the distance to the straight line and perform subsequent comparison operations.

[0115] Step four, determining a to-be-processed three-dimensional target frame corresponding to the to-be-processed two-dimensional target frame according to the first conversion point, the second conversion point and the third conversion point.

[0116] Specifically, the first conversion point, the second conversion point and the third conversion point can form an "L" shape, therefore, the size and position information of the detected target can be recovered according to the first conversion point, the second conversion point and the third conversion point, a new three-dimensional frame is obtained, and the three-dimensional frame is taken as the to-be-processed three-dimensional target frame corresponding to the to-be-processed two-dimensional target frame.

[0117] It should be noted that steps one to four are l-shape clustering methods, and specific details are not described herein. Of course, other clustering methods can also be selected to integrate and convert the target conversion points to obtain the to-be-processed three-dimensional target frame corresponding to the to-be-processed two-dimensional target frame.

[0118] S290, determining a fusion three-dimensional target frame corresponding to the initial point cloud data according to the initial three-dimensional target frame and the to-be-processed three-dimensional target frame.

[0119] The target detection method provided in this embodiment solves the problem of low accuracy of projecting the point cloud target frame into the image coordinate system by converting the eight corner points of the initial three-dimensional target frame into the image coordinate system to determine the corner point two-dimensional data, and constructing the conversion two-dimensional target frame corresponding to the initial three-dimensional target frame according to the corner point two-dimensional data. By determining the overlap degree between the initial two-dimensional target frame and each conversion two-dimensional target frame, the to-be-processed two-dimensional target frame that needs to be reconstructed into a three-dimensional target frame is determined, and by performing depth search on the sub-blocks obtained by dividing the to-be-processed two-dimensional target frame according to the preset search parameters and the grid center points, the target conversion points corresponding to each sub-block are determined, the target conversion points are integrated to obtain the to-be-processed three-dimensional target frame corresponding to the to-be-processed two-dimensional target frame, the problem that the image recognition result cannot be accurately fused and converted into the point cloud recognition result is solved, the accuracy of fusing the image recognition result and the point cloud recognition result is improved, and the effect of avoiding loss of the detected target is achieved.

[0120] Figure 3 The flowchart of another target detection method in the embodiments of the present disclosure is shown in FIG. 6. As shown in FIG. 6, the method specifically can include the following steps: Figure 3

[0121] S310, data set preparation and labeling.​

[0122] The predetermined point cloud data training set and the predetermined image data training set are labeled. For each frame of point cloud, the target class and three-dimensional position need to be labeled. For each frame of image, the target class and two-dimensional position need to be labeled.

[0123] S320, network model construction and training.

[0124] A three-dimensional perception network model is constructed according to the PointPillar network, which is used for subsequent identification of an initial three-dimensional target box. A two-dimensional perception network model is constructed according to the Centernet network, which is used for subsequent identification of an initial two-dimensional target box. The three-dimensional perception network model is trained by using the labeled point cloud data training set, and the two-dimensional perception network model is trained by using the labeled image data training set. The network loss function includes a classification loss and a regression loss.

[0125] S330, late fusion of network results.

[0126] After the three-dimensional perception network model and the two-dimensional perception network model are trained, the results of the two network models can be fused. The late fusion process is as shown in Figure 4

[0127] First, the time-synchronized initial point cloud data and initial image data are input into the three-dimensional perception network model and the two-dimensional perception network model, respectively, to obtain the corresponding initial three-dimensional target box and initial two-dimensional target box. Then, according to the camera calibration, the conversion relationship from the point cloud coordinate system to the image coordinate system is obtained, that is, the point cloud to image mapping table.

[0128] Further, the initial three-dimensional target box and the initial two-dimensional target box are associated. Specifically, the eight corner points of the initial three-dimensional target box are converted to the image coordinate system through the conversion relationship from the point cloud coordinate system to the image coordinate system, to obtain the two-dimensional data of the corner points. Then, according to the maximum horizontal coordinate (maximum horizontal axis), the minimum horizontal coordinate (minimum horizontal axis), the maximum vertical coordinate (maximum vertical axis), and the minimum vertical coordinate (minimum vertical axis) in the two-dimensional data of the corner points, a conversion two-dimensional target box corresponding to the initial three-dimensional target box is constructed. The overlap degree between each initial two-dimensional target box and each conversion two-dimensional target box is calculated. If the overlap degree meets a preset condition, it means that the two are associated, and the detection target is perceived in both the image and the point cloud. The initial two-dimensional target box that is not associated is taken as a to-be-processed two-dimensional target box.

[0129] For the to-be-processed two-dimensional target box, a block depth search of the to-be-processed two-dimensional target box is performed. Specifically, as shown in Figure 5 Figure 6 ​​The center point-based search mode shown performs a deep search, that is, first, the depth position information of the grid center point position is searched, and then the search radius is gradually expanded, and each radius is searched clockwise until the corresponding position has depth position information.

[0130] In order to obtain a more accurate to-be-processed three-dimensional target frame, the point cloud points (target conversion points) projected into the to-be-processed two-dimensional target frame are subjected to l-shape re-clustering. As shown in Figure 7 As shown, first, the point with the largest slope (the first conversion point) and the point with the smallest slope (the second conversion point) are determined. Further, a third conversion point farthest from the straight line of the two points is found, and the three points form an L shape, so that the size, position and other information of the detected target can be recovered, and the to-be-processed three-dimensional target frame is obtained. In actual application, constraints are also made according to the size prior of different categories.

[0131] Finally, the initial three-dimensional target frame and the to-be-processed three-dimensional target frame are fused to obtain a fused three-dimensional target frame, which is used to indicate the three-dimensional information of all detected targets.

[0132] Figure 8 FIG. 1 is a structural schematic diagram of a target detection device according to an embodiment of the present disclosure. As shown in Figure 8 The device includes an initial target frame determination module 410, a converted two-dimensional target frame determination module 420, a to-be-processed two-dimensional target frame determination module 430, a block processing module 440, a to-be-processed three-dimensional target frame recovery module 450 and a fused three-dimensional target frame determination module 460.

[0133] The initial target frame determination module 410 is configured to determine at least one initial three-dimensional target frame according to initial point cloud data, and determine at least one initial two-dimensional target frame according to initial image data that is time-synchronized with the initial point cloud data. The converted two-dimensional target frame determination module 420 is configured to determine a converted two-dimensional target frame corresponding to the initial three-dimensional target frame in an image coordinate system. The to-be-processed two-dimensional target frame determination module 430 is configured to associate the initial two-dimensional target frame with the converted two-dimensional target frame, and determine an initial two-dimensional target frame that fails in association as a to-be-processed two-dimensional target frame. The block processing module 440 is configured to perform block processing on the to-be-processed two-dimensional target frame to obtain a plurality of sub-blocks. The to-be-processed three-dimensional target frame recovery module 450 is configured to perform depth search on each sub-block to recover a to-be-processed three-dimensional target frame associated with the to-be-processed two-dimensional target frame. The fused three-dimensional target frame determination module 460 is configured to determine a fused three-dimensional target frame corresponding to the initial point cloud data according to the initial three-dimensional target frame and the to-be-processed three-dimensional target frame, wherein the fused three-dimensional target frame is used to indicate three-dimensional information of a detected target.

[0134] The target detection device provided by the embodiment performs target recognition on initial point cloud data and initial image data respectively, obtains an initial three-dimensional target frame and an initial two-dimensional target frame, further converts the initial three-dimensional target frame to an image coordinate system to obtain a converted two-dimensional target frame, and associates the initial two-dimensional target frame with the converted two-dimensional target frame, so that the initial two-dimensional target frame that fails in association is determined as a two-dimensional target frame to be processed, so as to perform subsequent three-dimensional recovery. Further, the two-dimensional target frame to be processed is processed in blocks to obtain a plurality of sub-blocks, and each sub-block is subjected to a depth search to recover a three-dimensional target frame to be processed associated with the two-dimensional target frame to be processed. Further, the initial three-dimensional target frame is fused with the three-dimensional target frame to be processed to obtain a fused three-dimensional target frame, so as to fuse the image recognition result and the point cloud recognition result. The problem that the point cloud recognition result and the image recognition result are prone to interference and have poor robustness when being fused is solved, accurate fusion of the point cloud target frame and the image target frame is achieved, and the robustness of target detection is improved.

[0135] On the basis of the above-mentioned embodiment, the converted two-dimensional target frame determination module 420 is further configured to determine, for each initial three-dimensional target frame, eight corner points of the initial three-dimensional target frame, determine, for each corner point, corner point two-dimensional data of the corner point in the image coordinate system, and determine, according to the corner point two-dimensional data, a converted two-dimensional target frame corresponding to the initial three-dimensional target frame.

[0136] On the basis of the above-mentioned embodiment, the converted two-dimensional target frame determination module 420 is further configured to determine, according to the corner point two-dimensional data, a horizontal axis maximum value, a horizontal axis minimum value, a vertical axis maximum value, and a vertical axis minimum value, and determine, according to the horizontal axis maximum value, the horizontal axis minimum value, the vertical axis maximum value, and the vertical axis minimum value, a converted two-dimensional target frame corresponding to the initial three-dimensional target frame.

[0137] On the basis of the above-mentioned embodiment, the two-dimensional target frame to be processed determination module 430 is further configured to determine, for each initial two-dimensional target frame, an overlap degree between the initial two-dimensional target frame and each converted two-dimensional target frame, determine that the initial two-dimensional target frame is successfully associated if the overlap degree is greater than a preset overlap threshold, and determine that the initial two-dimensional target frame fails in association and the initial two-dimensional target frame is the two-dimensional target frame to be processed if the overlap degrees are all not greater than the preset overlap threshold.

[0138] On the basis of the above-mentioned embodiment, the three-dimensional target frame to be processed recovery module 450 is further configured to determine, for each sub-block, a grid center point corresponding to the sub-block, determine, according to the grid center point and a preset search parameter, a target conversion point corresponding to the sub-block, and determine, according to the target conversion points, a three-dimensional target frame to be processed corresponding to the two-dimensional target frame to be processed.

[0139] On the basis of the above-mentioned embodiment, optionally, the three-dimensional target box to be processed recovery module 450 is further configured to, for each point cloud point in the initial point cloud data, determine a point cloud conversion point of the point cloud point in an image coordinate system; take the grid center point as a point to be matched; determine a matching result according to the point to be matched and each point cloud conversion point; if the matching result is a matching success, take the matched point cloud conversion point as a target conversion point corresponding to the sub-block; if the matching result is a matching failure, update the point to be matched according to the point to be matched and the preset search parameter, and return to perform the operation of determining a matching result according to the point to be matched and each point cloud conversion point.

[0140] On the basis of the above-mentioned embodiment, optionally, the three-dimensional target box to be processed recovery module 450 is further configured to, for each target conversion point, determine a slope to be compared of the target conversion point according to first coordinate information and second coordinate information of the target conversion point; determine a maximum slope value and a minimum slope value according to each slope to be compared, and take a target conversion point corresponding to the maximum slope value as a first conversion point and a target conversion point corresponding to the minimum slope value as a second conversion point; determine a third conversion point according to each target conversion point, the first conversion point and the second conversion point; and determine a three-dimensional target box to be processed corresponding to the two-dimensional target box to be processed according to the first conversion point, the second conversion point and the third conversion point.

[0141] The target detection apparatus provided by the embodiments of the present disclosure can execute the steps in the target detection method provided by the method embodiments of the present disclosure, and has the execution steps and beneficial effects which will not be repeated here.

[0142] Figure 9 FIG. 1 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. Figure 9 FIG. 1 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. Figure 9 The electronic device shown is merely an example, and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.

[0143] As shown in FIG. 1, the electronic device 500 can include a processor 510, a memory 520 and a communication interface 530. Figure 9As shown, the electronic device 500 can include a processing device (e.g., a central processor, a graphics processor, etc.) 501 that can perform various appropriate actions and processes to implement the methods of embodiments as described in the present disclosure according to programs stored in a read-only memory (ROM) 502 or loaded into a random access memory (RAM) 503 from a storage device 508. Various programs and data required by the electronic device 500 for operation are also stored in the RAM 503. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.

[0144] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts, thereby implementing the object detection method as described above. In such embodiments, the computer program can be downloaded and installed from a network through a communication device 509, or installed from the storage device 508, or installed from the ROM 502. When the computer program is executed by the processing device 501, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.

[0145] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or can exist separately without being assembled into the electronic device. The above-mentioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to:

[0146] According to the initial point cloud data, determine at least one initial three-dimensional target frame, and according to initial image data that is time-synchronized with the initial point cloud data, determine at least one initial two-dimensional target frame;

[0147] Determine a converted two-dimensional target frame corresponding to the initial three-dimensional target frame in an image coordinate system;

[0148] Associate the initial two-dimensional target frame with the converted two-dimensional target frame, and determine an initial two-dimensional target frame that fails in association as a two-dimensional target frame to be processed;

[0149] Perform block processing on the two-dimensional target frame to be processed to obtain a plurality of sub-blocks;

[0150] Respectively perform a depth search on each sub-block to recover a three-dimensional target frame to be processed associated with the two-dimensional target frame to be processed;

[0151] According to the initial three-dimensional target frame and the three-dimensional target frame to be processed, a fusion three-dimensional target frame corresponding to the initial point cloud data is determined, wherein the fusion three-dimensional target frame is used to indicate three-dimensional information of the detected target.

[0152] Optionally, when the one or more programs are executed by the electronic device, the electronic device can further perform other steps described in the above embodiments.

[0153] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above.

[0154] The above description is merely preferred embodiments of the present disclosure and a description of the principles of the technology applied. It should be understood by those skilled in the art that the disclosed scope of the present disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by any combinations of the above technical features or their equivalent features without departing from the above disclosed concept. For example, the above features are replaced with the technical features disclosed in the present disclosure (but not limited to) having similar functions to form technical solutions.

Claims

1. A target detection method characterized by, The method comprises: According to the initial point cloud data, at least one initial three-dimensional target frame is determined, and according to the initial image data time-synchronized with the initial point cloud data, at least one initial two-dimensional target frame is determined; Determine the corresponding conversion two-dimensional target frame of the initial three-dimensional target frame under the image coordinate system; Correlate the initial two-dimensional target frame with the conversion two-dimensional target frame, and determine the initial two-dimensional target frame that fails to be associated as a to-be-processed two-dimensional target frame; The to-be-processed two-dimensional target frame is processed by blocking to obtain a plurality of sub-blocks; For each sub-block, determine the grid center point corresponding to the sub-block; According to the grid center point and the preset search parameter, a target conversion point corresponding to the sub-block is determined; According to each target conversion point, a to-be-processed three-dimensional target frame corresponding to the to-be-processed two-dimensional target frame is determined; According to the initial three-dimensional target frame and the to-be-processed three-dimensional target frame, a fusion three-dimensional target frame corresponding to the initial point cloud data is determined, wherein the fusion three-dimensional target frame is used to indicate the three-dimensional information of the detected target.

2. The method of claim 1, wherein, The determination of the conversion two-dimensional target frame corresponding to the initial three-dimensional target frame under the image coordinate system comprises: For each initial three-dimensional target frame, eight corner points of the initial three-dimensional target frame are determined; For each corner point, corner point two-dimensional data of the corner point under the image coordinate system is determined; According to the corner point two-dimensional data, a conversion two-dimensional target frame corresponding to the initial three-dimensional target frame is determined.

3. The method of claim 2, wherein, The determination of the conversion two-dimensional target frame corresponding to the initial three-dimensional target frame according to the corner point two-dimensional data comprises: According to the corner point two-dimensional data, the maximum value of the horizontal axis, the minimum value of the horizontal axis, the maximum value of the vertical axis and the minimum value of the vertical axis are determined; According to the maximum value of the horizontal axis, the minimum value of the horizontal axis, the maximum value of the vertical axis and the minimum value of the vertical axis, a conversion two-dimensional target frame corresponding to the initial three-dimensional target frame is determined.

4. The method of claim 1, wherein, The correlation of the initial two-dimensional target frame with the conversion two-dimensional target frame to determine the initial two-dimensional target frame that fails to be associated as a to-be-processed two-dimensional target frame comprises: For each initial two-dimensional target frame, the degree of overlap between the initial two-dimensional target frame and each conversion two-dimensional target frame is determined; If the degree of overlap is greater than a preset overlap threshold, it is determined that the initial two-dimensional target frame is successfully associated; If none of the degrees of overlap is greater than the preset overlap threshold, it is determined that the initial two-dimensional target frame fails to be associated, and the initial two-dimensional target frame is determined as a to-be-processed two-dimensional target frame.

5. The method of claim 1, wherein, The determination of the target conversion point corresponding to the sub-block according to the grid center point and the preset search parameter comprises: For each point cloud point in the initial point cloud data, a point cloud conversion point of the point cloud point under the image coordinate system is determined; The grid center point is taken as a to-be-matched point; According to the to-be-matched point and each point cloud conversion point, a matching result is determined; If the matching result is a matching success, the matched point cloud conversion point is taken as the target conversion point corresponding to the sub-block; If the matching result is a matching failure, the matching point is updated according to the matching point and the preset search parameter, and the operation of determining a matching result according to the matching point and each point cloud conversion point is performed again.

6. The method of claim 1, wherein, The operation of determining a to-be-processed three-dimensional target frame corresponding to the to-be-processed two-dimensional target frame according to each target conversion point comprises the following operations. For each target conversion point, a to-be-compared slope of the target conversion point is determined according to the first coordinate information and the second coordinate information of the target conversion point. A slope maximum value and a slope minimum value are determined according to each to-be-compared slope, a target conversion point corresponding to the slope maximum value is taken as a first conversion point, and a target conversion point corresponding to the slope minimum value is taken as a second conversion point. A third conversion point is determined according to each target conversion point, the first conversion point, and the second conversion point. A to-be-processed three-dimensional target frame corresponding to the to-be-processed two-dimensional target frame is determined according to the first conversion point, the second conversion point, and the third conversion point.

7. A target detection apparatus characterized by comprising: The operation of determining a to-be-processed three-dimensional target frame corresponding to the to-be-processed two-dimensional target frame according to each target conversion point comprises the following operations. An initial target frame determination module is configured to determine at least one initial three-dimensional target frame according to initial point cloud data, and determine at least one initial two-dimensional target frame according to initial image data that is time-synchronized with the initial point cloud data. A conversion two-dimensional target frame determination module is configured to determine a conversion two-dimensional target frame corresponding to the initial three-dimensional target frame in an image coordinate system. A to-be-processed two-dimensional target frame determination module is configured to associate the initial two-dimensional target frame with the conversion two-dimensional target frame, and determine an initial two-dimensional target frame that fails in association as a to-be-processed two-dimensional target frame. A block processing module is configured to perform block processing on the to-be-processed two-dimensional target frame to obtain a plurality of sub-blocks. A to-be-processed three-dimensional target frame recovery module is configured to determine, for each sub-block, a grid center point corresponding to the sub-block, determine a target conversion point corresponding to the sub-block according to the grid center point and a preset search parameter, and determine a to-be-processed three-dimensional target frame corresponding to the to-be-processed two-dimensional target frame according to each target conversion point. A fusion three-dimensional target frame determination module is configured to determine a fusion three-dimensional target frame corresponding to the initial point cloud data according to the initial three-dimensional target frame and the to-be-processed three-dimensional target frame, wherein the fusion three-dimensional target frame is used to indicate three-dimensional information of a detected target.

8. An electronic device, comprising: The electronic device comprises: one or more processors; a storage device configured to store one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the target detection method according to any one of claims 1-6.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, The program is executed by the processor to implement the target detection method according to any one of claims 1-6.

Citation Information

Patent Citations

  • System and method for detecting position information of vehicles around target vehicle

    CN109270543A

  • Three-dimensional object detection method and device, computer equipment and storage medium

    CN111709923A

  • Target determination method and device, equipment and computer readable storage medium

    CN112633258A