Target detection post-processing method and device, equipment and storage medium

By calculating the occlusion ratio and two screening methods of the detected target detection frames in traffic scenes, the problem of the impact of vehicle normal driving caused by misdetection in traditional target detection methods is solved, and more accurate target detection results are achieved.

CN120014598APending Publication Date: 2025-05-16ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510091984.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

In traffic scenarios, traditional target detection methods may detect some targets that originally had no effect on normal driving, resulting in false detection, which in turn affects the normal driving of the vehicle.

Method used

By performing object detection on the input image, the occlusion ratio of each detection box is calculated and filtered twice. The first filtering filters out the detection boxes that are obscured by other detection boxes according to the occlusion ratio, and the second filtering filters out the detection boxes that are occasionally falsely detected based on the set of historical frame detection boxes.

Benefits of technology

It effectively avoids the impact of mis-checking targets on the normal driving of the vehicle and reduces emergency braking situations caused by mischecking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014598A_ABST
    Figure CN120014598A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection post-processing method and device, equipment and a storage medium, and relates to the field of target detection, and the method comprises the steps: carrying out the target detection of an input image, and obtaining an original detection frame set, the input image being a traffic scene image shot by a vehicle-mounted camera; determining the shielding proportion of a single detection frame in the original detection frame set, wherein the shielding proportion represents the shielding degree of a target object contained in the single detection frame by other objects; performing primary screening on the original detection frame set according to the shielding proportion to obtain a first target detection frame set; and performing secondary screening on the original detection frame set according to a historical frame detection frame set corresponding to the original detection frame set to obtain a second target detection frame set. According to the method, post-processing operation of two times of screening is carried out on the original detection frame set detected in the traffic scene, so that the situation that the vehicle is braked emergently due to the fact that the false detection target which does not affect normal driving of the vehicle intrudes into the driving track of the vehicle is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of target detection, and in particular to a target detection post-processing method, device, equipment and storage medium. Background Art

[0002] In the conventional object detection process, all objects in the scene are usually detected as much as possible to reduce the missed detection rate of object detection. However, in traffic scenes, if the focus is still on reducing the missed detection rate, some objects that originally have no impact on normal driving may be detected. Once these objects are mistakenly identified as objects with collision risks, they will affect the normal driving of the vehicle.

[0003] For example, at a certain moment, vehicle A is driving normally, and vehicle B is blocked by non-motor vehicles, pillars, etc. Therefore, at this moment, vehicle B will not cross the obstruction to affect the normal driving of vehicle A. However, if vehicle B is still used as the detection target at this time, vehicle A may mistakenly detect vehicle B as a target that is about to invade the driving track of vehicle A, causing vehicle A to take emergency braking, which in turn affects the normal driving of vehicle A. Therefore, the industry is in urgent need of a target detection post-processing method that can avoid the impact of misdetected targets on the normal driving of vehicles. Summary of the invention

[0004] The main purpose of the present application is to provide a target detection post-processing method, device, equipment and storage medium, aiming to provide a target detection post-processing method that can avoid the impact of false detection of targets on the normal driving of the vehicle.

[0005] To achieve the above object, the present application provides a target detection post-processing method, the method comprising the following steps:

[0006] Performing target detection on an input image to obtain a set of original detection frames, wherein the input image is a traffic scene image taken by a vehicle-mounted camera;

[0007] Determine an occlusion ratio of a single detection frame in the original detection frame set, where the occlusion ratio indicates the degree to which a target object contained in the single detection frame is occluded by other objects;

[0008] Screening the original detection frame set once according to the occlusion ratio to obtain a first target detection frame set;

[0009] The original detection frame set is secondary screened according to the historical frame detection frame set corresponding to the original detection frame set to obtain a second target detection frame set.

[0010] In one embodiment, the step of determining the occlusion ratio of a single detection frame in the original detection frame set includes:

[0011] Selecting a high confidence detection frame from the original detection frame set;

[0012] Determine a first point coordinate corresponding to the high-confidence detection frame, a second point coordinate corresponding to a single detection frame in the original detection frame set, and a third point coordinate corresponding to the vehicle-mounted camera;

[0013] Calculating a first distance between the high-confidence detection frame and the vehicle-mounted camera according to the first point coordinates and the third point coordinates, and calculating a second distance between the single detection frame and the vehicle-mounted camera according to the second point coordinates and the third point coordinates;

[0014] The occlusion ratio of the single detection frame is determined based on the first distance and the second distance.

[0015] In one embodiment, the step of determining the occlusion ratio of the single detection frame based on the first distance and the second distance includes:

[0016] If the sum of the first distance and the preset distance threshold is less than the second distance, determining the occlusion ratio of the single detection frame according to the occupation angles respectively corresponding to the high confidence detection frame and the single detection frame, wherein the occupation angle represents the angle range occupied by the detection frame in the field of view of the vehicle-mounted camera;

[0017] If the sum of the first distance and the preset distance threshold is greater than or equal to the second distance, it is determined that the occlusion ratio of the single detection frame is a preset minimum ratio.

[0018] In one embodiment, the step of determining the occlusion ratio of the single detection frame according to the occupation angles respectively corresponding to the high confidence detection frame and the single detection frame comprises:

[0019] Calculating a first occupancy angle corresponding to the high-confidence detection frame and a second occupancy angle corresponding to the single detection frame;

[0020] determining an angle overlap interval between the high-confidence detection frame and the single detection frame according to the first occupation angle and the second occupation angle;

[0021] The occlusion ratio of the single detection frame is calculated based on the angle overlap interval.

[0022] In one embodiment, the step of screening the original detection frame set according to the occlusion ratio to obtain the first target detection frame set includes:

[0023] Updating the initial confidence of the single detection frame according to the occlusion ratio to obtain an updated confidence;

[0024] The original detection frame set is screened based on the updated confidence to obtain a first target detection frame set.

[0025] In one embodiment, the step of performing secondary screening on the original detection frame set according to the historical frame detection frame set corresponding to the original detection frame set to obtain the second target detection frame set includes:

[0026] Selecting an occluded detection frame set from the original detection frame set, the occluded detection frame set being a set of detection frames whose occlusion ratio is greater than a preset ratio threshold;

[0027] Calculate a maximum intersection-over-union ratio set between any detection frame in the historical frame detection frame set corresponding to the original detection frame set and any detection frame in the obscured detection frame set;

[0028] The original detection frame set is screened for a second time based on the maximum intersection-over-union set to obtain a second target detection frame set.

[0029] In addition, to achieve the above-mentioned purpose, the present application also proposes a target detection post-processing device, the target detection post-processing device comprising:

[0030] Demand analysis module, used to select the target driving area from the map according to the user's driving needs;

[0031] a layer acquisition module, used to determine a first layer and a second layer corresponding to the target driving area, wherein the first layer is used to reflect the user takeover status of the vehicle in the target driving area when the driving assistance function is executed, and the second layer is used to reflect the usability status of the memory driving function of the vehicle in the target driving area;

[0032] A route recommendation module is used to generate a recommended driving route for the user in the target driving area based on the first layer and the second layer.

[0033] In addition, to achieve the above-mentioned objectives, the present application also proposes a target detection post-processing device, which includes: a memory, a processor, and a target detection post-processing program stored on the memory and executable on the processor, wherein the target detection post-processing program is configured to implement the steps of the target detection post-processing method described above.

[0034] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a target detection post-processing program is stored on the storage medium. When the target detection post-processing program is executed by the processor, the steps of the target detection post-processing method described above are implemented.

[0035] In addition, to achieve the above-mentioned purpose, the present invention also provides a computer program product, which includes a target detection post-processing program, and when the target detection post-processing program is executed by a processor, it implements the steps of the target detection post-processing method described above.

[0036] The present application performs target detection on an input image to obtain an original detection frame set, wherein the input image is an image of a traffic scene taken by a vehicle-mounted camera; determines the occlusion ratio of a single detection frame in the original detection frame set, wherein the occlusion ratio indicates the degree to which the target object contained in the single detection frame is occluded by other objects; screens the original detection frame set once according to the occlusion ratio to obtain a first target detection frame set; and screens the original detection frame set twice according to a historical frame detection frame set corresponding to the original detection frame set to obtain a second target detection frame set. Compared with the traditional target detection method, since the above method of the present application performs two post-processing operations of screening the original detection frame set detected in the traffic scene, the first time, based on the occlusion ratio of a single detection frame in the original detection frame set, the detection frame that is occluded to a certain extent by other detection frames is screened out, and the second time, based on the historical frame detection frame set, the detection frame that is occasionally misdetected is screened out, thereby avoiding the situation where the misdetected target that originally does not affect the normal driving of the vehicle invades the vehicle's driving trajectory to cause the vehicle to brake urgently. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 A schematic diagram of the structure of a target detection post-processing device in a hardware operating environment involved in an embodiment of the present application;

[0038] Figure 2 This is a flow chart of the first embodiment of the target detection post-processing method of the present application;

[0039] Figure 3 This is a flow chart of the second embodiment of the target detection post-processing method of the present application;

[0040] Figure 4 This is a schematic diagram of the occlusion of the target detection post-processing method of this application;

[0041] Figure 5 A schematic diagram of overlapping angle merging of the target detection post-processing method of this application;

[0042] Figure 6 This is a schematic diagram of occlusion ratio calculation of the target detection post-processing method of this application;

[0043] Figure 7 This is a flow chart of the third embodiment of the target detection post-processing method of the present application;

[0044] Figure 8 This is a structural block diagram of the first embodiment of the target detection post-processing device of the present application.

[0045] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0046] It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0047] Reference Figure 1 , Figure 1 This is a schematic diagram of the structure of the target detection post-processing device in the hardware operating environment involved in the embodiment of the present application.

[0048] like Figure 1 As shown, the target detection post-processing device may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. Among them, the communication bus 1002 is used to realize the connection and communication between these components. The user interface 1003 may include a display screen (Display), an input unit such as a keyboard (Keyboard), and the optional user interface 1003 may also include a standard wired interface and a wireless interface. The network interface 1004 may optionally include a standard wired interface and a wireless interface (such as a wireless fidelity (Wireless-Fidelity, Wi-Fi) interface). The memory 1005 may be a high-speed random access memory (Random Access Memory, RAM), or a stable non-volatile memory (Non-Volatile Memory, NVM), such as a disk storage. The memory 1005 may also be a storage device independent of the aforementioned processor 1001.

[0049] Those skilled in the art will understand that Figure 1 The structure shown in the figure does not constitute a limitation on the target detection post-processing device, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.

[0050] like Figure 1 As shown, the memory 1005 as a storage medium may include an operating system, a network communication module, a user interface module, and a target detection post-processing program.

[0051] exist Figure 1In the target detection post-processing device shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and the memory 1005 in the target detection post-processing device of the present application can be set in the target detection post-processing device, and the target detection post-processing device calls the target detection post-processing program stored in the memory 1005 through the processor 1001, and executes the target detection post-processing method provided in the embodiment of the present application.

[0052] The target detection post-processing method provided in the present application can be deployed and executed on a computer device, which can be a vehicle-mounted device (such as a vehicle-mounted controller, a vehicle-mounted processor or other vehicle-mounted units, etc.), or a terminal outside the vehicle, an independent server, a cloud, a server cluster, a distributed system, an Internet of Things device or an Internet of Vehicles device, etc. This embodiment is not limited to this, and those skilled in the art can set up a computer device to implement the target detection post-processing method according to actual needs.

[0053] The present application embodiment provides a target detection post-processing method, referring to Figure 2 , Figure 2 This is a flow chart of the first embodiment of the target detection post-processing method of the present application.

[0054] In this embodiment, the target detection post-processing method includes the following steps:

[0055] Step S1: performing target detection on an input image to obtain a set of original detection frames, wherein the input image is a traffic scene image taken by a vehicle-mounted camera.

[0056] It is understandable that the above input image may contain any elements that may exist in the traffic scene, such as motor vehicles, non-motor vehicles, pedestrians, pillars, fences, etc., and the above original detection frame set is a set of several detection frames that contain any of the above elements. Exemplarily, the above original detection frame set can be expressed as Input = {{x, y, z, w, h, l, sita, conf}, {}, {}…}, where each detection frame contains a center point (x, y, z), a detection size (w, h, l), a detection frame heading angle sita, and a detection frame confidence conf.

[0057] It should be understood that the above-mentioned target detection can be BEV (Bird's Eye View) target detection, and the corresponding specific process can be described as: feature extraction is performed on the above-mentioned input image through the network model to obtain a feature map of 1*C*W*H, where 1 represents the batch size, C represents the number of channels, W represents the width, and H represents the height; the depth of each pixel value of 1*C*W*H obtained by depth estimation is then passed through the LSS algorithm (a 3D target detection algorithm that constructs bird's-eye view features from the bottom up) to obtain a bird's-eye view, and then further feature extraction target detection Head (output layer) and the preset anchor (anchor frame) are performed, and finally decoding is performed to obtain a single detection frame in the above-mentioned original detection frame set.

[0058] Step S2: determining the occlusion ratio of a single detection frame in the original detection frame set, where the occlusion ratio indicates the degree to which the target object contained in the single detection frame is occluded by other objects.

[0059] In a specific implementation, the specific parameters of the part of any detection frame obscured by other detection frames (such as the angle range and area occupied by the detection frame in the field of view of the vehicle-mounted camera) can be divided by the specific parameters when the detection frame is not obscured. The quotient obtained can be used to represent the degree to which the target object contained in a single detection frame is obscured by other objects, that is, the above-mentioned occlusion ratio.

[0060] Step S3: screening the original detection frame set according to the occlusion ratio to obtain a first target detection frame set.

[0061] It is understandable that in the process of target detection in traffic scenes, the occluded target is allowed to be missed to a certain extent, because it is not the nearest target and there is no risk of collision. However, false detection should be avoided as much as possible, because false detection may cause the occluded target to have a very large expected speed, thereby invading the driving trajectory of the normal vehicle and causing the normal vehicle to take emergency braking measures. More importantly, since the full picture of the occluded target cannot be seen clearly, it is more prone to false detection. Among them, the occluded target refers to a target that is blocked to a certain extent.

[0062] Based on this, this embodiment proposes to filter the original detection frame set according to the occlusion ratio of each detection frame in the original detection frame set, thereby filtering out the detection frames in the original detection frame set that are occluded to a certain extent (which can be defined according to the actual scene, such as 70%), and obtaining the above-mentioned first target detection frame set.

[0063] Step S4: performing a secondary screening on the original detection frame set according to the historical frame detection frame set corresponding to the original detection frame set to obtain a second target detection frame set.

[0064] It should be noted that the above historical frame detection frame set refers to the detection frame set corresponding to the original detection frame set in the historical frame (such as the previous frame).

[0065] In the specific implementation, due to the existence of certain occasional missed detections in target detection, for example, there are large differences between the detection frame of the previous frame and the detection frame of the next frame, these differences may be caused by occasional missed detections. At the same time, since there may be certain missed detections of occluded targets in traffic scenes, a multi-frame continuous confirmation strategy can be adopted for occluded targets to eliminate the negative impact of occasional missed detections. Specifically, based on the multi-frame continuous confirmation strategy, all occluded detection frames can be selected from the above-mentioned original detection frame set, and then these occluded detection frames are secondary screened through the historical frame detection frame set, thereby filtering out the occluded detection frames in the original detection frame set that have large differences between the current frame and the historical frame, and obtaining the above-mentioned second target detection frame set.

[0066] It is understandable that all the detection frames included in the first target detection frame set and the second target detection frame set are the detection frames that the target detection post-processing method described in this embodiment ultimately needs to obtain. Compared with the detection frames in the original detection frame set, the detection frames in the first target detection frame set and the second target detection frame set increase the missed detection of occluded targets, but greatly eliminate false detections, which can prevent the target from generating a large predicted speed to invade the vehicle's driving trajectory, and are therefore more suitable for traffic scenes.

[0067] This embodiment performs target detection on an input image to obtain an original detection frame set, wherein the input image is an image of a traffic scene taken by a vehicle-mounted camera; determines the occlusion ratio of a single detection frame in the original detection frame set, wherein the occlusion ratio indicates the degree to which the target object contained in the single detection frame is occluded by other objects; screens the original detection frame set once according to the occlusion ratio to obtain a first target detection frame set; and screens the original detection frame set twice according to a historical frame detection frame set corresponding to the original detection frame set to obtain a second target detection frame set. Compared with the traditional target detection method, since the above method of this embodiment performs two post-processing operations of screening the original detection frame set detected in the traffic scene, the first time, based on the occlusion ratio of a single detection frame in the original detection frame set, the detection frame that is occluded to a certain extent by other detection frames is screened out, and the second time, based on the historical frame detection frame set, the detection frame that is occasionally misdetected is screened out, thereby avoiding the situation where the misdetected target that originally does not affect the normal driving of the vehicle invades the vehicle's driving trajectory to cause the vehicle to brake urgently.

[0068] refer to Figure 3 , Figure 3 This is a flow chart of the second embodiment of the target detection post-processing method of the present application.

[0069] In a feasible implementation manner, the step S2 may include:

[0070] Step S21: Select a high-confidence detection frame from the original detection frame set.

[0071] It should be noted that the above-mentioned high-confidence detection frame may refer to a detection frame whose confidence is greater than a preset confidence threshold (eg, 0.7).

[0072] Step S22: Determine the first point coordinates corresponding to the high-confidence detection frame, the second point coordinates corresponding to a single detection frame in the original detection frame set, and the third point coordinates corresponding to the vehicle-mounted camera.

[0073] It should be noted that the first point coordinates, the second point coordinates and the third point coordinates are all in the same reference coordinate system, for example, a vehicle coordinate system with the center point of the vehicle where the vehicle-mounted camera is located as the coordinate origin.

[0074] Step S23: Calculate the first distance between the high confidence detection frame and the vehicle-mounted camera according to the first point coordinates and the third point coordinates, and calculate the second distance between the single detection frame and the vehicle-mounted camera according to the second point coordinates and the third point coordinates.

[0075] It is understandable that, since the detection frame is a rectangle, the detection frame generally has four points (i.e., the four vertices of the rectangle). Based on this, the first point coordinates mentioned above may refer to the coordinates of the point closest to the vehicle camera among the four points of the high-confidence detection frame in the vehicle coordinate system, and the second point coordinates mentioned above may refer to the coordinates of the point closest to the vehicle camera among the four points of a single detection frame in the original detection frame set in the vehicle coordinate system.

[0076] Exemplarily, assuming that the coordinates of the first point or the second point are (X1, Y1), and the coordinates of the third point are (XC, YC), then the first distance or the second distance can be calculated based on the following formula 1: dis=sqrt((X1-XC)*(X1-XC)+(Y1-YC)*(Y1-YC)), where dis represents the first distance or the second distance.

[0077] Step S24: determining the occlusion ratio of the single detection frame based on the first distance and the second distance.

[0078] In a specific implementation, a first specific parameter corresponding to the portion of a single detection frame obscured by other detection frames (such as the angle range and area occupied by the detection frame in the field of view of the vehicle-mounted camera) can be determined based on the first distance and the second distance, and then the first specific parameter is compared with the second specific parameter when the single detection frame is not obscured to obtain the above-mentioned occlusion ratio.

[0079] In a feasible implementation manner, step S24 may be implemented by the following steps S241 and S242:

[0080] Step S241: If the sum of the first distance and the preset distance threshold is less than the second distance, the occupancy ratio of the single detection frame is determined according to the occupancy angles corresponding to the high-confidence detection frame and the single detection frame respectively, and the occupancy angle represents the angle range occupied by the detection frame in the field of view of the vehicle-mounted camera.

[0081] It should be noted that the above preset distance threshold can be set according to the size of the target object in the actual scene, for example, 0.3m.

[0082] Step S242: If the sum of the first distance and the preset distance threshold is greater than or equal to the second distance, it is determined that the occlusion ratio of the single detection frame is a preset minimum ratio.

[0083] In a specific implementation, if the sum of the above-mentioned first distance and the preset distance threshold is less than the above-mentioned second distance, it can be indicated that the above-mentioned single detection frame is blocked by the above-mentioned high-confidence detection frame, and therefore the occlusion ratio of the single detection frame can be determined according to the angle ranges respectively occupied by the high-confidence detection frame and the single detection frame in the field of view of the vehicle-mounted camera; if the sum of the above-mentioned first distance and the preset distance threshold is greater than or equal to the above-mentioned second distance, it can be indicated that the above-mentioned single detection frame is not blocked by the above-mentioned high-confidence detection frame (or even if the single detection frame is blocked by the high-confidence detection frame, the detection effect is the same as when there is no occlusion), and therefore it can be determined that the occlusion ratio of the single detection frame at this time is the preset minimum ratio (for example, 1%).

[0084] In a feasible implementation manner, the step S241 may include:

[0085] Step S2411: Calculate a first occupancy angle corresponding to the high-confidence detection frame and a second occupancy angle corresponding to the single detection frame.

[0086] It should be noted that the first occupied angle represents the angle range occupied by the high confidence detection frame in the field of view of the vehicle-mounted camera, and the second occupied angle represents the angle range occupied by the single detection frame in the field of view of the vehicle-mounted camera. Specifically, the occupied angle (including the first occupied angle and the second occupied angle) can be expressed in the form of [a, a+b], where a represents the starting angle at which the edge of the detection frame begins to be captured by the vehicle-mounted camera, and a+b represents the ending angle when the detection frame completely leaves the field of view of the vehicle-mounted camera.

[0087] In the specific implementation, taking the calculation of the first occupied angle corresponding to the high confidence detection box as an example, assuming that the first distance obtained based on the above formula 1 is dis1, and the horizontal coordinates of the four points of the high confidence detection box in the vehicle coordinate system are X1, X2, X3, and X4 respectively, then the above first occupied angle can be calculated based on the following formula:

[0088] angle1=arccos((X1-XC) / dist1);

[0089] angle2=arccos((X2-XC) / dist1);

[0090] angle3=arccos((X3-XC) / dist1);

[0091] angle4=arccos((X4-XC) / dist1);

[0092] a=min(angle1,angle2,angle3,angle4);

[0093] b=max(angle1,angle2,angle3,angle4).

[0094] Among them, angle1, angle2, angle3, and angle4 respectively represent the angles between the straight line between the corresponding points of X1, X2, X3, and X4 and the corresponding points of the vehicle camera and the reference direction (for example, the north direction), a represents the starting angle at which the edge of the high-confidence detection frame begins to be captured by the vehicle camera, and b represents the angle from the starting angle to when the high-confidence detection frame completely leaves the field of view of the vehicle camera. For example, you can refer to Figure 4 , Figure 4 This is a schematic diagram of the occlusion of the target detection post-processing method of this application. Figure 4 It can be seen that B is the detection frame blocked by the non-motor vehicle and the pillar, and B corresponds to the occupancy angle of the camera of the ego vehicle (i.e., the normal driving vehicle) and can be expressed as [a, a+b].

[0095] Step S2412: determining an angle overlap interval between the high-confidence detection frame and the single detection frame according to the first occupancy angle and the second occupancy angle.

[0096] Step S2413: Calculate the occlusion ratio of the single detection frame based on the angle overlap interval.

[0097] It should be understood that, since the single detection frame is blocked by the high-confidence detection frame, there is an angular overlap interval between the single detection frame and the high-confidence detection frame. Figure 5 , Figure 5 This is a schematic diagram of overlapping angle merging of the target detection post-processing method of the present application. Assuming that a single detection box is occluded by three high-confidence detection boxes, the corresponding occluded target set is Occbox1 = {[a1, b1], [a2, b2], [a3, b3]}, where [a1, b1], [a2, b2], [a3, b3] represent the occupancy angles corresponding to the three high-confidence detection boxes (i.e., the first occupancy angles mentioned above), and since [a1, b1] and [a2, b2] have overlapping intervals, they can be merged into [a1, b2]; and [a3, b3] has no overlap with other intervals and can be directly retained. Therefore, the first occupancy angles corresponding to the three high-confidence detection boxes can be merged and output as Occbox1 = {[a1, b2], [a3, b3]}.

[0098] Furthermore, the above-mentioned occlusion ratio can be calculated by the following steps: first, the angle overlap interval between the first occupancy angle and the second occupancy angle is calculated, then the total occlusion angle is calculated according to the angle overlap interval, and finally, the total occlusion angle is divided by the second occupancy angle to obtain the above-mentioned occlusion ratio.

[0099] In order to more clearly present the calculation principle of occlusion ratio, please refer to Figure 6 , Figure 6 This is a schematic diagram of the occlusion ratio calculation of the target detection post-processing method of this application. Figure 6 In , Occbox1 (i.e. [a1, b2], [a3, b3]) represents the first occupied angle corresponding to the high confidence detection box, and box1 (i.e. [a, b]) represents the second occupied angle corresponding to a single detection box. It can be seen that Figure 6 The occlusion ratio ratio in can be calculated based on the following formula:

[0100] ratio=(b2-a) / (ba)+(b-a3) / (ba).

[0101] In this embodiment, a high-confidence detection frame is selected from the original detection frame set; a first point coordinate corresponding to the high-confidence detection frame, a second point coordinate corresponding to a single detection frame in the original detection frame set, and a third point coordinate corresponding to the vehicle-mounted camera are determined; a first distance between the high-confidence detection frame and the vehicle-mounted camera is calculated according to the first point coordinate and the third point coordinate, and a second distance between the single detection frame and the vehicle-mounted camera is calculated according to the second point coordinate and the third point coordinate; if the sum of the first distance and a preset distance threshold is less than the second distance, a first occupancy angle corresponding to the high-confidence detection frame and a second occupancy angle corresponding to the single detection frame are calculated; an angle overlap interval between the high-confidence detection frame and the single detection frame is determined according to the first occupancy angle and the second occupancy angle; an occlusion ratio of the single detection frame is calculated based on the angle overlap interval; if the sum of the first distance and the preset distance threshold is greater than or equal to the second distance, the occlusion ratio of the single detection frame is determined to be a preset minimum ratio. The above method of this embodiment determines the angle overlap interval between the high-confidence detection frame and the single detection frame according to the first occupancy angle corresponding to the high-confidence detection frame and the second occupancy angle corresponding to the single detection frame when the single detection frame is occluded by the high-confidence detection frame, and then calculates the occlusion ratio of the single detection frame based on the angle overlap interval, thereby improving the calculation accuracy of the occlusion ratio, and further providing a data basis for the subsequent step of screening the original detection frame set according to the occlusion ratio.

[0102] refer to Figure 7 , Figure 7 This is a flowchart of the third embodiment of the target detection post-processing method of the present application.

[0103] In a feasible implementation manner, step S3 may be implemented by the following steps S31 and S32:

[0104] Step S31: updating the initial confidence of the single detection frame according to the occlusion ratio to obtain an updated confidence.

[0105] In a specific implementation, the initial confidence of the above single detection box can be updated according to the following formula:

[0106] conf=(conf'–conf_th)*ratio+conf_th.

[0107] Among them, conf represents the above-mentioned updated confidence, conf' represents the above-mentioned initial confidence, conf_th represents the hyperparameter threshold (generally set to about 0.7), and ratio represents the above-mentioned occlusion ratio.

[0108] Step S32: Screening the original detection frame set based on the updated confidence to obtain a first target detection frame set.

[0109] In a specific implementation, the original detection frame set can be screened according to the above-mentioned update confidence. When the update confidence is less than a certain threshold (e.g., 80%), the detection frame corresponding to the update confidence is directly deleted. In addition, after deleting the detection frame whose update confidence is less than a certain threshold, the overlapping detection frames in the original detection frame set can be filtered out by NMS (Non-Maximum Suppression) operation.

[0110] In a feasible implementation manner, step S4 can be implemented by the following steps S41, S42 and S43:

[0111] Step S41: selecting an occluded detection frame set from the original detection frame set, wherein the occluded detection frame set is a set of detection frames whose occlusion ratio is greater than a preset ratio threshold.

[0112] It should be noted that the above preset ratio threshold can be flexibly set according to actual scenarios, for example, it can be set to 70%.

[0113] Step S42: Calculate the maximum intersection-over-union ratio set between any detection frame in the historical frame detection frame set corresponding to the original detection frame set and any detection frame in the occluded detection frame set.

[0114] In a specific implementation, any detection frame in the above-mentioned obscured detection frame set can be converted from the vehicle coordinate system to the world coordinate system according to the positioning signal. Among them, the positioning signal is a 3*3 rotation matrix R and a 1*3 translation matrix T. According to the rotation matrix R, the yaw angle yaw rotated around Z can be obtained. Assuming that the detection frame of the vehicle coordinate system is represented by box1(x, y, z, l, w, h, sita), the converted global coordinate system detection frame can be represented by box2(x1, y1, z1, l, w, h, sita1), and the conversion formula is as follows:

[0115] [x1,y1,z1] T =R*[x,y,z] T +T

[0116] sita1=sita+yaw.

[0117] Based on the same steps, any detection frame in the historical frame detection frame set can be converted as above, and the maximum intersection-and-union ratio set between any detection frame in the historical frame detection frame set and any detection frame in the occluded detection frame set can be calculated based on the converted detection frame. For example, assuming that OccBox (i.e., the occluded detection frame set after conversion) contains {detection frame 1, detection frame 2}, and BoxH (i.e., the historical frame detection frame set after conversion) contains {detection frame A, detection frame B, detection frame C}, firstly, the detection frame 1 is compared with the detection frame A, detection frame B, and detection frame C respectively by the intersection-and-union ratio, and the maximum intersection-and-union ratio corresponding to the detection frame 1 is selected; the maximum intersection-and-union ratio acquisition process of the detection frame 2 is the same as that of the detection frame 1, which will not be repeated here. Therefore, in this example, the maximum intersection-and-union ratio set finally obtained contains 2 maximum intersection-and-union ratios.

[0118] Step S43: performing a secondary screening on the original detection frame set based on the maximum intersection-over-union set to obtain a second target detection frame set.

[0119] In a specific implementation, the maximum IoU contained in the maximum IoU set can be compared with a preset IoU threshold: if the maximum IoU is greater than the preset IoU threshold, the detection frame corresponding to the maximum IoU is retained; if the maximum IoU is less than or equal to the preset IoU threshold, the detection frame corresponding to the maximum IoU is deleted. Finally, the retained detection frame can be inversely transformed according to the conversion formula in step S42 above to transform the world coordinate system corresponding to the retained detection frame into the vehicle coordinate system, and the set of detection frames finally obtained is the second target detection frame set above.

[0120] This embodiment updates the initial confidence of the single detection frame according to the occlusion ratio to obtain an updated confidence; based on the updated confidence, the original detection frame set is screened once to obtain a first target detection frame set; an occluded detection frame set is selected from the original detection frame set, and the occluded detection frame set is a set of detection frames whose occlusion ratio is greater than a preset ratio threshold; the maximum intersection-and-union ratio set between any detection frame in the historical frame detection frame set corresponding to the original detection frame set and any detection frame in the occluded detection frame set is calculated; based on the maximum intersection-and-union ratio set, the original detection frame set is screened twice to obtain a second target detection frame set. Compared with the traditional target detection post-processing method, the above method of this embodiment updates the initial confidence of a single detection frame according to the occlusion ratio, and screens the original detection frame set based on the updated confidence obtained after the update, and at the same time, the original detection frame set is screened for a second time according to the maximum intersection-and-union ratio set between any detection frame in the historical frame detection frame set and any detection frame in the occluded detection frame set, thereby reducing the confidence of the occluded detection frame, and further reducing false detection in the target detection process and the missed detection of other vehicles caused by false detection.

[0121] In addition, an embodiment of the present application further proposes a storage medium, on which a target detection post-processing program is stored. When the target detection post-processing program is executed by a processor, the steps of the target detection post-processing method described above are implemented.

[0122] Reference Figure 8 , Figure 8 This is a structural block diagram of the first embodiment of the target detection post-processing device of the present application.

[0123] like Figure 8 As shown, the target detection post-processing device proposed in the embodiment of the present application includes:

[0124] The target detection module 801 is used to perform target detection on an input image to obtain an original detection frame set, wherein the input image is a traffic scene image taken by a vehicle-mounted camera;

[0125] An occlusion calculation module 802 is used to determine an occlusion ratio of a single detection frame in the original detection frame set, where the occlusion ratio indicates the degree to which the target object contained in the single detection frame is occluded by other objects;

[0126] A primary screening module 803, configured to perform a primary screening on the original detection frame set according to the occlusion ratio to obtain a first target detection frame set;

[0127] The secondary screening module 804 is used to perform secondary screening on the original detection frame set according to the historical frame detection frame set corresponding to the original detection frame set to obtain a second target detection frame set.

[0128] This embodiment performs target detection on an input image to obtain an original detection frame set, wherein the input image is an image of a traffic scene taken by a vehicle-mounted camera; determines the occlusion ratio of a single detection frame in the original detection frame set, wherein the occlusion ratio indicates the degree to which the target object contained in the single detection frame is occluded by other objects; screens the original detection frame set once according to the occlusion ratio to obtain a first target detection frame set; and screens the original detection frame set twice according to a historical frame detection frame set corresponding to the original detection frame set to obtain a second target detection frame set. Compared with the traditional target detection method, since the above method of this embodiment performs two post-processing operations of screening the original detection frame set detected in the traffic scene, the first time, based on the occlusion ratio of a single detection frame in the original detection frame set, the detection frame that is occluded to a certain extent by other detection frames is screened out, and the second time, based on the historical frame detection frame set, the detection frame that is occasionally misdetected is screened out, thereby avoiding the situation where the misdetected target that originally does not affect the normal driving of the vehicle invades the vehicle's driving trajectory to cause the vehicle to brake urgently.

[0129] Based on the first embodiment of the target detection post-processing device of the present application, a second embodiment of the target detection post-processing device of the present application is proposed.

[0130] In this embodiment, the occlusion calculation module 802 is also used to select a high-confidence detection frame from the original detection frame set; determine the first point coordinates corresponding to the high-confidence detection frame, the second point coordinates corresponding to the single detection frame in the original detection frame set, and the third point coordinates corresponding to the vehicle-mounted camera; calculate the first distance between the high-confidence detection frame and the vehicle-mounted camera according to the first point coordinates and the third point coordinates, and calculate the second distance between the single detection frame and the vehicle-mounted camera according to the second point coordinates and the third point coordinates; determine the occlusion ratio of the single detection frame based on the first distance and the second distance.

[0131] Furthermore, the occlusion calculation module 802 is also used to determine the occlusion ratio of the single detection frame according to the occupancy angles corresponding to the high-confidence detection frame and the single detection frame respectively if the sum of the first distance and the preset distance threshold is less than the second distance, and the occupancy angle represents the angular range occupied by the detection frame in the field of view of the vehicle-mounted camera; if the sum of the first distance and the preset distance threshold is greater than or equal to the second distance, then the occlusion ratio of the single detection frame is determined to be a preset minimum ratio.

[0132] Furthermore, the occlusion calculation module 802 is also used to calculate a first occupancy angle corresponding to the high-confidence detection frame and a second occupancy angle corresponding to the single detection frame; determine an angle overlap interval between the high-confidence detection frame and the single detection frame according to the first occupancy angle and the second occupancy angle; and calculate the occlusion ratio of the single detection frame based on the angle overlap interval.

[0133] Furthermore, the primary screening module 803 is also used to update the initial confidence of the single detection frame according to the occlusion ratio to obtain an updated confidence; and perform a primary screening on the original detection frame set based on the updated confidence to obtain a first target detection frame set.

[0134] Furthermore, the secondary screening module 804 is also used to select an occluded detection frame set from the original detection frame set, the occluded detection frame set being a set of detection frames whose occlusion ratio is greater than a preset ratio threshold; calculate the maximum intersection-and-union ratio set between any detection frame in the historical frame detection frame set corresponding to the original detection frame set and any detection frame in the occluded detection frame set; and perform secondary screening on the original detection frame set based on the maximum intersection-and-union ratio set to obtain a second target detection frame set.

[0135] Other embodiments or specific implementation methods of the target detection post-processing device of the present application can refer to the above-mentioned method embodiments and will not be repeated here.

[0136] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or system including the element.

[0137] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0138] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as a read-only memory / random access memory, a disk, or an optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.

[0139] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A target detection post-processing method, characterized in that: The method comprises the following steps: Performing target detection on an input image to obtain a set of original detection frames, wherein the input image is a traffic scene image taken by a vehicle-mounted camera; Determine an occlusion ratio of a single detection frame in the original detection frame set, where the occlusion ratio indicates the degree to which a target object contained in the single detection frame is occluded by other objects; Screening the original detection frame set once according to the occlusion ratio to obtain a first target detection frame set; The original detection frame set is secondary screened according to the historical frame detection frame set corresponding to the original detection frame set to obtain a second target detection frame set.

2. The target detection post-processing method according to claim 1, characterized in that: The step of determining the occlusion ratio of a single detection frame in the original detection frame set includes: Selecting a high confidence detection frame from the original detection frame set; Determine a first point coordinate corresponding to the high-confidence detection frame, a second point coordinate corresponding to a single detection frame in the original detection frame set, and a third point coordinate corresponding to the vehicle-mounted camera; Calculating a first distance between the high-confidence detection frame and the vehicle-mounted camera according to the first point coordinates and the third point coordinates, and calculating a second distance between the single detection frame and the vehicle-mounted camera according to the second point coordinates and the third point coordinates; The occlusion ratio of the single detection frame is determined based on the first distance and the second distance.

3. The target detection post-processing method according to claim 2, characterized in that: The step of determining the occlusion ratio of the single detection frame based on the first distance and the second distance includes: If the sum of the first distance and the preset distance threshold is less than the second distance, determining the occlusion ratio of the single detection frame according to the occupation angles respectively corresponding to the high confidence detection frame and the single detection frame, wherein the occupation angle represents the angle range occupied by the detection frame in the field of view of the vehicle-mounted camera; If the sum of the first distance and the preset distance threshold is greater than or equal to the second distance, it is determined that the occlusion ratio of the single detection frame is a preset minimum ratio.

4. The target detection post-processing method according to claim 3, characterized in that: The step of determining the occlusion ratio of the single detection frame according to the occupation angles respectively corresponding to the high confidence detection frame and the single detection frame comprises: Calculating a first occupancy angle corresponding to the high-confidence detection frame and a second occupancy angle corresponding to the single detection frame; determining an angle overlap interval between the high-confidence detection frame and the single detection frame according to the first occupation angle and the second occupation angle; The occlusion ratio of the single detection frame is calculated based on the angle overlap interval.

5. The target detection post-processing method according to claim 1, characterized in that: The step of screening the original detection frame set according to the occlusion ratio to obtain a first target detection frame set includes: Updating the initial confidence of the single detection frame according to the occlusion ratio to obtain an updated confidence; The original detection frame set is screened based on the updated confidence to obtain a first target detection frame set.

6. The target detection post-processing method according to claim 1, characterized in that: The step of performing secondary screening on the original detection frame set according to the historical frame detection frame set corresponding to the original detection frame set to obtain a second target detection frame set includes: Selecting an occluded detection frame set from the original detection frame set, the occluded detection frame set being a set of detection frames whose occlusion ratio is greater than a preset ratio threshold; Calculate a maximum intersection-over-union ratio set between any detection frame in the historical frame detection frame set corresponding to the original detection frame set and any detection frame in the obscured detection frame set; The original detection frame set is screened for a second time based on the maximum intersection-over-union set to obtain a second target detection frame set.

7. A target detection post-processing device, characterized in that: The target detection post-processing device comprises: The target detection module is used to perform target detection on the input image to obtain a set of original detection frames, wherein the input image is a traffic scene image taken by a vehicle-mounted camera; an occlusion calculation module, used to determine an occlusion ratio of a single detection frame in the original detection frame set, wherein the occlusion ratio indicates the degree to which the target object contained in the single detection frame is occluded by other objects; A primary screening module, used for screening the original detection frame set according to the occlusion ratio to obtain a first target detection frame set; The secondary screening module is used to perform secondary screening on the original detection frame set according to the historical frame detection frame set corresponding to the original detection frame set to obtain a second target detection frame set.

8. A target detection post-processing device, characterized in that: The device comprises: a memory, a processor, and a target detection post-processing program stored in the memory and executable on the processor, wherein the target detection post-processing program is configured to implement the steps of the target detection post-processing method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a target detection post-processing program is stored on the storage medium. When the target detection post-processing program is executed by a processor, the steps of the target detection post-processing method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a target detection post-processing program, and when the target detection post-processing program is executed by a processor, the steps of the target detection post-processing method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Target detection post-processing method, apparatus and device, and storage medium

    WO2026157557A1