Filtering method and device of object detection frame and related equipment

By calculating the offset parameters of the object detection frame to determine the degree of consistency of its position and size, the problem of object detection frame misfiltering in the existing technology is solved, and the recognition accuracy and completeness of three-dimensional solid objects are improved.

CN120689854APending Publication Date: 2025-09-23BEIJING AUTONAVI YUNMAP TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410294899.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-14
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

When modeling physical objects based on point cloud data, existing technologies have the problem of incorrect filtering of object detection frames, especially when three-dimensional physical objects are nested, resulting in incomplete recognition results.

Method used

By obtaining the feature parameters of the object detection frame and calculating the offset parameters, the position and size consistency of the object detection frame can be determined, and duplicate object detection frames can be filtered out to ensure the accuracy and completeness of the recognition results.

Benefits of technology

It effectively avoids the mis-filtering of object detection frames, improves the recognition quality of 3D solid objects, and ensures the completeness of recognition results, especially when nesting phenomena exist.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689854A_ABST
    Figure CN120689854A_ABST
Patent Text Reader

Abstract

The invention discloses a filtering method and device of an object detection frame and related equipment, and aims to avoid error filtering of the object detection frame and improve the recognition quality of a three-dimensional entity object. The method comprises the steps that feature parameters of object detection frames of a plurality of three-dimensional entity objects of the same type are obtained, and the feature parameters and the types of the object detection frames are obtained in advance based on point cloud data recognition; according to the characteristic parameters of the object detection frames, a plurality of offset parameters are obtained, and one offset parameter represents the consistency degree of the positions and the sizes of the two object detection frames in the multiple object detection; and according to the offset parameter, filtering out one of two repeated object detection frames until the offset parameter of each of the plurality of object detection frames is calculated, and outputting the reserved object detection frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of visual technology, and in particular to a method, device and related equipment for filtering an object detection frame. Background Art

[0002] In many fields, it is necessary to identify and model real-world physical objects. For example, in the field of high-precision mapping, physical objects such as traffic signs and traffic light backplanes need to be identified and modeled to generate representations of these physical objects in high-precision maps. High-precision maps can assist intelligent driving systems with positioning and / or driving decisions. Furthermore, in fields such as virtual reality (VR) and augmented reality (AR), the identification and modeling of real-world physical objects is also required.

[0003] Currently, physical objects in the real world can be identified and modeled based on pre-collected point cloud data. Specifically, when a vehicle equipped with a lidar travels on the road, the lidar emits laser signals around the vehicle. When the laser signals contact physical objects in the real world, the objects reflect the laser signals, and point cloud data is generated based on the collected reflected laser signals. Because point cloud data is generated based on the laser signals reflected by physical objects, the point cloud data records the shape of the corresponding physical objects. Therefore, existing technologies can generate digital modeling results of physical objects in the real world based on point cloud data.

[0004] The existing process of modeling physical objects based on point cloud data requires generating an object detection frame based on the point cloud data. This object detection frame represents relevant information, such as the relative position and shape of the physical object in the point cloud data. Because multiple object detection frames may be identified for the same physical object in real-world scenarios, existing technologies require filtering these object detection frames to reduce redundancy or duplication.

[0005] The inventors of the present application have discovered that when the entity object is a three-dimensional entity object that occupies a certain spatial volume, the object detection frame can be filtered based on the volume intersection and union ratio of the object detection frame. The volume intersection and union ratio refers to the ratio of the intersection volume to the union volume of two object detection frames. If the volume intersection and union ratio of two object detection frames is too large, it can be considered that the two object detection frames express the same three-dimensional entity object, so that one of the object detection frames can be removed. However, three-dimensional entity objects in the real world are nested. The so-called nesting means that a three-dimensional entity object includes other three-dimensional entity objects. At this time, if the volume intersection and union ratio is used to filter the object detection frame of the three-dimensional entity object, the object detection frame will be misfiltered, resulting in the lack of modeling results for certain three-dimensional entity objects. Summary of the Invention

[0006] The present application provides a method, apparatus, and related equipment for filtering object detection frames, which utilize the degree of consistency between object detection frames to filter object detection frames, thereby avoiding misfiltering of object detection frames and improving the recognition quality of three-dimensional solid objects.

[0007] In a first aspect, the present application provides a method for filtering an object detection frame, characterized in that the method comprises:

[0008] Obtaining feature parameters of object detection frames of multiple three-dimensional solid objects of the same type, wherein the feature parameters and types of the object detection frames are pre-identified based on point cloud data;

[0009] Obtaining a plurality of offset parameters according to the characteristic parameters of the object detection frames, wherein an offset parameter represents a degree of consistency between positions and sizes of two object detection frames in the plurality of object detections;

[0010] According to the offset parameter, one of the two repeated object detection frames is filtered out until the offset parameters of the multiple object detection frames are all calculated, and the retained object detection frame is output.

[0011] In a second aspect, the present application provides a device for filtering a physical object detection frame, the device comprising:

[0012] An acquiring unit, configured to acquire feature parameters of object detection frames of a plurality of three-dimensional solid objects of the same type, wherein the feature parameters and types of the object detection frames are pre-identified based on point cloud data;

[0013] an offset parameter calculation unit, configured to obtain a plurality of offset parameters based on the feature parameters of the object detection frames, wherein an offset parameter represents a degree of consistency between positions and sizes of two object detection frames in the plurality of object detections;

[0014] The filtering unit is configured to filter out one of the two repeated object detection frames according to the offset parameter until the offset parameters are calculated for all the multiple object detection frames, and output the retained object detection frame.

[0015] In a third aspect, the present application provides a device, comprising: a processor, the processor being coupled to a memory, the memory storing at least one computer program instruction, the at least one computer program instruction being loaded and executed by the processor so that the device implements a method as described in any one of the first aspects above.

[0016] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the storage medium stores at least one instruction, and when the instruction is executed on a computer, the computer executes the method as described in any one of the aforementioned first aspects.

[0017] In a fifth aspect, the present application provides a computer program product, characterized in that the computer program product includes one or more computer program instructions, which, when loaded and executed by a computer, enable the computer to execute the method as described in any one of the aforementioned first aspects.

[0018] It can be seen that this application has the following beneficial effects:

[0019] The solution provided by this application first obtains feature parameters for multiple object detection frames of the same type. Based on the feature parameters of the multiple object detection frames, multiple offset parameters are calculated. Finally, duplicate object detection frames are filtered out from the multiple object detection frames based on the multiple offset parameters and the retained object detection frames are output. The object detection frames of this application are obtained based on point cloud data that records three-dimensional solid objects. Each object detection frame corresponds to a three-dimensional solid object recorded in the point cloud data. Since an offset parameter is used to characterize the degree of consistency in the position and size of two object detection frames, and the position and size of an object detection frame essentially reflect the position and size of the three-dimensional solid object corresponding to the object detection frame, the offset parameter of this application reflects the degree of consistency between the three-dimensional solid objects corresponding to the two object detection frames. Even if two different three-dimensional solid objects are nested, they will not appear in the same position and have the same size. Therefore, the offset parameters of this application can accurately determine duplicate object detection frames for the same three-dimensional solid object. Moreover, even if there are nested three-dimensional solid objects among the multiple three-dimensional solid objects recorded in the point cloud data, this application can avoid incorrect filtering of object detection frames, ensure the completeness of the final three-dimensional solid object recognition results, and improve the recognition quality of three-dimensional solid objects. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments provided in the present application. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0021] Figure 1-a A schematic diagram of a three-dimensional solid object provided in an embodiment of the present application;

[0022] Figure 1-b A schematic diagram of another three-dimensional solid object provided in an embodiment of the present application;

[0023] Figure 1-c A schematic diagram of an object detection frame provided in an embodiment of the present application;

[0024] Figure 2 A schematic diagram of a flow chart of a method for filtering an object detection frame provided in an embodiment of the present application;

[0025] Figure 3 Another flowchart of the object detection frame filtering method provided in an embodiment of the present application;

[0026] Figure 4 A schematic diagram of a process for determining an offset parameter provided in an embodiment of the present application;

[0027] Figure 5 A schematic diagram of the structure of an object detection frame filtering device provided in an embodiment of the present application;

[0028] Figure 6 A schematic diagram of the structure of a device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0029] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0030] The following is an explanation of some terminology concepts involved in the embodiments of this application.

[0031] (1) Three-dimensional point cloud data.

[0032] Three-dimensional point cloud data, which can be called point cloud data, refers to data recorded in the form of points through scanning and other methods. Since the amount of point data is very large, it is called a point cloud. The information of each point in the point cloud data includes at least: three-dimensional position coordinates, and some may contain color information (RGB) or reflection intensity information (Intensity). Taking the high-precision map scene as an example, the point cloud data used to make high-precision maps is collected by the laser radar carried by the collection vehicle when the collection vehicle is driving on the road. Specifically: When the collection vehicle equipped with the laser radar is driving on the road, the laser radar emits a laser signal. The laser signal contacts the road and the physical objects on both sides of the road and will generate reflection. By receiving the reflected laser signal, point cloud data recording the physical objects on the road and on both sides of the road can be obtained. The physical objects on the road and on both sides of the road can include traffic signs, lane lines, driving direction indicators, etc.

[0033] (2) Coordinate system.

[0034] A coordinate system is a reference established to describe the position and motion of an object. In the process of identifying physical objects based on three-dimensional point cloud data, multiple coordinate systems can be established based on different principles. For example, a coordinate system can be established based on an object detection frame, or a coordinate system can be established based on the ground. In the embodiment of the present application, the coordinate system established based on the object detection frame can be referred to as the first coordinate system, and the coordinate system established based on the ground can be referred to as the second coordinate system or the geographic coordinate system.

[0035] (3) Object detection box.

[0036] An object detection box can be referred to as a physical object detection box. In this application, an object detection box is a polygonal box, such as a rectangular box, that is identified based on point cloud data and encloses the physical object recorded in the point cloud data. The size, shape, and position of the object detection box are related to the size, shape, and position of the physical object recorded in the point cloud data. If the point cloud data records a three-dimensional physical object, the object detection box for the three-dimensional physical object can be a cuboid.

[0037] (4) Nesting.

[0038] Nesting means that an entity object contains another entity object. Figure 1-a , as shown in the traffic sign 10, since the traffic sign 10 is a solid object occupying a certain spatial volume, the traffic sign can be called a three-dimensional solid object. Figure 1-a The traffic sign 10 shown in the figure shows two instructions: "Place A 300km" and "Place B 150km". When processing the 3D point cloud data recording the traffic sign 10, the gray body of the traffic sign 10, "Place A 300km" and "Place B 150km" will be treated as three 3D entity objects. Therefore, the gray body of the traffic sign 10 has the phenomenon of "Place A 300km" and "Place B 150km" nested. Similarly, Figure 1-b The body of the traffic sign 30 shown is embedded with three-dimensional entity objects such as “Slow” and “No Honking”.

[0039] The following is combined with Figure 1-a and attached Figure 1-b The existing volume intersection and union ratio schemes are introduced.

[0040] like Figure 1-aAs shown in FIG, after the three-dimensional point cloud data recording the traffic sign 10 is recognized, it is assumed that object detection frame 21, object detection frame 22, object detection frame 23 and object detection frame 24 are obtained, wherein object detection frame 21 and object detection frame 22 both correspond to the gray body of traffic sign 10, and there is duplication between the two. Therefore, according to the volume intersection and union ratio of object detection frame 21 and object detection frame 22 and the confidence of the object detection frame, object detection frame 22 with low confidence is filtered out, and object detection frame 21 is retained. Further, the volume intersection and union ratio of object detection frame 21 and object detection frame 23 is obtained. Since object detection frame 23 is completely contained in object detection frame 21, the intersection volume of object detection frame 21 and object detection frame 23 is the volume of object detection frame 23, and the union volume is the volume of object detection frame 21. If the volume intersection and union ratio obtained based on this is greater than the preset intersection and union ratio threshold, object detection frame 23 is filtered out, which will result in the final recognition result not including the recognition result of "A place 300km".

[0041] For example, see Figure 1-b The traffic sign 30 shown in the figure shows four instructions, namely "Vehicles entering the factory area, please slow down", "Slow down sign" and "No honking sign". After the three-dimensional point cloud data recording the traffic sign 30 is recognized, it is assumed that the following is obtained: Figure 1-b As shown, object detection frame 41, object detection frame 42, object detection frame 43, object detection frame 44, and object detection frame 45 are shown. For object detection frame 42, the volume of its intersection with object detection frame 41 is the volume of object detection frame 42, and the volume of its union is the volume of object detection frame 42. Because object detection frame 42 occupies a larger area, it may be filtered out due to its volume intersection-over-union ratio being greater than the intersection-over-union ratio threshold, resulting in the absence of the recognition result "Vehicles entering the factory area, please slow down" in the final recognition result.

[0042] To solve the above problems, an embodiment of the present application provides a method for filtering an object detection frame. The method provided by the present application is described in detail below in conjunction with the drawings in the specification.

[0043] See also Figure 2 , this figure is a flow chart of a method for filtering an object detection frame provided in an embodiment of the present application, which includes the following steps.

[0044] S201: Acquire feature parameters of object detection boxes of multiple three-dimensional solid objects of the same type.

[0045] When filtering multiple object detection frames, the method provided herein first obtains feature parameters of these object detection frames. The feature parameters of the object detection frames characterize the characteristics of the three-dimensional physical object corresponding to the object detection frames. The feature parameters may include one or more of the following: the center position parameter, orientation parameter, and size parameter of the object detection frames. The multiple parameters may be two or more.

[0046] It should be noted that the feature parameters of the object detection frame of the 3D solid object obtained in this application are pre-acquired based on 3D point cloud data. Those skilled in the art can, based on their own techniques, obtain the feature parameters of the object detection frame of the 3D solid object based on point cloud data, and this application will not elaborate on them. In this application, one object detection frame corresponds to one 3D solid object, and different object detection frames may correspond to the same 3D solid object.

[0047] In the embodiment of the present application, object detection frames can be classified according to the 3D solid objects they correspond to, and then object detection frames of the same category can be filtered. In other words, if multiple object detection frames are obtained based on 3D point cloud data, the multiple object detection frames can first be classified according to the type of 3D solid objects, and then each type of object detection frame can be filtered separately.

[0048] Taking high-precision maps as an example, the types of three-dimensional entity objects can be traffic signs, traffic light backboards, etc. This application processes the object detection frames of three-dimensional entity objects of the same type. For example, the object detection frames of the category "traffic signs" are classified into one category for processing. The category of three-dimensional entity objects can also be pre-acquired based on point cloud data, which will not be repeated in this application.

[0049] S202: Obtain multiple offset parameters according to the feature parameters of the object detection frame.

[0050] After obtaining the feature parameters of multiple object detection frames, multiple offset parameters can be obtained based on the feature parameters. Among them, an offset parameter can be calculated based on the offset parameters of two object detection frames. The offset parameter represents the degree of consistency between the two object detection frames in position and size. For a detailed introduction to the offset parameter, please refer to the following Figure 4 The introduction of the corresponding embodiments will not be repeated here.

[0051] In some possible implementations, offset parameters can be calculated between any two object detection frames. That is, assuming the number of acquired feature parameters for object detection frames of the same type of 3D solid objects is N, then N*(N-1) / 2 offset parameters can be calculated. Based on these offset parameters, it can be determined whether any two object detection frames overlap.

[0052] In some other possible implementations, the object detection frames can be filtered by selecting a reference object detection frame, calculating an offset parameter based on the reference object detection frame, filtering the object detection frame when there is a duplicate, and continuing to determine the next reference object detection frame. In this way, the number of offset parameters that need to be calculated can be reduced, the amount of calculation can be reduced, and the filtering efficiency can be improved. For more information about this embodiment, please refer to the following Figure 3 The introduction of the corresponding embodiments will not be repeated here.

[0053] S203: Filter out one of the two repeated object detection frames according to the offset parameter, until the offset parameters of the multiple object detection frames are all calculated, and output the retained object detection frame.

[0054] After determining the offset parameters, the object detection frames can be filtered according to the offset parameters. Specifically, the offset parameters can be used to determine whether the object detection frames corresponding to the offset parameters are repeated. If they are repeated, one object detection frame is filtered out from the two object detection frames and the other object detection frame is retained. For example, if a certain offset parameter indicates that the degree of consistency between the positions and sizes of the two object detection frames is greater than a threshold, then the two object detection frames can be considered to be duplicate object detection frames. Duplicate object detection frames refer to the three-dimensional solid objects corresponding to the two object detection frames being the same three-dimensional solid objects. Optionally, when filtering object detection frames, the object detection frames to be filtered out and the object detection frames to be retained can be determined based on the credibility of the object detection frames. In an embodiment of the present application, the object detection frames to be filtered out can be referred to as filtered object detection frames, and the object detection frames to be retained can be referred to as retained object detection frames. Generally, the confidence of the retained object detection frame can be higher than the confidence of the filtered object detection frame.

[0055] The above is the object detection frame filtering method provided by the present application. When a three-dimensional solid object A is nested in a three-dimensional solid object B, it can be understood that although the volume intersection and union ratio of the three-dimensional solid object A and the three-dimensional solid object B will cause the object detection frame to be misfiltered, there must be differences between the sizes and positions of the three-dimensional solid object A and the three-dimensional solid object B. Therefore, the size and position of the object detection frame surrounding the three-dimensional solid object must also be different. Therefore, the solution of the present application can ensure that the object detection frame of the nested three-dimensional solid object will not be misfiltered, thereby ensuring the comprehensiveness of the recognition results.

[0056] According to the above introduction, assuming that there are N object detection frames of the same type obtained in step S201, the solution can be to first calculate the offset parameters between all the object detection frames in the N object detection frames, and then repeatedly judge the object detection frames and perform filtering. According to this method, the present application calculates a total of N*(N-1) / 2 offset parameters, that is, the present application needs to perform N*(N-1) / 2 offset parameter calculations. When the number N is large, the number of offset parameters is large, and too many computing resources are occupied. To this end, in some possible implementation methods, the object detection frames can be filtered for multiple rounds to reduce the number of object detection frames in the next round of calculating offset parameters, thereby saving computing resources and improving efficiency. The following is combined with Figure 4 Provide a detailed introduction.

[0057] See also Figure 3 , which is another flowchart of the object detection frame filtering method provided in an embodiment of the present application, specifically comprising the following steps:

[0058] S301: Acquire feature parameters of object detection boxes of multiple three-dimensional solid objects of the same type.

[0059] For the introduction of step S301, please refer to the above and will not be repeated here.

[0060] For ease of introduction, Figure 3 The corresponding embodiment is introduced based on filtering N object detection frames (N is a positive integer greater than 1). Accordingly, in step S301, N feature parameters can be obtained.

[0061] S302: Establish an initial set T and a target set R.

[0062] In order to facilitate multiple rounds of filtering, Figure 3 In a corresponding implementation, an initial set T and a target set R can be established. The initial set T and the target set R are used to store object detection frames. Specifically, the initial set T can be used to store object detection frames whose retention has not yet been determined. When step S302 is first executed, the initial set T is used to store the object detection frames obtained in step S301. The target set R can be used to store the retained object detection frames. Storing object detection frames in the set can be implemented by recording the identifiers of the object detection frames in the set.

[0063] That is to say, the target set R can be used to store the object detection frames retained in step S203. It can be understood that before the first execution of step 303, the identifiers of N object detection frames are stored in the initial set T, and the target set R is empty. In an embodiment of the present application, the object detection frames in the initial set T can be arranged in sequence. Accordingly, the above-mentioned initial set T can also be referred to as the initial queue T, and the above-mentioned target set R can also be referred to as the target queue R. Specifically, the object detection frames can be arranged according to the confidence of the object detection frames. For example, the N object detection frames in the initial set T can be arranged in order from high to low confidence. The confidence of the object detection frame is generated when the object detection frame of the three-dimensional solid object is identified in advance based on point cloud data. The confidence is used to characterize the accuracy of the object detection frame. The higher the confidence, the higher the accuracy of the object detection frame.

[0064] S303: Determine an object detection frame with the highest confidence in the initial set T as a reference object detection frame, and transfer the reference object detection frame from the initial set T to the target set R.

[0065] At the beginning of each round of filtering, the present application determines the object detection frame with the highest confidence in the initial set T as the reference object detection frame, and transfers the reference object detection frame from the initial set T to the target set R. The reference object detection frame transferred to the target set R is determined as the retained object detection frame.

[0066] It can be understood that transferring an object detection frame from the initial set T to the target set R means that the reference object detection frame will be removed from the initial set T. That is, each round of filtering reduces at least one object detection frame from the initial set T. Furthermore, the confidence of any object detection frame remaining in the initial set T is lower than the confidence of the object detection frame remaining in the target set R.

[0067] Because the reference object detection frame selected in each round of filtering is the object detection frame with the highest confidence in the initial set T, even if one (or more) object detection frames in the initial set T correspond to the same 3D solid object as the reference object detection frame, the confidence level of the 3D solid object determined using the reference object detection frame is higher than that of the solid object determined using other reference object detection frames, due to the highest confidence level. Therefore, retaining the reference object detection frame and removing it from the initial set T ensures accurate 3D solid object recognition while reducing the computational effort required for offset parameters.

[0068] After the reference object detection frame is determined, the object detection frames sorted after the reference object detection frame in the initial set T can be filtered according to the reference object detection frame. Specifically, if the N object detection frames in the initial set T are arranged in order of confidence from high to low, the object detection frames can be filtered by judging the degree of consistency between each object detection frame sorted after the reference object detection frame in the initial set T and the reference object detection frame one by one. Optionally, the first object detection frame sorted after the reference object detection frame in the initial set T can be first determined as the object detection frame to be calculated, and a first round of filtering is performed. The object detection frame to be calculated is also the object detection frame with the highest confidence in the initial set T. Then, the second object detection frame sorted after the reference object detection frame in the initial set T (that is, the object detection frame with the second highest confidence in the initial set T) is determined as the object detection frame to be calculated, and a second round of filtering is performed, and so on. Please refer to the following steps for the filtering process:

[0069] S304: Obtain an offset parameter between the reference object detection frame and the object detection frame to be calculated according to feature parameters of the reference object detection frame and the object detection frame to be calculated.

[0070] S305: Based on the offset parameters between the reference object detection frame and the detection frame of the object to be calculated, determine whether there is any duplication between the reference object detection frame and the detection frame to be calculated. If not, execute step S306. If so, first filter the detection frame of the object to be calculated from the initial set T, and then execute step S306.

[0071] After obtaining the offset parameter between the reference object detection frame and the object detection frame to be calculated, it can be determined whether there is duplication between the reference object detection frame and the object detection frame to be calculated based on the offset parameter. If there is no duplication between the reference object detection frame and the object detection frame to be calculated, it means that the reference object detection frame and the object detection frame to be calculated correspond to different three-dimensional solid objects. Then, the object detection frame to be calculated is retained in the initial set T. If there is duplication between the reference object detection frame and the object detection frame to be calculated, it means that the reference object detection frame and the object detection frame to be calculated correspond to the same three-dimensional solid object, and then one object detection frame can be filtered out from the reference object detection frame and the object detection frame to be calculated. Since the object detection frames in the initial set T of this application are sorted from high to low according to confidence, the confidence of the object detection frame to be calculated must not be higher than that of the reference object detection frame. Therefore, the object detection frame to be calculated can be determined as the filtered object detection frame, the object detection frame to be calculated can be removed from the initial set T, and then step S306 is executed.

[0072] S306: Determine whether there is an object detection frame with uncalculated offset parameters in the initial set T. If so, use the object detection frame with uncalculated offset parameters as the object detection frame to be calculated and return to step S304. If not, execute step S307.

[0073] If there is an object detection frame in the initial set T for which the offset parameters have not been calculated, the object detection frame can be used as a new object detection frame to be calculated and the process returns to step S304 to filter the object detection frames for which the offset parameters have not been calculated. If the offset parameters have been calculated for each object detection frame in the initial set T with respect to the reference object detection frame, step S307 can be executed.

[0074] S307: Determine whether the number of object detection frames remaining in the initial set T is greater than 1. If it is greater than 1, return to execute step S303; if it is equal to 1, it means that there is 1 object detection frame remaining in the initial set T, then add the object detection frame to the target set R, and then output the object detection frame retained in the target set R; if it is less than 1, it means that the initial set T is empty, then directly output the object detection frame retained in the target set R.

[0075] After completing one round of filtering, the present application determines whether it is necessary to perform the next round of filtering on the remaining object detection frames in the initial set T. Specifically, it can be determined whether the number of remaining object detection frames in the initial set T is greater than 1 to determine whether it is necessary to perform the next round of filtering.

[0076] If the number of remaining object detection frames in the initial set T is greater than 1, it indicates that at least two object detection frames remain in the initial set T, and there is a possibility that the two object detection frames may be duplicated. Therefore, the initial set T can be further filtered. Specifically, the process returns to step S303 and selects the object detection frame with the highest confidence from the remaining object detection frames in the initial set T as the reference object detection frame for the next round.

[0077] If there are no remaining object detection frames in the initial set T, it means that the filtering of the initial set T is completed, the filtering can be ended, and the target set R is output. Subsequently, the three-dimensional solid objects can be identified based on the retained object detection frames stored in the target set R.

[0078] If there is only one object detection frame remaining in the initial set T, then there is no duplication between this object detection frame and any of the retained object detection frames in the target set R. That is, the 3D solid object corresponding to this object detection frame is different from the 3D solid object corresponding to any of the retained object detection frames in the target set R. Therefore, this object detection frame can be determined as a retained object detection frame, added to the target set R, and then the target set R is output.

[0079] according to Figure 3 Corresponding implementation method, if there are N object detection frames in the initial set T, except for the reference object detection frame in the first round of filtering, which needs to calculate the offset parameters with the other N-1 object detection frames, the number of object detection frames in the initial set T will inevitably decrease with each round of filtering. Moreover, as the number of filtering rounds increases, the number of object detection frames retained in the initial set T will become smaller and smaller, so that the total amount of offset parameter calculation is less than N*(N-1) / 2, which improves the computational efficiency.

[0080] The following is combined with the actual application scenario. Figure 3 The implementation method is introduced in this paper.

[0081] Assume that there are five object detection frames, namely object detection frame K1, object detection frame K2, object detection frame K3, object detection frame K4, and object detection frame K5, and the confidence of K1, K2, K3, K4, and K5 decreases from high to low. K1 and K3 correspond to 3D solid object A, K2 and K4 correspond to 3D solid object B, and K5 corresponds to 3D solid object C.

[0082] When filtering the object detection frame, K1, K2, K3, K4 and K5 can be added to the initial set T and multiple rounds of filtering can be performed.

[0083] In the first round of filtering, the object detection frame with the highest confidence (i.e., K1) can be selected from the initial set T as the reference object detection frame, and K1 can be transferred from the initial set T to the target set R. At this time, the object detection frames in the initial set T include: K2, K3, K4, and K5. Then, the object detection frame with the highest confidence (i.e., K2) can be selected from the initial set T as the object detection frame to be calculated, and the offset parameters of K1 and K2 are calculated. Since the three-dimensional solid object corresponding to K1 is different from the three-dimensional solid object corresponding to K2, it can be determined based on the offset parameters that K1 and K2 are not repeated, so K2 is retained in the initial set T. At this time, the object detection frames included in the initial set T are still K2, K3, K4, and K5. Among them, K3, K4, and K5 have not calculated the offset parameters with K1, so K3 is determined as the new object detection frame to be calculated, and then K1 and K3 are calculated. Since the 3D solid object corresponding to K1 is the same as the 3D solid object corresponding to K3, it can be determined based on the offset parameters that K1 and K3 are duplicates, and K3 is filtered out. At this time, the object detection frames included in the initial set T are K2, K4, and K5. Among them, K4 and K5 have not had their offset parameters calculated with K1. Then the offset parameters of K1 and K4 are continued to be calculated. Since the 3D solid object corresponding to K1 is different from the 3D solid object corresponding to K4, K4 is retained in the initial set T. Similarly, the offset parameters of K1 and K5 are continued to be calculated. Since the 3D solid object corresponding to K1 is different from the 3D solid object corresponding to K5, K5 can be retained in the initial set T. That is, after filtering according to K1, three object detection frames K2, K4, and K5 remain in the initial set T.

[0084] Since the number of remaining object detection boxes in the initial set T is greater than 1, a second round of filtering can be performed.

[0085] In the second round of filtering, the object detection frame with the highest confidence (i.e., K2) can be selected from the initial set T as the reference object detection frame, and K2 can be transferred from the initial set T to the target set R. At this time, there are still two object detection frames K4 and K5 in the initial set T. The object detection frame with the highest confidence (i.e., K4) can be selected from the initial set T as the object detection frame to be calculated, and the offset parameters of K2 and K4 are calculated. Since the three-dimensional solid object corresponding to K2 is the same as the three-dimensional solid object corresponding to K4, K4 can be filtered out by the same token, and the offset parameters of K2 and K5 can be calculated. Since the three-dimensional solid object corresponding to K2 is different from the three-dimensional solid object corresponding to K5, K5 is retained in the initial set T. After filtering according to K2, one object detection frame K5 remains in the initial set T.

[0086] Since the number of remaining object detection frames in the initial set T is 1, K5 can be determined as a retained object detection frame and added to the target set R. Next, the target set R can be output. The target set R contains three object detection frames: K1, K2, and K5. This completes the filtering of object detection frames.

[0087] It's easy to see that in the above implementation, the offset parameters are calculated 4+2=6 times. If the offset parameters between any two object detection frames are calculated, the offset parameters need to be calculated 5*(5-1) / 2=10 times. This shows that multiple rounds of filtering can reduce the amount of offset parameter calculations and improve filtering efficiency.

[0088] The above describes the method for filtering the object detection frame provided in the embodiment of the present application. The following describes an embodiment of determining the offset parameter provided in the present application with reference to the accompanying drawings.

[0089] First, the characteristic parameters are introduced.

[0090] In the embodiment of the present application, the feature parameters of the object detection frame may include at least a center position parameter of the object detection frame, an orientation parameter of the object detection frame, and a size parameter of the object detection frame, which are used to describe the center point, orientation, and size of the object detection frame, respectively. The following uses the i-th object detection frame as an example to explain the meaning of each parameter. The i-th object detection frame may be the i-th object detection frame in the initial set T. i is a positive integer less than n.

[0091] First, the center position parameters. The center position parameters of the object detection box can represent the geometric center point of the object detection box. For example, if the shape of the i-th object detection box is a cuboid, then the center position parameters of the i-th object detection box are the position parameters of the geometric center point of the cuboid. The center position parameters include the horizontal coordinate, the vertical coordinate, and the height coordinate.

[0092] For ease of introduction, we can use the vector P i Represents the center position parameters of the i-th object detection box. Vector P i =[x i ,y i ,z i ]. Among them, x i is the horizontal coordinate of the center of the cuboid corresponding to the i-th object detection frame in the second coordinate system, y i is the vertical coordinate of the center of the cuboid corresponding to the i-th object detection frame in the second coordinate system, z i is the height coordinate of the center of the cuboid corresponding to the i-th object detection box in the second coordinate system.

[0093] Second, orientation parameters.

[0094] The orientation parameter of an object detection box describes the orientation of a 3D object. In real-world scenarios, some 3D objects are directional. For example, in HD maps, the orientation of a 3D object like a traffic sign or traffic light backplane can be determined by the direction of the side with the information drawn on it (i.e., the side facing the driver).

[0095] To this end, when performing three-dimensional physical object recognition, the orientation of the object detection frame can be set, and then the orientation parameters of the object detection frame can be determined according to the relative position relationship between the object detection frame and the second coordinate system.

[0096] For example, the x-axis direction of the first coordinate system can be used as the orientation of the object detection frame. Accordingly, the orientation parameter of the object detection frame may include the orientation angle and / or direction vector of the object detection frame. The orientation angle of the object detection frame may be the angle between the x-axis of the first coordinate system and the x-axis of the second coordinate system. This angle may be referred to as the orientation angle θ i It can be understood that the angle θ i Equivalent to the Euler angle between the first coordinate system and the second coordinate system.

[0097] See also Figure 1-c ,exist Figure 1-c In the implementation shown, the first coordinate system O lwh and the second coordinate system (also called geographic coordinate system) Q xyz . Among them, the x-axis of the first coordinate system is d l , corresponding to the long side direction of the object detection frame, the y-axis of the first coordinate system is d w , corresponding to the short side direction of the object detection frame, the z axis of the first coordinate system is d l , corresponds to the vertical direction of the object detection frame. Accordingly, the orientation angle θ i The angle between the x-axis of the first coordinate system and the x-axis of the second coordinate system (that is, the geographic coordinate system) represents the angle of rotation of the x-axis of the second coordinate system to the x-axis of the first coordinate system.

[0098] Third, size parameters.

[0099] The size parameter is used to describe the size of the object detection frame. If the object detection frame is a cuboid, the size of the object detection frame can be described by the length, width and height of the object detection frame. Accordingly, the size parameters of the i-th object detection frame can include the length of the long side l i , short side length w i and the vertical side length h i (also known as height h i ).

[0100] Optionally, the above center position parameters, orientation parameters and size parameters can be obtained by the set t i Represents. Set t i =[x i ,y i ,z i ,θ i ,l i ,w i ,h i ].

[0101] The above introduces the characteristic parameters of the feature detection frame. The following takes the i-th object detection frame and the j-th object detection frame as examples to introduce some implementation methods of calculating the offset parameters based on the characteristic parameters. Among them, i and j are both positive integers less than N. Figure 3 In the corresponding implementation, the i-th object detection frame may be the reference object detection frame, and the j-th object detection frame may be the object detection frame to be calculated. For ease of introduction, the offset parameter of the i-th object detection frame and the j-th object detection frame is referred to as the offset parameter O below. ij .

[0102] In the embodiment of the present application, the calculation of the offset parameter can be directional. That is, the offset parameter between two object detection frames is the offset parameter of one object detection frame relative to the other object detection frame. Figure 3 In the corresponding implementation, the offset parameter may be an offset parameter of the object detection frame to be calculated relative to the reference object detection frame. For ease of introduction, the offset parameter is hereinafter referred to as ij The offset parameter of the j-th object detection frame relative to the j-th object detection frame is taken as an example for introduction.

[0103] See also Figure 4 , which is a flow chart of a method for determining an offset parameter provided in an embodiment of the present application, specifically comprising the following steps:

[0104] S401: Calculate the center offset of the i-th object detection frame and the j-th object detection frame according to the center position parameters of the i-th object detection frame and the center position parameters of the j-th object detection frame.

[0105] In calculating the offset parameter O ij The process includes first calculating the center position offset between the i-th object detection frame and the j-th object detection frame. The center position offset between the i-th object detection frame and the j-th object detection frame represents the deviation between the center point of the i-th object detection frame and the center point of the j-th object detection frame. Optionally, the center position offset offset can be obtained by the following formula (1).

[0106] Formula (1): offset=[x i ,yi ,z i ]-[x j ,y j ,z j ]

[0107] Since all the parameters in the above formula are position parameters in the second coordinate system, the center position offset offset represents the relative position deviation between the i-th object detection frame and the j-th object detection frame in the second coordinate system.

[0108] S402: Determine a direction vector of the i-th object detection frame according to the orientation parameter of the i-th object detection frame.

[0109] In addition to the center position offset, the direction vector of the i-th object detection frame can also be obtained based on the orientation parameter of the i-th object detection frame. The direction vector of the i-th object detection frame represents the orientation of each edge of the i-th object detection frame in the geographic coordinate system.

[0110] Among them, the direction vector of the i-th object detection frame may include the long side direction vector, short side direction vector and vertical direction vector of the i-th object detection frame, which respectively represent the orientation of the long side of the i-th object detection frame in the geographic coordinate system, the orientation of the short side of the i-th object detection frame in the geographic coordinate system and the orientation of the vertical side of the i-th object detection frame in the geographic coordinate system.

[0111] It should be noted that Figure 4 The long side direction, short side direction and vertical direction in the corresponding embodiment refer to the long side direction, short side direction and vertical side direction of the reference object detection frame (ie, the i-th object detection frame).

[0112] Specifically, if the orientation parameters of the i-th object detection frame include the orientation angle θ i , then the direction vector of the i-th object detection box can be obtained by the following formula (2).

[0113] Formula (2):

[0114] in, is the long side direction vector of the i-th object detection box, is the short side direction vector of the i-th object detection box, is the vertical direction vector of the i-th object detection box.

[0115] S403: Obtain the offset parameter O according to the center position offset, the direction vector of the i-th object detection frame, the size parameter of the i-th object acquisition frame, and the size parameter of the j-th object detection frame. ij .

[0116] After obtaining the direction vector of the i-th object detection frame, the offset parameter O can be calculated by combining the center position offset, the direction vector, the size parameters of the i-th object acquisition frame, and the size parameters of the j-th object detection frame. ij .

[0117] Optionally, the offset parameter O ij The offset parameters of the i-th object detection frame and the j-th object detection frame on each side (long side, short side and vertical side) can be included. ij =[O ij l ,O ij w ,O ij h ]. Among them, O ij l It can be called the long side offset parameter, which represents the offset between the i-th object detection frame and the j-th object detection frame on the long side of the i-th object detection frame. ij w It can be called the short side offset parameter, which represents the offset between the i-th object detection box and the j-th object detection box on the short side of the i-th object detection box. ij h It can be called a vertical offset parameter, which represents the offset between the i-th object detection box and the j-th object detection box on the vertical side of the i-th object detection box.

[0118] Accordingly, when calculating the offset parameter O ij The long side offset parameter O can be calculated separately ij l , short side offset parameter O ij w and vertical edge offset parameter O ij h .

[0119] When calculating the long side offset parameters O ij l On the one hand, the length of the long side of the i-th object detection frame and the length of the long side of the j-th object detection frame can be summed. Half of the sum represents the maximum length of the projection of the distance between the center point of the i-th object detection frame and the center point of the j-th object detection frame on the long side of the i-th object detection frame when the i-th object detection frame and the j-th object detection frame just intersect in the long side direction.

[0120] On the other hand, the projection length of the center position offset in the long side direction of the i-th object detection frame can be calculated. The center position offset represents the relative position relationship between the i-th object detection frame and the j-th object detection frame in the second coordinate system. Therefore, the projection length in the long side direction can be determined by the sum of the center position offset and the long side direction vector of the i-th object detection frame. Specifically, the long side direction vector of the i-th object detection frame can be converted to the center position offset after transposition (i.e., offset). T ) to obtain the projection length in the long side direction.

[0121] Then, the long side offset parameter O can be obtained based on the projection length in the long side direction and the summation result. ij l .

[0122] Specifically, the difference between half of the sum and the projection length in the long side direction can be used to represent the overlap distance between the i-th object detection frame and the j-th object detection frame. Alternatively, the difference can be called the overlap amount in the long side direction between the i-th object detection frame and the j-th object detection frame. ij l . Overlap in the long side direction ij l It is obtained by the following formula (3).

[0123] Formula (3):

[0124] In the first possible implementation, after obtaining the overlap in the long side direction, ij l After that, the overlap in the long side direction can be ij l As the long side offset parameter O ij l The larger the overlap in the long side direction, the higher the degree of overlap between the i-th object detection frame and the j-th object detection frame in the long side direction.

[0125] However, the first implementation does not take into account the size of the object detection frame itself, and there may be a certain error. Therefore, in the second possible implementation, the overlap amount in the long side direction can be combined with the overlap amount in the long side direction. ij l and the related parameters in the i-th object detection frame and the j-th object detection frame to obtain the long side offset parameter O ij l For example, the overlap in the long side direction can be ijl The length of the long side of the reference object detection box (i.e. l i ) and use the ratio as the long side offset parameter O ij l Alternatively, you can also set the overlap in the long side direction to ij l The length of the long side of the object detection box to be calculated (i.e. l j ) and use the ratio as the long side offset parameter O ij l .

[0126] However, the offset parameters calculated by the second implementation method may have errors in some scenarios. Therefore, in the third possible implementation method, the overlap amount in the long side direction can be ij l The minimum value among the long side length of the reference object detection frame and the long side length of the object detection frame to be calculated is used as the numerator, and the maximum value among the three is used as the denominator, and the ratio of the numerator and denominator is used as the long side offset parameter O ij l That is, the long side offset parameter O ij l It can be obtained by the following formula (4).

[0127] Formula (4):

[0128] In some scenarios, the confidence of the object detection frame with a longer side length can be higher than the confidence of the object detection frame with a shorter side length. Therefore, when calculating the maximum value, the object detection frame with a smaller confidence level can be ignored. When calculating the maximum value, the longer side length (i.e., l j When calculating the maximum value, the length of the long side of the object detection box with a smaller confidence level (i.e., l i ). Accordingly, the above formula (4) can be simplified to the following formula 5.

[0129] Formula (5):

[0130] Thus, if the center point of the i-th object detection frame and the center point of the j-th object detection frame coincide or differ slightly in the long side direction, the overlap in the long side direction is ij l is greater than the long side length of the j-th object detection frame. Then the long side offset parameter O in the above formula (5) is ij lThe numerator is the shorter length of the long side of the j-th object detection frame, and the denominator is the length of the long side of the i-th object detection frame. If the i-th object detection frame and the j-th object detection frame have a higher degree of consistency in the long side direction, the long side offset parameter O calculated according to formula (5) is ij l The closer it is to 1. If the consistency between the i-th object detection frame and the j-th object detection frame in the long side direction is lower, the long side offset parameter O calculated according to formula (5) is ij l The closer to 0.

[0131] Alternatively, if the center point of the i-th object detection frame and the center point of the j-th object detection frame coincide or differ slightly in the long side direction, the overlap is ij l is less than the long side length of the j-th object detection frame. Then the long side offset parameter O in the above formula (5) is ij l The numerator is the overlap in the long side direction ij l , the denominator is the length of the long side of the i-th object detection frame. If the i-th object detection frame and the j-th object detection frame have a higher degree of consistency in the long side direction, the long side offset parameter O calculated according to formula (5) will be ij l The closer it is to 1. If the consistency between the i-th object detection frame and the j-th object detection frame in the long side direction is lower, the long side offset parameter O calculated according to formula (5) is ij l The closer to 0.

[0132] In this way, the long side offset parameter O ij l The size of determines the similarity between the i-th object detection box and the j-th object detection box in the long side direction.

[0133] Similarly, the overlap in the short side direction can be calculated by the following formulas (6) and (7) respectively: ij w and the vertical overlap ij h .

[0134] Formula (6):

[0135] Formula (7):

[0136] Furthermore, the short side offset parameter O can be calculated by the following formulas (8) and (9) respectively: ijw and vertical edge offset parameter O ij h .

[0137] Formula (8):

[0138] Formula (9):

[0139] By using the above formulas (1) to (9), the offset parameter O between the i-th object detection frame and the j-th object detection frame can be calculated: ij .

[0140] Optionally, depending on whether the offset parameter satisfies the conditions for filtering the object detection frame, the method may include determining whether the offset parameter is greater than a preset upper offset threshold. If the offset parameter is greater than the upper offset threshold, it indicates that the size and position of the two object detection frames corresponding to the offset parameter are highly consistent, and therefore, it can be determined that the two object detection frames correspond to the same three-dimensional solid object. Optionally, if the offset parameter is calculated using the above formulas (1) to (9), then the upper offset threshold may be less than 1.

[0141] If the offset parameters include offset parameters in multiple directions, then when filtering object detection frames using the offset parameters, you can determine whether the offset parameter in each direction is greater than the upper offset threshold. If the offset parameters in all directions are greater than the upper offset threshold, it means that the position and size of the two object detection frames are highly consistent in all directions. Therefore, the two object detection frames can be considered to correspond to the same 3D solid object.

[0142] That is to say, when judging whether the i-th object detection frame and the j-th object detection frame are repeated, we can judge O ij l , O ij w and O ij h Is it greater than the upper limit of the offset? ij l , O ij w and O ij h If both are greater than the upper threshold of the offset, it is determined that the i-th object detection frame and the j-th object detection frame correspond to the same 3D entity object. ij l , O ij w and O ij h If any one of them is not greater than the offset upper limit threshold, it is determined that the i-th object detection frame and the j-th object detection frame correspond to different three-dimensional solid objects.

[0143] The above is an embodiment of calculating the offset parameters provided by this application. Furthermore, when using the volume intersection-union ratio to filter the object detection frame, not only may the object detection frame be incorrectly filtered due to the nesting between three-dimensional solid objects, but the object detection frame may also be missed due to the shape and volume of the three-dimensional solid object, and repeated object detection frames cannot be filtered out, resulting in duplicate identified three-dimensional solid objects.

[0144] For example, for 3D solid objects like traffic signs, their thickness (i.e., along their short sides) is relatively small. As a result, the point cloud data collected from them may be relatively scattered. Object detection frames derived from the point cloud data may include two (or more) object detection frames that do not overlap or have very low overlap.

[0145] For example. Assume that the actual coordinates of the traffic sign in the short side direction are [2, 5] (the unit can be centimeters, for example), but because the traffic sign is thin, the coordinates of the short side direction of the object detection frame identified based on the point cloud may be distributed in the range of [0, 7]. Assume that based on the point cloud, three object detection frames are determined: object detection frame 1, object detection frame 2, and object detection frame 3. Among them, the coordinates of object detection frame 1 in the short side direction are [0, 4], the coordinates of object detection frame 2 in the short side direction are [1, 4], and the coordinates of object detection frame 3 in the short side direction are [4, 7]. If the volume intersection and union ratio is used for filtering, since object detection frame 1 and object detection frame 2 have an intersection volume, filtering can be performed by the volume intersection and union ratio. However, since the intersection volume of object detection frame 3 and object detection frame 1 cannot be obtained, object detection frame 3 and object detection frame 2 cannot be filtered. In this way, two three-dimensional solid objects will be identified according to object detection frame 2 and object detection frame 3 respectively, resulting in the final identified three-dimensional solid object being inconsistent with the actual situation, resulting in repeated identification of the three-dimensional identified object.

[0146] In the object detection frame filtering method provided in the embodiment of the present application, the offset parameter can represent the degree of consistency of the position and size of the object detection frame. In this way, even if the two object detection frames do not overlap or the intersection volume is very small, if the two object detection frames correspond to the same three-dimensional solid object, the degree of consistency of the position and size of the two object detection frames will still be high. In this way, the problem that the technology of filtering object detection frames based on the intersection-over-union ratio cannot filter out duplicate object detection frames is solved, and the phenomenon of repeated recognition of three-dimensional objects is avoided.

[0147] For example, suppose the offset parameter is passed Figure 4If the offset parameter in a certain direction is less than zero, it means that the overlap used to calculate the offset parameter is less than zero. The projection length of the center point of the reference object detection frame and the center point of the object detection frame being calculated in that direction is greater than the maximum distance between the center points of the reference object detection frame and the center points of the object detection frame being calculated when they are exactly touching. The reference object detection frame and the object detection frame being calculated do not overlap in this direction.

[0148] Take the long side direction as an example to illustrate. ij l In formula (2), the minuend is The longest distance between the center points of the i-th object detection frame and the j-th object detection frame on the long side when they overlap. The subtrahend is the projection length of the center position offset on the long side, indicating the projection distance between the center point of the i-th object detection frame and the center point of the j-th object detection frame on the long side of the i-th object detection frame.

[0149] When the i-th object detection frame and the j-th object detection frame have the same orientation and are exactly touching (i.e., the faces perpendicular to the long side exactly coincide), the projection of the distance between the center point of the i-th object detection frame and the center point of the j-th object detection frame on the long side is Therefore, if the overlap in the long side direction is ij l If the overlap in the long side direction is less than zero, it means that even if the i-th object detection frame and the j-th object detection frame have the same orientation, there is no overlap between the i-th object detection frame and the j-th object detection frame. ij l (or long side offset parameter O ij l ) is less than 0, indicating that the i-th object detection box and the j-th object detection box do not overlap on the long side.

[0150] Therefore, if the offset upper threshold is less than 0, then when the offset parameter is less than 0 but greater than the offset upper threshold, it is still possible to determine that the reference object detection frame and the object detection frame to be calculated correspond to the same 3D solid object based on the offset parameter. In other words, if the reference object detection frame and the object detection frame to be calculated do not overlap, resulting in the calculated offset parameter being less than 0, it is still possible to determine whether the reference object detection frame and the object detection frame to be calculated correspond to the same 3D solid object using the offset upper threshold parameter being less than 0.

[0151] Alternatively, in some other possible implementations, a minimum distance threshold between object detection frames may be pre-set. Before filtering based on the offset parameter, it may be determined whether the shortest distance between the object detection frames is greater than the shortest distance threshold. If the shortest distance between two object detection frames is greater than the shortest distance threshold, it indicates that the shortest distance between the two object detection frames is far, and the two object detection frames correspond to different three-dimensional physical objects. If the shortest distance between the two object detection frames is close, it is considered that there is a possibility that the two object detection frames correspond to the same three-dimensional physical object, and the offset parameter is then used to determine whether the two object detection frames are duplicates.

[0152] Optionally, the “shortest distance between two object detection frames” may include the shortest distance between the two object detection frames in the short side direction.

[0153] Based on the above method embodiment, the present application embodiment also provides a physical object detection frame filtering device. Figure 5 , Figure 5 A schematic diagram of a physical object detection frame filtering device provided in an embodiment of the present application.

[0154] exist Figure 5 In the illustrated implementation, the filtering device 500 for the entity object detection frame includes:

[0155] An acquiring unit 510 is configured to acquire feature parameters of object detection frames of a plurality of three-dimensional solid objects of the same type, wherein the feature parameters and types of the object detection frames are pre-identified based on point cloud data;

[0156] an offset parameter calculation unit 520 for obtaining a plurality of offset parameters based on the feature parameters of the object detection frames, wherein an offset parameter represents a degree of consistency between positions and sizes of two object detection frames in the plurality of object detections;

[0157] The filtering unit 530 is configured to filter out one of the two repeated object detection frames according to the offset parameter until the offset parameters are calculated for all the multiple object detection frames, and output the retained object detection frames.

[0158] In some possible implementations, the filtering unit 530 is specifically used to establish an initial set T and a target set R, wherein the initial set T is used to store the multiple object detection frames, and the multiple object detection frames in T are sorted in order from high to low according to confidence; the target set R is used to store the retained object detection frames.

[0159] In some possible implementations, the filtering unit 530 is specifically configured to determine an object detection frame with the highest confidence in the initial set T as a reference object detection frame, and transfer the reference object detection frame from the initial set T to the target set R; the offset parameter calculation unit 520 is specifically configured to calculate an offset parameter based on feature parameters of the reference object detection frame and an object detection frame to be calculated, where the object detection frame to be calculated is an object detection frame in the initial set T that is sorted after the reference object detection frame.

[0160] In some possible implementations, the filtering unit 530 is specifically configured to determine, based on the offset parameters of the reference object detection frame and the object detection frame to be calculated, whether there is duplication between the reference object detection frame and the detection frame to be calculated; if there is duplication, filtering out the object detection frame to be calculated from the initial set T, and then determining whether there is an object detection frame in the initial set T that is sorted after the reference object detection frame and whose offset parameters are not calculated with the reference object detection frame; if there is no duplication, determining whether there is an object detection frame in the initial set T that is sorted after the reference object detection frame and whose offset parameters are not calculated with the reference object detection frame; if there is For an object detection frame that is sorted after the reference object detection frame and has not calculated an offset parameter with the reference object detection frame, the object detection frame is used as the object detection frame to be calculated, and the step of calculating the offset parameter according to the feature parameters of the reference object detection frame and the object detection frame to be calculated is performed; if there is no object detection frame in the initial set T that is sorted after the reference and has not calculated an offset parameter with the reference object detection frame, and the number of remaining object detection frames in the initial set T is greater than 1, the step of determining the object detection frame with the highest confidence in the initial set T as the reference object detection frame and transferring the reference object detection frame from the initial set T to the target set R is performed.

[0161] In some possible implementations, the feature parameters of the object detection frame include a center position parameter, an orientation parameter, and a size parameter of the object detection frame; the offset parameter calculation unit 520 is specifically configured to calculate a center position offset based on the center position parameters of the reference object detection frame and the object detection frame to be calculated; determine a direction vector of the reference object detection frame based on the orientation parameter of the reference object detection frame; and obtain the offset parameter based on the center position offset, the direction vector of the reference object detection frame, the size parameter of the reference object detection frame, and the size parameter of the object detection frame to be calculated.

[0162] In some possible implementations, the direction vector of the object detection frame includes a long side direction vector, a short side direction vector, and a vertical direction vector; the size parameters include: a length in the long side direction, a width in the short side direction, and a height in the vertical direction; and the one offset parameter includes: a long side offset parameter, a short side offset parameter, and a vertical offset parameter; the offset parameter calculation unit 520 is specifically configured to obtain, based on the center position offset, the long side direction vector, the short side direction vector, and the vertical direction vector, a projection length of the center position offset in the long side direction of the reference object detection frame, a projection length of the center position offset in the short side direction of the reference object detection frame, and a projection length of the center position offset in the vertical direction of the reference object detection frame; obtain the long side offset parameter based on the projection length in the long side direction and the lengths of the two object detection frames; and / or obtain the short side offset parameter based on the projection length in the short side direction and the widths of the two object detection frames; and / or obtain the vertical offset parameter based on the projection length in the vertical direction and the heights of the two object detection frames.

[0163] In some possible implementations, the offset parameter calculation unit 520 is specifically used to calculate the sum of the length of the reference object detection frame and the length of the object detection frame to be calculated; obtain the long side offset parameter based on the difference between half of the sum of the lengths and the projection length in the long side direction; and / or calculate the sum of the width of the reference object detection frame and the width of the object detection frame to be calculated; obtain the short side offset parameter based on the difference between half of the sum of the widths and the projection length in the short side direction; and / or calculate the sum of the height of the reference object detection frame and the height of the object detection frame to be calculated; obtain the vertical offset parameter based on the difference between half of the sum of the heights and the projection length in the vertical direction.

[0164] In some possible implementations, the filtering unit 510 is specifically used to determine whether the offset parameters of the reference object detection frame and the object detection frame to be calculated, including the long side offset parameter, the short side offset parameter, and the vertical offset parameter, are all greater than an offset upper limit threshold, and the offset upper limit threshold is less than 0; if greater, it is determined that there is duplication between the reference object detection frame and the detection frame to be calculated.

[0165] Based on the above method embodiment and apparatus embodiment, the present application also provides a device, which will be described below with reference to the accompanying drawings.

[0166] See also Figure 6 , Figure 6 This is a schematic diagram of a device provided in an embodiment of the present application. The device 600 includes: a processor 601, a memory 602, and a bus 603.

[0167] The processor 601 and the memory 602 communicate with each other via the bus 603 .

[0168] The processor 601 is configured to call program instructions in the memory 602 so as to implement the travel planning method described in any of the above embodiments in the device 600. The electronic device of the present application may be a server, a server cluster, a mobile phone, a PC, a PAD, a smart wearable device, etc.

[0169] The present application also provides a computer-readable storage medium, wherein the storage medium stores at least one instruction. When the instruction is executed on a computer, the computer executes the method described in any of the aforementioned embodiments.

[0170] The present application also provides a computer program product, characterized in that the computer program product includes one or more computer program instructions, and when the computer program instructions are loaded and executed by a computer, the computer is caused to execute the method as described in any of the aforementioned embodiments.

[0171] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

[0172] Here, “A refers to B” or “A sees B” means that A is the same as B or A is a simple variation of B.

[0173] The terms "first" and "second" in the description and claims of the embodiments of this application are used to distinguish different objects, not to describe a specific order of objects, and should not be construed as indicating or implying relative importance. For example, the terms "first road segment" and "second road segment" are used to distinguish different road segments, not to describe a specific order of road segments, and should not be construed as implying that the first road segment is more important than the second road segment.

[0174] The information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the road information involved in this application was obtained with full authorization.

[0175] In the embodiments of the present application, unless otherwise specified, "at least one" means one or more, and "a plurality" means two or more. For example, a plurality of road segments refers to two or more road segments.

[0176] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in accordance with the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

[0177] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for filtering an object detection frame, characterized in that: The method comprises: Obtaining feature parameters of object detection frames of multiple three-dimensional solid objects of the same type, wherein the feature parameters and types of the object detection frames are pre-identified based on point cloud data; Obtaining a plurality of offset parameters according to the characteristic parameters of the object detection frames, wherein an offset parameter represents a degree of consistency between positions and sizes of two object detection frames in the plurality of object detections; According to the offset parameter, one of the two repeated object detection frames is filtered out until the offset parameters of the multiple object detection frames are all calculated, and the retained object detection frame is output.

2. The method according to claim 1, characterized in that The method further comprises: An initial set T and a target set R are established, wherein the initial set T is used to store the multiple object detection frames obtained, and the multiple object detection frames in the initial set T are sorted in order from high to low according to confidence; and the target set R is used to store the retained object detection frames.

3. According to claim 2, it is characterized in that The method further comprises: Determine an object detection frame with the highest confidence in the initial set T as a reference object detection frame; Transferring the reference object detection box from the initial set T to the target set R; Obtaining the offset parameter according to the feature parameter of the object detection frame specifically includes: An offset parameter is calculated according to feature parameters of the reference object detection frame and a to-be-calculated object detection frame, where the to-be-calculated object detection frame is an object detection frame in the initial set T that is sorted after the reference object detection frame.

4. The method according to claim 3, characterized in that The step of filtering out one of the two repeated object detection frames according to the offset parameter until the offset parameters of the multiple object detection frames are all calculated includes: Determining whether there is overlap between the reference object detection frame and the detection frame to be calculated based on offset parameters of the reference object detection frame and the detection frame to be calculated; If there is duplication, filtering out the object detection frame to be calculated from the initial set T, and then determining whether there is an object detection frame in the initial set T that is sorted after the reference object detection frame and whose offset parameter is not calculated with the reference object detection frame; If there is no duplication, determining whether there is an object detection frame in the initial set T that is sorted after the reference object detection frame and whose offset parameter is not calculated with the reference object detection frame; If there is an object detection frame in the initial set T that is sorted after the reference object detection frame and whose offset parameters have not been calculated with the reference object detection frame, then use the object detection frame as the object detection frame to be calculated, and perform the step of calculating the offset parameters based on the feature parameters of the reference object detection frame and the object detection frame to be calculated; If there is no object detection frame in the initial set T that is sorted after the reference and has not calculated an offset parameter with the reference object detection frame, and the number of remaining object detection frames in the initial set T is greater than 1, then the step of determining the object detection frame with the highest confidence in the initial set T as the reference object detection frame and transferring the reference object detection frame from the initial set T to the target set R is performed.

5. The method according to claim 3 or 4, characterized in that The characteristic parameters include center position parameters, orientation parameters and size parameters; Calculating the offset parameter according to the feature parameters of the reference object detection frame and the object detection frame to be calculated includes: Calculating a center position offset according to center position parameters of the reference object detection frame and the object detection frame to be calculated; Determining a direction vector of the reference object detection frame according to an orientation parameter of the reference object detection frame; The offset parameter is obtained according to the center position offset, the direction vector of the reference object detection frame, the size parameter of the reference object detection frame, and the size parameter of the object detection frame to be calculated.

6. The method according to claim 5, characterized in that The direction vector of the reference object detection frame includes a long side direction vector, a short side direction vector, and a vertical direction vector; the size parameters include: the length of the long side, the width of the short side, and the height of the vertical direction; the offset parameters include: a long side offset parameter, a short side offset parameter, and a vertical offset parameter; The obtaining of the offset parameter according to the center position offset, the direction vector of the reference object detection frame, the size parameter of the reference object detection frame, and the size parameter of the object detection frame to be calculated includes: Obtaining, according to the center position offset, the long side direction vector, the short side direction vector, and the vertical direction vector, a projection length of the center position offset in the long side direction of the reference object detection frame, a projection length of the center position offset in the short side direction of the reference object detection frame, and a projection length of the center position offset in the vertical direction of the reference object detection frame; Obtaining the long side offset parameter according to the projection length in the long side direction, the length of the reference object detection frame, and the length of the reference object detection frame to be calculated; and / or, Obtaining the short side offset parameter according to the projection length in the short side direction, the width of the reference object detection frame, and the width of the reference object detection frame to be calculated; and / or, The vertical offset parameter is obtained according to the projection length in the vertical direction, the height of the reference object detection frame, and the height of the reference object detection frame to be calculated.

7. The method according to claim 6, characterized in that Obtaining the long side offset parameter according to the projection length in the long side direction, the length of the reference object detection frame, and the length of the reference object detection frame to be calculated includes: Calculating the sum of the length of the reference object detection frame and the length of the object detection frame to be calculated; Obtaining the long side offset parameter according to a difference between half of the sum of the lengths and the projection length in the long side direction; The short side offset parameter is obtained according to the projection length in the short side direction, the width of the reference object detection frame, and the width of the reference object detection frame to be calculated, including: Calculating the sum of the width of the reference object detection frame and the width of the object detection frame to be calculated; Obtaining the short side offset parameter according to a difference between half of the sum of the widths and the projection length in the short side direction; Obtaining the vertical offset parameter according to the projection length in the vertical direction, the height of the reference object detection frame, and the height of the reference object detection frame to be calculated includes: Calculating the sum of the height of the reference object detection frame and the height of the object detection frame to be calculated; The vertical offset parameter is obtained according to the difference between half of the sum of the heights and the projection length in the vertical direction.

8. The method according to claim 6, characterized in that The determining, based on the offset parameters of the reference object detection frame and the object detection frame to be calculated, whether there is duplication between the reference object detection frame and the detection frame to be calculated comprises: Determine whether the offset parameters of the reference object detection frame and the object detection frame to be calculated, including the long side offset parameter, the short side offset parameter, and the vertical offset parameter, are all greater than an offset upper limit threshold, and the offset upper limit threshold is less than 0; if greater than, determine that there is duplication between the reference object detection frame and the detection frame to be calculated.

9. A filtering device for an object detection frame, characterized in that: The device comprises: An acquiring unit, configured to acquire feature parameters of object detection frames of a plurality of three-dimensional solid objects of the same type, wherein the feature parameters and types of the object detection frames are pre-identified based on point cloud data; an offset parameter calculation unit, configured to obtain a plurality of offset parameters based on the feature parameters of the object detection frames, wherein an offset parameter represents a degree of consistency between positions and sizes of two object detection frames in the plurality of object detections; The filtering unit is configured to filter out one of the two repeated object detection frames according to the offset parameter until the offset parameters are calculated for all the multiple object detection frames, and output the retained object detection frame.

10. A device, characterized in that The device includes: a processor, the processor is coupled to a memory, the memory stores at least one computer program instruction, and the at least one computer program instruction is loaded and executed by the processor, so that the device implements the method according to any one of claims 1 to 8.