Precision test method and device of target detection model, electronic equipment and medium
By projecting the target labeling information and detection results in the two-dimensional space into the three-dimensional space, and using vision sensors and camera calibration information, the performance evaluation problem of the target detection model in the three-dimensional space is solved, the cost of lidar truth value labeling is reduced, and efficient three-dimensional space accuracy testing is achieved.
Patent Information
- Application Number
- CN202510519942.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-08-08
AI Technical Summary
The lack of performance evaluation of the target detection model in three-dimensional space in the prior art has affected the implementation of automatic parking function and the cost of lidar truth value labeling is high.
By projecting the target labeling information and target detection results in the two-dimensional space into the three-dimensional space, the three-dimensional data projection is used to use the two-dimensional image collected by the visual sensor, and combining the camera calibration information, the performance evaluation of the target detection model in the three-dimensional space is achieved.
The accuracy test of the target detection model in three-dimensional space is realized, which reduces the cost of lidar truth value annotation and improves the efficiency and accuracy of performance evaluation.
Smart Images

Figure CN120451713A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of target detection technology, and in particular to a method, device, electronic device, and medium for testing the accuracy of a target detection model. Background Art
[0002] With the rapid development of advanced driver assistance technology, automatic parking (APA) has garnered increasing attention as a core feature of the technology. Object detection algorithms are one of the most critical algorithms in APA, used to identify obstacles such as people, vehicles, and cones during parking and patrolling. Therefore, the accuracy of object detection algorithms plays a crucial role in the implementation of APA.
[0003] At present, in related technologies, the performance evaluation of corresponding target detection algorithms usually evaluates the detection rate, recall rate, regression accuracy of the detection frame, etc. of the target detection algorithm in the image plane, but lacks performance evaluation in three-dimensional space. Summary of the Invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, at least one embodiment of the present disclosure provides a method, device, electronic device and medium for testing the accuracy of a target detection model.
[0005] In a first aspect, the present disclosure provides a method for testing the accuracy of an object detection model, comprising:
[0006] Obtaining a test image and target annotation information of the test image;
[0007] Inputting the test image into the target detection model to be tested to perform target detection, and obtaining target detection information corresponding to the test image;
[0008] transforming the target annotation information into three-dimensional annotation information, and transforming the target detection information into three-dimensional detection information;
[0009] An accuracy test result of the target detection model is determined based on the target annotation information, the target detection information, the three-dimensional annotation information, and the three-dimensional detection information.
[0010] In a second aspect, the present disclosure provides an accuracy testing device for an object detection model, comprising:
[0011] A first acquisition module is used to acquire a test image and target annotation information of the test image;
[0012] A second acquisition module is used to input the test image into the target detection model to be tested to perform target detection and obtain target detection information corresponding to the test image;
[0013] a conversion module, configured to convert the target annotation information into three-dimensional annotation information, and to convert the target detection information into three-dimensional detection information;
[0014] A determination module is used to determine an accuracy test result of the target detection model based on the target annotation information, the target detection information, the three-dimensional annotation information and the three-dimensional detection information.
[0015] In a third aspect, the present disclosure provides an electronic device, comprising: a processor and a memory;
[0016] The processor is used to execute the accuracy testing method of the target detection model as described in the first aspect by calling the program or instruction stored in the memory.
[0017] In a fourth aspect, the present disclosure provides a computer-readable storage medium, which stores a program or instruction, and the program or instruction enables a computer to execute the accuracy testing method of the target detection model as described in the first aspect.
[0018] In a fifth aspect, the present disclosure provides a computer program product, which is used to execute the accuracy testing method of the target detection model as described in the first aspect.
[0019] The technical solution provided by the embodiments of the present disclosure has at least the following advantages compared with the prior art:
[0020] In the disclosed embodiment, a test image and target annotation information of the test image are obtained, and the test image is input into the target detection model to be tested for target detection, and target detection information corresponding to the test image is obtained. Then, the target annotation information is transformed into three-dimensional annotation information, and the target detection information is transformed into three-dimensional detection information. Finally, based on the target annotation information, target detection information, three-dimensional annotation information and three-dimensional detection information, the accuracy test result of the target detection model is determined. By adopting the above technical solution, the target annotation information of the test image and the target detection information obtained by using the target detection model are respectively transformed into three-dimensional annotation information and three-dimensional detection information, the projection of two-dimensional information to three-dimensional space is realized, which solves the problem of lack of three-dimensional data in the accuracy test of the target detection model, and realizes the accuracy test of the target detection model in two-dimensional space and three-dimensional space simultaneously based on the test image in two-dimensional space, and it can be completed without relying on the laser radar to collect the true value of three-dimensional information, thereby reducing the cost of evaluating the performance of the target detection model in three-dimensional space. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0022] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0023] Figure 1 A flowchart of a method for testing the accuracy of an object detection model provided by an exemplary embodiment of the present disclosure;
[0024] Figure 2 A flowchart of a method for testing the accuracy of an object detection model provided by another exemplary embodiment of the present disclosure;
[0025] Figure 3 A flowchart of a method for testing the accuracy of an object detection model provided by another exemplary embodiment of the present disclosure;
[0026] Figure 4 A schematic diagram of the accuracy test process of the target detection model according to an exemplary embodiment of the present disclosure is shown;
[0027] Figure 5 A schematic diagram of the structure of an accuracy testing device for an object detection model provided in one embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the present disclosure is further described in detail below with reference to the accompanying drawings and embodiments. It is understood that the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. The specific embodiments described herein are only used to explain the present disclosure, rather than to limit the present disclosure. In the absence of conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field fall within the scope of protection of the present disclosure.
[0029] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.
[0030] The performance evaluation of target detection models includes two dimensions, two-dimensional and three-dimensional. For performance evaluation in two-dimensional space, the detection rate, recall rate, and regression accuracy of the detection box of the target detection algorithm in the image plane are usually evaluated. During the evaluation, only the relevant information of the target in the image needs to be labeled, which is low-cost. For performance evaluation in three-dimensional space, the position accuracy and heading accuracy of the target detection algorithm in the world coordinate system are usually evaluated. This requires the use of lidar to collect the true value of three-dimensional information, which is relatively costly.
[0031] In response to the above problems, the present disclosure provides a method for testing the accuracy of a target detection model. By projecting the target annotation information and target detection results in the two-dimensional space into the three-dimensional space, the projected data is used to evaluate the performance of the target detection model in the three-dimensional space, thereby solving the problem of lack of three-dimensional data. The method can realize the test of the position accuracy of the target detection model in the three-dimensional space. This process only relies on visual sensors and only requires the addition of camera calibration information. It does not need to rely on lidar, and the cost is low, which solves the high cost problem of lidar true value annotation.
[0032] The following describes in detail the specific implementations of the target detection model accuracy testing method, device, electronic device, and medium disclosed in the present invention in conjunction with the accompanying drawings.
[0033] Figure 1 A flowchart of a method for testing the accuracy of a target detection model provided in an exemplary embodiment of the present disclosure is provided. The method can be performed by a device for testing the accuracy of a target detection model provided in an embodiment of the present disclosure. The device for testing the accuracy of a target detection model can be implemented using software and / or hardware and can be integrated into an electronic device.
[0034] like Figure 1 As shown, the accuracy testing method of the target detection model may include the following steps:
[0035] Step 101: Acquire a test image and target annotation information of the test image.
[0036] There are multiple test images, and target annotation is performed for each test image, and annotation information of all targets on the test image (i.e., target annotation information) is annotated. The annotation information includes, but is not limited to, part or all of the target type (such as vehicle, pedestrian, obstacle, etc.), target annotation box information (used to indicate the location of the target), the location of the target ground point (annotated coordinates), and target occlusion attributes (including whether it is occluded, occlusion level, etc.). When the target occlusion attribute is included in the annotation information, whether the corresponding target participates in the matching of the target detection value can be determined based on the target occlusion attribute. For example, if the occlusion level of the target is high (the area of the target occluded is large), the target does not participate in the matching of its detection value.
[0037] Optionally, test images can be collected from multiple scenes, such as above ground, underground, sunny, cloudy, rainy, cement, brick, asphalt, with distant and near targets, etc. By enriching the collection scenes of test images, the performance of the target detection model in different scenes can be comprehensively evaluated to improve the accuracy of the test results.
[0038] Step 102: Input the test image into the target detection model to be tested to perform target detection, and obtain target detection information corresponding to the test image.
[0039] Among them, the target detection model is pre-trained and can identify targets such as vehicles and pedestrians in the input image and output the detection information of the target. The detection information may include but is not limited to the type of detected target, target detection box information used to indicate the location of the target, and the coordinates of the grounding point of the target that is in contact with the ground.
[0040] In this embodiment, for each acquired test image, the test image can be input into the target detection model to be tested, and the target detection model performs target detection on the input test image and outputs target detection information corresponding to all targets detected in the test image.
[0041] Step 103: transform the target annotation information into three-dimensional annotation information, and transform the target detection information into three-dimensional detection information.
[0042] In this embodiment, the target annotation information and target detection information corresponding to each test image can be transformed into three-dimensional space to obtain three-dimensional annotation information corresponding to the target annotation information and three-dimensional detection information corresponding to the target detection information. When performing the coordinate system conversion, all coordinates in the target annotation information and target detection information can be transformed, or only the coordinates of the target ground point can be converted, which is not limited by the present disclosure. Through this transformation, the corresponding three-dimensional information of the target can be obtained by relying solely on the two-dimensional image collected by the visual sensor, without relying on laser radar to collect data, thereby reducing costs.
[0043] As an example, when transforming two-dimensional target labeling information and target detection information into three-dimensional space data, the target labeling information and target detection information can be first converted from the pixel coordinate system to the camera coordinate system, and then converted from the camera coordinate system to the three-dimensional vehicle coordinate system to obtain three-dimensional labeling information and three-dimensional detection information.
[0044] Step 104 : Determine an accuracy test result of the target detection model based on the target annotation information, the target detection information, the 3D annotation information, and the 3D detection information.
[0045] In this embodiment, the accuracy test results of the target detection model in two dimensions can be determined based on the target annotation information and the target detection information, and the accuracy test results of the target detection model in three dimensions can be determined based on the three-dimensional annotation information and the three-dimensional detection information. Thus, the accuracy of the target detection model in both two and three dimensions can be tested solely using two-dimensional images captured by the visual sensor, without relying on three-dimensional data collected by the lidar. This ensures the acquisition of three-dimensional spatial data while reducing data acquisition costs.
[0046] Exemplarily, the accuracy indicators of two-dimensional space may include precision, recall, position accuracy, etc. By matching the target labeling information and target detection information of each target in the same test image, true positive samples, false positive samples, and false negative samples can be identified, wherein the target detection result that successfully matches the target labeling result is identified as a true positive sample, the target detection result that does not match the target labeling result is identified as a false positive sample, and the target labeling result that is not matched is identified as a false negative sample. Then, the precision and recall rates are determined based on the number of true positive samples, the number of false positive samples, and the number of false negative samples. For example, the precision rate can be expressed as the ratio of the true positive samples divided by the sum of the true positive samples and the false positive samples, and the recall rate can be expressed as the ratio of the sum of the true positive samples and the false positive samples divided by the sum of the true positive samples and the false negative samples. In addition, in the true positive samples, by comparing the two-dimensional position difference and the three-dimensional position difference between the labeling box and the detection box corresponding to the same target, the position accuracy of the target detection model in two-dimensional space and the position accuracy in three-dimensional space can be determined.
[0047] In an optional embodiment of the present disclosure, it is also possible to count the targets falling in a preset area of interest, for example, 5m*5m in front and left and right, and determine the detection accuracy of the target detection model in the corresponding area based on the target labeling information, target detection information, three-dimensional labeling information and three-dimensional detection information of these targets.
[0048] The accuracy test method of the target detection model of the embodiment of the present disclosure obtains a test image and target annotation information of the test image, and inputs the test image into the target detection model to be tested for target detection, obtains target detection information corresponding to the test image, then transforms the target annotation information into three-dimensional annotation information, and transforms the target detection information into three-dimensional detection information, and finally, determines the accuracy test result of the target detection model based on the target annotation information, target detection information, three-dimensional annotation information and three-dimensional detection information. By adopting the above technical solution, by transforming the target annotation information of the test image and the target detection information obtained by using the target detection model into three-dimensional annotation information and three-dimensional detection information respectively, the projection of two-dimensional information to three-dimensional space is realized, which solves the problem of lack of three-dimensional data in the accuracy test of the target detection model, and realizes the accuracy test of the target detection model in two-dimensional space and three-dimensional space simultaneously based on the test image in two-dimensional space, and can be completed without relying on the true value of three-dimensional information collected by the laser radar, thereby reducing the cost of evaluating the performance of the target detection model in three-dimensional space.
[0049] In an optional embodiment of the present disclosure, the target marking information includes the marking coordinates of the target grounding point, and the target detection information includes the detection coordinates of the target grounding point, such as Figure 2 As shown, based on the above embodiment, step 103 may include the following sub-steps:
[0050] Step 201 : Dedistort and transform the key point coordinates to obtain the dedistorted coordinates of the key point coordinates on the camera normalization plane, wherein the key point coordinates include the marked coordinates and the detected coordinates.
[0051] Typically, most cameras installed on vehicles are fisheye cameras. Test images captured by fisheye cameras are distorted. To obtain the true position of the target touchdown point in the vehicle coordinate system, in this embodiment, the detected coordinates and annotated coordinates of the target touchdown point can be dedistorted.
[0052] As an example, the currently commonly used fisheye camera dedistortion algorithm can be used to dedistort the labeled coordinates and detected coordinates of the target ground point respectively, and then the dedistorted labeled coordinates and detected coordinates are respectively converted to the camera normalized plane to obtain the dedistorted labeled coordinates and dedistorted detected coordinates of the target ground point on the camera normalized plane (collectively referred to as dedistorted coordinates).
[0053] As another example, the camera intrinsic parameters and distortion coefficients of the camera can be obtained, wherein the camera is a fisheye camera, the camera intrinsic parameters and distortion coefficients are pre-calibrated, and the camera intrinsic parameters include the camera focal length (f x , f y ) and the optical center (c x , c y), the distortion coefficients include radial distortion coefficients (k1, k2, and k3) and tangential distortion coefficients (p1 and p2). Next, the keypoint coordinates are converted to the camera normalization plane based on the camera intrinsic parameters to obtain the camera coordinates of the keypoint coordinates on the camera normalization plane. The distortion amount is then determined based on the camera coordinates and the distortion coefficients. Finally, the camera coordinates are dedistorted based on the distortion amount to obtain the dedistorted coordinates of the keypoint coordinates on the camera normalization plane.
[0054] The camera normalization plane is a virtual two-dimensional plane located in front of the camera, and the z coordinate of any coordinate point on the plane is 1. Assuming that the key point coordinates of the target ground point in the test image (fisheye original image) are (u, v), the key point coordinates are converted to the camera normalization plane based on the camera intrinsic parameters. The camera coordinates of the obtained camera normalization plane are marked as (x, y, 1). The camera coordinates can be converted using the following formula (1):
[0055]
[0056] Next, based on the obtained camera coordinates and the obtained distortion coefficients, the formula for determining the distortion amount is shown in formula (2):
[0057] Δx=x*(k1r 2 +k2r 4 +k3r 6 )+2p1xy+p2(r 2 +2x 2 )
[0058] Δy=y*(k1r 2 +k2r 4 +k3r 6 )+2p2xy+p1(r 2 +2y 2 )(2)
[0059] Among them, r 2 =x 2 +y 2
[0060] Finally, dedistortion is performed based on the obtained distortion amount and camera coordinates to obtain the dedistorted coordinates of the key point coordinates on the camera normalized plane. The dedistorted coordinates are marked as (x', y', 1), and the dedistorted coordinates can be expressed as shown in the following formula (3):
[0061]
[0062] Thus, through the above dedistortion and coordinate conversion operations, the dedistorted coordinates of the labeled coordinates of the target ground point on the camera normalized plane can be obtained, as well as the dedistorted coordinates of the detected coordinates of the target ground point on the camera normalized plane.
[0063] In step 202, the target touchdown point is located on the ground in the vehicle coordinate system, and the dedistorted coordinates are converted to the vehicle coordinate system to obtain the three-dimensional coordinates of the target touchdown point in the vehicle coordinate system, where the three-dimensional coordinates include the three-dimensional annotated coordinates corresponding to the annotated coordinates and the three-dimensional detected coordinates corresponding to the detected coordinates.
[0064] Among them, the vehicle body coordinate system is used to describe the relative position relationship between the objects around the vehicle and the vehicle. The vehicle body coordinate system of this embodiment can be defined using the currently commonly used vehicle body coordinate system definition method, and this disclosure does not impose any restrictions on this.
[0065] In this embodiment, after obtaining the undistorted coordinates of the target ground point in the camera normalization plane, the undistorted coordinates can be converted to the vehicle coordinate system. Since the undistorted coordinates are the coordinates of the target ground point, the target ground point can be located on the ground in the vehicle coordinate system. The depth of the target ground point (i.e., the coordinate value in the z-axis direction) is adjusted so that it touches the ground in the vehicle coordinate system to obtain the three-dimensional coordinates of the target ground point in the vehicle coordinate system. It can be understood that when the undistorted coordinates are obtained by undistorting and converting the labeled coordinates of the target ground point, the obtained three-dimensional coordinates are the three-dimensional coordinates corresponding to the labeled coordinates of the target ground point (referred to as the three-dimensional labeled coordinates); when the undistorted coordinates are obtained by undistorting and converting the detected coordinates of the target ground point, the obtained three-dimensional coordinates are the three-dimensional coordinates corresponding to the detected coordinates of the target ground point (referred to as the three-dimensional detected coordinates).
[0066] In an optional embodiment of the present disclosure, when converting the dedistorted coordinates into three-dimensional coordinates in the vehicle coordinate system, the camera extrinsic matrix of the camera can be first obtained, wherein the camera extrinsic matrix is calibrated after the camera is installed, and the camera extrinsic matrix is composed of parameters used to describe the position and posture of the fisheye camera in the vehicle coordinate system, and may include a rotation matrix and a translation vector. Then, a depth parameter is set for the dedistorted coordinates to obtain a first moving coordinate, and the first moving coordinate is converted to the vehicle coordinate system based on the camera extrinsic matrix to obtain a second moving coordinate. Finally, with the target ground point located on the ground in the vehicle coordinate system as the target, the depth parameter of the second moving coordinate is adjusted to obtain the three-dimensional coordinate of the target ground point in the vehicle coordinate system.
[0067] Assuming the dedistorted coordinates are (x', y', 1), the first moving coordinates can be expressed as (x'*d, y'*d, d), where d is the depth parameter. The first moving coordinates are converted to the vehicle coordinate system to obtain the second moving coordinates. Assuming the second moving coordinates are marked as (x", y", z"), the second moving coordinates can be expressed as shown in the following formula (4):
[0068]
[0069] In formula (4), T represents the camera extrinsic parameter matrix. By continuously adjusting the value of d so that z″ = 0, the depth parameter d can be obtained. Substituting the obtained depth parameter d into formula (4) above, the three-dimensional coordinates of the target ground point in the vehicle coordinate system can be obtained.
[0070] The accuracy testing method of the target detection model of the embodiment of the present disclosure dedistorts and transforms the key point coordinates to obtain the dedistorted coordinates of the key point coordinates on the normalized plane of the camera. The key point coordinates include the labeled coordinates and the detected coordinates. Then, with the target ground point located on the ground in the vehicle coordinate system as the target, the dedistorted coordinates are transformed into the vehicle coordinate system to obtain the three-dimensional coordinates of the target ground point in the vehicle coordinate system. The three-dimensional coordinates include the three-dimensional labeled coordinates corresponding to the labeled coordinates and the three-dimensional detected coordinates corresponding to the detected coordinates. Therefore, by dedistorting the labeled coordinates and the detected coordinates of the target ground point and then transforming them into the vehicle coordinate system, the accuracy of the converted three-dimensional detected coordinates and three-dimensional labeled coordinates is guaranteed.
[0071] In an optional embodiment of the present disclosure, Figure 3 As shown, based on the above embodiment, step 104 may include the following sub-steps:
[0072] Step 301 : Compare the target labeling information and the target detection information to determine a sample pair, where the sample pair includes matching target labeling information and target detection information.
[0073] In this embodiment, for each target in the test image, the target labeling information and target detection information corresponding to the same target can be compared to determine whether the two match. If they match, the successfully matched target labeling information and target detection information are combined into a sample pair, thereby obtaining multiple sample pairs.
[0074] In an optional embodiment of the present disclosure, the target labeling information and target detection information corresponding to the same target can be compared to determine whether the target labeling information and the target detection information meet the preset conditions. If it is determined that the preset conditions are met, the target labeling information and the target detection information currently being compared are determined to be a sample pair.
[0075] As an example, the preset condition may be that the target type is consistent and the difference in the grounding point coordinates is within a preset range. Thus, when determining a sample pair, the target detection information and the target annotation information are compared to see whether the target type of the same target is consistent, and whether the difference between the detected coordinates and the annotated coordinates of the target grounding point is within a preset range. If the target type is consistent and the difference between the detected coordinates and the annotated coordinates of the target grounding point is within a preset range, then the currently compared target detection information and target annotation information are determined to be a successful match and are determined to be a sample pair. Specifically, when determining whether the difference between the detected coordinates and the annotated coordinates is within a preset range, the horizontal coordinate difference and the vertical coordinate difference between the detected coordinates and the annotated coordinates can be determined first. If both the horizontal coordinate difference and the vertical coordinate difference are within the preset range, then the difference between the detected coordinates and the annotated coordinates is determined to be within the preset range.
[0076] As another example, the target annotation information also includes the annotation type and target annotation frame information of the target, and the target detection information also includes the detection type and target detection frame information of the target. The preset condition can be that the annotation type and detection type of the same target are consistent, and the difference between the target annotation frame information and the target detection frame information of the same target is within a preset error range, and the deviation between the annotation coordinates and the detection coordinates of the target ground point of the same target is less than a preset value. Thus, when determining a sample pair, the target detection information and the target annotation information are compared to see whether the target type of the same target is consistent, whether the difference between the target annotation frame information and the target detection frame information is within a preset error range, and whether the deviation between the detection coordinates and the annotation coordinates of the target ground point is less than a preset value. If, for the same target, its annotation type and detection type are consistent, and the difference between the target annotation frame information and the target detection frame information is within a preset error range, and the deviation between the detection coordinates and the annotation coordinates of the target ground point is less than a preset value, then it is determined that the currently compared target detection information and target annotation information match successfully and are determined as a sample pair. Among them, when judging whether the difference between the target annotation frame information and the target detection frame information of the same target is within a preset error range, the coordinate difference of the four corner points of the frame can be compared based on the target annotation frame information and the target detection frame information. If the coordinate difference of the four corner points is all within the preset error range, then it is determined that the difference between the target annotation frame information and the target detection frame information of the target is within the preset error range. When judging whether the deviation between the detected coordinates and the labeled coordinates of the target ground point of the same target is less than a preset value, the absolute value of the horizontal coordinate difference and the absolute value of the vertical coordinate difference of the detected coordinates and the labeled coordinates can be first calculated. If the absolute value of the horizontal coordinate difference and the absolute value of the vertical coordinate difference are both less than the preset value, then it is determined that the deviation between the detected coordinates and the labeled coordinates of the target ground point is less than the preset value.
[0077] Step 302 : Determine a two-dimensional position accuracy test result of the target detection model based on the labeled coordinates and the detected coordinates of the target ground point in the sample alignment.
[0078] In this embodiment, after determining the sample pairs whose target annotation information and target detection information match, the two-dimensional position accuracy test results of the target detection model can be determined based on the annotation coordinates and detection coordinates of the target ground point in the sample pairs.
[0079] As an example, the horizontal and vertical deviations of the detection coordinates of each sample pair relative to the labeled coordinates can be calculated, and the one with the larger deviation (larger index value, regardless of positive or negative signs) between the horizontal and vertical coordinates is selected to calculate regression evaluation indicators such as mean squared error (MSE) and mean absolute error, and the calculated regression evaluation index value is determined as the two-dimensional position accuracy test result of the target detection model.
[0080] As another example, the total number of sample pairs can be counted, and the number of target sample pairs whose deviation between the detection coordinates and the labeled coordinates is less than a preset deviation threshold can be counted, and the ratio of the number of target sample pairs divided by the total number of sample pairs can be calculated as the two-dimensional position accuracy test result of the target detection model.
[0081] Step 303 : Determine a three-dimensional position accuracy test result of the target detection model based on the three-dimensional annotated coordinates and the three-dimensional detected coordinates of the target ground point in the sample alignment.
[0082] In this embodiment, after determining the sample pairs whose target annotation information and target detection information match, the three-dimensional position accuracy test results of the target detection model can be determined based on the three-dimensional annotation coordinates and three-dimensional detection coordinates of the target ground point in the sample pairs.
[0083] It should be noted that the three-dimensional position accuracy test results of the target detection model can be calculated in the same way as the two-dimensional position accuracy test results of the target detection model. To avoid repetition, it will not be repeated here.
[0084] The accuracy testing method of the target detection model of the embodiment of the present invention determines a sample pair by comparing the target annotation information and the target detection information, the sample pair includes the matched target annotation information and the target detection information, and then determines the two-dimensional position accuracy test result of the target detection model based on the annotation coordinates and the detection coordinates of the target ground point in the sample pair, and determines the three-dimensional position accuracy test result of the target detection model based on the three-dimensional annotation coordinates and the three-dimensional detection coordinates of the target ground point in the sample pair. Therefore, the sample pairs in which the target annotation information and the target detection information are successfully matched are first screened out, and then the two-dimensional position accuracy of the target detection model is determined according to the annotation coordinates and the detection coordinates of the target ground point in the sample pair, and the three-dimensional position accuracy of the target detection model is determined according to the three-dimensional annotation coordinates and the three-dimensional detection coordinates. This can reduce the amount of calculation, improve the test efficiency, and quickly realize the position accuracy detection of the target detection model in two-dimensional space and three-dimensional space.
[0085] Figure 4 FIG. 4 shows a schematic diagram of the accuracy test process of the target detection model of an exemplary embodiment of the present disclosure. Figure 4 As shown, after the camera is installed, it is necessary to calibrate the camera parameters, including camera intrinsic parameters, camera extrinsic parameter matrix and distortion coefficient, and use the installed camera to capture test images, perform target annotation on the captured test images, annotate the target annotation information (two-dimensional) of all targets in the test images, and input the test images into the target detection model to be tested for target detection, and obtain the target detection information (two-dimensional) output by the target detection model. The target annotation information and target detection information are respectively reverse-projected and converted to the vehicle coordinate system to obtain annotation information with three-dimensional information (three-dimensional annotation information) and detection information with three-dimensional information (three-dimensional detection information). Finally, the two-dimensional target annotation information, target detection information, three-dimensional annotation information and three-dimensional detection information are combined to determine the test results of the two-dimensional test indicators and the test results of the three-dimensional test indicators of the target detection model.
[0086] In order to implement the above embodiments, the present disclosure further provides an accuracy testing device for a target detection model. The accuracy testing device for a target detection model can be implemented using software and / or hardware and can be integrated into an electronic device.
[0087] Figure 5 This is a schematic diagram of the structure of the accuracy testing device of the target detection model provided by an embodiment of the present disclosure, such as Figure 5 As shown, the target detection model accuracy testing device 40 may include: a first acquisition module 410 , a second acquisition module 420 , a transformation module 430 and a determination module 440 .
[0088] The first acquisition module 410 is used to acquire the test image and target annotation information of the test image;
[0089] The second acquisition module 420 is used to input the test image into the target detection model to be tested to perform target detection and obtain target detection information corresponding to the test image;
[0090] a transformation module 430 for transforming target annotation information into three-dimensional annotation information and transforming target detection information into three-dimensional detection information;
[0091] The determination module 440 is used to determine the accuracy test result of the target detection model based on the target annotation information, the target detection information, the three-dimensional annotation information and the three-dimensional detection information.
[0092] Optionally, the target annotation information includes the annotation coordinates of the target grounding point, and the target detection information includes the detection coordinates of the target grounding point; the transformation module 430 includes:
[0093] A first transformation unit is used to perform dedistortion and transformation processing on the key point coordinates to obtain dedistorted coordinates of the key point coordinates on the camera normalization plane, wherein the key point coordinates include marked coordinates and detected coordinates;
[0094] The second transformation unit is used to transform the dedistorted coordinates into the vehicle body coordinate system with the target ground point located on the ground in the vehicle body coordinate system as the target, thereby obtaining the three-dimensional coordinates of the target ground point in the vehicle body coordinate system, wherein the three-dimensional coordinates include the three-dimensional annotated coordinates corresponding to the annotated coordinates and the three-dimensional detected coordinates corresponding to the detected coordinates.
[0095] Further optionally, the first transform unit is further configured to:
[0096] Get the camera's intrinsic parameters and distortion coefficients;
[0097] Based on the camera intrinsic parameters, the key point coordinates are converted to the camera normalization plane to obtain the camera coordinates of the key point coordinates in the camera normalization plane;
[0098] Determine the amount of distortion based on the camera coordinates and the distortion coefficient;
[0099] The camera coordinates are dedistorted based on the distortion amount to obtain the dedistorted coordinates of the key point coordinates on the camera normalized plane.
[0100] Further optionally, the second transform unit is further configured to:
[0101] Get the camera extrinsic parameter matrix of the camera;
[0102] Setting the depth parameter for the dedistorted coordinate to obtain the first moving coordinate;
[0103] The first moving coordinates are converted to the vehicle coordinate system based on the camera extrinsic matrix to obtain the second moving coordinates;
[0104] Taking the target grounding point on the ground in the vehicle coordinate system as the target, the depth parameter of the second moving coordinate is adjusted to obtain the three-dimensional coordinates of the target grounding point in the vehicle coordinate system.
[0105] Optionally, the determination module 440 includes:
[0106] a matching unit, configured to compare the target labeling information and the target detection information to determine a sample pair, the sample pair including the matched target labeling information and the target detection information;
[0107] A first determining unit is configured to determine a two-dimensional position accuracy test result of a target detection model based on the marked coordinates and the detected coordinates of the target ground point in the sample alignment;
[0108] The second determining unit is used to determine the three-dimensional position accuracy test result of the target detection model based on the three-dimensional labeled coordinates and the three-dimensional detection coordinates of the target grounding point in the sample center.
[0109] Optionally, the matching unit is further configured to:
[0110] Compare the target labeling information and target detection information corresponding to the same target to determine whether the target labeling information and target detection information meet the preset conditions;
[0111] When it is determined that the preset conditions are met, the target labeling information and the target detection information currently being compared are determined to be a sample pair.
[0112] Further optionally, the target annotation information further includes target annotation type and target annotation frame information, and the target detection information further includes target detection type and target detection frame information; the preset conditions include:
[0113] The annotation type and detection type of the same target are consistent, the difference between the target annotation box information and the target detection box information of the same target is within the preset error range, and the deviation between the annotation coordinates and the detection coordinates of the target ground point of the same target is less than the preset value.
[0114] The accuracy testing device for an object detection model applicable to an electronic device provided in the embodiments of the present disclosure can execute the accuracy testing method for an object detection model provided in the embodiments of the present disclosure, and has the corresponding functional modules and beneficial effects of the execution method. Any details not fully described in the embodiments of the present disclosure may be referred to in the description of any method embodiment of the present disclosure.
[0115] An embodiment of the present disclosure also provides an electronic device including a processor and a memory; the processor calls a program or instruction stored in the memory to execute the steps of each embodiment of the accuracy testing method of the target detection model as described above. To avoid repeated description, they will not be repeated here.
[0116] The embodiments of the present disclosure also provide a computer-readable storage medium, which is non-transitory and stores programs or instructions. The programs or instructions enable a computer to execute the steps of each embodiment of the accuracy testing method of the target detection model as described above. To avoid repeated description, they will not be repeated here.
[0117] The embodiments of the present disclosure also provide a computer program product, which is used to execute the steps of each embodiment of the accuracy testing method of the target detection model as described above.
[0118] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0119] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.
Claims
1. A method for testing the accuracy of a target detection model, characterized in that: The method comprises: Obtaining a test image and target annotation information of the test image; Inputting the test image into the target detection model to be tested to perform target detection, and obtaining target detection information corresponding to the test image; transforming the target annotation information into three-dimensional annotation information, and transforming the target detection information into three-dimensional detection information; An accuracy test result of the target detection model is determined based on the target annotation information, the target detection information, the three-dimensional annotation information, and the three-dimensional detection information.
2. The method according to claim 1, characterized in that The target marking information includes the marking coordinates of the target grounding point, and the target detection information includes the detection coordinates of the target grounding point; The converting the target annotation information into three-dimensional annotation information and converting the target detection information into three-dimensional detection information includes: Performing dedistortion and transformation processing on the key point coordinates to obtain dedistorted coordinates of the key point coordinates on a camera normalized plane, wherein the key point coordinates include the marked coordinates and the detected coordinates; Taking the target ground point as being located on the ground in the vehicle body coordinate system, the dedistorted coordinates are converted to the vehicle body coordinate system to obtain the three-dimensional coordinates of the target ground point in the vehicle body coordinate system, wherein the three-dimensional coordinates include the three-dimensional annotated coordinates corresponding to the annotated coordinates and the three-dimensional detected coordinates corresponding to the detected coordinates.
3. The method according to claim 2, characterized in that The dedistortion and conversion processing is performed on the key point coordinates to obtain the dedistorted coordinates of the key point coordinates on the camera normalization plane, including: Get the camera's intrinsic parameters and distortion coefficients; Converting the key point coordinates to a camera normalized plane based on the camera intrinsic parameters to obtain the camera coordinates of the key point coordinates on the camera normalized plane; Determining a distortion amount based on the camera coordinates and the distortion coefficients; The camera coordinates are dedistorted based on the distortion amount to obtain dedistorted coordinates of the key point coordinates on the camera normalization plane.
4. The method according to claim 2, characterized in that The step of converting the dedistorted coordinates to the vehicle body coordinate system with the target touchdown point located on the ground in the vehicle body coordinate system as the target to obtain the three-dimensional coordinates of the target touchdown point in the vehicle body coordinate system includes: Get the camera extrinsic matrix of the camera; Setting a depth parameter for the dedistorted coordinate to obtain a first moving coordinate; Converting the first moving coordinates to the vehicle coordinate system based on the camera extrinsic parameter matrix to obtain second moving coordinates; Taking the target grounding point located on the ground in the vehicle body coordinate system as the target, the depth parameter of the second movement coordinate is adjusted to obtain the three-dimensional coordinates of the target grounding point in the vehicle body coordinate system.
5. The method according to any one of claims 2 to 4, characterized in that: The determining, based on the target labeling information, the target detection information, the three-dimensional labeling information, and the three-dimensional detection information, an accuracy test result of the target detection model includes: Comparing the target labeling information and the target detection information to determine a sample pair, wherein the sample pair includes matching target labeling information and target detection information; Determining a two-dimensional position accuracy test result of the target detection model based on the marked coordinates and the detected coordinates of the target ground point in the sample pair; Based on the three-dimensional annotated coordinates and the three-dimensional detected coordinates of the target ground point in the sample pair, a three-dimensional position accuracy test result of the target detection model is determined.
6. The method according to claim 5, characterized in that The comparing the target labeling information and the target detection information to determine a sample pair includes: Comparing the target labeling information and the target detection information corresponding to the same target, and determining whether the target labeling information and the target detection information meet a preset condition; When it is determined that the preset condition is met, the target annotation information and the target detection information currently being compared are determined to be a sample pair.
7. The method according to claim 6, characterized in that The target annotation information also includes the target annotation type and target annotation frame information, and the target detection information also includes the target detection type and target detection frame information; The preset conditions include: The annotation type and the detection type of the same target are consistent, and the difference between the target annotation box information and the target detection box information of the same target is within a preset error range, and the deviation between the annotation coordinates and the detection coordinates of the target ground point of the same target is less than a preset value.
8. A target detection model accuracy test device, characterized in that: include: A first acquisition module is used to acquire a test image and target annotation information of the test image; A second acquisition module is used to input the test image into the target detection model to be tested to perform target detection and obtain target detection information corresponding to the test image; a conversion module, configured to convert the target annotation information into three-dimensional annotation information, and to convert the target detection information into three-dimensional detection information; A determination module is used to determine an accuracy test result of the target detection model based on the target annotation information, the target detection information, the three-dimensional annotation information and the three-dimensional detection information.
9. An electronic device, characterized in that: include: processor and memory; The processor is used to execute the accuracy testing method of the target detection model as described in any one of claims 1 to 7 by calling the program or instruction stored in the memory.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program or instruction, which enables a computer to execute the accuracy testing method of the target detection model according to any one of claims 1 to 7.