A method and device for evaluating a visual perception algorithm

By acquiring ground truth and detection information from point cloud data frames and image frames, and combining this with evaluation rules, the shortcomings of performance evaluation for visual perception algorithms are addressed, enabling comprehensive evaluation of visual perception algorithms and improving detection accuracy and stability.

CN114170448BActive Publication Date: 2026-02-27MOMENTA (SUZHOU) TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202010841528.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-08-20
Publication Date
2026-02-27
Estimated Expiration
2040-08-20

AI Technical Summary

Technical Problem

The lack of effective methods in existing technologies to evaluate the performance of visual perception algorithms affects the accuracy of autonomous driving systems, facial recognition systems, and identity verification systems.

Method used

By obtaining ground truth and detection information from point cloud data frames and image frames in the evaluation dataset, and combining them with preset evaluation rules for result accuracy and algorithm stability, the detection accuracy and algorithm stability of the visual perception algorithm are determined, including automatically annotating the 3D information of objects and matching ground truth and detection information, and error curves are plotted for evaluation.

Benefits of technology

It enables a comprehensive evaluation of visual perception algorithms, improves the accuracy of detection results and the stability assessment of algorithms, and provides a more intuitive visualization of evaluation results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114170448B_ABST
    Figure CN114170448B_ABST
Patent Text Reader

Abstract

The embodiment of the application discloses a visual perception algorithm evaluation method and device, the method comprises the following steps: obtaining the true value information corresponding to each object determined based on each point cloud data frame in the evaluation data set; obtaining the detection information corresponding to each detected object based on a preset visual detection algorithm and the detection of each image frame in the evaluation data set; based on the preset result accuracy evaluation rule, the preset algorithm stability evaluation rule, the labeled pose information and object motion information in the true value information corresponding to each object, and the detection pose information and detection motion information in the detection information corresponding to each detected object, determining the evaluation information corresponding to the preset visual perception algorithm, wherein the evaluation information comprises: the first evaluation information representing the detection result accuracy of the preset visual perception algorithm and the second evaluation information representing the algorithm stability, so as to comprehensively evaluate the performance of the visual perception algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of algorithm evaluation, in particular to a visual perception algorithm evaluation method and device. BACKGROUND

[0002] The visual perception algorithm is a core component in the fields of automatic driving systems, face recognition systems and identity verification systems, and the accuracy of the perception result of the visual perception algorithm affects the accuracy of the output result of the above systems to a certain extent. Accordingly, in order to ensure the performance of the above systems, the performance of the visual perception algorithm needs to be evaluated before the visual perception algorithm is actually applied.

[0003] Therefore, how to provide a method for evaluating the performance of the visual perception algorithm becomes a problem to be solved. SUMMARY

[0004] The present application provides a visual perception algorithm evaluation method and device to comprehensively evaluate the performance of the visual perception algorithm. The specific technical solutions are as follows:

[0005] In a first aspect, the present application provides a visual perception algorithm evaluation method, which comprises:

[0006] Obtaining true value information corresponding to each object determined based on each point cloud data frame in an evaluation data set, wherein the true value information at least includes labeled pose information and object motion information of the corresponding object, and each evaluation data includes point cloud data frames and image frames that exist in a corresponding relationship;

[0007] Obtaining detection information corresponding to each detection object detected based on a preset visual perception algorithm and each image frame in the evaluation data set, wherein the detection information at least includes detection pose information and detection motion information of the corresponding detection object;

[0008] Based on a preset result accuracy evaluation rule, a preset algorithm stability evaluation rule, labeled pose information and object motion information in the true value information corresponding to each object, and detection pose information and detection motion information in the detection information corresponding to each detection object, determining evaluation information corresponding to the preset visual perception algorithm, wherein the evaluation information includes first evaluation information representing the detection result accuracy of the preset visual perception algorithm and second evaluation information representing the algorithm stability.

[0009] Optionally, the process of obtaining the true value information corresponding to each object determined based on each point cloud data frame in the evaluation data set comprises:

[0010] Obtaining an evaluation data set;

[0011] annotating each point cloud data frame in the evaluation data set based on the pre-trained three-dimensional data perception model to annotate the annotation box information of each object corresponding to each point cloud data frame, to determine the annotation position information and the annotation pose information of each object corresponding to each point cloud data frame, and to obtain the annotation pose information of each object corresponding to each point cloud data frame;

[0012] Based on the annotation position information and the annotation pose information of each object corresponding to each point cloud data frame in the evaluation data set, and the time sequence information between each point cloud data frame in the evaluation data set, the annotation speed information and the annotation acceleration information of each object corresponding to each point cloud data frame are determined to obtain the object motion information of each object corresponding to each point cloud data frame, and the true value information corresponding to each object corresponding to each point cloud data frame is obtained.

[0013] Optionally, the step of obtaining the detection information corresponding to each object detected based on the preset visual perception algorithm and each image frame in the evaluation data set comprises:

[0014] Based on the preset visual perception algorithm, each image frame in the evaluation data set is detected to obtain the detection box information of each detected object corresponding to each image frame, to determine the detection position information and the detection pose information of each detected object corresponding to each image frame, and to obtain the detection pose information of each detected object corresponding to each image frame;

[0015] Based on the preset visual perception algorithm and the detection position information and the detection pose information of each detected object corresponding to each image frame, the detection speed information and the detection acceleration information of each detected object corresponding to each image frame are determined to obtain the detection motion information of each detected object corresponding to each image frame, and the detection information corresponding to each detected object corresponding to each image frame is obtained.

[0016] Optionally, the true value information includes the annotation box information of each object corresponding to each point cloud data frame, the detection information includes the detection box information of each detected object corresponding to each image frame, and the detection box information includes the two-dimensional position information of the corresponding detected object in the image frame.

[0017] The step of determining the evaluation information corresponding to the preset visual perception algorithm based on the preset result accuracy evaluation rule, the preset algorithm stability evaluation rule, the annotation pose information and the object motion information in the true value information corresponding to each object, and the detection pose information and the detection motion information in the detection information corresponding to each detected object comprises:

[0018] For each object corresponding to each point cloud data frame, based on the annotation box information corresponding to the object, the position conversion relationship between the point cloud data frame acquisition device and the image frame acquisition device, and the intrinsic information of the image frame acquisition device, the projection position information of the projection frame of the annotation box corresponding to the object projected into the image frame corresponding to the point cloud data frame is determined as the projection frame position information corresponding to the object;

[0019] For each object corresponding to each point cloud data frame, based on the projection frame position information corresponding to each object and the two-dimensional position information of each detected object in the image frame corresponding to the point cloud data frame, the matching projection frame position information and two-dimensional position information are determined to determine the matching ground truth information and detection information corresponding to each point cloud data frame and its corresponding image frame, wherein the matching projection frame position information and two-dimensional position information are the projection frame position information and two-dimensional position information whose intersection over union value of the corresponding frame exceeds the preset intersection over union threshold;

[0020] Based on the preset result accuracy evaluation rule, the matching ground truth information and detection information corresponding to each point cloud data frame and its corresponding image frame, the detection information that does not match the ground truth information, and the ground truth information that does not match the detection information, the first evaluation information representing the detection result accuracy of the preset visual perception algorithm is determined.

[0021] Based on the preset algorithm stability evaluation rule and the matching ground truth information and detection information corresponding to each point cloud data frame and its corresponding image frame, the second evaluation information representing the algorithm stability of the preset visual perception algorithm is determined.

[0022] Optionally, the detection pose information includes: the detection position information and the detection attitude information of each detected object in each image frame determined by its detection frame information; the detection motion information includes: the detection speed information and the detection acceleration information of each detected object in each image frame; the annotation pose information includes: the annotation position information and the annotation attitude information of each object corresponding to each point cloud data frame determined by its annotation frame information; and the object motion information includes: the annotation speed information and the annotation acceleration information of each object corresponding to each point cloud data frame.

[0023] The step of determining the first evaluation information representing the detection result accuracy of the preset visual perception algorithm based on the preset result accuracy evaluation rule, the matching ground truth information and detection information corresponding to each point cloud data frame and its corresponding image frame, the detection information that does not match the ground truth information, and the ground truth information that does not match the detection information, comprises:

[0024] determine, based on the matching true value information and the detection information corresponding to each point cloud data frame and the image frame corresponding thereto, a precision information and a recall information of a detection result corresponding to the preset visual perception algorithm;

[0025] determine, based on the matching true value information and the detection information corresponding to each point cloud data frame and the image frame corresponding thereto, a detection position error value between the annotation position information included in the matching true value information and the detection position information included in the detection information;

[0026] determine, based on the matching true value information and the detection information corresponding to each point cloud data frame and the image frame corresponding thereto, a detection posture error value between the annotation posture information included in the matching true value information and the detection posture information included in the detection information;

[0027] determine, based on the matching true value information and the detection information corresponding to each point cloud data frame and the image frame corresponding thereto, a detection speed error value between the annotation speed information included in the matching true value information and the detection speed information included in the detection information;

[0028] determine, based on the matching true value information and the detection information corresponding to each point cloud data frame and the image frame corresponding thereto, a detection acceleration error value between the annotation acceleration information included in the matching true value information and the detection acceleration information included in the detection information;

[0029] determine, based on the matching true value information and the detection information corresponding to each point cloud data frame and the image frame corresponding thereto, a length-width error value of a detection box between the annotation box information included in the matching true value information and the detection box information included in the detection information;

[0030] draw an error curve corresponding to the target error value based on the target error value between the matching true value information and the detection information and a preset error threshold value corresponding to the target error value, wherein an abscissa axis of the error curve is the preset error threshold value, an ordinate axis of the error curve is a ratio of a number of target error values smaller than each preset error threshold value in the target error value to a total amount of data in an evaluation data set, and the target error value is the detection position error value, the detection posture error value, the detection speed error value, the detection acceleration error value, or the length-width error value of the detection box;

[0031] sort the target error value between the matching true value information and the detection information according to the size of the numerical value to obtain a sorting sequence corresponding to the target error value, determine a first target error value with the largest numerical value in a first percentage of the target error value in the sorting sequence corresponding to the target error value, and a second target error value with the largest numerical value in a second percentage of the target error value;

[0032] determine first evaluation information representing detection result accuracy of the preset visual perception algorithm based on precision information and recall information of detection results corresponding to the preset visual perception algorithm, the error curve corresponding to the target error information, the first target error value and the second target error value, and / or the target error value.

[0033] Optionally, the step of determining second evaluation information representing algorithm stability of the preset visual perception algorithm based on the preset algorithm stability evaluation rule, the matching true value information and the detection information corresponding to each point cloud data frame and its corresponding image frame, comprises:

[0034] determine target error values corresponding to the same object from the target error values based on time sequence information between point cloud data frames or image frames in the evaluation data set;

[0035] For target error values corresponding to different objects, based on time sequence information between point cloud data frames or image frames in the evaluation data set, the target error values corresponding to the objects, and a preset curve fitting algorithm, a fitting error curve corresponding to the target error values of the object is fitted, wherein the fitting error curve comprises fitting errors corresponding to the object at each point cloud data frame or image frame corresponding acquisition time.

[0036] For target error values corresponding to different objects, based on the target error values corresponding to the objects, and fitting errors corresponding to the object at each point cloud data frame or image frame corresponding acquisition time contained in the fitting error curve corresponding to the target error values of the object, second evaluation information representing algorithm stability of the preset visual perception algorithm is determined.

[0037] Optionally, the step of determining second evaluation information representing algorithm stability of the preset visual perception algorithm based on the target error values corresponding to different objects, the target error values corresponding to the objects, and fitting errors corresponding to the object at each point cloud data frame or image frame corresponding acquisition time contained in the fitting error curve corresponding to the target error values of the object, comprises:

[0038] For target error values corresponding to different objects, based on the target error values corresponding to the objects, and fitting errors corresponding to the object at each point cloud data frame or image frame corresponding acquisition time contained in the fitting error curve corresponding to the target error values of the object, the difference between the target error and the fitting error corresponding to the same acquisition time is calculated.

[0039] For the target error values corresponding to different objects, a difference error curve corresponding to the target error of the object is drawn based on the difference between the target error and the fitting error of each collection time corresponding to the object and a preset difference threshold corresponding to the target error value of the object, wherein the horizontal axis of the difference error curve corresponding to the target error of the object is the preset difference threshold corresponding to the target error value of the object, and the vertical axis of the difference error curve corresponding to the target error of the object is the ratio of the number of differences less than each preset difference threshold to the total number of target error values corresponding to the target error of the object.

[0040] For the target error values corresponding to different objects, the differences corresponding to the target error values of the object are sorted according to the size of the values, the first difference with the largest value in the first third of the sequence is determined, and the second difference with the largest value in the first fourth of the sequence is determined.

[0041] Based on the first difference and the second difference corresponding to the target error values of each object, and / or the difference error curve corresponding to the target error values of each object, second evaluation information representing the algorithm stability of the preset visual perception algorithm is determined.

[0042] Optionally, the method further comprises:

[0043] The first evaluation information, the second evaluation information, the precision rate information and the recall rate information of the detection result corresponding to the preset visual perception algorithm, the error curve corresponding to the target error information, the first target error value and the second target error value, the target error value, the target error value corresponding to the same object determined and / or the fitting error curve corresponding to the target error value of each object fitted are displayed.

[0044] In a second aspect, an embodiment of the present application provides a visual perception algorithm evaluation device, the device comprising:

[0045] A first obtaining module configured to obtain true value information corresponding to each object determined based on each point cloud data frame in an evaluation data set, wherein the true value information at least includes labeled pose information and object motion information of the corresponding object, and each evaluation data includes point cloud data frames and image frames with corresponding relationship;

[0046] A second obtaining module configured to obtain detection information corresponding to each detection object detected based on a preset visual perception algorithm and each image frame in the evaluation data set, wherein the detection information at least includes detection pose information and detection motion information of the corresponding detection object;

[0047] The determining module is configured to determine evaluation information corresponding to the preset visual perception algorithm based on the preset result accuracy evaluation rule, the preset algorithm stability evaluation rule, the labeled pose information and the object motion information in the true value information corresponding to each object, and the detected pose information and the detected motion information in the detection information corresponding to each detected object, wherein the evaluation information includes first evaluation information representing detection result accuracy of the preset visual perception algorithm and second evaluation information representing algorithm stability.

[0048] Optionally, the first obtaining module is specifically configured to obtain evaluation data sets.

[0049] Each point cloud data frame in the evaluation data sets is labeled based on the pre-trained three-dimensional data perception model, and the labeled box information of each object corresponding to each point cloud data frame is labeled to determine the labeled position information and the labeled pose information of each object corresponding to each point cloud data frame, so as to obtain the labeled pose information of each object corresponding to each point cloud data frame.

[0050] Based on the labeled position information and the labeled pose information of each object corresponding to each point cloud data frame in the evaluation data sets and the time sequence information between each point cloud data frame in the evaluation data sets, the labeled speed information and the labeled acceleration information of each object corresponding to each point cloud data frame are determined to obtain the object motion information of each object corresponding to each point cloud data frame, and the true value information corresponding to each object corresponding to each point cloud data frame is obtained.

[0051] Optionally, the second obtaining module is specifically configured to detect each image frame in the evaluation data sets based on the preset visual perception algorithm to obtain the detected box information of each object corresponding to each image frame to determine the detected position information and the detected pose information of each detected object corresponding to each image frame, and obtain the detected pose information of each detected object corresponding to each image frame.

[0052] Based on the preset visual perception algorithm and the detected position information and the detected pose information of each detected object corresponding to each image frame, the detected speed information and the detected acceleration information of each detected object corresponding to each image frame are determined to obtain the detected motion information of each detected object corresponding to each image frame, and the detection information corresponding to each detected object corresponding to each image frame is obtained.

[0053] Optionally, the true value information includes the labeled box information of each object corresponding to each point cloud data frame, the detection information includes the detected box information of each detected object corresponding to each image frame, and the detected box information includes two-dimensional position information of the corresponding detected object in the image frame.

[0054] The determining module includes:

[0055] The first determination unit is configured to determine, for each object corresponding to each point cloud data frame, projection position information of a projection frame of a label frame of the object in an image frame corresponding to the point cloud data frame as projection frame position information of the object, based on the label frame information of the object, a position conversion relationship between a point cloud data frame acquisition device and an image frame acquisition device, and intrinsic parameter information of the image frame acquisition device;

[0056] The second determination unit is configured to determine matched projection frame position information and two-dimensional position information for each object corresponding to each point cloud data frame, based on the projection frame position information of each object and two-dimensional position information of each detected object in an image frame corresponding to the point cloud data frame, to determine matched ground truth information and detection information corresponding to each point cloud data frame and the image frame corresponding thereto, wherein the matched projection frame position information and two-dimensional position information are projection frame position information and two-dimensional position information whose intersection over union value of the corresponding frames exceeds a preset intersection over union threshold.

[0057] The third determination unit is configured to determine first evaluation information representing detection result accuracy of the preset visual perception algorithm based on a preset result accuracy evaluation rule, matched ground truth information and detection information corresponding to each point cloud data frame and the image frame corresponding thereto, detection information that does not match the ground truth information, and ground truth information that does not match the detection information.

[0058] The fourth determination unit is configured to determine second evaluation information representing algorithm stability of the preset visual perception algorithm based on a preset algorithm stability evaluation rule and matched ground truth information and detection information corresponding to each point cloud data frame and the image frame corresponding thereto.

[0059] Optionally, the detection pose information includes detection position information and detection attitude information of each detected object in each image frame determined by detection frame information of the detected object; the detection motion information includes detection speed information and detection acceleration information of each detected object in each image frame; the label pose information includes label position information and label attitude information of each object corresponding to each point cloud data frame determined by label frame information of the object; and the object motion information includes label speed information and label acceleration information of each object corresponding to each point cloud data frame.

[0060] The third determination unit is specifically configured to determine precision rate information and recall rate information of a detection result corresponding to the preset visual perception algorithm based on matched ground truth information and detection information corresponding to each point cloud data frame and the image frame corresponding thereto, detection information that does not match the ground truth information, and ground truth information that does not match the detection information.

[0061] determine a detection position error value between the matched ground truth information and the detection information based on the labeled position information included in the matched ground truth information and the detection position information included in the detection information corresponding to each point cloud data frame and the image frame corresponding thereto;

[0062] determine a detection pose error value between the matched ground truth information and the detection information based on the labeled pose information included in the matched ground truth information and the detection pose information included in the detection information corresponding to each point cloud data frame and the image frame corresponding thereto;

[0063] determine a detection speed error value between the matched ground truth information and the detection information based on the labeled speed information included in the matched ground truth information and the detection speed information included in the detection information corresponding to each point cloud data frame and the image frame corresponding thereto;

[0064] determine a detection acceleration error value between the matched ground truth information and the detection information based on the labeled acceleration information included in the matched ground truth information and the detection acceleration information included in the detection information corresponding to each point cloud data frame and the image frame corresponding thereto;

[0065] determine a length-width error value of a detection bounding box between the matched ground truth information and the detection information based on the labeled bounding box information included in the matched ground truth information and the detection bounding box information included in the detection information corresponding to each point cloud data frame and the image frame corresponding thereto;

[0066] draw an error curve corresponding to the target error value based on the target error value between the matched ground truth information and the detection information and a preset error threshold value corresponding to the target error value, wherein the horizontal axis of the error curve is the preset error threshold value, the vertical axis of the error curve is a ratio of a number of target error values smaller than each preset error threshold value in the target error value to a total amount of data in the evaluation dataset, and the target error value is the detection position error value, the detection pose error value, the detection speed error value, the detection acceleration error value, or the length-width error value of the detection bounding box;

[0067] sort the target error value between the matched ground truth information and the detection information according to the size of the numerical value to obtain a sorting sequence corresponding to the target error value, determine a first target error value with the largest numerical value in a first percentage of target error values in the sorting sequence corresponding to the target error value, and a second target error value with the largest numerical value in a second percentage of target error values;

[0068] determine first evaluation information representing the accuracy of the detection result of the preset visual perception algorithm based on the precision information and the recall information of the detection result corresponding to the preset visual perception algorithm, the error curve corresponding to the target error information, the first target error value and the second target error value, and / or the target error value.

[0069] Optionally, the fourth determining unit is specifically configured to determine, from the target error values, a target error value corresponding to a same object based on timing information between point cloud data frames or image frames in the evaluation data set;

[0070] For target error values corresponding to different objects, based on timing information between point cloud data frames or image frames in the evaluation data set, the target error value of the object, and a preset curve fitting algorithm, a fitting error curve corresponding to the target error value of the object is fitted, wherein the fitting error curve comprises fitting errors corresponding to the object at each point cloud data frame or image frame corresponding acquisition time.

[0071] For target error values corresponding to different objects, based on the target error value of the object, and fitting errors corresponding to the object at each point cloud data frame or image frame corresponding acquisition time contained in the fitting error curve corresponding to the target error value of the object, second evaluation information characterizing algorithm stability of the preset visual perception algorithm is determined.

[0072] Optionally, the fourth determining unit is specifically configured to, for target error values corresponding to different objects, based on the target error value of the object, and fitting errors corresponding to the object at each point cloud data frame or image frame corresponding acquisition time contained in the fitting error curve corresponding to the target error value of the object, calculate the difference value of the target error and the fitting error corresponding to the same acquisition time.

[0073] For target error values corresponding to different objects, based on the difference value of the target error and the fitting error of each acquisition time of the object, and a preset difference value threshold corresponding to the target error value of the object, a difference error curve corresponding to the target error of the object is drawn, wherein the horizontal axis of the difference error curve corresponding to the target error of the object is the preset difference value threshold corresponding to the target error value of the object, and the vertical axis of the difference error curve corresponding to the target error of the object is the ratio of the number of difference values less than each preset difference value threshold to the total number of target error values of the object.

[0074] For target error values corresponding to different objects, the difference values corresponding to the target error values of the object are sorted according to the size of the values, the first difference value with the largest value in the first third of the difference values in the sorting sequence is determined, and the second difference value with the largest value in the first fourth of the difference values is determined.

[0075] Based on the first difference value and the second difference value corresponding to the target error values of each object, and / or the difference error curve corresponding to the target error values of each object, second evaluation information characterizing algorithm stability of the preset visual perception algorithm is determined.

[0076] Optionally, the apparatus further comprises:

[0077] The display module is configured to display the first evaluation information, the second evaluation information, the precision rate information and the recall rate information of the detection result corresponding to the preset visual perception algorithm, the error curve corresponding to the target error information, the first target error value and the second target error value, the target error value, the target error value corresponding to the same object determined and / or the fitting error curve corresponding to the target error value corresponding to each object obtained by fitting.

[0078] From the above, the embodiment of the application provides a visual perception algorithm evaluation method and device, the true value information corresponding to each object is obtained based on each point cloud data frame in the evaluation data set, wherein the true value information at least includes the labeled pose information and the object motion information of the corresponding object, each evaluation data includes the point cloud data frame and the image frame having a corresponding relationship; the detection information corresponding to each detection object is obtained based on the preset visual perception algorithm and the detection of each image frame in the evaluation data set, wherein the detection information at least includes the detection pose information and the detection motion information of the corresponding detection object; based on the preset result accuracy evaluation rule, the preset algorithm stability evaluation rule, the labeled pose information and the object motion information in the true value information corresponding to each object, and the detection pose information and the detection motion information in the detection information corresponding to each detection object, the evaluation information corresponding to the preset visual perception algorithm is determined, wherein the evaluation information includes the first evaluation information representing the detection result accuracy of the preset visual perception algorithm and the second evaluation information representing the algorithm stability.

[0079] By applying the embodiment of the application, the first evaluation information of the accuracy of the detection result of the preset visual perception algorithm and the second evaluation information of the algorithm stability can be determined based on the preset result accuracy evaluation rule, the preset algorithm stability evaluation rule, the labeled pose information and the object motion information in the true value information corresponding to each object, and the detection pose information and the detection motion information in the detection information corresponding to each detection object, the preset visual perception algorithm is evaluated from the accuracy of the detection result and the stability of the detection result of the algorithm, and the performance of the visual perception algorithm is comprehensively evaluated. Of course, any product or method implementing the application does not necessarily need to achieve all the advantages described above.

[0080] The innovation points of the embodiment of the application include:

[0081] 1. The first evaluation information of the accuracy of the multi-aspect detection result of the preset visual perception algorithm and the second evaluation information of the algorithm stability can be determined based on the preset result accuracy evaluation rule, the preset algorithm stability evaluation rule, the labeled pose information and object motion information in the true value information corresponding to each object, and the detection pose information and detection motion information in the detection information corresponding to each detected object, so as to evaluate the preset visual perception algorithm from the accuracy of the multi-aspect detection result and the stability of the detection result of the algorithm, and realize comprehensive evaluation of the performance of the visual perception algorithm.

[0082] 2. Each point cloud data frame in the evaluation data set is automatically labeled based on the pre-trained three-dimensional data perception model, and the labeling box information of each object corresponding to each point cloud data frame is obtained, and then the labeling position information and the labeling pose information of each object corresponding to each point cloud data frame are calculated. The speed information and the acceleration information of each object corresponding to each point cloud data frame are determined by combining the time sequence information between each point cloud data frame in the evaluation data, and the three-dimensional information including the labeling box information, the labeling position information, the labeling pose information, the labeling speed information and the labeling acceleration information of each object corresponding to each point cloud data frame is obtained, so as to realize automatic labeling of the three-dimensional information of each labeled object and save labor cost.

[0083] 3. The true value information and the detection information are matched based on the labeling box information and the detection box information in the true value information, and then the first evaluation information representing the accuracy of the detection result of the preset visual perception algorithm and the second evaluation information representing the algorithm stability are determined based on the preset result accuracy evaluation rule, the preset algorithm stability evaluation rule, the matching true value information and detection information corresponding to each point cloud data frame and its corresponding image frame, the detection information not matched to the true value information, and the true value information not matched to the detection information, so as to realize comprehensive evaluation of the result accuracy and the algorithm stability of the preset visual perception algorithm.

[0084] 4. In the process of evaluating the accuracy of the detection results of the preset visual perception algorithm, in addition to evaluating the accuracy of the detection results from the perspectives of precision and recall, a new evaluation index is added. Based on the target error value between the matched true value information and the detection information, and the preset error threshold corresponding to the target error value, an error curve corresponding to the target error value is plotted. The first target error value with the largest value among the first percentage of target error values ​​in the sorted sequence corresponding to the target error value, and the second target error value with the largest value among the first and second percentages of target error values ​​are determined. Then, combined with the precision and recall information of the detection results, the error curve corresponding to the target error information, the first target error value, the second target error value and / or the target error value, the first evaluation information characterizing the accuracy of the detection results of the preset visual perception algorithm is determined, so as to achieve a more comprehensive evaluation of the accuracy of the detection results.

[0085] 5. To achieve a comprehensive evaluation of the preset visual perception algorithm, an evaluation of the algorithm's stability has been added. First, the target error value corresponding to the same object is determined from the target error values. Then, a preset curve fitting algorithm is used to fit the fitting error curve corresponding to the target error value of the object. Based on the target error value corresponding to the object and the fitting error of the object at the corresponding acquisition time in each point cloud data frame or image frame in the fitting error curve, a second evaluation information characterizing the algorithm stability of the preset visual perception algorithm is determined, thus achieving the evaluation of the algorithm stability of the preset visual perception algorithm. For example: calculate the difference between the target error value corresponding to the object and the fitting error at the same acquisition time, and then plot the difference error curve with the horizontal axis as the preset difference threshold and the vertical axis as the ratio of the number of differences corresponding to the target error value of the object that are less than each preset difference threshold to the total number of target error values ​​corresponding to the object; and / or sort the differences corresponding to the target error value of the object according to the size of the values, determine the first difference with the largest value in the third percentile of the differences in the sorted sequence, and the second difference with the largest value in the fourth percentile of the differences, and determine the second evaluation information through the first difference and the second difference, and / or the difference error curve.

[0086] 6. Visualize the intermediate and final evaluation results of the visual perception algorithm evaluation process, providing users with a more intuitive evaluation process of the accuracy and stability of the visual perception algorithm detection results. Attached Figure Description

[0087] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only represent some of the embodiments of the present application. For those of ordinary skill in the art, other drawings can also be obtained from these drawings without any creative effort.

[0088] Figure 1 A flowchart of a visual perception algorithm evaluation method provided by an embodiment of the present application;

[0089] Figure 2 An example diagram of an error curve corresponding to a target error value provided by an embodiment of the present application;

[0090] Figure 3 A structural diagram of a visual perception algorithm evaluation device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0091] The technical solutions in the embodiments of the present application will be described clearly and completely in the following with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments only represent some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without any creative effort fall within the scope of protection of the present application.

[0092] It should be noted that the terms "comprising" and "having" and any variations thereof in the embodiments of the present application and the drawings are intended to cover non-exclusive inclusion. For example, the processes, methods, systems, products or devices comprising a series of steps or units are not limited to the listed steps or units, but can optionally further comprise steps or units not listed, or can optionally further comprise other steps or units inherent to these processes, methods, products or devices.

[0093] The present application provides a visual perception algorithm evaluation method and device to comprehensively evaluate the performance of a visual perception algorithm. The embodiments of the present application will be described in detail in the following.

[0094] Figure 1 A flowchart of a visual perception algorithm evaluation method provided by an embodiment of the present application. The method can include the following steps:

[0095] S101: Obtain the true value information corresponding to each object determined based on each point cloud data frame in the evaluation data set.

[0096] The true value information at least includes labeled pose information of the corresponding object and object motion information, and each evaluation data includes point cloud data frames and image frames that have a corresponding relationship. The point cloud data frames can be data frames collected by a laser radar sensor, and the image frames can be image frames collected by an image collection device.

[0097] The evaluation method of the visual perception algorithm provided in the embodiments of the present application can be applied to any electronic device with computing capability, which can be a terminal or a server.

[0098] The labeled pose information of the corresponding object and the object motion information included in the true value information can be information based on a three-dimensional space, for example, can be pose information and motion information in a device coordinate system of a device that collects the point cloud data frames, or can be pose information and motion information in a preset space rectangular coordinate system, which can be a world coordinate system or an image collection device coordinate. The labeled pose information can include labeled position information and labeled attitude information. The object motion information can include, but is not limited to, speed information and acceleration information of the labeled object. For the sake of clarity, the speed information included in the object motion information determined based on each point cloud data frame in the evaluation data set can be referred to as labeled speed information, and the acceleration information can be referred to as labeled acceleration information.

[0099] In one case, the visual perception algorithm can be a visual perception algorithm applied in an automatic driving system, and correspondingly, each evaluation data included in the evaluation data set can be evaluation data collected by a target vehicle during driving, and each evaluation data includes point cloud data frames and image frames that have a corresponding relationship. The corresponding relationship can refer to point cloud data frames and image frames collected in the same collection period. Correspondingly, the laser radar sensor and the image collection device can be arranged in the target vehicle.

[0100] In the case that the visual perception algorithm is a visual perception algorithm applied in an automatic driving system, the objects can include, but are not limited to, vehicles and pedestrians, etc. In one case, when the object is a vehicle, the labeled position information in the labeled pose information of the corresponding object included in the true value information can refer to position information of a center point of the vehicle, position information of a center point of a vehicle tail, or position information of a center point of a vehicle head, which are all possible. The labeled attitude information in the labeled pose information of the corresponding object included in the true value information can refer to various angle information of the vehicle relative to coordinate axes of a coordinate system in which the vehicle is located during driving, including pitch angle information, roll angle information, and yaw angle information. In one case, the pitch angle information and the roll angle information generated by the vehicle during driving on the ground surface can be ignored, i.e., the pitch angle information and the roll angle information generated by the vehicle during driving on the ground surface are considered to be zero.

[0101] In an implementation, the evaluation data in the evaluation data set can include evaluation data collected for normal driving scenarios, or evaluation data collected for large vehicle or special vehicle scenarios, or evaluation data collected for pedestrians, complex intersections, and specific weather conditions, all of which are possible.

[0102] In an implementation, the electronic device can directly obtain the true value information corresponding to each object determined based on each point cloud data frame in the evaluation data set sent by the other device.

[0103] In an implementation of the present application, S101 can include the following steps 011-013:

[0104] 011: Obtain an evaluation data set;

[0105] 012: Label each point cloud data frame in the evaluation data set based on a pre-trained three-dimensional data perception model, label the bounding box information of each object corresponding to each point cloud data frame, to determine the label position information and label pose information of each object corresponding to each point cloud data frame, and obtain the label pose information of each object corresponding to each point cloud data frame;

[0106] 013: Based on the label position information and label pose information of each object corresponding to each point cloud data frame in the evaluation data set, and the time sequence information between each point cloud data frame in the evaluation data set, determine the label speed information and label acceleration information of each object corresponding to each point cloud data frame, to obtain the object motion information of each object corresponding to each point cloud data frame, and obtain the true value information corresponding to each object corresponding to each point cloud data frame.

[0107] In the present implementation, the electronic device can directly obtain an evaluation data set, wherein the evaluation data set includes a plurality of evaluation data; the electronic device inputs the point cloud data frame included in each evaluation data in the evaluation data set into a pre-trained three-dimensional data perception model, detects each object in each point cloud data frame through the pre-trained three-dimensional data perception model, and labels through a bounding box to obtain the bounding box information of each object corresponding to each point cloud data frame, wherein the bounding box can be a cube. The bounding box information of each object includes information that can represent the length, width and height of the object, and information that can represent the pose information of the object.

[0108] Subsequently, the electronic device converts the label box information of each object corresponding to each point cloud data frame output by the pre-trained three-dimensional data perception model to obtain the label position information and the label pose information of each object corresponding to each frame of point cloud data. The pre-trained three-dimensional data perception model can be a neural network model trained based on sample point cloud data frames and their corresponding calibration information including the calibration box information of each object in the sample point cloud data frames. For details of the model training process, please refer to the related art, which will not be described here.

[0109] The evaluation data in the evaluation data set is generally continuous data obtained by continuous acquisition, that is, the point cloud data frames in the evaluation data set are continuous frames, and the image frames are continuous frames. Correspondingly, the electronic device can determine the label speed information and the label acceleration information of each object corresponding to each point cloud data frame based on the label position information and the label pose information of each object corresponding to each point cloud data frame in the evaluation data set, and the time sequence information between each point cloud data frame in the evaluation data set.

[0110] In one implementation, the label position information of each object can include lateral position information, longitudinal position information, and radial position information. Then, based on the lateral position information of each object corresponding to each point cloud data frame and the time sequence information between each point cloud data frame, the label lateral speed information and the label lateral acceleration information of each object can be determined; based on the longitudinal position information of each object corresponding to each point cloud data frame and the time sequence information between each point cloud data frame, the label longitudinal speed information and the label longitudinal acceleration information of each object can be determined; and based on the radial position information of each object corresponding to each point cloud data frame and the time sequence information between each point cloud data frame, the label radial speed information and the label radial acceleration information of each object can be determined.

[0111] S102: Obtain detection information corresponding to each detected object detected based on a preset visual perception algorithm and each image frame in the evaluation data set.

[0112] The detection information at least includes detection pose information and detection motion information of the corresponding detected object.

[0113] The preset visual perception algorithm can be used to detect the detection information of each object in the image from the image frame. In order to describe clearly, the object detected from the image frame by using the preset visual perception algorithm can be called a detected object.

[0114] The detection information can include two-dimensional information and three-dimensional information corresponding to the object. The two-dimensional information corresponding to the object can include two-dimensional position information and two-dimensional velocity information of the object in the image frame. The three-dimensional information corresponding to the object can include, but is not limited to, detection pose information of the object in a specified space rectangular coordinate system, detection motion information including, but not limited to, detection velocity information and detection acceleration information of the corresponding object.

[0115] In an implementation, the electronic device can directly obtain the detection information corresponding to each detected object detected by the other device based on the preset visual perception algorithm and each image frame in the evaluation data set.

[0116] In an implementation of the present application, S102 can include the following steps 021-022.

[0117] 021: Based on the preset visual perception algorithm, each image frame in the evaluation data set is detected to obtain the detection box information corresponding to each detected object in each image frame, so as to determine the detection position information and the detection pose information of each detected object in each image frame, and obtain the detection pose information of each detected object in each image frame.

[0118] 022: Based on the preset visual perception algorithm and the detection position information and the detection pose information of each detected object in each image frame, the detection velocity information and the detection acceleration information of each detected object in each image frame are determined to obtain the detection motion information of each detected object in each image frame, and the detection information corresponding to each detected object in each image frame is obtained.

[0119] In the implementation, the electronic device or the connected storage device pre-stores the preset visual perception algorithm. After obtaining the evaluation data set, the electronic device can detect each image frame in the evaluation data set based on the preset visual perception algorithm to obtain the detection box information corresponding to each detected object in each image frame. The detection box information includes information representing the length, width and height of the corresponding detected object, information representing the pose information of the corresponding detected object, and information representing the two-dimensional position information of the corresponding detected object in the corresponding image frame.

[0120] The electronic device determines, based on the preset visual perception algorithm and the detected bounding box information corresponding to each object in each image frame, detection position information and detection posture information of each detected object corresponding to each image frame. And based on the preset visual perception algorithm and the detection position information and detection posture information of each detected object corresponding to each image frame, the detection speed information and the detection acceleration information of each detected object corresponding to each image frame are determined to obtain the detection motion information of each detected object corresponding to each image frame.

[0121] In one implementation, the detection position information of each detected object can include lateral position information, longitudinal position information and radial position information. Then, based on the lateral position information of each detected object and the time sequence information between each image frame, the detection lateral speed information and the detection lateral acceleration information of each detected object can be determined; based on the longitudinal position information of each detected object and the time sequence information between each image frame, the detection longitudinal speed information and the detection longitudinal acceleration information of each detected object can be determined; based on the radial position information of each detected object and the time sequence information between each image frame, the detection radial speed information and the detection radial acceleration information of each detected object can be determined.

[0122] S103: Based on the preset result accuracy evaluation rule, the preset algorithm stability evaluation rule, the labeled posture information and object motion information in the true value information corresponding to each object, and the detection posture information and detection motion information in the detection information corresponding to each detected object, determine the evaluation information corresponding to the preset visual perception algorithm.

[0123] The evaluation information includes first evaluation information representing the detection result accuracy of the preset visual perception algorithm and second evaluation information representing the algorithm stability.

[0124] In this step, the electronic device can process the true value information and the detection information that have a corresponding relationship based on the preset result accuracy evaluation rule to obtain the first evaluation information representing the detection result accuracy of the preset visual perception algorithm; and process the true value information and the detection information that have a corresponding relationship based on the preset algorithm stability evaluation rule to obtain the second evaluation information representing the algorithm stability of the preset visual perception algorithm.

[0125] The preset result accuracy evaluation rule can include but is not limited to a specific detection result accuracy evaluation index and a process of determining a result corresponding to the specific detection result accuracy evaluation index based on the true value information and the detection information; and the preset algorithm stability evaluation rule can include but is not limited to a specific algorithm stability evaluation index and a process of determining a result corresponding to the specific algorithm stability evaluation index based on the true value information and the detection information.

[0126] According to the embodiment of the present application, the first evaluation information of the accuracy of the multi-aspect detection result of the preset visual perception algorithm and the second evaluation information of the algorithm stability can be determined based on the preset result accuracy evaluation rule, the preset algorithm stability evaluation rule, the labeled pose information and object motion information in the true value information corresponding to each object, and the detection pose information and detection motion information in the detection information corresponding to each detection object. The preset visual perception algorithm is evaluated from the accuracy of the multi-aspect detection result and the stability of the detection result of the algorithm, and the performance of the visual perception algorithm is comprehensively evaluated.

[0127] In another embodiment of the present application, the true value information includes the labeled box information of each object corresponding to each point cloud data frame, and the detection information includes the detection box information of each detection object detected in each image frame corresponding to the image frame, and the detection box information includes the two-dimensional position information of the corresponding detection object in the image frame.

[0128] The S103 can include the following steps 031-034.

[0129] 031: For each object corresponding to each point cloud data frame, based on the labeled box information corresponding to the object, the position conversion relationship between the point cloud data frame acquisition device and the image frame acquisition device, and the intrinsic information of the image frame acquisition device, the projection position information of the projection box of the labeled box corresponding to the object in the image frame corresponding to the point cloud data frame is determined as the projection box position information corresponding to the object.

[0130] 032: For each object corresponding to each point cloud data frame, based on the projection box position information corresponding to each object and the two-dimensional position information of each detection object in the image frame corresponding to the point cloud data frame, the matching projection box position information and two-dimensional position information are determined to determine the matching true value information and detection information corresponding to each point cloud data frame and its corresponding image frame.

[0131] The matching projection box position information and two-dimensional position information are the projection box position information and two-dimensional position information whose intersection over union value exceeds a preset intersection over union threshold.

[0132] 033: Based on the preset result accuracy evaluation rule, the matching true value information and detection information corresponding to each point cloud data frame and its corresponding image frame, the detection information that does not match the true value information, and the true value information that does not match the detection information, the first evaluation information representing the accuracy of the detection result of the preset visual perception algorithm is determined.

[0133] 034: Based on the preset algorithm stability evaluation rule, the matching true value information and detection information corresponding to each point cloud data frame and its corresponding image frame, the second evaluation information representing the algorithm stability of the preset visual perception algorithm is determined.

[0134] In this implementation, before determining the evaluation information corresponding to the preset visual perception algorithm based on the true value information and the detection information, the electronic device needs to match the true value information and the detection information, so as to determine the evaluation information corresponding to the preset visual perception algorithm through the mutually matched true value information and the detection information. Correspondingly, the electronic device can, for each object corresponding to a point cloud data frame, based on the label box information corresponding to the object and the position conversion relationship between the point cloud data frame acquisition device and the image frame acquisition device, convert the label box corresponding to the label box information of the object from the coordinate system of the point cloud data frame acquisition device to the coordinate system of the image frame acquisition device to obtain the position information of the label box corresponding to the label box information of the object in the coordinate system of the image frame acquisition device; and then, based on the position information of the label box corresponding to the label box information of the object in the coordinate system of the image frame acquisition device and the intrinsic parameter information of the image frame acquisition device, project the label box corresponding to the object into the image frame corresponding to the point cloud data frame to determine the projection position information of the projection box of the label box corresponding to the object projected into the image frame corresponding to the point cloud data frame as the projection box position information corresponding to the object.

[0135] After projecting the label box corresponding to each object into the image frame corresponding to the point cloud data frame for each point cloud data frame, the intersection over union between the label box corresponding to each object and the two-dimensional detection box of each detected object in the image frame corresponding to the point cloud data frame can be calculated based on the projection box position information corresponding to each object and the two-dimensional position information of each detected object in the image frame corresponding to the point cloud data frame, that is, the ratio between the intersection area between the label box corresponding to each object and the two-dimensional detection box of each detected object in the image frame corresponding to the point cloud data frame and the union area between the label box corresponding to each object and the two-dimensional detection box of each detected object in the image frame corresponding to the point cloud data frame. For each ratio, the ratio is compared with a preset intersection over union threshold to determine the size of the ratio and the preset intersection over union threshold. If the ratio exceeds the preset intersection over union threshold, it is determined that the object corresponding to the projection box position information corresponding to the ratio and the detected object corresponding to the two-dimensional position information corresponding to the ratio are the same object, and correspondingly, the projection box position information corresponding to the ratio and the two-dimensional position information corresponding to the ratio are matched projection box position information and two-dimensional position information.

[0136] For example, the corresponding objects in the point cloud data frame A include object 1, object 2 and object 3; the corresponding detection objects in the image frame a corresponding to the point cloud data frame A include detection object 1, detection object 2, detection object 3 and detection object 4; for the object 1 corresponding to the point cloud data frame A, based on the projection frame position information corresponding to the object 1 and the two-dimensional position information of the detection object 1, the intersection over union between the projection frame corresponding to the object 1 and the two-dimensional detection frame corresponding to the detection object 1 is calculated; based on the projection frame position information corresponding to the object 1 and the two-dimensional position information of the detection object 2, the intersection over union between the projection frame corresponding to the object 1 and the two-dimensional detection frame corresponding to the detection object 2 is calculated; based on the projection frame position information corresponding to the object 1 and the two-dimensional position information of the detection object 3, the intersection over union between the projection frame corresponding to the object 1 and the two-dimensional detection frame corresponding to the detection object 3 is calculated; based on the projection frame position information corresponding to the object 1 and the two-dimensional position information of the detection object 4, the intersection over union between the projection frame corresponding to the object 1 and the two-dimensional detection frame corresponding to the detection object 4 is calculated.

[0137] Similarly, for the object 2 corresponding to the point cloud data frame A, the intersection over union between the projection frame corresponding to the object 2 and the two-dimensional detection frame corresponding to the detection object 1 is calculated; the intersection over union between the projection frame corresponding to the object 2 and the two-dimensional detection frame corresponding to the detection object 2 is calculated; the intersection over union between the projection frame corresponding to the object 2 and the two-dimensional detection frame corresponding to the detection object 3 is calculated; the intersection over union between the projection frame corresponding to the object 2 and the two-dimensional detection frame corresponding to the detection object 4 is calculated. And for the object 3 corresponding to the point cloud data frame A, the intersection over union between the projection frame corresponding to the object 3 and the two-dimensional detection frame corresponding to the detection object 1 is calculated; the intersection over union between the projection frame corresponding to the object 3 and the two-dimensional detection frame corresponding to the detection object 2 is calculated; the intersection over union between the projection frame corresponding to the object 3 and the two-dimensional detection frame corresponding to the detection object 3 is calculated; the intersection over union between the projection frame corresponding to the object 3 and the two-dimensional detection frame corresponding to the detection object 4 is calculated.

[0138] The sizes of each intersection over union, i.e. the ratio, and the preset intersection over union threshold value are respectively judged; for example: if the intersection over union between the projection frame corresponding to the object 1 and the two-dimensional detection frame corresponding to the detection object 3 exceeds the preset intersection over union threshold value, it is determined that the projection frame position information corresponding to the object 1 and the two-dimensional position information corresponding to the detection object 3 are matched projection frame position information and two-dimensional position information, and correspondingly, the true value information corresponding to the object 1 and the detection information corresponding to the detection object 3 are matched true value information and detection information.

[0139] If the cross-union ratio (CUI) between the projection box corresponding to object 3 and the 2D detection box corresponding to object 1, the cross-union ratio between the projection box corresponding to object 3 and the 2D detection box corresponding to object 2, the cross-union ratio between the projection box corresponding to object 3 and the 2D detection box corresponding to object 3, and the cross-union ratio between the projection box corresponding to object 3 and the 2D detection box corresponding to object 4 all do not exceed the preset CUI threshold, then it is determined that there is no detection object among the detection objects 1-4 that is the same physical object as object 3. That is, the truth information corresponding to object 3 is the truth information of unmatched detection information. Accordingly, object 3 can be called a missed detection object.

[0140] If the cross-union ratio (CUI) between the projection box of object 1 and the 2D detection box of object 4, the cross-union ratio between the projection box of object 2 and the 2D detection box of object 4, and the cross-union ratio between the projection box of object 3 and the 2D detection box of object 4 all do not exceed the preset CUI threshold, then it is determined that there is no object among objects 1-3 that is the same physical object as the detected object 4. That is, the detection information corresponding to the detected object 4 is the detection information that has not matched the true value information. Accordingly, the detected object 4 can be called a false detection object.

[0141] Subsequently, the electronic device can determine the first evaluation information representing the accuracy of the detection results of the preset visual perception algorithm based on the preset result accuracy evaluation rules, the ground truth information and detection information corresponding to each point cloud data frame and its corresponding image frame, the detection information that does not match the ground truth information, and the ground truth information that does not match the detection information. Based on the preset algorithm stability evaluation rules, the ground truth information and detection information corresponding to each point cloud data frame and its corresponding image frame, the second evaluation information representing the algorithm stability of the preset visual perception algorithm can be determined.

[0142] In another embodiment of the present invention, the detection pose information includes: the detection position information and detection pose information of each detected object corresponding to each image frame, determined by its detection box information; the detection motion information includes: the detection velocity information and detection acceleration information of each detected object corresponding to each image frame.

[0143] The annotation pose information includes: the annotation position information and annotation pose information of each object corresponding to each point cloud data frame, determined by its annotation box information; the object motion information includes: the annotation velocity information and annotation acceleration information of each object corresponding to each point cloud data frame.

[0144] The 033 may include the following steps 0331-0339:

[0145] 0331: Determine the precision information and recall information of the detection result corresponding to the preset visual perception algorithm based on the matched true value information and detection information corresponding to each point cloud data frame and its corresponding image frame, the detection information not matched to the true value information, and the true value information not matched to the detection information.

[0146] 0332: Determine the detection position error value between the matched true value information and detection information based on the annotation position information included in the matched true value information and the detection position information included in the detection information corresponding to each point cloud data frame and its corresponding image frame.

[0147] 0333: Determine the detection posture error value between the matched true value information and detection information based on the annotation posture information included in the matched true value information and the detection posture information included in the detection information corresponding to each point cloud data frame and its corresponding image frame.

[0148] 0334: Determine the detection speed error value between the matched true value information and detection information based on the annotation speed information included in the matched true value information and the detection speed information included in the detection information corresponding to each point cloud data frame and its corresponding image frame.

[0149] 0335: Determine the detection acceleration error value between the matched true value information and detection information based on the annotation acceleration information included in the matched true value information and the detection acceleration information included in the detection information corresponding to each point cloud data frame and its corresponding image frame.

[0150] 0336: Determine the length-width error value of the detection box between the matched true value information and detection information based on the annotation box information included in the matched true value information and the detection box information included in the detection information corresponding to each point cloud data frame and its corresponding image frame.

[0151] 0337: Draw an error curve corresponding to the target error value based on the target error value between the matched true value information and detection information and the preset error threshold value corresponding to the target error value.

[0152] Wherein, the horizontal axis of the error curve is the preset error threshold value, the vertical axis of the error curve is the ratio of the number of target error values less than each preset error threshold value in the target error value to the total amount of data in the evaluation data set, and the target error value is: detection position error value, detection posture error value, detection speed error value, detection acceleration error value, or length-width error value of the detection box.

[0153] 0338: Sort the target error values between the matched true value information and the detection information according to the size of the numerical values to obtain a sorting sequence corresponding to the target error values; determine a first target error value with the largest numerical value in the first percentage of the target error values in the sorting sequence corresponding to the target error values, and a second target error value with the largest numerical value in the second percentage of the target error values.

[0154] 0339: Based on the precision information and the recall information of the detection result corresponding to the preset visual perception algorithm, the error curve corresponding to the target error information, the first target error value and the second target error value and / or the target error value, determine the first evaluation information representing the accuracy of the detection result of the preset visual perception algorithm.

[0155] In the present implementation, the preset visual perception algorithm can detect 2D information and 3D information of each detected object in the image frame, including two-dimensional position information of each detected object in the image frame, detection position information of each detected object in the specified space rectangular coordinate system, detection attitude information, detection speed information and detection acceleration information.

[0156] In order to realize the multi-dimensional evaluation of the accuracy of the detection result of the preset visual perception algorithm and the stability of the algorithm, the true value information includes a plurality of dimensional annotation parameters of the corresponding object, which can include but not limited to annotated position information, annotated attitude, annotated speed information and annotated acceleration information of the corresponding object.

[0157] The preset result accuracy evaluation rule can include a rule indicating the determination of the precision information and the recall information of the detection result. Accordingly, the electronic device can determine the precision information and the recall information of the detection result corresponding to the preset visual perception algorithm based on the matched true value information and the detection information, the detection information not matched to the true value information, and the true value information not matched to the detection information of a point cloud data frame and its corresponding image frame according to the preset precision information determination manner and the preset recall rate determination manner. The precision information and the recall information of the detection result corresponding to the preset visual perception algorithm are used as an evaluation index for evaluating the accuracy of the detection result corresponding to the preset visual perception algorithm. The preset precision information determination manner and the preset recall rate determination manner can refer to the precision information determination manner and the recall rate determination manner in related technologies, which will not be described here.

[0158] In the present implementation, a new detection result accuracy evaluation index corresponding to a preset visual perception algorithm is added, a special form of error curve is drawn, and the detection result accuracy corresponding to the preset visual perception algorithm is evaluated through the error curve. Another newly added detection result accuracy evaluation index corresponding to a preset visual perception algorithm is: counting the number of error values corresponding to different proportions in the same dimension error values, and then evaluating the detection result accuracy corresponding to the preset visual perception algorithm based on the number of error values corresponding to different proportions in the same dimension error values counted.

[0159] Specifically, first, the error values corresponding to different dimensions are calculated based on the matching true value information and the detection information:

[0160] Based on the annotation position information included in the matching true value information and the detection position information included in the detection information corresponding to each point cloud data frame and its corresponding image frame, the detection position error value between the matching true value information and the detection information is determined. That is, for each matching true value information and detection information, the coordinate systems of the annotation position information and the detection position information are unified, and then the detection position error value between the unified annotation position information and the unified detection position information is calculated; wherein the detection position error value can be an absolute error value and / or a relative error value between the annotation position information and the detection position information.

[0161] The detection position error value can include but is not limited to a position error value in a horizontal direction, a position error value in a vertical direction, and a combined position error value. The combined position error value is a combined value of the position error values in each direction. The vertical direction can refer to the driving direction of the target vehicle, and the horizontal direction can refer to the vertical direction of the driving direction of the target vehicle.

[0162] Based on the annotation pose information included in the matching true value information and the detection pose information included in the detection information corresponding to each point cloud data frame and its corresponding image frame, the detection pose error value between the matching true value information and the detection information is determined. That is, for each matching true value information and detection information, the coordinate systems of the annotation pose information and the detection pose information are unified, and then the detection pose error value between the unified annotation pose information and the unified detection pose information is calculated; wherein the detection pose error value can be an absolute error value and / or a relative error value between the annotation pose information and the detection pose information.

[0163] The detection pose error value can include but is not limited to a pose error value in a horizontal direction, a pose error value in a vertical direction, and a combined pose error value. The combined pose error value is a combined value of the pose error values in each direction.

[0164] Based on the matching true value information corresponding to each point cloud data frame and its corresponding image frame, the labeled speed information and the detected speed information included in the detection information are determined, that is, for each matching true value information and detection information, the coordinate system of the labeled speed information and the detected speed information is unified, and then the detection speed error value between the labeled speed information and the detected speed information after the unified coordinate system is calculated; wherein the detection speed error value can be the absolute error value and / or the relative error value between the labeled speed information and the detected speed information.

[0165] The detection speed error value can include but is not limited to the speed error value in the lateral direction, the speed error value in the longitudinal direction, and the synthetic speed error value, wherein the synthetic speed error value is the synthetic value of the speed error values in each direction.

[0166] Based on the matching true value information corresponding to each point cloud data frame and its corresponding image frame, the labeled speed information and the detected speed information included in the detection information are determined, that is, for each matching true value information and detection information, the coordinate system of the labeled speed information and the detected speed information is unified, and then the detection speed error value between the labeled speed information and the detected speed information after the unified coordinate system is calculated; wherein the detection speed error value can be the absolute error value and / or the relative error value between the labeled speed information and the detected speed information.

[0167] The detection speed error value can include but is not limited to the speed error value in the lateral direction, the speed error value in the longitudinal direction, and the synthetic speed error value, wherein the synthetic speed error value is the synthetic value of the speed error values in each direction.

[0168] Based on the matching true value information corresponding to each point cloud data frame and its corresponding image frame, the labeled speed information and the detected speed information included in the detection information are determined, that is, for each matching true value information and detection information, the coordinate system of the labeled speed information and the detected speed information is unified, and then the detection speed error value between the labeled speed information and the detected speed information after the unified coordinate system is calculated; wherein the detection speed error value can be the absolute error value and / or the relative error value between the labeled speed information and the detected speed information.

[0169] In one case, the electronic device can sequentially take the determined detection position error value, the detection posture error value, the detection speed error value, the detection acceleration error value, and the length-width error value of the detection frame as target error values, respectively; draw an error curve corresponding to the target error values based on the target error values between the matching true value information and the detection information and the preset error threshold corresponding to the target error values. Specifically, for each preset error threshold, the number of target error values less than the preset error threshold in the target error values can be counted, and the ratio of the number of target error values less than the preset error threshold in the target error values to the total amount of data in the evaluation dataset can be calculated; then, the preset error threshold is taken as the horizontal axis of the error curve corresponding to the target error values, and the number of target error values less than each preset error threshold in the target error values is taken as the vertical axis of the error curve corresponding to the target error values, to draw the error curve corresponding to the target error values. The preset error threshold includes a plurality of preset error thresholds, which can be set from 0 and increase sequentially.

[0170] As shown in Figure 2 , it is an example diagram of the error curve corresponding to the target error values.

[0171] In another case, the electronic device sorts the target error values between the matching true value information and the detection information according to the size of the values to obtain a sorting sequence corresponding to the target error values; determines a first target error value with the largest value in the first percentage of the target error values in the sorting sequence, and a second target error value with the largest value in the second percentage of the target error values. The first target error value can be referred to as 1sigma, and the second target error value can be referred to as 2sigma.

[0172] For example, the first percentage can be the first 68.26% of the sorting sequence, and the second percentage can be the first 95.44% of the sorting sequence.

[0173] Subsequently, the electronic device can determine first evaluation information representing the accuracy of the detection result of the preset visual perception algorithm based on the precision information and the recall information of the detection result corresponding to the preset visual perception algorithm, the error curve corresponding to the target error information, the first target error value and the second target error value, and / or the target error value.

[0174] It can be understood that the higher the precision information and the recall information of the detection result corresponding to the preset visual perception algorithm, the higher the accuracy of the detection result corresponding to the preset visual perception algorithm.

[0175] For the error curve corresponding to the target error values, as shown in Figure 2As shown, the greater the area under the error curve (AUC: Area Under Curve) corresponding to the target error value, that is, the closer the area value to 1, the smaller the error of the corresponding dimension of the matching true value information and the detection information as a whole, that is, the better the performance of the preset visual perception algorithm. For example, the target error value is the detection position error value, which measures the position judgment ability of the preset visual perception algorithm. The greater the area under the curve in the error curve corresponding to the detection position error value, the better the position judgment ability of the preset visual perception algorithm. For another example, the target error value is the detection posture error value, which measures the posture judgment ability of the preset visual perception algorithm. The greater the area under the curve in the error curve corresponding to the detection posture error value, the better the position judgment ability of the preset visual perception algorithm.

[0176] For the evaluation index of the number of error values corresponding to different proportions in the same dimension error value, the smaller the values of the first target error value and the second target error value determined, the better the performance of the preset visual perception algorithm, that is, the higher the accuracy of the detection result.

[0177] In one implementation, the evaluation index of the accuracy of the detection result corresponding to the preset visual perception algorithm can also include, but is not limited to, P-R curve and the proportion of false detection and missed detection of objects in the detection result in the whole detection result, and the like.

[0178] In order to improve the comprehensiveness of the evaluation of the preset visual perception algorithm, in addition to evaluating the accuracy of the detection result corresponding to the preset visual perception algorithm, the algorithm stability of the preset visual perception algorithm is also evaluated in the embodiments of the present application. In another embodiment of the present application, the034may include the following steps:

[0179] 0341: determining the target error value corresponding to the same object from the target error value based on the time sequence information between the point cloud data frames or image frames in the evaluation data set;

[0180] 0342: fitting the target error curve corresponding to the target error value of the object based on the time sequence information between the point cloud data frames or image frames in the evaluation data set, the target error value of the object, and a preset curve fitting algorithm,

[0181] wherein the fitted error curve includes the fitted error corresponding to the object at each acquisition time corresponding to the point cloud data frame or image frame;

[0182] 0343: For the target error value corresponding to different objects, based on the target error value corresponding to the object and the fitting error curve corresponding to the target error value corresponding to the object at the acquisition time corresponding to each point cloud data frame or image frame, determine the second evaluation information that characterizes the algorithm stability of the preset visual perception algorithm.

[0183] In this implementation, the point cloud data frames and image frames in the evaluation dataset exhibit temporal continuity. Therefore, the electronic device can determine the target error value corresponding to the same object from the target error values ​​based on the temporal information between the point cloud data frames or image frames in the evaluation dataset. The target error values ​​corresponding to the same object are arranged according to the temporal information between their corresponding point cloud data frames or image frames.

[0184] Accordingly, for different objects, the electronic device, based on the temporal information between point cloud data frames or image frames in the evaluation dataset, the target error value corresponding to the object, and a preset curve fitting algorithm, fits the target error value corresponding to the object to obtain a fitting error curve. The fitting error curve corresponding to the target error value of the object may include: the fitting error of the object at the corresponding acquisition time of each point cloud data frame or image frame. The preset curve fitting algorithm can be any type of curve fitting algorithm in related technologies, and this embodiment of the invention does not limit it.

[0185] Subsequently, the electronic device determines second evaluation information characterizing the algorithm stability of the preset visual perception algorithm based on the target error value corresponding to different objects and the fitting error curve corresponding to the target error value of the object at the acquisition time corresponding to each point cloud data frame or image frame. In another embodiment of the present invention, 0343 includes:

[0186] 03431: For different objects, based on the target error value corresponding to the object and the fitting error of the object at the acquisition time corresponding to each point cloud data frame or image frame contained in the fitting error curve corresponding to the target error value of the object, calculate the difference between the target error and the fitting error corresponding to the same acquisition time.

[0187] 03432: For the target error value corresponding to different objects, based on the difference between the target error and the fitting error at each acquisition time corresponding to the object, and the preset difference threshold corresponding to the target error value of the object, the difference error curve corresponding to the target error of the object is plotted.

[0188] The horizontal axis of the difference error curve corresponding to the target error of the object is a preset difference threshold corresponding to the target error value of the object. The vertical axis of the difference error curve corresponding to the target error of the object is the ratio of the number of differences less than each preset difference threshold to the total number of target error values of the object.

[0189] 03433: For the target error values corresponding to different objects, the differences corresponding to the target error values of the object are sorted according to the size of the values, the first difference with the largest value in the first third of the sorted sequence is determined, and the second difference with the largest value in the first fourth of the sorted sequence is determined.

[0190] 03434: Based on the first difference and the second difference corresponding to the target error values of each object, and / or the difference error curve corresponding to the target error values of each object, second evaluation information representing the algorithm stability of the preset visual perception algorithm is determined.

[0191] In the present implementation, the target error values of the object correspond to different point cloud data frames or image frames, and each point cloud data frame or image frame corresponds to a collection time. Correspondingly, in one implementation, the target error values of the object correspond to different collection times. Therefore, the electronic device can calculate the difference between the target error and the fitting error corresponding to the same collection time based on the target error values of the object and the fitting errors corresponding to the collection times of the object in each point cloud data frame or image frame included in the fitting error curve corresponding to the target error values of the object.

[0192] Further, for different preset difference thresholds corresponding to the target error values of the object, the number of differences less than the preset difference threshold is counted among the differences between the target error and the fitting error of each collection time of the object. The ratio of the number of differences less than the preset difference threshold to the total number of target error values of the object is calculated. The difference error curve corresponding to the target error of the object is drawn with the ratio of the number of differences less than each preset difference threshold to the total number of target error values of the object as the vertical axis and the preset difference threshold corresponding to the target error value of the object as the horizontal axis.

[0193] In another implementation, for the target error values corresponding to different objects, the electronic device can sort the differences corresponding to the target error values of the object according to the size of the values, determine the first difference with the largest value in the first third of the sorted sequence, and determine the second difference with the largest value in the first fourth of the sorted sequence. The third percentage can be the same as or different from the first percentage, and the fourth percentage can be the same as or different from the second percentage.

[0194] Subsequently, the electronic device can determine second evaluation information representing algorithm stability of the preset visual perception algorithm based on the first difference value and the second difference value corresponding to the target error value of each object and / or the difference error curve corresponding to the target error of each object. The smaller the value between the first difference value and the second difference value corresponding to the target error value of each object, the smaller the difference between the target error and the fitting error of each collection time corresponding to the object, and accordingly, the better the algorithm stability of the preset visual perception algorithm in the dimension corresponding to the target error value.

[0195] The larger the area under the difference error curve corresponding to the target error value of each object, that is, the closer the area value to 1, the smaller the error between the target error and the fitting error of each collection time corresponding to the object as a whole, that is, the better the performance of the preset visual perception algorithm, and the higher the detection stability of the preset visual perception algorithm in the dimension corresponding to the target error value.

[0196] In order to provide user experience and help users understand the evaluation information, the embodiment of the present application can visually display the evaluation information and the intermediate information generated in the evaluation process. In another embodiment of the present application, the method further comprises:

[0197] displaying the first evaluation information, the second evaluation information, the precision rate information and the recall rate information of the detection result corresponding to the preset visual perception algorithm, the error curve corresponding to the target error information, the first target error value and the second target error value, the target error value, the determined target error value corresponding to the same object, and / or the fitting error curve corresponding to the target error value corresponding to each object fitted.

[0198] The electronic device can visually display the evaluation information and the intermediate information generated in the evaluation process in a web-based visualization manner or a 3D rendering-based visualization manner, and can provide a function for interacting with the user.

[0199] Corresponding to the method embodiment described above, the embodiment of the present application provides an evaluation device for visual perception algorithm, as shown in Figure 3 The device can include:

[0200] The first obtaining module 310 is configured to obtain the true value information corresponding to each object determined based on each point cloud data frame in the evaluation data set, wherein the true value information at least includes the labeled pose information of the corresponding object and the object motion information, and each evaluation data includes point cloud data frame and image frame having a corresponding relationship;

[0201] The second obtaining module 320 is configured to obtain detection information corresponding to each detected object detected based on the preset visual perception algorithm and each image frame in the evaluation data set, wherein the detection information at least includes detection pose information and detection motion information of the corresponding detected object;

[0202] The determining module 330 is configured to determine evaluation information corresponding to the preset visual perception algorithm based on the preset result accuracy evaluation rule, the preset algorithm stability evaluation rule, the labeled pose information and the object motion information in the true value information corresponding to each object, and the detection pose information and the detection motion information in the detection information corresponding to each detected object, wherein the evaluation information includes first evaluation information representing detection result accuracy of the preset visual perception algorithm and second evaluation information representing algorithm stability.

[0203] According to the embodiment of the present application, the first evaluation information representing the accuracy of the detection result of the preset visual perception algorithm and the second evaluation information representing the stability of the algorithm can be determined based on the preset result accuracy evaluation rule, the preset algorithm stability evaluation rule, the labeled pose information and the object motion information in the true value information corresponding to each object, and the detection pose information and the detection motion information in the detection information corresponding to each detected object. The preset visual perception algorithm is evaluated from the accuracy of the detection result and the stability of the algorithm, so that the performance of the visual perception algorithm is comprehensively evaluated.

[0204] In another embodiment of the present application, the first obtaining module 310 is specifically configured to obtain an evaluation data set;

[0205] Each point cloud data frame in the evaluation data set is labeled based on the pre-trained three-dimensional data perception model, and the labeled box information of each object corresponding to each point cloud data frame is labeled to determine the labeled position information and the labeled pose information of each object corresponding to each point cloud data frame, so as to obtain the labeled pose information of each object corresponding to each point cloud data frame.

[0206] The labeled speed information and the labeled acceleration information of each object corresponding to each point cloud data frame are determined based on the labeled position information and the labeled pose information of each object corresponding to each point cloud data frame in the evaluation data set and the time sequence information between each point cloud data frame in the evaluation data set, so as to obtain the object motion information of each object corresponding to each point cloud data frame, and obtain the true value information corresponding to each object corresponding to each point cloud data frame.

[0207] In another embodiment of the present application, the second obtaining module 320 is specifically configured to detect each image frame in the evaluation data set based on a preset visual perception algorithm to obtain detection box information of each detected object corresponding to each image frame, so as to determine detection position information and detection posture information of each detected object corresponding to each image frame, and obtain detection position information of each detected object corresponding to each image frame.

[0208] Based on the preset visual perception algorithm and the detection position information and the detection posture information of each detected object corresponding to each image frame, detection speed information and detection acceleration information of each detected object corresponding to each image frame are determined, so as to obtain detection motion information of each detected object corresponding to each image frame, and obtain detection information corresponding to each detected object corresponding to each image frame.

[0209] In another embodiment of the present application, the true value information includes label box information of each object corresponding to each point cloud data frame, and the detection information includes detection box information of each detected object corresponding to each image frame, and the detection box information includes two-dimensional position information of the corresponding detection object in the image frame.

[0210] The determining module 330 includes:

[0211] A first determining unit (not shown in the figure) is configured to, for each object corresponding to each point cloud data frame, determine, as projection position information of a projection box corresponding to the object, projection position information of a projection of a label box corresponding to the object in an image frame corresponding to the point cloud data frame based on the label box information corresponding to the object, a position conversion relationship between a point cloud data frame acquisition device and an image frame acquisition device, and intrinsic parameter information of the image frame acquisition device.

[0212] A second determining unit (not shown in the figure) is configured to, for each object corresponding to each point cloud data frame, determine matching projection box position information and two-dimensional position information based on the projection box position information corresponding to each object and two-dimensional position information of each detected object in an image frame corresponding to the point cloud data frame, so as to determine matching true value information and detection information corresponding to each point cloud data frame and the image frame corresponding thereto, wherein the matching projection box position information and the two-dimensional position information are projection box position information and two-dimensional position information whose intersection over union value exceeds a preset intersection over union threshold.

[0213] A third determining unit (not shown in the figure) is configured to determine first evaluation information representing detection result accuracy of the preset visual perception algorithm based on a preset result accuracy evaluation rule, matching true value information and detection information corresponding to each point cloud data frame and the image frame corresponding thereto, detection information that does not match true value information, and true value information that does not match detection information.

[0214] A fourth determining unit (not shown in the figure) is configured to determine second evaluation information representing algorithm stability of the preset visual perception algorithm based on a preset algorithm stability evaluation rule, matching true value information and detection information corresponding to each point cloud data frame and its corresponding image frame.

[0215] In another embodiment of the present application, the detected pose information includes detection position information and detection posture information of each detected object in each image frame determined by its detection box information; the detected motion information includes detection speed information and detection acceleration information of each detected object in each image frame; the labeled pose information includes labeled position information and labeled posture information of each object corresponding to each point cloud data frame determined by its labeled box information; and the object motion information includes labeled speed information and labeled acceleration information of each object corresponding to each point cloud data frame.

[0216] The third determining unit is specifically configured to determine precision rate information and recall rate information of a detection result corresponding to the preset visual perception algorithm based on matching true value information and detection information corresponding to each point cloud data frame and its corresponding image frame, detection information not matched to true value information, and true value information not matched to detection information.

[0217] Based on labeled position information included in the matching true value information and detection position information included in the detection information corresponding to each point cloud data frame and its corresponding image frame, a detection position error value between the matching true value information and the detection information is determined.

[0218] Based on labeled posture information included in the matching true value information and detection posture information included in the detection information corresponding to each point cloud data frame and its corresponding image frame, a detection posture error value between the matching true value information and the detection information is determined.

[0219] Based on labeled speed information included in the matching true value information and detection speed information included in the detection information corresponding to each point cloud data frame and its corresponding image frame, a detection speed error value between the matching true value information and the detection information is determined.

[0220] Based on labeled acceleration information included in the matching true value information and detection acceleration information included in the detection information corresponding to each point cloud data frame and its corresponding image frame, a detection acceleration error value between the matching true value information and the detection information is determined.

[0221] Determine a length-width error value of a detection box between the matched ground truth information and the detection information based on the annotation box information included in the matched ground truth information and the detection box information included in the detection information corresponding to each point cloud data frame and its corresponding image frame;

[0222] Draw an error curve corresponding to the target error value based on the target error value between the matched ground truth information and the detection information and a preset error threshold corresponding to the target error value, wherein the horizontal axis of the error curve is the preset error threshold, the vertical axis of the error curve is a ratio of a number of target error values smaller than each preset error threshold in the target error value to a total amount of data in the evaluation data set, and the target error value is a detection position error value, a detection posture error value, a detection speed error value, a detection acceleration error value, or a length-width error value of a detection box.

[0223] Sort the target error values between the matched ground truth information and the detection information according to the size of the numerical value to obtain a sorting sequence corresponding to the target error value, and determine a first target error value with the largest numerical value in the first percentage of the target error values in the sorting sequence corresponding to the target error value and a second target error value with the largest numerical value in the second percentage of the target error values.

[0224] Determine first evaluation information representing the detection result accuracy of the preset visual perception algorithm based on the precision information and the recall information of the detection result corresponding to the preset visual perception algorithm, the error curve corresponding to the target error information, the first target error value and the second target error value, and / or the target error value.

[0225] In another embodiment of the application, the fourth determination unit is specifically configured to determine the target error value corresponding to the same object from the target error value based on the time sequence information between the point cloud data frames or the image frames in the evaluation data set.

[0226] For the target error values corresponding to different objects, a fitting error curve corresponding to the target error value of the object is fitted based on the time sequence information between the point cloud data frames or the image frames in the evaluation data set, the target error value of the object, and a preset curve fitting algorithm, wherein the fitting error curve includes the fitting error corresponding to the object at each acquisition time corresponding to each point cloud data frame or image frame.

[0227] For the target error values corresponding to different objects, the second evaluation information representing the algorithm stability of the preset visual perception algorithm is determined based on the target error value of the object and the fitting error corresponding to the object at each acquisition time corresponding to each point cloud data frame or image frame included in the fitting error curve corresponding to the target error value of the object.

[0228] In another embodiment of the present application, the fourth determining unit is specifically configured to, for the target error values corresponding to different objects, calculate the difference between the target error and the fitting error corresponding to the same acquisition time based on the target error values corresponding to the objects, the target error values corresponding to the objects, and the fitting error curves containing the fitting error of the object at each point cloud data frame or image frame corresponding to the acquisition time of the object.

[0229] For the target error values corresponding to different objects, the difference value error curve corresponding to the target error of the object is drawn based on the difference between the target error and the fitting error corresponding to each acquisition time of the object and the preset difference value threshold corresponding to the target error value of the object, wherein the horizontal axis of the difference value error curve corresponding to the target error of the object is the preset difference value threshold corresponding to the target error value of the object, and the vertical axis of the difference value error curve corresponding to the target error of the object is the ratio of the number of difference values smaller than each preset difference value threshold to the total number of target error values of the object.

[0230] For the target error values corresponding to different objects, the difference values corresponding to the target error values of the object are sorted according to the size of the values, the first difference value with the largest value in the first third of the sequence is determined, and the second difference value with the largest value in the first fourth of the sequence is determined.

[0231] Based on the first difference value and the second difference value corresponding to the target error values of each object and / or the difference value error curve corresponding to the target error values of each object, the second evaluation information representing the algorithm stability of the preset visual perception algorithm is determined.

[0232] In another embodiment of the present application, the device further comprises:

[0233] The display module (not shown in the figure) is configured to display the first evaluation information, the second evaluation information, the precision rate information and the recall rate information of the detection result corresponding to the preset visual perception algorithm, the error curve corresponding to the target error information, the first target error value and the second target error value, the target error value, the target error value corresponding to the same object determined, and / or the fitting error curve corresponding to the target error value of each object fitted.

[0234] The system and device embodiments described above correspond to the system embodiments and have the same technical effects as the method embodiments. For specific descriptions, refer to the method embodiments. The device embodiments are based on the method embodiments, and specific descriptions can be referred to the method embodiment part, which will not be described here. Those skilled in the art can understand that the drawings are only a schematic diagram of an embodiment, and the modules or flows in the drawings are not necessarily necessary for implementing the present application.

[0235] Those skilled in the art can understand that the modules in the device in the embodiments can be distributed in the device in the embodiments as described in the embodiments, or can be changed to be located in one or more devices different from the embodiments. The modules in the above embodiments can be combined into one module, or can be further split into multiple sub-modules.

[0236] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for evaluating a visual perception algorithm, the method comprising: The method comprises: obtaining true value information corresponding to each object determined based on each point cloud data frame in an evaluation data set, wherein the true value information at least includes labeled pose information and object motion information of the corresponding object, and each evaluation data includes point cloud data frames and image frames having a corresponding relationship; obtaining detection information corresponding to each detection object detected based on a preset visual perception algorithm and each image frame in the evaluation data set, wherein the detection information at least includes detection pose information and detection motion information of the corresponding detection object; determining evaluation information corresponding to the preset visual perception algorithm based on a preset result accuracy evaluation rule, a preset algorithm stability evaluation rule, the labeled pose information and the object motion information in the true value information corresponding to each object, and the detection pose information and the detection motion information in the detection information corresponding to each detection object, wherein the evaluation information includes first evaluation information representing detection result accuracy of the preset visual perception algorithm and second evaluation information representing algorithm stability of the preset visual perception algorithm; determining the second evaluation information representing algorithm stability of the preset visual perception algorithm based on the preset algorithm stability evaluation rule, the true value information corresponding to each object, and the detection information corresponding to each detection object, comprising: determining a target error value corresponding to the same object from target error values between matched true value information and detection information based on time sequence information between point cloud data frames or image frames in the evaluation data set, wherein the matched true value information and detection information are matched true value information and detection information corresponding to each point cloud data frame and its corresponding image frame, and the target error value is a detection position error value, a detection attitude error value, a detection speed error value, a detection acceleration error value, or a length-width error value of a detection box; for each target error value corresponding to an object, fitting a fitting error curve corresponding to the target error value of the object based on time sequence information between point cloud data frames or image frames in the evaluation data set, the target error value of the object, and a preset curve fitting algorithm, wherein the fitting error curve includes fitting errors corresponding to the object at each point cloud data frame or image frame corresponding to the collection time; for each target error value corresponding to an object, calculating a difference value between the target error and the fitting error corresponding to the same collection time based on the target error value of the object and fitting errors of the object at each point cloud data frame or image frame corresponding to the collection time included in the fitting error curve corresponding to the target error value of the object. For each object corresponding to the target error value, based on the difference between the target error and the fitting error of each collection time corresponding to the object, and the preset difference threshold corresponding to the target error value of the object, a difference error curve corresponding to the target error of the object is drawn, wherein the horizontal axis of the difference error curve corresponding to the target error of the object is the preset difference threshold corresponding to the target error value of the object, and the vertical axis of the difference error curve corresponding to the target error of the object is the ratio of the number of differences less than each preset difference threshold to the total number of target error values corresponding to the target error of the object. For each object corresponding to the target error value, the difference corresponding to the target error value of the object is sorted according to the size of the value, and the first difference with the largest value in the first third of the sequence is determined, and the second difference with the largest value in the first fourth of the sequence is determined. Based on the first difference and the second difference corresponding to the target error value of each object, and / or the difference error curve corresponding to the target error value of each object, second evaluation information representing the algorithm stability of the preset visual perception algorithm is determined.

2. The method of claim 1, wherein, The process of obtaining the true value information of each object determined based on each point cloud data frame in the evaluation data set includes: obtaining an evaluation data set; annotating each point cloud data frame in the evaluation data set based on a pre-trained three-dimensional data perception model to annotate the bounding box information of each object corresponding to each point cloud data frame to determine the annotation position information and the annotation pose information of each object corresponding to each point cloud data frame, and obtain the annotation pose information of each object corresponding to each point cloud data frame; based on the annotation position information and the annotation pose information of each object corresponding to each point cloud data frame in the evaluation data set, and the time sequence information between each point cloud data frame in the evaluation data set, the annotation speed information and the annotation acceleration information of each object corresponding to each point cloud data frame are determined to obtain the object motion information of each object corresponding to each point cloud data frame, and the true value information corresponding to each object corresponding to each point cloud data frame is obtained.

3. The method of claim 1, wherein, The step of obtaining the detection information corresponding to each detection object detected based on the preset visual perception algorithm and each image frame in the evaluation data set includes: based on the preset visual perception algorithm, each image frame in the evaluation data set is detected to obtain the detection box information of each object corresponding to each image frame to determine the detection position information and the detection pose information of each detection object corresponding to each image frame, and obtain the detection pose information of each detection object corresponding to each image frame; based on the preset visual perception algorithm and the detection position information and the detection pose information of each detection object corresponding to each image frame, the detection speed information and the detection acceleration information of each detection object corresponding to each image frame are determined to obtain the detection motion information of each detection object corresponding to each image frame, and the detection information corresponding to each detection object corresponding to each image frame is obtained.

4. The method according to any one of claims 1 to 3, characterized in that, The true value information includes bounding box information of each object corresponding to each point cloud data frame, and the detection information includes detection box information of each detected object corresponding to each image frame, and the detection box information includes two-dimensional position information of the corresponding detection object in the image frame; The step of determining the evaluation information corresponding to the preset visual perception algorithm based on the preset result accuracy evaluation rule, the preset algorithm stability evaluation rule, the labeled pose information and the object motion information in the true value information corresponding to each object, and the detection pose information and the detection motion information in the detection information corresponding to each detection object, comprises: For each object corresponding to each point cloud data frame, based on the labeled box information corresponding to the object, the position conversion relationship between the point cloud data frame acquisition device and the image frame acquisition device, and the intrinsic information of the image frame acquisition device, the projection position information of the projection of the labeled box corresponding to the object to the projection box in the image frame corresponding to the point cloud data frame is determined as the projection box position information corresponding to the object; For each object corresponding to each point cloud data frame, based on the projection box position information corresponding to each object and the two-dimensional position information of each detection object in the image frame corresponding to the point cloud data frame, the matching projection box position information and two-dimensional position information are determined to determine the matching true value information and detection information corresponding to each point cloud data frame and its corresponding image frame, wherein the matching projection box position information and two-dimensional position information are: the projection box position information and the two-dimensional position information of the corresponding box whose intersection over union value exceeds the preset intersection over union threshold; Based on the preset result accuracy evaluation rule, the matching true value information and detection information corresponding to each point cloud data frame and its corresponding image frame, the detection information that does not match the true value information, and the true value information that does not match the detection information, the first evaluation information representing the detection result accuracy of the preset visual perception algorithm is determined. Based on the preset algorithm stability evaluation rule, the matching true value information and detection information corresponding to each point cloud data frame and its corresponding image frame, the second evaluation information representing the algorithm stability of the preset visual perception algorithm is determined.

5. The method of claim 4, wherein, The detection pose information includes: the detection position information and the detection posture information of each detected object corresponding to each image frame determined by its detection box information; the detection motion information includes: the detection speed information and the detection acceleration information of each detected object corresponding to each image frame; the labeled pose information includes: the labeled position information and the labeled posture information of each object corresponding to each point cloud data frame determined by its labeled box information; and the object motion information includes: the labeled speed information and the labeled acceleration information of each object corresponding to each point cloud data frame; The step of determining the first evaluation information representing the detection result accuracy of the preset visual perception algorithm based on the preset result accuracy evaluation rule, the matching true value information and detection information corresponding to each point cloud data frame and its corresponding image frame, the detection information that does not match the true value information, and the true value information that does not match the detection information, comprises: determine, based on the matching ground truth information and the detection information corresponding to each point cloud data frame and the image frame corresponding thereto, a precision information and a recall information of the detection result corresponding to the preset visual perception algorithm; determine, based on the matching ground truth information and the detection information corresponding to each point cloud data frame and the image frame corresponding thereto, a detection position error value between the matching ground truth information and the detection information; determine, based on the matching ground truth information and the detection information corresponding to each point cloud data frame and the image frame corresponding thereto, a detection posture error value between the matching ground truth information and the detection information; determine, based on the matching ground truth information and the detection information corresponding to each point cloud data frame and the image frame corresponding thereto, a detection speed error value between the matching ground truth information and the detection information; determine, based on the matching ground truth information and the detection information corresponding to each point cloud data frame and the image frame corresponding thereto, a detection acceleration error value between the matching ground truth information and the detection information; determine, based on the matching ground truth information and the detection information corresponding to each point cloud data frame and the image frame corresponding thereto, a length-width error value of a detection bounding box between the matching ground truth information and the detection information; draw an error curve corresponding to the target error value based on the target error value between the matching ground truth information and the detection information and a preset error threshold value corresponding to the target error value, wherein an abscissa axis of the error curve is the preset error threshold value, and an ordinate axis of the error curve is a ratio of a number of target error values smaller than each preset error threshold value in the target error value to a total amount of data in the evaluation data set; sort the target error value between the matching ground truth information and the detection information according to the size of the numerical value to obtain a sorting sequence corresponding to the target error value; determine a first target error value with the largest numerical value in a first percentage of the target error value in the sorting sequence corresponding to the target error value, and a second target error value with the largest numerical value in a second percentage of the target error value; determine a first evaluation information representing the accuracy of the detection result of the preset visual perception algorithm based on the precision information and the recall information of the detection result corresponding to the preset visual perception algorithm, the error curve corresponding to the target error value, the first target error value and the second target error value, and / or the target error value.

6. The method of claim 5, wherein, The method further comprises: displaying the first evaluation information, the second evaluation information, the precision information and the recall information of the detection result corresponding to the preset visual perception algorithm, the error curve corresponding to the target error value, the first target error value and the second target error value, the target error value, the determined target error value corresponding to the same object, and / or the fitting error curve corresponding to the fitting target error value corresponding to each object.

7. A device for evaluating visual perception algorithms, characterized in that, The device comprises: The first obtaining module is configured to obtain true value information corresponding to each object determined based on each point cloud data frame in the evaluation dataset, wherein the true value information at least includes labeled pose information of the corresponding object and object motion information, and each evaluation data includes point cloud data frames and image frames having a corresponding relationship; The second obtaining module is configured to obtain detection information corresponding to each detected object based on a preset visual perception algorithm and each image frame in the evaluation dataset, wherein the detection information at least includes detection pose information of the corresponding detected object and detection motion information. The determining module is configured to determine evaluation information corresponding to the preset visual perception algorithm based on the preset result accuracy evaluation rule, the preset algorithm stability evaluation rule, the labeled pose information and the object motion information in the true value information corresponding to each object, and the detection pose information and the detection motion information in the detection information corresponding to each detected object, wherein the evaluation information includes first evaluation information representing detection result accuracy of the preset visual perception algorithm and second evaluation information representing algorithm stability of the preset visual perception algorithm; wherein the second evaluation information representing the algorithm stability of the preset visual perception algorithm is determined based on the preset algorithm stability evaluation rule, the true value information corresponding to each object, and the detection information corresponding to each detected object, including: determining target error values corresponding to the same object from target error values between matching true value information and detection information based on the time sequence information between the point cloud data frames or the image frames in the evaluation data set, the matching true value information and detection information being the matching true value information and detection information corresponding to each point cloud data frame and its corresponding image frame; for each target error value corresponding to an object, a fitting error curve corresponding to the target error value of the object is fitted based on the time sequence information between the point cloud data frames or the image frames in the evaluation data set, the target error value of the object, and a preset curve fitting algorithm, wherein the fitting error curve includes the fitting error corresponding to the object at the collection time corresponding to each point cloud data frame or image frame; for each target error value corresponding to an object, the difference between the target error and the fitting error corresponding to the same collection time is calculated based on the target error value of the object and the fitting error corresponding to the collection time corresponding to each point cloud data frame or image frame included in the fitting error curve corresponding to the target error value of the object; for each target error value corresponding to an object, a difference value error curve corresponding to the target error of the object is drawn based on the difference between the target error and the fitting error of each collection time corresponding to the object and the preset difference value threshold corresponding to the target error value of the object, wherein the horizontal axis of the difference value error curve corresponding to the target error of the object is the preset difference value threshold corresponding to the target error value of the object, and the vertical axis of the difference value error curve corresponding to the target error of the object is the ratio of the number of difference values less than each preset difference value threshold to the total number of target error values of the object; for each target error value corresponding to an object, the target error value corresponding to the object is sorted according to the size of the numerical value, and the first difference value with the largest numerical value in the first third of the sorted sequence and the second difference value with the largest numerical value in the first fourth of the sorted sequence are determined.Determine second evaluation information representing algorithm stability of the preset visual perception algorithm based on the first difference value and the second difference value corresponding to the target error value corresponding to each object, and / or the difference error curve corresponding to the target error value corresponding to each object, the target error value being a detection position error value, a detection posture error value, a detection speed error value, a detection acceleration error value, or a length-width error value of a detection frame.

8. The apparatus of claim 7, wherein, The first obtaining module is specifically configured to obtain an evaluation dataset; label each point cloud data frame in the evaluation dataset based on a pre-trained three-dimensional data perception model, and label bounding box information of each object corresponding to each point cloud data frame to determine labeled position information and labeled pose information of each object corresponding to each point cloud data frame, and obtain labeled pose information of each object corresponding to each point cloud data frame; determine labeled speed information and labeled acceleration information of each object corresponding to each point cloud data frame based on the labeled position information and the labeled pose information of each object corresponding to each point cloud data frame in the evaluation dataset and time sequence information between each point cloud data frame in the evaluation dataset, to obtain object motion information of each object corresponding to each point cloud data frame, and obtain true value information corresponding to each object corresponding to each point cloud data frame.

Citation Information

Patent Citations

  • Obstacle sensing method, system, computer equipment and computer storage medium

    CN109163707A

  • A method and a device for acquiring information

    CN109903308A

  • High-speed automatic driving scene obstacle perception evaluation method and device

    CN110287832A

  • Image labeling method, device and system and host

    CN111127422A

  • Labeling result processing method, device and equipment and storage medium

    CN111368927A