Systems and methods for object detection models

Through the collaborative visualization and interactive tools of the object detection model visual analysis platform, the problem of lack of contextual information in model performance evaluation is solved, and the rapid identification of model weaknesses and performance improvement are achieved.

CN113128329BActive Publication Date: 2025-09-30ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202011624636.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-12-31
Filing Date
2020-12-31
Publication Date
2025-09-30
Estimated Expiration
2040-12-31

AI Technical Summary

Technical Problem

In existing technologies, performance evaluation methods for object detection models lack contextual information, making it difficult to track the root causes of errors, resulting in difficulties in improving model performance. In addition, evaluation tools lack interactive analysis and recommendation mechanisms.

Method used

This paper provides a visual analysis platform for object detection models, which generates recommended graphical user interface elements through collaborative visualization and interactive tools to help users identify and update model weaknesses, including data extraction, aggregation, visualization and model updating.

Benefits of technology

Through detailed data analysis and user interaction, it can quickly identify and improve the weaknesses of object detection models, improve model performance, and provide quantitative model evaluation and update methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113128329B_ABST
    Figure CN113128329B_ABST
Patent Text Reader

Abstract

A visual analytics platform for updating object detection models in autonomous driving applications is provided. A visual analytics tool for updating object detection models in autonomous driving applications. In one embodiment, an object detection model analysis system includes a computer and an interface device. The interface device includes a display device. The computer includes an electronic processor configured to: extract object information from image data using a first object detection model; extract characteristics of the object from metadata associated with the image data; generate a summary of the object information and characteristics; generate a collaborative visualization based on the summary and the characteristics; generate a recommended graphical user interface element based on the collaborative visualization and a first one or more user inputs; and update the first object detection model based at least in part on a classification of one or more individual objects as actual weaknesses in the first object detection model to generate a second object detection model for autonomous driving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to a visual analytics platform for object detection models and more specifically to a visual analytics platform for object detection models in autonomous driving applications. Background Art

[0002] Object detection is an important component in autonomous driving. It is also appropriate to evaluate and analyze the performance of object detection models implemented by one or more electronic processors to perform object detection in autonomous driving.

[0003] The mean average precision metric (referred to herein as "mAP") has been used to evaluate the accuracy of object detection models. The mAP metric is typically aggregated over multiple cases with different thresholds for all object categories and the intersection over union (IoU) value (referred to herein as "IoU") over the overlap area between the ground truth and predicted bounding boxes. The mAP metric provides a quantitative measure of model performance by object detection models and allows for comparison between different object detection models. Summary of the Invention

[0004] However, as a single, aggregated metric, mAP does not provide context or details beyond a high-level value of model performance. There are many factors that may affect model performance, such as object class, size, image background, or other suitable factors. Without context or details about the various factors that directly affect model performance, it is difficult to track down the root cause of errors in object detection models. Therefore, improving the model performance of object detection models is difficult because the root cause of errors may be unknown and difficult to track down.

[0005] Among other things, conventional evaluation methods and tools still face the following challenges: 1) mAP aggregate metric that does not consider contextual information of object characteristics, 2) no intuitive mechanism to understand the overall characteristics of the dataset and model performance, 3) no interactive exploration and guidance tools to analyze model performance with customized context, and 4) a lot of manual effort is required to guide and narrow down the root cause. To address these challenges, the present disclosure also includes, among other things, a visual analysis platform with collaborative visualization to perform multi-faceted performance analysis of object detection models in autonomous driving applications.

[0006] For example, in one embodiment, the present disclosure includes an object detection model analysis system. The object detection model analysis system includes a computer and an interface device. The interface device includes a display device. The computer includes: a communication interface configured to communicate with the interface device; a memory including an object detection model visual analysis platform for autonomous driving and a first object detection model; and an electronic processor communicatively connected to the memory. The electronic processor is configured to: extract object information from image data using a first object detection model; extract characteristics of the objects from metadata associated with the image data; generate a summary of the extracted object information and characteristics; generate a collaborative visualization based on the summary and the extracted characteristics; output the collaborative visualization for display on a display device; receive a first or more user inputs selecting a portion of information to be included in the collaborative visualization; generate a recommended graphical user interface element based on the collaborative visualization and the first or more user inputs, the recommended graphical user interface element summarizing one or more potential weaknesses of the first object detection model; output the recommended graphical user interface element for display on the display device; receive a second user input, the second user input selecting one or more individual objects from the object information extracted and included in the one or more potential weaknesses; output an image based on the image data and the second user input, the image highlighting the one or more individual objects; receive a third user input, the third user input classifying the one or more individual objects as actual weaknesses with respect to the first object detection model; and update the first object detection model based at least in part on the classification of the one or more individual objects as actual weaknesses with respect to the first object detection model to generate a second object detection model for autonomous driving.

[0007] Additionally, in another embodiment, the present disclosure includes a method for updating an object detection model for autonomous driving using an object detection model visual analytics platform. The method includes extracting object information from image data using an electronic processor and a first object detection model for autonomous driving. The method includes extracting, using the electronic processor, object characteristics from metadata associated with the image data. The method includes generating, using the electronic processor, a summary of the extracted object information and characteristics. The method includes generating, using the electronic processor, a collaborative visualization based on the summary and the extracted characteristics. The method includes outputting, using the electronic processor, the collaborative visualization for display on a display device. The method includes receiving, using the electronic processor, a first or more user inputs from an interface device via a communication interface, the first or more user inputs selecting a portion of information to be included in the collaborative visualization. The method includes generating, using the electronic processor, a recommended graphical user interface element based on the collaborative visualization and the first or more user inputs, the recommended graphical user interface element summarizing one or more potential weaknesses of the first object detection model. The method includes outputting, using the electronic processor, the recommended graphical user interface element for display on the display device. The method includes receiving, using the electronic processor, a second user input selecting one or more individual objects from the extracted object information included in the one or more potential weaknesses. The method includes outputting, using an electronic processor, an image for display on a display device based on image data and a second user input, the image highlighting the one or more individual objects. The method includes receiving, using the electronic processor, a third user input that classifies the one or more individual objects as actual weaknesses with respect to a first object detection model. The method also includes updating, using the electronic processor, the first object detection model based at least in part on the classification of the one or more individual objects as actual weaknesses with respect to the first object detection model to generate a second object detection model for autonomous driving.

[0008] Additionally, in yet another embodiment, the present disclosure includes a non-transitory computer-readable medium containing instructions that, when executed by an electronic processor, cause the electronic processor to perform a set of operations. The set of operations includes extracting object information from image data using a first object detection model for autonomous driving. The set of operations includes extracting characteristics of the object from metadata associated with the image data. The set of operations includes generating a summary of the extracted object information and characteristics. The set of operations includes generating a collaborative visualization based on the summary and the extracted characteristics. The set of operations includes outputting the collaborative visualization for display on a display device. The set of operations includes receiving first or more user inputs from an interface device via a communication interface, the first or more user inputs selecting a portion of information to be included in the collaborative visualization. The set of operations includes generating a recommended graphical user interface element based on the collaborative visualization and the first or more user inputs, the recommended graphical user interface element summarizing one or more potential weaknesses of the first object detection model. The set of operations includes outputting the recommended graphical user interface element for display on the display device. The set of operations includes receiving second user input, the second user input selecting one or more individual objects from the object information extracted and included in the one or more potential weaknesses. The set of operations includes outputting an image for display on a display device based on the image data and the second user input, the image highlighting the one or more individual objects. The set of operations includes receiving a third user input that classifies the one or more individual objects as actual weaknesses with respect to the first object detection model. The set of operations includes updating the first object detection model based at least in part on the classification of the one or more individual objects as actual weaknesses with respect to the first object detection model to generate a second object detection model for autonomous driving.

[0009] Before explaining any embodiment in detail, it should be understood that the embodiments are not limited in their application to the details of the configuration and arrangement of the components set forth in the following description or illustrated in the accompanying drawings. These embodiments can be practiced or implemented in various ways. In addition, it should be understood that the embodiments may include hardware, software, and electronic components or modules, which, for the purposes of discussion, may be illustrated and described as if most components are implemented only in hardware. However, those skilled in the art and based on a reading of this detailed description will recognize that, in at least one embodiment, the electronic-based aspects may be implemented in software (e.g., stored on a non-transitory computer-readable medium) that can be executed by one or more processing units, such as microprocessors and / or application-specific integrated circuits ("ASICs"). As such, it should be noted that the embodiments may be implemented using multiple hardware- and software-based devices and multiple different structural components. For example, the "computer" and "interface device" described in the specification may include one or more processing units, one or more computer-readable media modules, one or more input / output interfaces, and various connections (e.g., a system bus) that connect the components together. It should also be understood that although some embodiments depict components as logically separated, such depictions are for illustrative purposes only. In some embodiments, the components shown may be combined or divided into separate software, firmware, and / or hardware components. Regardless of how they are combined or divided, these components may be executed or located on the same computing device, or may be distributed among different computing devices connected via one or more networks or other suitable connections.

[0010] Other aspects of the various embodiments will become apparent by consideration of the detailed description and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 is a block diagram illustrating an object detection model analysis system according to various aspects of the present disclosure.

[0012] Figure 2 It is a diagram showing various aspects of the present disclosure. Figure 1 Diagram of the workflow of the various components in the visual analytics platform for object detection models.

[0013] Figure 3 is a diagram illustrating a size distribution visualization with a first user input selection, an IoU distribution visualization with a second user input selection, and an area under the curve (AUC) score visualization according to various aspects of the present disclosure.

[0014] Figure 4 is a block diagram illustrating an exemplary embodiment of an analysis recommendation graphical user interface (GUI) element regarding an object detection model according to various aspects of the present disclosure.

[0015] Figure 5 and 6 is a diagram illustrating two different analyses of an object detection model according to various aspects of the present disclosure.

[0016] Figure 7A and 7B is a diagram illustrating various aspects of the present disclosure. Figure 1 Flowchart of an exemplary method 700 performed by the object detection model visual analysis system 10.

[0017] Figure 8 is a diagram illustrating a score distribution visualization according to various aspects of the present disclosure.

[0018] Figure 9 is a diagram illustrating a class distribution visualization according to various aspects of the present disclosure. DETAILED DESCRIPTION

[0019] Before explaining any embodiments of the present disclosure in detail, it is to be understood that the present disclosure is not limited in its application to the details of construction and the arrangement of components set forth in the following description or illustrated in the following drawings. The present disclosure is capable of other embodiments and of being practiced or carried out in various ways.

[0020] Figure 1 is a block diagram illustrating an object detection model visual analysis system 10. Figure 1 In the example of , the object detection visual analysis system 10 includes a computer 100 and an interface device 120. The interface device 120 includes a display device and can be a personal desktop computer, laptop computer, tablet computer, personal digital assistant, mobile phone, or other suitable computing device.

[0021] However, it should be understood that in some embodiments, there is Figure 1 For example, the functionality described herein with respect to computer 100 can be extended to a different number of computers to provide distributed processing. Additionally, the functionality described herein with respect to computer 100 can be implemented solely by interface device 120, such that Figure 1 The "cloud-type" systems described in also apply to "personal computing" systems.

[0022] The computer 100 includes an electronic processor 102 (e.g., a microprocessor or other suitable processing device), a memory 104 (e.g., a non-transitory computer-readable storage medium), and a communication interface 114. It should be understood that in some embodiments, the computer 100 may include a computer having a configuration different from that of the computer 102. Figure 1Furthermore, the computer 100 may perform additional functionality beyond that described herein. Furthermore, the functionality of the computer 100 may be incorporated into other computers. Figure 1 As illustrated in , the electronic processor 102, memory 104, and communication interface 114 are electrically coupled by one or more control or data buses, enabling communications between the components.

[0023] In one example, electronic processor 102 executes machine-readable instructions stored in memory 104. For example, electronic processor 102 may execute instructions stored in memory 104 to perform the functionality described herein.

[0024] The memory 104 may include a program storage area (e.g., read-only memory (ROM)) and a data storage area (e.g., random access memory (RAM) and other non-transitory machine-readable media). In some examples, the program storage area may store instructions for the object detection model visual analysis platform 106 (also referred to as the "object detection model visual analysis tool 106") and the object detection model 108. Additionally, in some examples, the data storage area may store image data 110 and metadata 112 associated with the image data 110.

[0025] The object detection model visual analysis tool 106 has machine-readable instructions that cause the electronic processor 102 to process (e.g., retrieve) image data 110 and associated metadata 112 from the memory 104 using the object detection model 108 and generate different visualizations based on the image data 110 and the associated metadata 112. In particular, the electronic processor 102 generates different visualizations based on the image data 110 and the associated metadata 112 when processing the image data 110 and the metadata 112 associated with the object detection model 108 (e.g., as shown below). Figure 2 ), uses machine learning to generate and output developer insights, data visualizations, and enhancement recommendations for the object detection model 108.

[0026] Machine learning generally refers to the ability of a computer program to learn without being explicitly programmed. In some embodiments, a computer program (e.g., a learning engine) is configured to build an algorithm based on input. Supervised learning involves presenting a computer program with example inputs and their desired outputs. The computer program is configured to learn general rules that map inputs to outputs based on the training data it receives. Example machine learning engines include decision tree learning, association rule learning, artificial neural networks, classifiers, inductive logic programming, support vector machines, clustering, Bayesian networks, reinforcement learning, representation learning, similarity and metric learning, sparse dictionary learning, and genetic algorithms. By using one or more of the above methods, a computer program can ingest, parse, and understand data, and gradually improve the algorithms used for data analysis.

[0027] Communication interface 114 receives data from and provides data to devices external to computer 100, such as interface device 120. For example, communication interface 114 may include a port or connection for receiving a wired connection (e.g., an Ethernet cable, a fiber optic cable, a telephone cable, etc.), a wireless transceiver, or a combination thereof. In some examples, communication interface 114 may communicate with interface device 120 via the Internet.

[0028] In some examples, computer 100 includes one or more graphical user interfaces (as described in more detail below and in Figure 2 -6 ) and one or more user interfaces (not shown). The one or more graphical user interfaces (e.g., one or more web pages) include graphical elements that allow a user of the interface device 120 to interface with the computer 100. The one or more graphical user interfaces may include, or be a portion of, a display screen that displays developer insights, data visualizations, and enhancement recommendations and outputs derived by the electronic processor 102 from executing the object detection model visual analysis tool 106.

[0029] In some examples, the computer 100 includes one or more user interfaces (not shown). The one or more user interfaces include one or more input mechanisms (e.g., a touch screen, a keypad, buttons, knobs, etc.), one or more output mechanisms (e.g., a display, a printer, a speaker, etc.), or a combination thereof. The one or more optional user interfaces receive input from a user, provide output to a user, or a combination thereof. In some embodiments, as an alternative to or in addition to managing input and output through one or more optional user interfaces, the computer 100 can receive user input, provide user output, or both by communicating with an external device (e.g., an interface device 120 or a workstation (not shown)) over a wired or wireless connection.

[0030] Figure 2 It's a picture Figure 1FIG2 is a diagram of a workflow 200 of various components 202-210 in the object detection model visual analysis tool 106. Figure 2 In the example of the object detection model visual analysis tool 106, the object detection model visual analysis tool 106 includes a data extraction component 202, a data aggregation component 204, a collaborative visualization component 206, and an evaluation component 207. Figure 1 The object detection model 108 analysis recommendation component 208, and update Figure 1 The model update component 210 of the object detection model 108 is configured.

[0031] exist Figure 2 In the example, the object detection model visualization analysis tool 106 uses Figure 1 The object detection model 108 is used to detect traffic lights in the image data 110. Figure 2 As illustrated in FIG, a data extraction component 202 extracts object characteristics and model performance metrics from a dataset (e.g., image data 110 and associated metadata 112) and one or more object detection models (e.g., object detection model 108). A data aggregation component 204 aggregates the characteristics of these datasets. A collaborative visualization component 206 presents the data aggregation using a collaborative interactive visualization set that enables interactive analysis of the relationships between different metrics. Finally, an analysis recommendation component 208 recommends important error patterns based on user interests. The workflow 200 begins with a data source that includes an object image 110 (e.g., a driving scene with a traffic light) and associated metadata 112 (e.g., a label including classification information, a bounding box, or other suitable metadata). In summary, the object detection model visualization analysis tool 106 has four components: a data extraction component 202, a data (metric and metadata) aggregation component 204, a collaborative visualization component 206, and an analysis recommendation component 208.

[0032] As illustrated with respect to the data extraction component 202, there are two types of data that are extracted to understand the impact of object characteristics on model performance. The first type of data extracted is the characteristics of the object. The characteristics of the object are descriptive information including object type (e.g., green, yellow, red light, arrow, other suitable object types), object size, aspect ratio (which can be derived from bounding box information), and occlusion. The characteristics of the object can be extracted from the associated metadata 112. In addition, complex visual features in the image data 110 can also be extracted using a feature extractor, such as a feature map from a deep convolutional neural network, which is referred to as a "visual feature extractor."

[0033] The second type of data extracted is the detection results. These can include detection bounding boxes, detection classes, and detection scores. Using image data 110 as input, the detection results can be derived from the inference results from the object detection model 108. Using the ground truth and the detection bounding boxes, the object detection model 108 can also calculate the IoU value.

[0034] As illustrated with respect to the data summary component 204, the data summary component 204 summarizes the data distribution of both object characteristics and detection metrics (e.g., IoU and confidence scores.) The data summary component 204 also calculates an area under the curve (AUC) score based on user selections utilizing interactions with the collaborative visualization, as described in more detail below.

[0035] As illustrated with respect to the collaborative visualization component 206, the collaborative visualization component 206 is a collection of visualizations generated from the data aggregation. For example, Figure 3 is a diagram illustrating a size distribution visualization 302 with a first user input selection 304, and an IoU distribution visualization 306 and an area under the curve (AUC) score visualization 310 with a second user input selection 308. Figure 3 As illustrated in , the first user input selection 304 selects a size range between 3.5 and 4.5, and the second user input selection 308 selects an IoU range between approximately 0.6 and 0.7, and then an AUC score visualization 310 for large-sized objects is calculated by the data aggregation component 204 and displayed by the collaborative visualization component 206.

[0036] By calculating the AUC score after selecting a size range and an IoU range, the data aggregation component 204 and the collaborative visualization component 206 provide the user with contextual AUC. Conventionally, contextual AUC is not feasible because the mAP metric is aggregated by averaging AUC over many IoU ranges.

[0037] In addition, the collaborative visualization set can include the distribution of object classes and sizes, as well as detection metrics of IoU and scores, to allow users to interactively select objects of interest and understand the relationship between the metrics of interest. Specifically, the impact of different object characteristics on the IoU score, or how the IoU score is related to the confidence score. A scatter plot visualization of all objects is provided to reveal the overall pattern between object characteristics and model performance metrics. Each object is assigned a position in the two-dimensional scatter plot based on the dimensionality reduction result of its object characteristics and multivariate values ​​of the detection results. Dimensionality reduction techniques such as PCA, MDS, t-SNET, or other suitable dimensionality reduction techniques can be used here.

[0038] As illustrated with respect to the analysis recommendation component 208, when executed by the electronic processor 102, the analysis recommendation component 208 generates a graphical user interface element that identifies at least one of a strength or weakness of the object detection model 108 that a user can investigate with minimal manual exploration. Additionally, as Figure 2 As illustrated in , the analytics recommendation component 208 includes three subcomponents: 1) metric profiles, 2) clustering, and 3) analytics recommender.

[0039] Regarding the metric profile, the analysis recommendation component 208 has predefined metric profiles, including false alarms (zero-sized objects with no ground truth bounding box but with medium to high confidence scores), mislabeled data (zero-sized objects with very high confidence scores), and other suitable predefined metric profiles. In some embodiments, users of the object detection model visualization analysis tool 106 can also define their own metric profiles of interest, such as large size with small IoU, small size with large IoU, or other suitable metric profiles of interest.

[0040] Regarding clustering, the analysis recommendation component 208 includes metric pattern discovery to help users identify useful performance metric patterns across all objects by clustering object characteristics and model metrics. Finally, regarding the analysis recommender, the analysis recommendation component 208 can summarize features in clusters and guide users to clusters of interest. To summarize clusters, the analysis recommendation component 208 can use feature selection methods to highlight the highest priority features to the user of the object detection model visual analysis tool 106.

[0041] For example, Figure 4 is a block diagram illustrating an exemplary embodiment of an analysis recommendation graphical user interface (GUI) element 400 for an object detection model 108 that identifies three weaknesses: a plurality of potential false alarms portion 402, a plurality of potential missing labels portion 404, and a missing detection group portion 406. The missing detection group portion 406 further includes a first sub-portion dedicated to small size and dark backgrounds for objects with low IOU, and a second sub-portion dedicated to yellow light and high brightness. In this way, portions 402-406 of the analysis recommendation GUI element 400 allow a user to quickly locate these objects in a two-dimensional scatter plot visualization and allow further analysis without the user manually exploring each individual detected object.

[0042] For example, Figure 5 and 6is a diagram illustrating two different analyses 500 and 600 of the object detection model 108. Analysis 500 is connected to the potential false alarm portion 402 of the analysis recommendation GUI element 400 and allows the user to drill down into a segment of the two-dimensional scatter plot 502 to locate one type of model error in the object detection model 108, specifically false positives (e.g., images 504 and 506). Analysis 600 is connected to the potential missing labels portion 404 of the analysis recommendation GUI element 400 and allows the user to drill down into a segment of the two-dimensional scatter plot 602 to locate another type of model error in the object detection model 108, specifically data with missing labels (e.g., images 604 and 606). In short, the analysis recommendation GUI element 400 provides the user with a means of analyzing segments of the two-dimensional scatter plot visualization (i.e., the two-dimensional scatter plots 502 and 602, respectively) to address deficiencies in the object detection model 108.

[0043] With respect to the model update component 210 , after the user analyzes a segment of the two-dimensional scatter plot to address deficiencies in the object detection model 108 , the user may identify whether deficiencies exist in the object detection model 108 and how the deficiencies should be classified.

[0044] After classifying some or all of the defects highlighted by the analysis recommendation GUI element 400, the model update component 210, when executed by the electronic processor 102, updates the object detection model 108 to generate a second object detection model that does not include the defects identified in the object detection model 108. In short, the model update component 210 provides a means for the user to update the object detection model 108 to generate a new object detection model after some or all of the defects in the object detection model 108 have been resolved by the user using the analysis recommendation GUI element 400. Figure 4-6 As shown in the figure.

[0045] Figure 7A and 7B It is a diagram of Figure 1 Flowchart of an exemplary method 700 performed by the object detection model visual analysis system 10. Figure 1 –6 describes Figure 7.

[0046] 7 , method 700 includes extracting object information from image data using a first object detection model for autonomous driving by the electronic processor 102 of the computer 100 (at block 702). For example, the electronic processor 102 of the computer 100 extracts object information from the image data 110 using the object detection model 108 for autonomous driving.

[0047] The method 700 includes the electronic processor 102 extracting characteristics of the object from metadata associated with the image data (at block 704). For example, the electronic processor 102 extracts characteristics of the object from the associated metadata 112 associated with the image data 110 as Figure 2 Part of the data extraction component 202.

[0048] The method 700 includes the electronic processor 102 generating a summary of the extracted object information and characteristics (at block 706). For example, the electronic processor 102 generates a data distribution, an area under the curve (AUC) value, and a mAP value as Figure 2 Part of the data aggregation component 204.

[0049] The method 700 includes the electronic processor 102 generating a collaborative visualization based on the summary and the extracted characteristics (at block 708). The method 700 includes the electronic processor 102 outputting the collaborative visualization for display on a display device (at block 710). For example, the electronic processor 102 generates and outputs the collaborative visualization as Figure 2 Part of the collaborative visualization component 206.

[0050] The method 700 includes the electronic processor 102 receiving, via the communication interface, a first one or more user inputs from the interface device, the first one or more user inputs selecting a portion of information to be included in the collaborative visualization (at block 712). For example, the electronic processor 102 receives, via the communication interface 114, the first one or more user inputs from the interface device 120, and the first one or more user inputs include Figure 3 The first user input 304 and the second user input 308 are shown.

[0051] The method 700 includes the electronic processor 102 generating a recommended graphical user interface element based on the collaborative visualization and the first one or more user inputs, the recommended graphical user interface element summarizing one or more potential weaknesses of the first object detection model (at block 714). The method 700 includes the electronic processor 102 outputting the recommended graphical user interface element for display on a display device (at block 716). For example, the electronic processor 102 generates and outputs Figure 4 Recommended graphical user interface elements 400.

[0052] The method 700 includes the electronic processor 102 receiving a second user input that selects one or more individual objects from the object information extracted and included in the one or more potential weaknesses (at block 718). For example, the second user input is a selection of Figure 4 The selection of one of the potential false alarm portion 402, the potential missing labels portion 404, or the missing detection group portion 406, and the two-dimensional scatter plot (ie, Figure 5 and6 Selection of one or more individual objects in the two-dimensional scatter plot 502 or 602).

[0053] Method 700 includes the electronic processor 102 outputting an image for display on a display device based on the image data and the second user input, the image highlighting one or more individual objects (at block 720). For example, the electronic processor 102 outputs images 504 and 506 that highlight the individual objects selected in the two-dimensional scatter plot 502, such as Figure 5 Alternatively, for example, electronic processor 102 outputs images 604 and 606 that highlight individual objects selected in two-dimensional scatter plot 602, such as Figure 6 As shown in the figure.

[0054] The method 700 includes the electronic processor 102 receiving a third user input that classifies one or more individual objects as actual weaknesses with respect to the first object detection model at block 722. For example, the electronic processor 102 receives the third user input that classifies one or more individual objects in the two-dimensional scatter plot 502 as false positives in the object detection model 108. Alternatively, for example, the electronic processor 102 receives the third user input that classifies one or more individual objects in the two-dimensional scatter plot 602 as missing data labels in the object detection model 108.

[0055] Method 700 also includes the electronic processor 102 updating the first object detection model based at least in part on the classification of the one or more individual objects as actual weaknesses with respect to the first object detection model to generate a second object detection model for autonomous driving (at block 724). For example, the electronic processor 102 updates the object detection model 108 based at least in part on the false positives or missing data labels identified in the third user input to generate the second object detection model for autonomous driving, the second object detection model excluding the weaknesses identified in the object detection model 108 and increasing the object detection performance of the second object detection model over that of the object detection model 108.

[0056] In some embodiments, method 700 may further include extracting visual features based on the image data, and generating a collaborative visualization based on the aggregated and extracted visual features. In these embodiments, the visual features may include feature type, feature size, feature texture, and feature shape.

[0057] In some embodiments, the collaborative visualization may include a size distribution visualization, an IoU distribution visualization, an area under the curve (AUC) visualization, a score distribution, and a class distribution. In these embodiments, when the first one or more user inputs include a size range input for the size distribution visualization and an IoU value range input for the IoU distribution visualization, the AUC visualization may be based on the size range input and the IoU value range input.

[0058] In some embodiments, an individual object of the one or more individual objects is an example weakness of the one or more potential weaknesses of the first object detection model.In some embodiments, the object information may include a detection result, a score, and an intersection-over-union (IoU) value.

[0059] Figure 8 is a diagram illustrating a score distribution visualization 800. The score distribution visualization 800 is part of the collaborative visualization described above. Figure 8 In the example of , score distribution visualization 800 illustrates the distribution of mAP scores from 0.0 to 1.0 for all objects detected in image data 110. A user can select a range of mAP scores to further analyze detected objects that fall within the range of mAP scores with respect to object detection model 108.

[0060] Figure 9 is a diagram illustrating a class distribution visualization 900. The class distribution visualization 900 is also part of the collaborative visualization described above. Figure 9 In the example of FIG, class distribution visualization 900 illustrates the distribution of classifiers for all objects detected in image data 110. In particular, class distribution visualization 900 includes a no detection classifier, a yellow classifier, a red classifier, a green classifier, and an off classifier. A user can select one or more classifiers to further analyze the detected objects classified within the one or more classifiers.

[0061] The following examples illustrate example systems, methods, and non-transitory computer-readable media described herein.

[0062] Example 1: An object detection model visual analysis system, comprising: an interface device, including a display device; and a computer, including a communication interface configured to communicate with the interface device; a memory, including an object detection model visual analysis platform for autonomous driving and a first object detection model; and an electronic processor communicatively connected to the memory, the electronic processor being configured to: extract object information from image data using the first object detection model; extract characteristics of the object from metadata associated with the image data; generate a summary of the extracted object information and characteristics; generate a collaborative visualization based on the summary and the extracted characteristics; output the collaborative visualization for display on the display device; receive a first one or more user inputs selecting a portion of information to be included in the collaborative visualization; generate a collaborative visualization based on the collaborative visualization and the first one or more user inputs Recommending a graphical user interface element that summarizes one or more potential weaknesses of a first object detection model; outputting the recommended graphical user interface element for display on a display device; receiving a second user input that selects one or more individual objects from the object information extracted and included in the one or more potential weaknesses; outputting an image based on the image data and the second user input, the image highlighting the one or more individual objects; receiving a third user input that classifies the one or more individual objects as actual weaknesses with respect to the first object detection model; and updating the first object detection model based at least in part on the classification of the one or more individual objects as actual weaknesses with respect to the first object detection model to generate a second object detection model for autonomous driving.

[0063] Example 2: The object detection model visual analysis system of Example 1, wherein the electronic processor is further configured to extract visual features based on the image data and generate a collaborative visualization based on the aggregated and extracted visual features.

[0064] Example 3: The object detection model visual analysis system of Example 2, wherein the visual features include feature type, feature size, feature texture, and feature shape.

[0065] Example 4: The object detection model visual analysis system of any of Examples 1-3, wherein the collaborative visualization includes size distribution visualization, IoU distribution visualization, area under the curve (AUC) visualization, score distribution, and class distribution.

[0066] Example 5: The object detection model visual analysis system of Example 4, wherein the first one or more user inputs include a size range input for size distribution visualization and an IoU value range input for IoU distribution visualization, and wherein the AUC visualization is based on the size range input and the IoU value range input.

[0067] Example 6: The object detection model visual analysis system of any of Examples 1-5, wherein an individual object of the one or more individual objects is an example of one or more potential weaknesses of the first object detection model.

[0068] Example 7: The object detection model visual analysis system of any of Examples 1-6, wherein the object information includes detection results, scores, and intersection-over-union (IoU) values.

[0069] Example 8: A method for updating an object detection model for autonomous driving using an object detection model visual analysis platform, the method comprising: extracting object information from image data using an electronic processor and a first object detection model for autonomous driving; extracting characteristics of the object from metadata associated with the image data using the electronic processor; generating a summary of the extracted object information and characteristics using the electronic processor; generating a collaborative visualization based on the summary and the extracted characteristics using the electronic processor; outputting the collaborative visualization for display on a display device using the electronic processor; receiving first one or more user inputs from an interface device via a communication interface using the electronic processor, the first one or more user inputs selecting a portion of information to be included in the collaborative visualization; generating a recommended graphical user interface element using the electronic processor based on the collaborative visualization and the first one or more user inputs, the recommended graphical user interface The invention also provides a method for summarizing one or more potential weaknesses of a first object detection model; outputting, using an electronic processor, a recommended graphical user interface element for display on a display device; receiving, using the electronic processor, a second user input that selects one or more individual objects from the object information extracted and included in the one or more potential weaknesses; outputting, using the electronic processor, an image based on the image data and the second user input for display on the display device, the image highlighting the one or more individual objects; receiving, using the electronic processor, a third user input that classifies the one or more individual objects as actual weaknesses with respect to the first object detection model; and updating, using the electronic processor, the first object detection model based at least in part on the classification of the one or more individual objects as actual weaknesses with respect to the first object detection model to generate a second object detection model for autonomous driving.

[0070] Example 9: The method of Example 8, further comprising: extracting visual features based on the image data; and generating a collaborative visualization based on the aggregated and extracted visual features.

[0071] Example 10: The method of Example 9, wherein the visual features include feature type, feature size, feature texture, and feature shape.

[0072] Example 11: The method of any of Examples 9 and 10, wherein the collaborative visualization includes size distribution visualization, IoU distribution visualization, area under the curve (AUC) visualization, score distribution, and class distribution.

[0073] Example 12: The method of Example 11, wherein the first one or more user inputs include a size range input for the size distribution visualization and an IoU value range input for the IoU distribution visualization, and wherein the AUC visualization is based on the size range input and the IoU value range input.

[0074] Example 13: The method of any of Examples 8-12, wherein an individual object of the one or more individual objects is an example of one or more potential weaknesses of the first object detection model.

[0075] Example 14: The method of any of Examples 8-13, wherein the object information includes detection results, scores, and intersection-over-union (IoU) values.

[0076] Example 15: A non-transitory computer-readable medium containing instructions that, when executed by an electronic processor, cause the electronic processor to perform a set of operations, the set of operations comprising: extracting object information from image data using a first object detection model for autonomous driving; extracting characteristics of the object from metadata associated with the image data; generating a summary of the extracted object information and characteristics; generating a collaborative visualization based on the summary and the extracted characteristics; outputting the collaborative visualization for display on a display device; receiving first one or more user inputs from an interface device via a communication interface, the first one or more user inputs selecting a portion of information to be included in the collaborative visualization; generating a recommended graphical user interface element based on the collaborative visualization and the first one or more user inputs, the recommended graphical user interface element An element summarizes one or more potential weaknesses of a first object detection model; outputs a recommended graphical user interface element for display on a display device; receives a second user input, the second user input selecting one or more individual objects from the object information extracted and included in the one or more potential weaknesses; outputs an image for display on the display device based on the image data and the second user input, the image highlighting the one or more individual objects; receives a third user input, the third user input classifying the one or more individual objects as actual weaknesses with respect to the first object detection model; and updates the first object detection model based at least in part on the classification of the one or more individual objects as actual weaknesses with respect to the first object detection model to generate a second object detection model for autonomous driving.

[0077] Example 16: The non-transitory computer-readable medium of Example 15, wherein the set of operations further comprises extracting visual features based on the image data; and generating a collaborative visualization based on the aggregated and extracted visual features.

[0078] Example 17: The non-transitory computer-readable medium of Example 15 or 16, wherein the collaborative visualization comprises a size distribution visualization, an IoU distribution visualization, an area under the curve (AUC) visualization, a score distribution, and a class distribution.

[0079] Example 18: The non-transitory computer-readable medium of Example 17, wherein the first one or more user inputs include a size range input for the size distribution visualization and an IoU value range input for the IoU distribution visualization, and wherein the AUC visualization is based on the size range input and the IoU value range input.

[0080] Example 19: The non-transitory computer-readable medium of any of Examples 15-18, wherein an individual object of the one or more individual objects is an example of one or more potential weaknesses of the first object detection model.

[0081] Example 20: The non-transitory computer-readable medium of any of Examples 15-19, wherein the object information includes detection results, scores, and intersection-over-union (IoU) values.

[0082] Thus, the present disclosure provides, among other things, a visual analytics platform for updating object detection models in autonomous driving applications. Various features and advantages of the present invention are set forth in the appended claims.

Claims

1. A visual analysis system for an object detection model, comprising: An electronic processor configured to: extracting object information from the image data using a first object detection model, wherein the object information includes detection results, scores, and intersection-over-union values; extracting characteristics of the object from metadata associated with the image data, wherein the characteristics of the object are descriptive information including object type, object size, aspect ratio, and occlusion; generating a summary of the extracted object information and characteristics; Generate collaborative visualizations based on the summaries and extracted features; outputting the collaborative visualization for display on a display device; receiving user input selecting a portion of the information to be included in the collaborative visualization; generating a recommended graphical user interface element based on the collaborative visualization and the first one or more user inputs, the recommended graphical user interface element summarizing one or more potential weaknesses of the first object detection model, and wherein the one or more potential weaknesses include one or more weaknesses from the group consisting of: a potential false alarm weakness, a potential missing label weakness, and a missing detection group weakness; outputting the recommended graphical user interface element for display on a display device; receiving a second user input selecting one or more individual objects from the object information extracted and included in the one or more potential vulnerabilities; outputting an image based on the image data and the second user input, the image highlighting the one or more individual objects; receiving a third user input that classifies the one or more individual objects as actual weaknesses with respect to the first object detection model; as well as The first object detection model is updated based on the classification of the one or more individual objects as actual weaknesses with respect to the first object detection model to generate a second object detection model for autonomous driving.

2. The object detection model visual analysis system of claim 1 , wherein the electronic processor is further configured to extract visual features based on the image data, and Generate collaborative visualizations based on the aggregated and extracted visual features.

3. The object detection model visual analysis system according to claim 2, wherein the visual features include feature type, feature size, feature texture and feature shape.

4. The object detection model visual analysis system according to claim 1, wherein collaborative visualization includes size distribution visualization, intersection-over-union distribution visualization, area under the curve visualization, score distribution and class distribution.

5. The object detection model visual analysis system of claim 4 , wherein the first one or more user inputs include a size range input for visualizing the size distribution and an IoU value range input for visualizing the IoU distribution, and wherein the area under the curve visualization is based on the size range input and the IoU value range input.

6. The object detection model visual analysis system of claim 1, wherein an individual object of the one or more individual objects is an example of one or more potential weaknesses of the first object detection model.

7. A method for updating an object detection model for autonomous driving using an object detection model visual analysis platform, the method comprising: extracting object information from the image data using an electronic processor and a first object detection model for autonomous driving, wherein the object information includes detection results, scores, and intersection-over-union values; extracting, using an electronic processor, characteristics of the object from metadata associated with the image data, wherein the characteristics of the object are descriptive information including object type, object size, aspect ratio, and occlusion; generating, using an electronic processor, a summary of the extracted object information and characteristics; generating collaborative visualizations based on the aggregated and extracted features using electronic processors; outputting the collaborative visualization for display on a display device using an electronic processor; receiving, with the electronic processor, from the interface device via the communication interface, a first one or more user inputs, the first one or more user inputs selecting a portion of information to be included in the collaborative visualization; generating, with an electronic processor, a recommended graphical user interface element based on the collaborative visualization and the first one or more user inputs, the recommended graphical user interface element summarizing one or more potential weaknesses of the first object detection model, and wherein the one or more potential weaknesses include one or more weaknesses from the group consisting of: a potential false alarm weakness, a potential missing label weakness, and a missing detection group weakness; outputting, using an electronic processor, the recommended graphical user interface element for display on a display device; receiving, with the electronic processor, a second user input selecting one or more individual objects from the object information extracted and included in the one or more potential weaknesses; outputting, with an electronic processor, an image for display on a display device based on the image data and the second user input, the image highlighting the one or more individual objects; receiving, with the electronic processor, a third user input classifying the one or more individual objects as actual weaknesses with respect to the first object detection model; as well as The first object detection model is updated, using an electronic processor, based on the classification of the one or more individual objects as actual weaknesses with respect to the first object detection model to generate a second object detection model for autonomous driving.

8. The method according to claim 7, further comprising: Extract visual features based on image data; as well as Generate collaborative visualizations based on the aggregated and extracted visual features.

9. The method of claim 8, wherein the visual features include feature type, feature size, feature texture, and feature shape.

10. The method of claim 7, wherein the collaborative visualization comprises size distribution visualization, intersection-over-union distribution visualization, area under the curve visualization, score distribution, and class distribution.

11. The method of claim 10, wherein the first one or more user inputs include a size range input for the size distribution visualization and an IoU value range input for the IoU distribution visualization, and wherein the area under the curve visualization is based on the size range input and the IoU value range input.

12. The method of claim 7, wherein an individual object of the one or more individual objects is an example of one or more potential weaknesses of the first object detection model.

13. A non-transitory computer-readable medium containing instructions that, when executed by an electronic processor, cause the electronic processor to perform a set of operations comprising: extracting object information from the image data using a first object detection model for autonomous driving, wherein the object information includes detection results, scores, and intersection-over-union values; extracting characteristics of the object from metadata associated with the image data, wherein the characteristics of the object are descriptive information including object type, object size, aspect ratio, and occlusion; generating a summary of the extracted object information and characteristics; Generate collaborative visualizations based on the summaries and extracted features; outputting the collaborative visualization for display on a display device; receiving, via the communication interface, from the interface device, a first one or more user inputs, the first one or more user inputs selecting a portion of the information to be included in the collaborative visualization; generating a recommended graphical user interface element based on the collaborative visualization and the first one or more user inputs, the recommended graphical user interface element summarizing one or more potential weaknesses of the first object detection model, and wherein the one or more potential weaknesses include one or more weaknesses from the group consisting of: a potential false alarm weakness, a potential missing label weakness, and a missing detection group weakness; outputting the recommended graphical user interface element for display on a display device; receiving a second user input selecting one or more individual objects from the object information extracted and included in the one or more potential vulnerabilities; outputting an image for display on a display device based on the image data and the second user input, the image highlighting the one or more individual objects; receiving a third user input that classifies the one or more individual objects as actual weaknesses with respect to the first object detection model; as well as The first object detection model is updated based on the classification of the one or more individual objects as actual weaknesses with respect to the first object detection model to generate a second object detection model for autonomous driving.

14. The non-transitory computer-readable medium of claim 13, wherein the set of operations further comprises: Extract visual features based on image data; as well as Generate collaborative visualizations based on the aggregated and extracted visual features. 15 . The non-transitory computer-readable medium of claim 13 , wherein the collaborative visualization comprises a size distribution visualization, an intersection-over-union distribution visualization, an area under the curve visualization, a score distribution, and a class distribution.

16. The non-transitory computer-readable medium of claim 15, wherein the first one or more user inputs include a size range input for the size distribution visualization and an IoU value range input for the IoU distribution visualization, and wherein the area under the curve visualization is based on the size range input and the IoU value range input.

17. The non-transitory computer-readable medium of claim 13, wherein an individual object of the one or more individual objects is an example of one or more potential weaknesses of the first object detection model.