Automatic Identification of Training Data Candidates for a Perception System

The automated identification and relabeling of perception weaknesses in machine learning systems address inefficiencies in training data, enhancing accuracy and reducing manual labor, thereby improving the performance of perception systems.

JP7701687B2Active Publication Date: 2025-07-02EDGE CASE RESEARCH INC
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
JP2022549825
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-21
Filing Date
2021-02-22
Publication Date
2025-07-02
Estimated Expiration
2041-02-22

AI Technical Summary

Technical Problem

Existing machine learning-based perception systems face inefficiencies in identifying and addressing biases in training data, leading to reduced accuracy in object classification, particularly due to insufficient training data coverage and manual labor-intensive processes for data labeling and retraining.

Method used

A computer-implemented method that automates the identification of perception weaknesses by processing input data through a defect detection engine, relabeling scenes with identified weaknesses, and retraining the perception system to improve labeling performance and reduce manual effort.

Benefits of technology

Enhances the efficiency and effectiveness of machine learning perception systems by reducing manual analysis time and cost, improving the quality and diversity of training data, and optimizing the labeling process to enhance model performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701687000001
    Figure 0007701687000001
  • Figure 0007701687000002
    Figure 0007701687000002
  • Figure 0007701687000003
    Figure 0007701687000003
Patent Text Reader

Abstract

A method is described for automatically identifying perceptual weaknesses in training data used in improving the performance of a perception system. Defects in simulated data are also identified. The method, which can be incorporated into a system or into instructions placed on a storage medium, includes comparing the results of the perception system between baseline results and results with augmented input, and identifying perceptual weaknesses in response to the comparison. The perception system is retrained using the relabeled data, thereby improving.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] <Cross - Reference to Related Applications> This international patent application claims priority to U.S. Provisional Patent Application No. 62 / 979,776, “Automated Identification of Training Data Candidates for Perception Systems,” filed on February 21, 2020, the entire specification of which is incorporated herein by reference.

[0002] <Technical Field> The present invention relates to a novel attempt to identify data useful for training a system that learns using examples (e.g., a system using a type of machine learning). More particularly, the present invention relates to a method and system for identifying valuable training data used when retraining a learning system such as a system implemented using machine learning. Visual - based perception is exemplified, but as will be apparent to those skilled in the art, other systems will also be suitable for this attempt.

Background Art

[0003] Autonomous systems, such as self - driving vehicles, use a perception system to convert sensor data (e.g., data from a video camera) into a model of the world around the system (e.g., a list of objects and their positions). This is a more general example of a system that operates based on training data (e.g., implemented using machine learning such as deep learning based on a convolutional neural network).

[0004] An important factor in the performance of such a system is whether the training data covers all the important features of the operation to be learned with sufficient breadth and a sufficient number of examples of each object type to be learned. Each object type usually has multiple subtypes, and these subtypes need to be included in the training data for proper coverage. In a simple example, a pedestrian is an object type, and the subtypes could be whether they are wearing a red coat, a black coat, or a blue coat. "Subtype" in this context does not necessarily mean the classification information required by the end application, but rather various presentations of the object that each require sufficient training data and can be generalized to all relevant objects of that type. In practice, insufficient training data can bias the output of the machine learning process and reduce the probability of correct classification of objects to an unacceptable level. This is because some features are sufficiently different from other instances of that object type and may expose defects in the learning results.

[0005] For example, consider a perception system that has been carefully trained on pedestrians wearing red coats, black coats, and blue coats. If the perception system learns through training data that yellow is associated with objects other than pedestrians (such as traffic warning signs or traffic barriers), then a pedestrian wearing a yellow raincoat may have a lower probability of being detected as a pedestrian or may not be recognized as a pedestrian at all. This situation can be said to be caused by a bias in the training data that affects the system to undervalue pedestrians wearing yellow coats.

[0006] One attempt to identify biases affecting such a system is based on comparing object detection and / or classification results between a baseline sensor data stream and a sensor data stream augmented (e.g., with Gaussian noise). Such an attempt is described in International Patent Application PCT / US19 / 59619, “Systems and Methods for Evaluating Perception System Quality,” filed November 4, 2019, which is hereby incorporated by reference in its entirety. The attempt described in that patent application is generally referred to as “weakness detections” for a perception system. For example, the comparison between a baseline video data stream and a video data stream with Gaussian noise injected is an example of a “perception weakness detector” (note that such a detector can also function with other sensor modalities such as lidar and radar). An instance where a sensor data sample detects a weakness is called a “detection.”

[0007] One approach using a perception weakness detector is to present examples of weaknesses to a human analyst. The human analyst can look at a number of detections, identify patterns, and determine that the corresponding weaknesses exist in the perception system. Once a weakness is identified, the perception designer may search for more examples of more similar data samples and use them to retrain the system, which may improve performance. If the perception system is not improved by retraining alone, even after adding a sufficient number of data samples, the perception designer may need to reconstruct the system in some way to further improve the performance against the identified weakness. Using manual analysis to determine perception weaknesses can be effective, but it can also be time-consuming, even for highly skilled experts. The object of the present invention is to reduce the manual analysis effort required to improve the performance of a perception system. In an embodiment, the improvement in the performance of the perception system is indicated by a reduction in the number of perception weaknesses or defects identified by the system.

[0008] As described above, once a perception weakness is identified, additional training data must be identified. This may require a significant amount of manual work, such as scanning archived videos of pedestrians wearing yellow coats along the lines of the previous examples. It is a further object of the present invention to reduce the time and effort required to identify such training data.

[0009] After the weaknesses of the perception system are identified, the training data must be labeled before being input into the retraining process of machine learning. Since it requires a significant amount of manual labor, labeling can be very costly. For example, it has been reported that it can take up to 800 hours of manual labor to label one hour of driving data. Another objective of the present invention is to make the labeling cost more efficient by increasing the proportion of data that can actually contribute to improving the perception performance by being used as retraining data among the data to be labeled. In this specification, "training data" means both training data and validation data, such as those used in a machine learning approach of supervised learning. The "training data" identified by the disclosed embodiments can be divided into all training data, all validation data, or a combination thereof, which are considered advantageous by the perception system designer and perception verification activities.

[0010] These and other features and advantages will become apparent by reading the following detailed description and examining the accompanying drawings. It should be understood that the foregoing summary, the following detailed description, and the accompanying drawings are for illustrative purposes only and do not limit the various aspects as described in the claims.

Summary of the Invention

[0011] In a first aspect, a computer-implemented method for improving the labeling performance of a machine learning (ML) perception system for an input data set is provided, the method including processing the input data set through the perception system. After the input data system is processed, the method of the present invention further includes identifying one or more perception weaknesses, separating at least one scene from the input data set including the one or more perception weaknesses, relabeling the one or more perception weaknesses in the at least one scene to obtain at least one relabeled scene, and retraining the perception system with the at least one relabeled scene. In this way, the labeling performance of the ML perception system is improved.

[0012] In a second aspect, a computer-implemented method for evaluating the quality of simulated data used to train an ML perception system is disclosed. In some embodiments, the method includes processing an input data set of simulated data through the perception system to obtain one or more defect candidates, extracting a simulated object rendering associated with the one or more defect candidates, comparing the simulated object rendering with other examples of objects, and flagging simulated objects with low reliability matches with the other examples.

[0013] In a third aspect, a computer-implemented method for determining trigger conditions for perception weaknesses identified during a perception engine test using simulated data is disclosed. In some embodiments, the method includes processing an input data set including simulated data through the perception system to obtain a first result set, processing an input data set including augmented simulated data through the perception system to obtain a second result set, using a defect detection engine to compare the first result set and the second result set to identify perception weaknesses caused by simulated objects, where the simulated objects are defined by a plurality of parameters, and creating a modified simulated object by changing at least one parameter. In some embodiments, the third aspect further includes repeating the processing step and the comparison step using the modified simulated object to determine whether the modified simulated object has an impact on the perception weakness.

[0014] In some embodiments, these computer-implemented methods can be enabled and executed via a computer system having at least a memory or other data storage mechanism, one or more processors, and a non-transitory computer-readable medium storing instructions that, when enabled by the one or more processors, perform one or more of the foregoing aspects. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] For a complete understanding of the features and advantages of the present invention, reference should be made to the following detailed description taken in conjunction with the accompanying drawings.

[0016]

Figure 1

[0017]

Figure 2

[0018]

Figure 3

[0019] The detailed description provided below in connection with the accompanying drawings is intended as a description of the examples and is not intended to show the only forms in which the examples can be constructed or utilized. In the following description, the functions of the examples and a series of procedures that constitute and operate the examples will be described. However, it is also possible to realize the same or equivalent functions and procedures in different examples.

[0020] References to "one embodiment", "an embodiment", "exemplary embodiments", "some embodiments", "one example", "an example", "an instance", and "a certain example" etc. mean that the described embodiments, examples or instances can include specific features, structures or characteristics, but all embodiments, examples or instances do not necessarily have to include specific features, structures or characteristics. Also, these expressions do not necessarily refer to the same embodiment, example, or instance. Further, when a specific feature, structure or characteristic is described in connection with an embodiment, example or instance, it should be understood that such feature, structure or characteristic can be implemented in connection with other embodiments, examples or instances whether or not explicitly described.

[0021] Figure 1 shows a portion of the dataset used by the defect detection engine 10, which is used to identify the presence of perceptual weaknesses (hereinafter also referred to as "defects") within a baseline dataset 20, which is a given set of input data. The defect detection engine 10 corresponds to a preferred embodiment described in International Application PCT / US2019 / 59619 "Systems and Methods for Evaluating Perception System Quality", the entire content of which is incorporated herein by reference, and identifies weak perceptual performance.

[0022] In some embodiments, the defect detection engine 10 is parameterized by the following attributes: baseline dataset 20, extension 30, detector 40, and optional comparison dataset 50. The baseline dataset 20 is the dataset on which the defect detection engine 10 operates to identify defects. In some embodiments, it may occur that the engine 10 needs to modify the baseline dataset 20 in some way to facilitate defect detection. The extension 30 is a specification indicating how the data should be modified to achieve this purpose. In alternative embodiments, the engine may be provided with a comparison dataset 50 that is used in cooperation with the baseline dataset 20 to execute the defect detection process.

[0023] The engine 10 uses one or more computational components to perform the desired detection process. The detector 40 is a specification indicating the components used to achieve this purpose. The engine 10 executes a defect detection process on the baseline dataset 20 (and in some embodiments, the comparison dataset 50) of a particular "system under test" (SUT) 60 and evaluates the performance against the test criteria created by a given detector 40. In some embodiments, the SUT 60 is a perceptual system such as a deep convolutional neural network. Perceptual weaknesses, i.e., "defects", are placed in a defect database 70 for further processing or label inspection according to some embodiments of the present invention.

[0024] An embodiment of the workflow system 200 of the present invention is shown in FIG. 2. This embodiment assumes that the work executed in the workflow system 200 is represented by a sequence of job specifications 205. In this specification, the job specification 205 is a description of data input, configuration metadata, processing steps, and output for each unit of work processed by the system 200. These job specifications 205 are processed by an orchestrator 210, and the orchestrator 210 organizes and executes the workflow components necessary to satisfy each submitted job specification 205. It can be assumed that all the parameters necessary for the components of the workflow system 200 to execute their stages in the overall workflow are provided by the orchestrator 210.

[0025] Labeled data is required for the training of machine learning models. Labeled data can be created with sufficient quality and diversity only by humans involved in the labeling in some way and / or by quality checks of the automated labeling support. Hiring, training, and managing a team of human labelers requires cost and time. In order to maximize the ROI of the labeler team, it is necessary to optimize the quality, unit price, and value of the labels produced by the team. The quality of the labels is optimized by placing well-trained labelers and implementing a strong quality control process for the labels they produce. The unit price of the labels is optimized by maximizing the amount of labels generated per labeler while maintaining quality and minimizing the effort required to determine which items among all the items that can be labeled are the most suitable for labeling. The value of the labels is optimized by ensuring that the labels created by the labelers result in the highest possible improvement in the models trained using those labels.

[0026] The use and efficiency of labeling resources used in both the training and evaluation of a particular perception system are improved by automatically identifying highly valuable defects in the training values used in training. Some embodiments represent examples where scenes containing objects identified as perception defects are not sufficiently trained for the machine learning model being evaluated to process, and thus, when added to the training / evaluation dataset, may exploit the property of having a positive impact on the operation of the model. In practice, labels generated from scenes identified by a defect detection engine such as 10 have a higher expected value for the development process than labels created from scenes selected via other means such as random extraction or subjective human decision-making, as they reveal the sensitivity within the perception system. Further, since the defect detection engine is an automated tool that does not require human intervention to identify scenes that reveal sensitive parts within the perception system, it reduces the time and effort required to determine items to be labeled, and as a result, lowers the cost per label created. In other words, automated defect identification is more scalable than defect identification performed by humans. Note that in some embodiments, the labeling activity is automatically invoked in response to automated defect identification, rather than being used by a human analyst to hypothesize about the root cause of the perception defect.

[0027] Some embodiments provide a method for automatically identifying data items that are expected to have a positive effect on subsequent model learning processes that consume those labels when labeled by a human labeling team. As shown in FIG. 2, an autonomous vehicle (e.g., a test vehicle) under the control of an autonomous stack is tasked with driving in the world (track or road). Road data 210, including both sensor data and behavior data regarding what the vehicle encountered and how it behaved, is acquired during those drives. Eventually, the road data 210 is uploaded to a storage device 215 and stored for later access by a workflow system 200. In some embodiments described below, this same workflow can be applied equally well to data obtained from simulated drives, including pure software simulations and / or hybrid Hardware in Loop simulations.

[0028] The following steps outline the stages of an exemplary embodiment shown in FIG. 2. A job specification 205 is sent to an orchestrator 207 to request the execution of labeling for a portion of the available driving data. The orchestrator 207, in some embodiments, may be an algorithm represented by software that controls a series of steps that need to be executed to achieve a particular task, and uses a set of parameters supplied from the job specification 205 to instruct a defect detection engine 220 to perform the labeling execution. In a beneficial example, the specified input data may be selected from historical data found in a sensor data storage device 215.

[0029] In an embodiment, the defect detection engine 220 is then instructed to perform one or more specified augmentations on the specified input data, and perform the specified detection using the baseline data and the augmented data as inputs. Next, as a result of the evaluation, defect candidates are identified and stored in the defect database 225. Optionally, operational data (e.g., the progress of data analysis) is provided to the operations dashboard, and the orchestrator 207 is notified that the analysis is complete.

[0030] Next, the orchestrator 207 tells the defect detection filter 235 that a new collection of defect candidates is available for transmission to the filtering and labeling mechanism 240. Next, the filter 235 extracts those defects that match the desired filtering criteria from the current job, generates a set of instructions for the labeling mechanism 240 indicating the scenes to be labeled, and sends those instructions to the labeling mechanism 240. As used herein, the term "scene" refers to one complete sample of the environment when a particular sensor is designed for that environment. In a non-limiting example, a scene may be data related to a single video frame or a single lidar sweep.

[0031] In some embodiments, the filter configurator 245 optionally provides a way to filter some of the defects to focus on specific defects of engineering interest (e.g., a perception weakness related to pedestrians taller than 250 pixels). The associated viewer displays examples of defects that meet the optional tuning filter criteria by the system user. However, the user does not need to view the results of the filtering, nor does the filter need to create a subset of the defect database. Therefore, this part of the workflow can be done with little human intervention.

[0032] The labeling mechanism 240 labels the requested scenes and artifacts, adds the results to the available set of labeled data in the training and validation database 250, and notifies the orchestrator 207 that the labeling is complete. The orchestrator 207 then instructs the system to resume perceptual system retraining 260. When the machine learning models associated with the newly labeled data are retrained, a newly trained perceptual engine 265 results in response to the training data identified in the previous step. This data provides labeled examples of data samples for which the performance of the perceptual engine 265 was determined to be weak, so training on such samples is expected to improve perceptual performance.

[0033] In some embodiments, those steps are iteratively performed as a technique that continuously identifies new perceptual weaknesses and automatically incorporates relabeled scenes into the data sets used to train the associated ML models. In embodiments, the performance dashboard displays current and past perceptual engine performance data (such as a reduction in defects due to successive executions of the workflow).

[0034] In some embodiments of the present invention, the evaluator 270 and the reengineer process 275 handle situations where brute force retraining reveals problems on systems that do not follow such retraining. The ML model evolves along two main axes of learning and reconstruction, and can acquire functionality and performance over time through the continuous expansion and curation of the training dataset; as the set of training data grows and its quality improves, the functionality and performance of a properly constructed model improve. However, the ML model becomes extremely large and, at various points in development, further learning will no longer enable new capabilities to be demonstrated or performance to be improved. When such a plateau is reached, model reconstruction is required. Examples of reconstruction include changing the model's architecture and hyperparameters, as known to those skilled in the art (e.g., increasing the number of layers in a deep network topology training approach).

[0035] Some embodiments of the present invention utilize signals specific to the results obtained from a series of executions of a defect detection engine on an ML model trained against a growing collection of training data (such as that generated via FIG. 2). Evaluating the results of a series of labeling runs is expected to reveal progress (indicating usefulness in continuing retraining with additional data towards the intended goal) or the lack thereof (indicating that the model needs to be reconstructed as it has reached a maximum and no sufficient progress is seen with additional labeled data).

[0036] As described above, some embodiments include monitoring the progress of an ML model over a series of training runs to identify when further training is no longer effective and the model needs to be rebuilt. In particular, the orchestrator 207 instructs the evaluator 270 to compare a newly created set of defects with a set of defects generated from a previous labeling run on the same source dataset. If it is determined that there has been significant progress by adding new labels, the evaluator 270 instructs the orchestrator 207 to continue labeling. If it is determined that there has been no significant progress by adding new labels, the evaluator 270 instructs the orchestrator 207 to start rebuilding with respect to the identified set of defects that have proven to be resistant to retraining the model. Significant progress in this context is measured by an improvement in the Precision-Recall (PR) curve of the ML perception system. In this curve, "precision" is the rate of false positives and "recall" is the rate of false negatives. In an embodiment, if the curve improves by more than 0.5%, it is still considered significant progress such that rebuilding is not yet necessary. In a non-limiting example of "significant improvement", in a run over run comparison of the results of the perception system, after relabeling, the precision improved by 0.5% without a decrease in recall. Nevertheless, one of ordinary skill in the art will understand that significant improvement can be characterized in various ways.

[0037] In an embodiment, the perception system rebuild 275 may include adjusting model hyperparameters, adding / changing layers, applying regularization techniques, and other techniques. Once completed, the perception system rebuild 275 submits a new job specification to the orchestrator 207 and labeling and other workflows resume using the newly adjusted perception engine 265.

[0038] In some embodiments, the system shown in FIG. 2 can also be used to search for training errors that affect the system according to the following steps.

[0039] The job specification 205 is sent to the orchestrator 207, and the execution of defect detection for a portion of the stored real-world operation data is requested. The orchestrator 207 instructs the defect detection engine 220 to perform an evaluation according to the following procedure. The parameters used in this example are for illustrative purposes only and would be sent from the job specification 205.

[0040] Data that has already been collected and has available ground truth is selected from the sensor data storage device 215. This data corresponds to data samples that are also in the training & validation database 250 and preferably has available labels. No augmentation is performed on the data of this embodiment. The specified detection is performed using the baseline data and the ground truth data, and the defect candidates are stored in the defect database 225. Next, the evaluator 270 compares the newly created set of defects with the set of defects generated from previous evaluation runs against the ground truth. If no significant progress as described above has been made by the model since the last evaluation against the ground truth, the orchestrator 207 starts the core reconstruction workflow described above. In some embodiments, this comparison may include the selection of specific object types and characteristics to be evaluated by reversing the execution order of the evaluator 270 and the filter 235, and the transmission of the results evaluated after being filtered for the perception system reconstruction 275.

[0041] In a further embodiment, the system and method shown in FIG. 2 may be used to search for classification errors as follows: The job specification 205 is sent to the orchestrator 207 and an evaluation execution on a portion of the stored previously collected real-world data is requested. The orchestrator 207 instructs the defect detection engine 220 to perform a defect detection execution. As before, the parameters used in this example are for illustrative purposes only and would be obtained from the job specification 205. First, data with available ground truths is selected from the historical data sensor storage 215 and some extensions are performed on a copy of the data to enhance the weakness detection as taught in PCT / US2019 / 59619. Then, specific detections are identified using the extended data and the ground truth data, and defect candidates are stored in the defect database 225. Next, the newly created set of defects is compared to the set of defects generated from the previous evaluation execution against the ground truth. If no significant improvement has been made in the model since the last evaluation against the ground truth, the orchestrator 207 starts the core reconstruction workflow outlined above. As explained before, the filter 235 and the evaluator 270 can optionally be executed in the reverse order. This embodiment is envisioned to be repeatedly executed as a means of constantly evaluating the current state of the model and whether further training and reconstruction are needed.

[0042] In yet another embodiment, the quality of the simulations used to train the ML model can be improved, and as a result, the trust placed in the simulation tool will increase. The autonomous system is tested via simulation as a means of expanding the scenario scope and exploring the existing test space. However, when evaluating the ML model against simulated data, there is always a question as to whether the behavior of the ML model is due to the ML model itself or the quality of the simulation used to train it. By providing a mechanism to quantitatively evaluate the quality of the simulation objects, both the producers and users of the simulation tool can build trust in the results of running the ML model against the simulated world.

[0043] The embodiment shown in FIG. 3 leverages potential signals embodied in the results captured from the execution of a series of defect detections on a given ML model trained with simulated data. When a defect detection engine recognizes a particular feature or object of a simulated scene as a weakness in an ML model, that scene (one or more data samples) may be automatically processed through an evaluation process against existing ground truth information to evaluate whether the simulated scene data itself is the cause of the problem. Put another way, this embodiment outlines a method of capturing the sensitivity of a particular ML model to simulated features and automatically invoking an evaluation of those features to determine the quality of those features against a given set of quantitative metrics.

[0044] The vehicle simulated under the control of the autonomous stack is tasked with being driven in the simulated world, and sensor data and behavioral data regarding what the vehicle encountered and how it behaved are captured during the drive and ultimately moved to a repository accessible to components within the workflow system.

[0045] FIG. 3 is a diagram showing some embodiments of the present invention adjusted to evaluate the quality of simulation objects. In this workflow 300, a job specification 305 is sent to an orchestrator 310, requesting the execution of labeling for a certain part of the stored simulated operation data 315. This is different from the previous embodiments using actual road data. The orchestrator 310 instructs a defect detection engine 320 to perform an evaluation in the following procedure. Similar to the previous, the parameters used in this example are for illustrative purposes only and would be obtained from the job specification 305.

[0046] Data with available ground truth is selected from the sensor data storage 315, but no extension is performed on the data. Specific detections are performed by the defect detection engine 320 using baseline data and ground truth data, and defect candidates are stored in a defect database 325. In some embodiments, the scenes related to the defect candidates are also stored in the database 325.

[0047] Next, the orchestrator 310 starts an evaluation simulation object engine 330 to determine the rendering quality of the objects in the scenes identified as defect candidates. In some embodiments, the workflow 300 extracts the rendering within each bounding box identifying the defect candidates, determines the type of the simulated objects included, and executes an automatic comparison function with other examples of those objects stored in an existing catalog of labeled actual sensor video images. In other embodiments, the learning model may be trained with a library of existing labeled objects. If the match is insufficient, a new simulation is created to check whether the new objects are highly reliable compared to what the model shows. The better, the higher. Note that since the simulator is generating sensor data, the ground truth is known without the need for human-assisted labeling.

[0048] If a subsequent quality assessment 335 determines that the quality of the rendered entity is unacceptable, the corresponding defect candidate is flagged in the defect database 325 as requiring an update to the simulation world model. The orchestrator 310 is notified that the quality assessment 335 has been completed. However, if there are remaining defect candidates that do not require an update to the simulation world model, the orchestrator 310 starts the remainder of the standard labeling workflow and proceeds to the next steps of the filter 340, the labeling 345, the storage of the labeled data in the training database 360, and the subsequent retraining 370 by the perception engine 380 as described above. If a defect candidate is flagged as requiring an update to the simulation world model, the orchestrator 310 provides a set of defect candidates and related metadata to the developer / maintainer of the simulation information base according to the updated world model step 350. The developer and maintainer of the simulation information base examines the defect candidate information and metadata, identifies the necessary improvements, and updates the affected world model. When the adjustment to the affected world model is complete, the execution of the new simulator 355 is started and new simulation entries are generated in the sensor data storage device 315 for use in future evaluation runs. This embodiment is envisioned to be repeatedly executed as a means of constantly evaluating the quality of the simulated features provided by the simulation tool.

[0049] In yet other embodiments, an automated process is provided for determining the trigger conditions of defects discovered during model testing using simulated data. In the development of autonomous systems, it is often necessary to identify the root cause of the observed behavior of the system, i.e., the "trigger conditions". Usually, this process is a manual process where highly trained and experienced human assessors examine individual scenes within a dataset, assert potential causes for the demonstrated behavior, and conduct tests to confirm or refute those assertions. Since the execution cost of this manual process is very high, finding a way to significantly or fully automate this process would bring great value.

[0050] Some embodiments of the disclosed systems and methods utilize potential signals embodied in the results captured from the execution of a series of defect detections on a given ML model trained with simulated data. When a defect detection engine identifies a particular feature in a simulated scene as generating sensitivity within a particular ML model, it can be considered to have automatically identified interesting conditions for further evaluation. Additionally, by identifying defects from the scenes that occurred in the simulated world, it has also identified the particular scenarios and world models that embody the sensitivity. Since the scenario configurations and world models are parameterized, they can be used to generate a series of new simulations using the original scenarios and / or world models with one or more parameters changed. Executing the defect detection engine against the outputs from the simulations generated using these changed scenarios and / or world models generates a new series of defects that can be compared to the defects generated from the simulation runs that triggered the original defects. If it is shown that changing a particular parameter of a scenario or world model has a significant impact on the series of defects generated by the defect detection engine over a series of executions, then depending on the class of parameters, it may be possible to infer the trigger conditions for the defects from the nature of the changed parameters.

[0051] This embodiment outlines a method for identifying defect candidates for root cause analysis of trigger conditions, and a mechanism that utilizes a defect detection engine to evaluate simulated operation data and automatically infer the trigger conditions for those defects.

[0052] In some embodiments, a simulated vehicle under the control of an autonomous stack is tasked with being driven in a simulated world, and sensor and behavior data regarding what the vehicle encountered and how it behaved is captured during operation and ultimately moved to a repository accessible to components within the workflow system.

[0053] Continuing to refer to FIG. 3, next, a modified version of the trigger identification process 300 will be described. The job specification 305 is sent to the orchestrator 310 to instruct the defect detection engine 320 to perform a search execution. As before, the parameters used in this example are for illustrative purposes only and would be obtained from the job specification 305. Data with available ground truth (e.g., data from a previously labeled learning and validation database 360, or results of previous simulation data, etc.) is selected. Next, specific detections are performed by the defect detection engine 320 using baseline data and ground truth data. Defect candidates are stored in the defect database 325 along with references to the scenarios used to generate the starting simulations (for the simulated data) or annotated scenario descriptions of the labeled road data. Next, the orchestrator 310 instructs that related simulations with one or more modified parameters be generated, enabling the exploration of the simulation space in the vicinity of the features identified by the set of defect candidates.

[0054] In some embodiments, the simulated objects within each bounding box for each defect candidate can be identified using the associated ground truth, and from past executions of the workflow, it can be determined whether parameter changes in the current simulation scenario or world model had a significant impact on the set of detections generated in the current execution. If there is a positive trend in the significant impact (i.e., showing a correlation), but it is not yet sufficient to infer a causal relationship, the next step is to generate one or more new simulation scenarios or world models that further change the same parameters associated with the entities identified by the current set of detections. However, if there is a negative trend in the significant impact, or if there is no significant impact, one or more new simulation scenarios or world models are generated and various parameters associated with the entities identified in the current detection set are changed.

[0055] For new simulation scenarios or world models, a new job specification 305 is sent for additional workflow execution for the new simulation instance. If a sufficient causal relationship can be determined between parameter changes in one or more scenes or world models, the set of parameter changes is recorded in the defect database 325 along with the set of defect candidates, and the entire workflow ends. If it is determined that the causal relationship between parameter changes in one or more scenes or world models is not sufficient, the set of defects is marked for human review, and the entire workflow ends.

[0056] Examples of modifiable parameters include object characteristics (e.g., color, shape, size, display angle or actors, vehicles, road equipment, etc.), background (e.g., color, texture, brightness), scene characteristics (e.g., contrast, glare), environmental characteristics (e.g., haze, fog, snow, sleet, rain), occlusion (slight occlusion, partial occlusion, heavy occlusion), and the like.

[0057] This embodiment is assumed to be repeatedly executed as a method for constantly searching for defects and identifying the conditions that trigger them. It will be apparent to those skilled in the art that with a sufficiently accurate automatic perception labeling function, the labeling can be fully automated or semi-automated.

[0058] Various embodiments of the present invention can be advantageously used to identify good candidates for perceptual retraining, identify weaknesses that affect the system of the perceptual system, and identify defects that affect the system in the simulator output. The present invention can be implemented not only via a physical device (e.g., an extended physical computer module comprising hardware and software that modifies a data stream transmitted to another perceptual system and a physical computer module that performs comparisons), but also via a pure software mechanism hosted on a general computer platform.

[0059] Some aspects of the embodiments described herein include process steps. The process steps of the embodiments can be embodied in software, firmware, or hardware, and when embodied in software, it should be noted that it can be downloaded and reside on various platforms used by various operating systems and be operated therefrom. The embodiments may also be made into a computer program product executable on a computer system.

[0060] Embodiments also relate to a system for performing the operations described herein. Such a system may be specially constructed for the purpose, or may include a general purpose computer selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored on a computer readable storage medium, which may include, but is not limited to, a floppy disk, optical disk, CD-ROM, magneto-optical disk, read only memory (ROM), random access memory (RAM), EPROM, EEPROM, magnetic or optical card, application specific integrated circuit (ASIC), or any type of media suitable for storing electronic instructions, each coupled to a computer system bus. The memory / storage device may be transient or non-transient. The memory may include any of the above and / or other devices capable of storing information / data / program. Further, the computer devices referred to herein may include a single processor or may be an architecture that employs multiple processor designs to enhance computing capabilities.

[0061] Although the invention has been described and illustrated with reference to preferred embodiments and specific examples thereof, it will be readily apparent to those of ordinary skill in the art that other embodiments and examples can perform similar functions and / or achieve similar results. All such equivalent embodiments and examples are within the spirit and scope of the invention, are contemplated thereby, and are intended to be covered by the claims.

Claims

1. A computer-implemented method for improving the labeling performance of a machine learning (ML) perception system for an input data set, comprising: processing, by a computer device, the input data set via the ML perception system; identifying, by the computer device, one or more perception weaknesses; separating, by the computer device, at least one scene including the one or more perception weaknesses from the input data set; relabeling, in the at least one scene, the one or more perception weaknesses to obtain at least one relabeled scene; retraining, by the computer device, the ML perception system using the at least one relabeled scene; thereby improving the labeling performance of the ML perception system; wherein the step of processing the input data set includes: creating, by the computer device, an extended data set; processing, by the computer device, the extended data set via the ML perception system; and comparing, by the computer device, the results of the input data set with the results of the extended data set.

2. The computer-implemented method according to claim 1, wherein the step of identifying the one or more perception weaknesses is performed by a defect detection engine.

3. The computer-implemented method according to claim 1, wherein the one or more perception weaknesses include labeled objects of a bounding box.

4. The computer-implemented method according to claim 1, wherein the step of relabeling the at least one scene includes sending the at least one scene to be labeled by a human.

5. The computer-implemented method according to claim 1, wherein the input data set is composed of one or more of baseline data, extended data, ground truth data, and simulation data.

6. The computer-implemented method according to claim 1, further comprising filtering, by the computer device, the results of the processing step to focus on specific perception weaknesses of engineering interest.

7. The computer-implemented method according to claim 1, further comprising repeating the method, comparing the perceptual weakness with the perceptual weakness identified in the step of processing that occurred before the relabeling, and determining an improvement amount in the performance of the ML perception system.

8. The computer-implemented method according to claim 7, further comprising repeating the step of processing, the step of identifying, the step of separating, the step of relabeling, and the step of retraining until the labeling performance no longer significantly improves even when the ML perception system is retrained using the at least one relabeled scene.

9. The computer-implemented method according to claim 8, further comprising reconstructing the ML perception system when the labeling performance is not significantly improved in the at least one relabeled scene.

10. A computer-implemented method for determining a trigger condition for a perceptual weakness identified during a perceptual engine test using simulated data, (a) a computer device processing an input data set including simulated data through a perception system to obtain a first result set; (b) the computer device processing an input data set including extended simulated data through the perception system to obtain a second result set; (c) the computer device comparing the first result set and the second result set using a defect detection engine to identify a perceptual weakness caused by a simulated object defined by a plurality of parameters; (d) the computer device changing at least one parameter to create a modified simulated object; and repeating the processing step (a), the processing step (b), and the comparing step (c) using the modified simulated object to determine whether the modified simulated object affected the perceptual weakness.

11. The computer-implemented method according to claim 10, further comprising repeating the processing step (a), the processing step (b), the comparing step (c), and the changing step (d) until the trigger condition is identified.

12. In a computer system for improving the labeling performance of a machine learning (ML) perception system with respect to an input data set, a memory or other data storage mechanism, one or more processors, a non-transitory computer-readable medium, and is provided with, when the non-transitory computer-readable medium is executed by the one or more processors, processing the input data set by the ML perception system to identify one or more perception weaknesses; separating at least one scene including the one or more perception weaknesses from the input data set; relabeling the one or more perception weaknesses in the at least one scene to obtain at least one relabeled scene; retraining the ML perception system using the at least one relabeled scene; stores instructions for improving the labeling performance of the perception system, processing the input data set by the ML perception system includes the computer device creating an extended data set, the computer device processing the extended data set via the ML perception system, and the computer device comparing the results of the input data set with the results of the extended data set. A computer system.

13. A non-transitory computer-readable medium for improving the labeling performance of a machine learning (ML) perception system with respect to an input data set, when executed by a processor, processing the input data set by the ML perception system to identify one or more perception weaknesses; separating at least one scene including the one or more perception weaknesses from the input data set; relabeling the one or more perception weaknesses in the at least one scene to obtain at least one relabeled scene; retraining the ML perception system using the at least one relabeled scene; stores instructions for performing improving the labeling performance of the ML perception system, Processing the input data set with the ML perception system includes the computer device creating an extended data set, the computer device processing the extended data set via the ML perception system, and the computer device comparing the results of the input data set with the results of the extended data set. A non-transitory computer-readable medium.

Citation Information

Patent Citations

  • Apparatus and method for information processing

    JP2015087903A

  • Identifier, identification program, and identification method

    JP2015095212A

  • Data processing device, data processing method, and computer program

    JP2016076073A

  • Generation of a virtual world to assess real-world video analysis performance

    JP2017151973A

  • Information processing device, information processing method, and information processing program

    JP2018156316A