Object detection performance evaluation method and device, controller, vehicle and medium

By using a surrogate model to evaluate the performance of the object detection model on different datasets, the problem of declining model performance on other datasets after customized training is solved, training efficiency is improved and performance on multiple datasets is optimized.

CN120877028APending Publication Date: 2025-10-31ROBERT BOSCH GMBH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410543601.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-30
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing object detection models often experience a decline in performance on other datasets after being customized for specific datasets, resulting in low training efficiency and wasted resources, making it difficult to effectively balance performance across multiple datasets.

Method used

A proxy model is used for performance evaluation. By training a proxy model corresponding to the original model, the reasoning behavior is imitated, and its performance on different datasets is evaluated. This allows for the acquisition of comprehensive performance information without calling the original model, enabling targeted adjustments to the original model to optimize its performance on multiple datasets.

Benefits of technology

It improves the training efficiency of the object detection model, reduces resource consumption, ensures that the model maintains excellent performance on multiple datasets, and achieves a performance trade-off in different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877028A_ABST
    Figure CN120877028A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to a performance evaluation method and device for object detection, equipment, a controller and a medium. The method includes training a first model for object detection based on a first sample in a first data set and a first truth tag corresponding to the first sample. The method further includes training a second model for evaluating the trained first model based on the first sample and the first truth tag, and a first output of the trained first model for the first sample. The method further includes determining performance of the trained first model for a second data set different from the first data set by verifying performance of the trained second model for the second data set. In this way, after the performance of the model is optimized for the specific data set, the performance of the model for other data sets can be evaluated by using the agent model corresponding to the model, so that comprehensive performance information of the original model is obtained under the condition that the original model is not called, and the training efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this disclosure generally relate to the field of computers, and more particularly to methods, apparatuses, devices, controllers, and media for performance evaluation of object detection. Background Technology

[0002] Object detection is a computer vision technique used to locate and identify objects of interest within visual information such as images or videos. Examples include vehicles, pedestrians, and road signs in driving scenarios. Object detection is a crucial component of driver assistance and autonomous driving technologies. This technology helps vehicle driving systems (such as Advanced Driver Assistance Systems (ADAS) and Autonomous Driving (AD) systems) better understand the driving environment, enabling them to make safer and more effective decisions and take appropriate actions.

[0003] Artificial intelligence (AI) / machine learning (ML) models for object detection have the ability to extract representational features from visual information and identify target objects after a certain period of training. With the continuous advancement of technologies such as computer vision and machine learning, the accuracy and real-time performance of object detection have been continuously improved. Summary of the Invention

[0004] Embodiments of this disclosure provide a method, apparatus, device, and medium for performance evaluation of object detection.

[0005] According to a first aspect of this disclosure, a method for performance evaluation of object detection is provided. The method includes training a first model for object detection based on first samples in a first dataset and first ground truth labels corresponding to the first samples. The method further includes training a second model for evaluating the trained first model based on the first samples, the first ground truth labels, and a first output of the trained first model for the first samples. The method also includes determining the object detection performance of the trained first model for the second dataset by verifying the object detection performance of the trained second model for the second dataset, which is different from the first dataset.

[0006] According to a second aspect of this disclosure, an apparatus for performance evaluation of object detection is provided. The apparatus includes a first training module configured to train a first model for object detection based on first samples in a first dataset and first ground truth labels corresponding to the first samples. The apparatus also includes a second training module configured to train a second model for evaluating the trained first model based on the first samples, the first ground truth labels, and a first output of the trained first model for the first samples. The apparatus further includes a performance determination module configured to determine the object detection performance of the trained first model for the second dataset by verifying the object detection performance of the trained second model for the second dataset, which is different from the first dataset.

[0007] According to a third aspect of this disclosure, an electronic device is provided. The electronic device includes at least one processor. The electronic device also includes a memory coupled to the at least one processor and having instructions stored thereon that, when executed by the at least one processor, cause the device to perform the steps of the method in the first aspect of this disclosure.

[0008] According to a fourth aspect of this disclosure, a vehicle is provided. The vehicle includes the electronic equipment described in the third aspect of this disclosure.

[0009] According to a fifth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions, which are executed by a computer's processor to implement the steps of the method in the first aspect of this disclosure. Attached Figure Description

[0010] The above and other objects, features, and advantages of this disclosure will become clearer through a more detailed description of exemplary embodiments of the present disclosure taken in conjunction with the accompanying drawings. In exemplary embodiments of the present disclosure, the same or similar reference numerals generally represent the same or similar parts, components, etc.

[0011] Figure 1 The illustration shows a schematic diagram of an example environment in which the methods and / or apparatus according to embodiments of the present disclosure may be implemented;

[0012] Figure 2 The illustration shows a flowchart of a method for performance evaluation of object detection according to an embodiment of the present disclosure;

[0013] Figure 3 The illustration shows a performance evaluation process for an object detection model according to an embodiment of the present disclosure;

[0014] Figure 4 The illustration shows a process for determining whether models meet consistency or correlation requirements according to an embodiment of the present disclosure;

[0015] Figure 5 The illustration shows another process for determining whether consistency or correlation requirements are met between models according to an embodiment of the present disclosure;

[0016] Figure 6 A diagram illustrating the model architecture of the original model and the proxy model according to embodiments of this disclosure is shown;

[0017] Figure 7 A schematic diagram of an apparatus for performance evaluation of object detection according to embodiments of the present disclosure is shown; and

[0018] Figure 8 A schematic block diagram of an example device suitable for implementing embodiments of the present disclosure is shown.

[0019] In the various figures, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation

[0020] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0021] In the description of embodiments of this disclosure, the term "comprising" and its variations should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects unless explicitly indicated otherwise.

[0022] Object detection is a computer vision technology that plays a crucial role in numerous application scenarios. It identifies the location of objects based on visual information (e.g., in the form of bounding boxes) and determines the identity, category, and other information of the identified objects. Object detection technology helps facilitate scene understanding and behavior analysis. For example, in intelligent driving scenarios, object detection can identify key elements such as pedestrians, other vehicles, and road signs, enabling safe and effective driving decisions. Accurate object detection functionality benefits the performance of Advanced Driver Assistance Systems (ADAS) and Autonomous Driving Systems (AD), improving features such as blind spot detection, Automatic Parking Assist (APA), Adaptive Cruise Control (ACC), Lane Keeping Assist (LKA), and Automatic Emergency Braking. Furthermore, in the field of intelligent security, object detection can monitor abnormal objects and behaviors in real time, such as intruder intrusions and fires, and issue timely alerts.

[0023] For example, an artificial intelligence (AI) / machine learning (ML) model for object detection is configured to extract key features from visual information and, after a certain amount of training, identify objects of interest. Building a well-performing and stable model requires a sufficient amount of data for training, which is crucial for both the model's accuracy and generalization ability. When data is insufficient, data-augmented semi-supervised learning or even unsupervised learning strategies can be used to train the AI / ML model.

[0024] After an AI / ML model for object detection has been largely trained with sufficient data, if its performance on one or more datasets is below expectations, or if it will be used for specific application needs, then customized retraining of the AI / ML model is required. To illustrate this non-restrictively, consider a driving scenario: if dataset 1 primarily contains data samples related to driving in rainy or snowy weather, and the object detection model performs poorly on this dataset, or if the object detection model will be developed primarily for performing object detection while driving in rainy or snowy weather, then retraining of the object detection model for dataset 1 is necessary.

[0025] After a custom-trained object detection model is optimized for a specific dataset to meet accuracy requirements, its performance on other datasets may be worse than before the customization. For example, after a custom-trained model on a dataset including data samples about driving in rain or snow, its performance on dataset 1 improves, but its performance on another dataset including data samples about driving in congested traffic deteriorates. Therefore, it is necessary to obtain comprehensive performance information for other datasets after the model has been customized for a specific dataset, in order to weigh performance across these datasets.

[0026] To achieve performance tradeoffs for object detection models across multiple datasets, some solutions might involve retraining a customized object detection model on each of the datasets other than the specific dataset. However, this is extremely time-consuming, as AI / ML model architectures for object detection are typically very complex, and model iteration is very slow. Therefore, significant effort is required to compare the training results for each dataset across these datasets for any side effects caused by the aforementioned customized training, and this may require many iterations to maintain the final performance tradeoffs across different datasets. Furthermore, for data-driven processes in AI / ML models, performance analysis of iterative data is quite frequent, making it particularly time-consuming and labor-intensive.

[0027] To address at least the aforementioned and other potential problems, embodiments of this disclosure provide a scheme for performance evaluation of object detection. The scheme for performance evaluation of object detection according to embodiments of this disclosure includes training a first model for object detection based on a first sample in a first dataset and a first ground truth label corresponding to the first sample. The scheme further includes training a second model for evaluating the trained first model based on the first sample, the first ground truth label, and a first output of the trained first model for the first sample. The scheme also includes determining the object detection performance of the trained first model for the second dataset by verifying the object detection performance of the trained second model for a second dataset different from the first dataset. In this way, after the model's performance is optimized for a dataset, a surrogate model corresponding to that model can be used to evaluate the model's performance for other datasets, thereby obtaining comprehensive performance information of the original model without calling the original model, thus improving training efficiency.

[0028] The following is for reference. Figures 1 to 8 The present disclosure is provided to illustrate the basic principles and several example implementations. It should be understood that these exemplary embodiments are given only to enable those skilled in the art to better understand and implement the embodiments of the present disclosure, and are not intended to limit the scope of the disclosure in any way.

[0029] Figure 1 The illustration shows a schematic diagram of an example environment 100 in which the methods and / or processes according to embodiments of the present disclosure may be implemented. For ease of understanding, object detection in a driving scenario will be described exemplarily below. It should be understood that this is not limiting, and the methods according to embodiments of the present disclosure can also be used in other different use cases, such as the intelligent security scenarios discussed above, where object detection in intelligent security scenarios can be applied to schools, banks, etc.

[0030] like Figure 1 As shown, the example environment 100 may include a vehicle 110, training data 120, a computing device 130, and a storage device 140, and these components may be coupled to each other for interaction, such as... Figure 1 As shown in the illustration. It should be understood that limited components are shown in the example environment 100 for implementing embodiments of this disclosure for purposes of ease of understanding and illustration only, and embodiments of this disclosure are not limited thereto. For example, example environment 100 may also include a display (not shown) configured to display various datasets, performance information of object detection models, etc.

[0031] According to embodiments of this disclosure, vehicle 110 can be any type of motorized or non-motorized vehicle capable of carrying people and / or goods and being movable. Vehicle 110 typically includes one or more wheels, one or more seats, one or more load-bearing structures (such as a cabin, compartment, etc.), one or more power systems (such as an engine, electric motor, etc.), one or more control systems (such as a steering wheel, accelerator pedal, etc.), and one or more safety systems (such as seat belts, airbags, etc.), etc.

[0032] like Figure 1 As shown, vehicle 110 is illustrated as an automobile. However, this is merely exemplary and not limiting. By way of example, vehicle 110 may include, but is not limited to, buses, trucks, SUVs, sports cars, motorcycles, etc. Furthermore, vehicle 110 may be based on fossil fuels, clean energy, or a combination thereof. Fossil fuel-based vehicles primarily refer to vehicles that use fossil fuels such as oil and natural gas as their power source, such as traditional gasoline vehicles and diesel vehicles. Clean energy-based vehicles refer to vehicles that use clean energy as their power source, such as electric vehicles, hydrogen fuel cell vehicles, solar-powered vehicles, etc.

[0033] According to embodiments of this disclosure, vehicle 110 may include multiple sensors (not shown), and these sensors may be configured to sense the driving environment in which vehicle 110 is located when parked or in motion, the position and category, state and behavior of objects and persons (such as other vehicles and pedestrians) in that driving environment, etc., for the purpose of enabling various functionalities of ADAS and AD systems. The multiple sensors included in vehicle 110 may include at least one vision sensor (such as a camera, infrared sensor, depth sensor, etc.), and it is configured to capture visual information indicating the driving environment in which vehicle 110 is located, the position and category, state and behavior of objects in that driving environment, etc. In some embodiments, examples of the captured visual information may include, but are not limited to, images and videos, and are sampled to form training data 120 for an object detection model.

[0034] According to embodiments of this disclosure, training data 120 may include sample data for training an object detection model to detect objects in a driving environment. This sample data is formed from visual information captured by onboard sensors that indicates the driving environment and objects therein, processed through a series of steps. Examples of these processes may include, but are not limited to, normalization, cropping, scaling, flipping / rotating, brightness / contrast adjustment, color jittering, noise addition, labeling, and data augmentation. By using training data 120 to train the object detection model, the location and category, state, and behavior of objects in the driving environment can be identified. It should be understood that the training data may include not only sample data captured by onboard sensors as described above, but also sample data acquired through other means, such as sample data formed based on visual information acquired by external vision capture devices, and classic sample data specific to driving scenarios.

[0035] like Figure 1 As shown, training data 120 may include multiple datasets, which can be collections of sample data with the same or similar characteristics that have been classified based on a clustering strategy. For ease of understanding and illustration, in... Figure 1 The diagram illustrates training data 120, which includes data samples for driving scenarios and comprises three datasets: dataset 121, dataset 122, and dataset 133. By way of example, and not limitation, dataset 121 may include data samples on driving in rain or snow, dataset 122 may include data samples on driving in congested traffic, and dataset 123 may include sample data on driving in tunnel sections. In other words, these datasets can be tailored to specific application scenarios or usage requirements.

[0036] In some embodiments, the object detection model can be trained first using all the sample data in the training data 120, so that it has a certain ability to handle object detection in the corresponding scenario (e.g., driving scenario). Then, it can be customized by using a specific dataset corresponding to a certain application scenario or usage requirement, so that the model has optimized performance when performing object detection in the application scenario or for the usage requirement.

[0037] Training data 120 may be stored in storage device 140 and accessed by computing device 130. In some embodiments, training data 120 may include sample data presented in the form of images, video including image frames, etc. Such visual information may be captured by visual sensors on vehicle 110 such as cameras, infrared sensors, depth sensors, etc., and may therefore include red-green-blue (RGB) images, infrared images, and depth images, etc. It should be understood that this is not limiting, and the sample data in training data 120 may also be obtained in other ways, such as through data augmentation processing.

[0038] According to embodiments of this disclosure, computing device 130 may include an in-vehicle computing device (meaning computing device 130 may be located within vehicle 100), an external computing device, or a combination thereof. Computing device 130 may have computing capabilities adapted to perform object detection, on which AI / ML models for object detection can run. Computing device 130 is configured to perform performance evaluations for object detection according to embodiments of this disclosure. During the performance evaluation for object detection, computing device 130 may process visual information from vehicle 100 to form sample data in training data 120 and may access storage device 140 to perform corresponding operations. It should be understood that computing device 130 in… Figure 1 The example is shown as a computing device, but this is only illustrative and not limiting; a greater number of computing devices may be present in the example environment 100. The operation of the computing device 130 will be described in further detail below.

[0039] By way of example and not limitation, computing device 130 may include, but is not limited to, in-vehicle computers, personal computers, laptop computers, server computers, mobile devices (such as smartphones, tablets, etc.), wearable electronic devices, multimedia players, personal digital assistants (PDAs), smart home devices, consumer electronics products, or distributed computing environments that include any one or more of the above devices.

[0040] According to embodiments of this disclosure, storage device 140 can be configured to store training data 120, models and their parameters to be run on computing device 130, object detection results, performance reports, etc. It should be understood that storage device 140... Figure 1 The device shown is a storage device, but this is only illustrative and not restrictive; there may be many more storage devices in example environment 100.

[0041] By way of example and not limitation, storage device 140 may include, but is not limited to, local storage devices, remote storage devices, and combinations thereof. In some embodiments, the plurality of storage devices in storage device 140 may include, but is not limited to, hard disk drives (HDDs), solid-state drives (SSDs), etc., and some of the plurality of storage devices may be located locally, while others may be located remotely, for example, coupled together via lines or networks.

[0042] The above combination Figure 1 An example environment 100 in which methods and / or processes according to embodiments of the present disclosure may be implemented is described. The following will be combined with... Figure 2 This document describes a flowchart of a method 200 for performance evaluation of object detection according to embodiments of the present disclosure. This method 200 enables the evaluation of the model's performance on other datasets using a corresponding proxy model after the model has been optimized for a given dataset. This allows for the acquisition of comprehensive performance information of the original model without invoking it. Based on this comprehensive performance information, the original model can be retrained to address potential performance weaknesses, achieving a performance trade-off across multiple datasets.

[0043] Figure 2 A flowchart of a method 200 for performance evaluation of object detection according to an embodiment of the present disclosure is illustrated. At block 210, a first model for object detection is trained based on a first sample in a first dataset and a first ground truth label corresponding to the first sample. As described above, training data 120 is divided into multiple datasets (e.g., based on predetermined clustering strategies and classification rules, etc.), each of which includes sample data with certain commonalities. By way of example and not limitation, training data 120 may generally be image data corresponding to driving scenarios, and the first dataset may be the aforementioned dataset 121, which may include data samples about driving in rainy or snowy weather. The AI / ML model for performing object detection in driving scenarios may be customized and trained using dataset 121, such that the customized trained model has optimized performance when performing object detection in rainy or snowy weather application scenarios for dataset 121. It should be understood that customized training using a single dataset (or for a single application scenario) is described herein by way of example, and performance optimization can certainly be performed based on multiple datasets (or for multiple application scenarios).

[0044] At box 220, a second model is trained to evaluate the trained first model based on the first sample and the first ground truth label, and the first output of the trained first model for the first sample. The architecture of AI / ML models for object detection is often very complex. If one or more datasets to be judged are fed separately into the model for training, and then it is necessary to determine whether previous customized performance optimizations (e.g., for a specific dataset 121) lead to performance degradation for other datasets, and then to implement performance trade-offs for potential defects, this process often involves high time and computational costs. After all, in addition to the complexity of the model architecture, the amount of samples used for performance trade-offs is quite large, which undoubtedly increases the difficulty of the optimization process.

[0045] According to embodiments of this disclosure, a proxy model corresponding to the object detection model can be used. In some embodiments, the proxy model may have a simpler model architecture than the original model and be faster in terms of model iteration speed, while maintaining a certain level of object detection accuracy. Here, a second model will be trained using a first sample used to train the first model and ground truth labels corresponding to the first sample, as well as a first output of the trained first model for the first sample, so that the second model can mimic the reasoning behavior of the trained first model. The proxy model training process according to embodiments of this disclosure will be described in further detail below.

[0046] In box 230, the object detection performance of the trained first model on the second dataset is determined by validating the object detection performance of the second model on a second dataset different from the first dataset. In box 220, the surrogate model trained can mimic the inference behavior of the original model optimized for a specific dataset. Therefore, by validating the performance of this surrogate model on one or more other datasets different from the specific dataset, it can be determined whether the optimized original model experiences performance degradation on other datasets, and then the original model can be adjusted accordingly to achieve a performance trade-off. This process is targeted and efficient in improving performance.

[0047] The performance evaluation method for object detection according to embodiments of this disclosure enables the evaluation of a model's performance on other datasets or in different application scenarios after the model has been optimized for performance on one or more specific datasets or for one or more application scenarios. This allows for the acquisition of comprehensive performance information of the original model without invoking it. Furthermore, performance trade-offs can be specifically implemented to address potential shortcomings, improving training efficiency while reducing unnecessary resource consumption and ensuring the model maintains excellent performance across multiple datasets. The implementation of the performance evaluation process for object detection models according to embodiments of this disclosure will be described in further detail below.

[0048] Figure 3 The illustration shows a performance evaluation process 300 for an object detection model according to an embodiment of the present disclosure. Figure 3 As exemplarily illustrated, the performance evaluation process 300 may include a customized training subprocess 310, a first verification subprocess 320, a surrogate model training subprocess 330, a correlation testing subprocess 340, a second verification subprocess 350, and a performance analysis subprocess 360. This performance evaluation process 300 and its subprocesses can be abstracted into performance evaluation units and corresponding sub-units for each subprocess (e.g., customized training sub-unit, first verification sub-unit, etc.). These units and sub-units can be software-based components or systems for performance evaluation of object detection and can run on a computing device (such as computing device 130).

[0049] According to embodiments of this disclosure, in the customized training sub-process 310, the performance of the object detection model can be optimized in a specified aspect by custom training the model using a specific dataset. For example, the object detection model can achieve improved performance when performing inference in application scenarios corresponding to that specific dataset. By way of example and not limitation, the object detection model can be customized-trained using a dataset 121 that includes data samples about driving in rainy or snowy weather, so that the trained model has improved accuracy and robustness when performing object detection in rainy or snowy weather application scenarios.

[0050] According to embodiments of this disclosure, in the first verification sub-process 320, the performance of the customized trained object detection model can be verified. In some embodiments, the object detection model is optimized for dataset 121. In response to the customized trained model's performance on dataset 121 being higher than or equal to a predetermined performance threshold, a surrogate model training sub-process 330 can be started, i.e., training a surrogate model corresponding to the customized trained model. Conversely, in response to the customized trained model's performance on dataset 121 being lower than the predetermined performance threshold, meaning the customization of the model for dataset 121 has not met expectations, the customization training sub-process 310 can be returned to perform retraining of the object detection model on dataset 121. During such retraining, the training weight of dataset 121 in the entire dataset can be increased.

[0051] In some embodiments, the performance of an object detection model can be judged based on the intersection-over-union ratio (IoU). IoU is a key performance indicator that can indicate positional deviation; a larger IoU value means worse object detection performance. Its calculation method is given by the following equation (1):

[0052]

[0053] According to embodiments of this disclosure, in the surrogate model training sub-process 330, a surrogate model corresponding to the customized trained model can be trained for evaluating the customized trained model. As mentioned above, the original model may have a complex architecture and slow model iteration. Validating the performance of the original model across all datasets would result in high time costs and wasted computational resources. The surrogate model is used to replace this complex process, for example, by training the surrogate model on each dataset in these datasets and then validating its performance, in order to provide a faster and more economical solution.

[0054] A surrogate model can be trained using samples used to custom-train an object detection model, corresponding ground truth labels for those samples, and the output of the custom-trained model for those samples. This allows the surrogate model to mimic the reasoning behavior of the custom-trained model. In some embodiments, the surrogate model can learn the deviation between the ground truth labels corresponding to the samples used during the custom training of the original model and the output of the custom-trained original model for those samples to generate a trained surrogate model. Using the output of the trained surrogate model (i.e., the inferred deviation from the ground truth labels), the performance of the corresponding custom-trained original model can be estimated.

[0055] According to embodiments of this disclosure, in the relevance testing sub-process 340, it can be determined whether the proxy model trained in the proxy model training sub-process 330 and the original model customized in the customization training sub-process 310 satisfy predetermined conditions of consistency or relevance. If such predetermined conditions are met, the trained proxy model can be considered to correspond to the customized original model, and the second verification sub-process 350 is initiated to input other datasets different from dataset 121 into the trained proxy model for verification. If the predetermined conditions are not met, it can be considered that the trained proxy model cannot yet replace the customized original model, and the proxy model training sub-process 330 is returned to strengthen the proxy model's learning of the reasoning behavior of the customized original model. This will be discussed in conjunction with... Figure 4 and Figure 5 The predetermined conditions for consistency or correlation among the above models are described in further detail.

[0056] Figure 4The illustration shows a process 400 for determining whether models meet consistency or correlation requirements according to an embodiment of the present disclosure. At 410, a predetermined number of datasets are input to a customized-trained original model and a trained surrogate model, respectively. In some embodiments, the predetermined number may be significantly less than the number of all datasets (e.g., all datasets in training data 120). At 420, the difference between the deviation between the output and the corresponding ground truth label of the customized-trained original model for each of the predetermined number of datasets (i.e., the output deviation of the customized-trained original model for each of the predetermined number of datasets) and the difference between the output of the trained surrogate model for each of the predetermined number of datasets is determined. At 430, in response to the determined difference being less than a predetermined difference threshold, it is determined that the customized-trained original model and the trained surrogate model meet predetermined conditions for consistency or correlation. At 440, in response to the determined difference being greater than or equal to the predetermined difference threshold, it is determined that the customized-trained original model and the trained surrogate model do not meet the predetermined conditions.

[0057] Figure 5 The illustration shows another process 500 for determining whether models meet consistency or correlation requirements according to an embodiment of the present disclosure. The original model is pre-trained on all training data (e.g., training data 120) and then customized-trained on a specific dataset (e.g., dataset 121). At 510, a predetermined number of datasets are input to the customized-trained original model and the trained surrogate model, respectively. In some embodiments, this predetermined number may be significantly less than the number of all datasets. At 520, a first performance pattern of the customized-trained original model relative to its pre-training performance on the predetermined number of datasets is determined, i.e., a performance trend of the customized-trained original model on these datasets is determined. At 530, a second performance pattern of the trained surrogate model relative to its pre-training performance on the predetermined number of datasets is determined, i.e., a performance trend of the trained surrogate model on these datasets is determined. At 540, in response to a positive correlation between the first performance pattern and the second performance pattern, it is determined that the customized-trained original model and the trained surrogate model meet predetermined conditions for consistency or correlation. At 550, in response to the negative correlation or no correlation between the first performance mode and the second performance mode, it is determined that the original model trained with customization and the surrogate model trained do not meet the predetermined condition.

[0058] In some embodiments, the correlation between patterns can be determined based on the Pearson correlation coefficient, which is used to evaluate the degree of linear correlation between two variables. Its calculation method is given by the following equation (2):

[0059]

[0060] return Figure 3 In the second validation sub-process 350, the performance of the trained surrogate model is validated on one or more datasets to be judged (different from the datasets used in the previous customized training of the original model). The detection performance of the trained surrogate model on each of the one or more datasets to be judged can be validated separately. In the performance analysis sub-process 360, the performance of the customized trained original model on the one or more datasets to be judged is determined by analyzing the performance of the trained surrogate model on the one or more datasets to be judged. Through this analysis, potential shortcomings can be identified, and performance trade-offs can be implemented to address these shortcomings, ensuring that the model maintains excellent performance across multiple datasets.

[0061] Figure 6 A diagram illustrating the model architecture of the original model 610 and the proxy model 620 according to embodiments of the present disclosure is shown. Figure 6 As shown, the sample input 601 from one or more datasets to be judged is input into the original model 610 on a sample-by-sample basis. Based on the feature extraction layer 602 (also called the backbone) of the original model 610, feature vectors 603 for each channel can be extracted from the samples in these sample sets. The extracted feature vectors 603 include key information that can characterize the corresponding samples. The extracted feature vectors 603 can be dimensionality reduced by average pooling, for example, from multi-dimensional vectors to one-dimensional scalar values ​​604. In this way, the input is more suitable for the surrogate model 620.

[0062] According to embodiments of this disclosure, by inputting the converted scalar value 604 into the trained surrogate model 620, the output 605 of the trained surrogate model 620 can be obtained. During the aforementioned consistency or correlation verification phase, the difference between the deviation between the output 606 of the customized-trained original model 610 and the corresponding ground truth label and the output 605 of the trained surrogate model 620 can be used to determine whether replacement is possible. For example, if such a difference is less than a predetermined difference threshold, replacement is possible; otherwise, replacement is not possible. Here, based on the output of the trained surrogate model 620, the performance of the trained surrogate model 620 on one or more datasets to be judged is verified. By verifying the performance of the surrogate model 620 on one or more datasets to be judged, it is possible to determine whether the customized-trained original model has defects on these datasets, thereby achieving targeted performance trade-offs and ensuring that the model maintains excellent performance across multiple datasets.

[0063] Figure 7A schematic diagram of an apparatus 700 for performance evaluation of object detection according to an embodiment of the present disclosure is shown. The apparatus 700 may include multiple units or modules for performing, for example... Figure 2 The corresponding steps in method 200 discussed herein. For example... Figure 7 As shown, the apparatus 700 includes: a first training module 710 configured to train a first model for object detection based on a first sample in a first dataset and a first ground truth label corresponding to the first sample; a second training module 720 configured to train a second model for evaluating the trained first model based on the first sample, the first ground truth label, and a first output of the trained first model for the first sample; and a performance determination module 730 configured to determine the object detection performance of the trained first model for the second dataset by verifying the object detection performance of the trained second model for the second dataset, which is different from the first dataset.

[0064] In some embodiments, the apparatus 700 may further include a first verification module configured to verify a first performance of the trained first model for the first dataset; to start training the second model in response to the verified first performance being higher than or equal to a predetermined performance threshold; and to perform retraining of the first model for the first dataset in response to the verified first performance being lower than the predetermined performance threshold.

[0065] In some embodiments, the apparatus 700 may further include a training weight control module configured to increase the training weight of the first dataset in the entire dataset during the retraining of the first model.

[0066] In some embodiments, training the second model based on the first sample and the first ground truth label, and the first output of the trained first model for the first sample, may include generating the trained second model by having the second model learn the deviation between the first ground truth label corresponding to the first sample and the first output of the first model for the first sample.

[0067] In some embodiments, the apparatus 700 may further include a consistency or correlation verification module configured to: determine whether the trained second model satisfies a predetermined condition with the trained first model; and in response to the trained second model satisfying the predetermined condition with the trained first model, input the second dataset into the trained second model to verify the object detection performance of the trained second model for the second dataset.

[0068] In some embodiments, determining whether the trained second model and the trained first model satisfy the predetermined condition may include: inputting a predetermined number of multiple datasets into the trained first model and the trained second model respectively; determining the difference between the deviation between the output error and the corresponding ground truth label of the trained first model for each of the multiple datasets (i.e., the output deviation of the trained first model for each of the multiple datasets) and the output of the trained second model for each of the multiple datasets; determining that the trained second model and the trained first model satisfy the predetermined condition in response to the determined difference being less than a predetermined difference threshold; and determining that the trained second model and the trained first model do not satisfy the predetermined condition in response to the determined difference being greater than or equal to the predetermined difference threshold.

[0069] In some embodiments, wherein the first model is a pre-trained model, and determining whether the trained second model satisfies the predetermined condition with the trained first model may include: inputting a predetermined number of multiple datasets into the trained first model and the trained second model respectively; determining a first performance pattern of the trained first model relative to the pre-training performance for the multiple datasets; determining a second performance pattern of the trained second model relative to the pre-training performance for the multiple datasets; determining that the trained second model and the trained first model satisfy the predetermined condition in response to a positive correlation between the first performance pattern and the second performance pattern; and determining that the trained second model and the trained first model do not satisfy the predetermined condition in response to a negative correlation or no correlation between the first performance pattern and the second performance pattern.

[0070] In some embodiments, the apparatus 700 may further include a feature extraction and dimensionality reduction module configured to extract feature vectors for each channel from a second sample in the second dataset based on a feature extraction layer in the first model; and to convert the extracted feature vectors for each channel into scalar values ​​via average pooling.

[0071] In some embodiments, verifying the object detection performance of the trained second model for the second dataset may include: obtaining a second output of the trained second model by inputting the transformed scalar value into the trained second model; and verifying the object detection performance of the trained second model for the second dataset based on the second output of the trained second model.

[0072] In some embodiments, the first model can be used to detect objects in a driving environment, and the object detection performance of the trained first model and the second model can be determined based on the intersection-union ratio of the bounding box locations of objects in the driving environment output by the respective models.

[0073] In some embodiments, the first sample set may include the first sample associated with a first application in a driving scenario, and the second sample set may include the second sample associated with a second application in the driving scenario that is different from the first application. The device 700 may also include a driving object detection module configured to perform object detection corresponding to the second application based on the trained first model in response to a verified second performance of the trained first model for the second dataset being higher than or equal to a predetermined performance threshold, so as to detect the location and category, state and behavior of objects in the driving environment.

[0074] Figure 8 A schematic block diagram of an example device 800 suitable for implementing embodiments of the present disclosure is shown. The controller described above can be implemented using device 800. As shown, device 800 includes a processor 801 that can perform various appropriate actions and processes based on computer program instructions loaded into random access memory (RAM) 803 according to computer program instructions stored in read-only memory (ROM) 802. Various programs and data required for the operation of device 800 may also be stored in RAM 803. The processor 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0075] The various processes and procedures described above, such as method 200 and processes 400 and 500, may be executed by processor 801. For example, in some embodiments, method 200 and processes 400 and 500 may be implemented as computer software programs tangibly contained in a machine-readable medium. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802. When the computer program is loaded into RAM 803 and executed by processor 801, one or more actions of method 200 and processes 400 and 500 described above may be performed. According to embodiments of this disclosure, a vehicle is provided that may include device 800 as described above for performing various aspects of this disclosure.

[0076] This disclosure can be a method, apparatus, electronic device, vehicle, computer-readable storage medium, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of this disclosure.

[0077] Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage media used herein are not to be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.

[0078] The computer-readable program instructions described herein can be downloaded from computer-readable storage media to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to the computer-readable storage media in the respective computing / processing device.

[0079] Computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized by utilizing the status information of the computer-readable program instructions to implement various aspects of this disclosure.

[0080] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0081] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0082] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0083] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0084] The various embodiments of this disclosure have been described above. These descriptions are exemplary and not exhaustive, and are not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method (200) for performance evaluation of object detection, comprising: Based on the first sample in the first dataset and the first ground truth label corresponding to the first sample, train (210) a first model for object detection; Based on the first sample and the first ground truth label, and the first output of the trained first model for the first sample, a second model is trained (220) to evaluate the trained first model; as well as The object detection performance of the trained first model on the second dataset is determined by verifying the object detection performance of the trained second model on a second dataset that is different from the first dataset.

2. The method (200) according to claim 1, further comprising: Verify the first performance of the trained first model on the first dataset; In response to the verified first performance being higher than or equal to a predetermined performance threshold, training of the second model begins; as well as In response to the verified first performance being lower than the predetermined performance threshold, the first model is retrained for the first dataset.

3. The method (200) according to claim 2, further comprising: During the retraining of the first model, the training weight of the first dataset in the entire dataset is increased.

4. The method (200) of claim 1, wherein training the second model based on the first sample and the first ground truth label, and the first output of the trained first model for the first sample comprises: The trained second model is generated by having the second model learn the deviation between the first ground truth label corresponding to the first sample and the first output of the first model for the first sample.

5. The method (200) according to claim 1, further comprising: Determine whether the trained second model satisfies predetermined conditions with the trained first model; as well as In response to the predetermined condition being met by the trained second model and the trained first model, the second dataset is input into the trained second model to verify the object detection performance of the trained second model for the second dataset.

6. The method (200) of claim 5, wherein determining whether the trained second model satisfies the predetermined condition with the trained first model comprises: A predetermined number of multiple datasets are respectively input into the trained first model and the trained second model; Determine the difference between the deviation between the output and the corresponding ground truth label of the first model trained for each of the plurality of datasets and the difference between the output of the second model trained for each of the plurality of datasets; In response to the determined difference being less than a predetermined difference threshold, it is determined that the trained second model and the trained first model satisfy the predetermined condition; as well as In response to the determined difference being greater than or equal to the predetermined difference threshold, it is determined that the trained second model and the trained first model do not satisfy the predetermined condition.

7. The method (200) according to claim 5, wherein the first model is a pre-trained model, and determining whether the trained second model satisfies the predetermined condition with the trained first model comprises: A predetermined number of multiple datasets are respectively input into the trained first model and the trained second model; Determine a first performance pattern of the trained first model relative to its pre-training performance for the plurality of datasets; Determine a second performance pattern of the trained second model relative to its pre-training performance for the plurality of datasets; In response to the positive correlation between the first performance mode and the second performance mode, it is determined that the trained second model and the trained first model satisfy the predetermined condition; as well as In response to the first performance mode being negatively correlated or uncorrelated with the second performance mode, it is determined that the trained second model and the trained first model do not satisfy the predetermined conditions.

8. The method (200) according to claim 1, further comprising: Based on the feature extraction layer in the first model, feature vectors for each channel are extracted from the second samples in the second dataset. The feature vector of each extracted channel is converted into a scalar value via average pooling.

9. The method (200) of claim 8, wherein verifying the object detection performance of the trained second model on the second dataset comprises: The second output of the trained second model is obtained by inputting the transformed scalar value into the trained second model; Based on the second output of the trained second model, the object detection performance of the trained second model on the second dataset is verified.

10. The method (200) according to claim 1, wherein: The first model was used to detect objects in the driving environment, and The object detection performance of the trained first and second models is determined based on the intersection-union ratio of the object bounding box locations in the driving environment, as output by the respective models.

11. The method (200) of claim 1, wherein the first dataset includes the first samples associated with a first application in a driving scenario, and the second dataset includes the second samples associated with a second application in the driving scenario that is different from the first application, and the method (200) further includes: In response to the verified second performance of the trained first model for the second dataset being higher than or equal to a predetermined performance threshold, object detection corresponding to the second application is performed based on the trained first model to detect the location and category, state and behavior of objects in the driving environment.

12. An apparatus (700) for performance evaluation of object detection, comprising: The first training module (710) is configured to train a first model for object detection based on a first sample in the first dataset and a first ground truth label corresponding to the first sample; The second training module (720) is configured to train a second model for evaluating the trained first model based on the first sample and the first ground truth label, and the first output of the trained first model for the first sample. as well as The performance determination module (730) is configured to determine the object detection performance of the trained first model for the second dataset by verifying the object detection performance of the trained second model for a second dataset different from the first dataset.

13. A controller, comprising: At least one processor; as well as A memory, coupled to the at least one processor and having instructions stored thereon, which, when executed by the at least one processor, cause the controller to perform the method according to any one of claims 1-11.

14. A vehicle comprising the controller according to claim 13.

15. A computer program product tangibly stored on a computer-readable medium and comprising computer-executable instructions that, when executed by a processor of a computer, cause the computer to perform the method according to any one of claims 1 to 11.