Data efficient ai / ML training

The data processing unit and method efficiently address the challenges of redundant and mislabeled data in AI/ML training by automatically identifying and removing noncompliant images, leading to cost-effective and accurate training.

WO2025124712A1PCT designated stage expired Publication Date: 2025-06-19HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2023/085671
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-13
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Existing AI/ML training methods face challenges in efficiently reducing data without incurring high computation costs, and they often rely on manual intervention for correcting mislabeled samples, which is time-consuming and expensive.

Method used

A data processing unit and method that computes soft labels for images, identifies noncompliant images based on entropy values and label discrepancies, and automatically modifies the dataset by removing redundant and mislabeled samples, thereby forming a modified image dataset for efficient training.

Benefits of technology

This approach enables efficient data pruning, automatic rectification of mislabeled samples, and reduction of outlier samples, resulting in cost-effective and energy-efficient AI/ML training with minimal loss in classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2023085671_19062025_PF_FP_ABST
    Figure EP2023085671_19062025_PF_FP_ABST
Patent Text Reader

Abstract

A data processing unit for data pruning in training a weighted processing network; the device comprising one or more processors configured to: receive an image dataset comprised of a plurality of annotated images comprising one or more image and one or more associated image labels; compute a soft label for each of the plurality of annotated images in the image dataset, the soft labels representing an object class information within the plurality of annotated images; identify one or more noncompliant images of the plurality of annotated images in the image data set, based on the computed soft label of each of the plurality annotated images in the received image dataset; modify the image dataset based on the identified one or more noncompliant image to form a modified image dataset comprised of fewer annotated images than the plurality of annotated images in the received image dataset.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DATA EFFICIENT AI / ML TRAINING

[0002] TECHNICAL FIELD

[0003] This present disclosure relates to a data processing unit for data pruning in training a weighted processing network. The present disclosure also relates to a method of processing one or more pass for heterogeneous pipeline processing of data.

[0004] BACKGROUND

[0005] Computer Vision (CV) is a field of Artificial Intelligence (Al) that enables computers and systems to process and derive meaningful information from digital images, videos, other visual inputs - and take actions or make recommendations based on that input. Common tasks in CV include image classification, object detection, semantic segmentation, among others.

[0006] Neural network-based models for CV have evolved rapidly in recent years, driven by the advances in computing power, the development of new algorithms, and availability of large datasets. The increase of computing power prompted the development of deeper and more complex models - commonly referred to as Deep Neural Networks (DNNs).

[0007] Image classification is the task of assigning a label to an image from a predefined set of categories / labels. Convolutional Neural Networks (CNNs) and Transformer-based architectures specialized for images (e.g., Vision Transformers or ViTs) are types of DNNs that have achieved state-of-the-art performance on a variety of image classification benchmarks. Popular CNNs for image classification include Residual Networks or ResNets.

[0008] Efficient artificial intelligence (AI) / machine learning (ML) training is the process of training AI / ML models with minimal resources, such as with less data, time, energy, and computing power. Data provides AI / ML systems with examples of the real world. AI / ML systems learn by identifying patterns in the data - a process referred to as model training or training. Large-scale training of DNNs for image classification often involves millions to billions of annotated samples. However, not all samples contribute to the training of the deep learning model . Instead, a significant portion of the training samples are redundant, and their removal or pruning would not result in any noticeable loss in the classification accuracy. In addition to the redundant samples, datasets can often contain training samples with incorrect labels. These errors in labels can arise from the subjective nature of human annotation or due to the confusing or uncharacteristic nature of the sample. AI / ML models trained on such incorrect labels can lead to erroneous prediction during test time. Other techniques for efficient AI / ML training include Model Pruning and Quantization. These techniques are used to reduce the size and complexity of AI / ML models making the model more efficient to train and deploy. However, for training, these techniques are mostly limited by their inability to efficiently reduce samples without incurring a significantly high computation cost. Other solutions are limited by the choice of hardware.

[0009] Some work has been done to address the above issues, in a first example it is known to provide a method for identifying mislabelled samples (i.e., samples with incorrect labels, confusing samples sharing characteristic with multiple classes) that are present in the training data for AI / ML deep neural network models. Using mislabelled training samples fortraining of AI / ML models is undesirable as it can cause the models to generate erroneous results. The flow diagram shown in Fig. 1 demonstrates a known method for identifying mislabelled samples. During training, an AI / ML model is trained on the available data. Adversarial attack is a standard technique in ML which is used to obtain incorrect output from an AI / ML model by designing the input in a specific way. The system of Fig. 1 performs adversarial attacks on the training samples using varying levels of strength (controlled by a parameter). Following the adversarial attack, misguided samples in the dataset (i.e., the AI / ML model outputs an incorrect class label) are identified and sorted based on the strength of the attack. Among these, the samples for which the strength of the adversarial attack is below a pre-defined threshold are identified as mislabelled samples.

[0010] A second known method for identifying label errors made during annotation is illustrated in the flow chart of Fig. 2. The system uses an AI / ML model which is first trained on an annotated dataset which might contain labelling errors. Next, the probability level at the output for each training sample is compared with two user-specified thresholds - first and second confidence levels. Samples are identified to be potentially mislabelled (i) if the predicted label probability is higher than the first confidence level and the predict label is different from the annotator-assigned label, or (ii) if the predicted label probabilities are at a level less than a second confidence level and the predicted label is identical to the annotator-assigned label. These mis-labelled samples are first reviewed and re-labelled manually by annotators. Post label-rectification, these samples are re-introduced for training.

[0011] There is also a third known technique for using two metrics to quantify the importance of each training sample. These metrics are the Gradient Normed (GraNd) and the Error L2-Norm (EL2N) scores, used to identify important training samples and thereby remove redundant samples for the image classification task. These scores, when averaged over several weight initializations, can be used to prune a considerable fraction of training samples early (i.e., after a few epochs) in training. The GraNd score is well-approximated by the EL2N score and provides a stronger signal for pruning. Consequently, the EL2N score is used to rank and prune samples. Samples with low EL2N scores are identified as redundant samples and are pruned and not used for the training of the AI / ML model. A flow diagram demonstrating this approach is shown in Figure. 3. In this flow diagram it can be seen that multiple models are trained for many epochs in order to compute the EL2N scores which may be found by the expression:

[0012] E||p (wt,x) - y||2

[0013] Where (x, y) refer to an image-label pair, w_t is a vector representing weights of the neural network at t>0, p(w_t,x) refers to the output of the neural network. The EL2N score (Error L2 Norm) is defined by the norm of the error vector under the cross-entropy loss.

[0014] The flow diagram of Fig. 3 then prunes the samples that are below a threshold.

[0015] A fourth known method of quantifying the importance of a training sample using a “forgetting score”. A “forgetting event” for a training sample is a point in training when the Al image classification model switches the output label from a correct classification label to an incorrect one. The “forgetting score” for each training sample is then defined as the number of times it went through a forgetting event during training. The forgetting scores are typically computed at the end of the training process. In this method it is observed that (i) a large number of training samples are never forgotten once learnt, i.e., low forgetting scores, and (ii) training samples with noisy labels and visually complicated images (i.e., images with uncommon or rare features) often demonstrate high forgetting scores. Training samples with low forgetting scores, can be completely omitted during training without any noticeable degradation of the classification accuracy. The complete flow diagram for this method is shown in Fig. 4, in which a model is trained, forgetting scores are computed in order to remove samples and then train a model with the remaining samples.

[0016] These approaches can be categorised in two ways, firstly into systems that identify and remove mislabelled samples from the training dataset as in the first and second approaches, and secondly into systems that identify and remove redundant samples from the training dataset as in the third and fourth approaches. Each category has associated disadvantages.

[0017] In the first and second approaches identification and removal of mislabelled samples does not target efficient AI / ML model training. Instead, the primary focus of these approaches is to mitigate the negative impacts (i.e., incorrect classification during test-time) caused by training using incorrect labels. Furthermore, the first known approach relies on evaluating the logits at the end of a training cycle. Overparameterized networks (i.e., DNNs with large number of model parameters) can often memorize mislabelled samples making them not susceptible to adversarial perturbations. The mislabelled sample thus remains undetected. Rectification of mislabelled samples is either not present or done via manual review in the first and second approaches described above. Manual review is time-consuming and expensive as expert annotators might need to re-visit large amount of data. Similarly, in the approaches three and four discussed above, the GraNd and EL2N scores are unreliable at the early stages of training from a single training process. Accordingly, multiple instantiations (i.e., 10-20 models) of the training process are needed for reliable scores which in turn increases the cost of training multiplicatively (i.e., N models lead to N-times cost). In the fourth approach it is observed that samples with noisy labels tend to have high forgetting scores. However, this approach does not propose a means of correcting / re-labelling such samples. Instead, these training samples are used as-is which might negatively impact the training process. Only redundant samples with low forgetting scores are removed. Although the “forgetting score” accounts for the training dynamics (i.e., evolution of logit values during training), it does so at discrete steps when “forgetting events” occur. Consequently, this results in poor performance when it comes to data pruning. Furthermore, forgetting scores are computed at the end of the complete training process. This results in high training and GPU compute cost. Finally, known the third and fourth known approaches incur significantly high computational overhead, either via multiple instantiations for reliable score estimations or due to the requirement of a longer training window when pruning samples. Overall, the inability of prior systems to automatically rectify incorrectly labelled samples and to efficiently prune redundant samples hinder their applicability.

[0018] Existing systems targeting data-efficient training of AI / ML systems do not fully leverage the training dynamics (i.e., the logit trajectories) for the task of pruning. This inefficient utilization results in existing solutions incurring significantly higher computation costs and / or parallel instantiations. Furthermore, existing solutions that target efficient training are hardware specific (2:4 sparsity) and do not generalize to all CPU / GPU resources.

[0019] The present cost of training a single large Al model is currently expensive and the cost of training a model on a large dataset can be even more expensive, into the tens of millions of dollars. OpenAI estimates that the cost of training a model on a large dataset will increase into the hundreds of millions by 2030. This is due to the increasing size of datasets, as well as the need for more computing power to train larger models.

[0020] SUMMARY

[0021] According to a first aspect of this disclosure there is provided a data processing unit for data pruning in training a weighted processing network; the device comprising one or more processors configured to: receive an image dataset comprised of a plurality of annotated images comprising one or more image and one or more associated image labels; compute a soft label for each of the plurality of annotated images in the image dataset, the soft labels representing an object class information within the plurality of annotated images; identify one or more noncompliant images of the plurality of annotated images in the image data set, based on the computed soft label of each of the plurality annotated images in the received image dataset; modify the image dataset based on the identified one or more noncompliant image to form a modified image dataset comprised of fewer annotated images than the plurality of annotated images in the received image dataset. This allows the removal of easy samples, automatic rectification of mis-labelled samples, and the removal of outlier samples all in an end-to-end framework.

[0022] The data processing unit as described above, wherein the one or more processors are further configured to: in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, compute an entropy value for each of the plurality annotated image based on the computed soft label; and determine that the computed entropy value is equal to or below a predetermined threshold value; wherein the entropy value represents the confidence of assigning an image to one of a plurality of classes; and in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, remove the identified one or more noncompliant image from the received data set when forming the modified data set. This allows uninformative images to be removed from the image dataset.

[0023] The data processing unit as described above, wherein the one or more processors are further configured to: in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, determine whether the soft label for each of the plurality of annotated images is the same as the associated image label; and in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, automatically re-label the associated image label to be the soft label when the associated image label does not match the computed soft label. This allows for correction of mislabelled images in the image dataset.

[0024] The data processing unit as described above, wherein the one or more processors are further configured to: in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, determine a maximum class identification for each of the annotated images in the image data set by computing a maximum class identification to a threshold value, where the maximum class identification is based on the computed soft label for the annotated image; determine whether the maximum class identification is equal to or below a predetermined threshold value, where the image is determined to be noncompliant if the maximum class value is equal to or below the predetermined threshold; and in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, remove the identified one or more noncompliant image from the received data set when forming the modified data set. This allows for images which may be assigned to multiple classes and are not determined to be assignable to a single class to be removed from the image dataset.

[0025] The data processing unit as described above, wherein the entropy value is determined based on the following expression:

[0026] Where H is the entropy value, y is the soft label, i is the image class, c is the number of image classes and p is probability that an object in the image is in an image class. This provides a means of identifying an entropy of the images in an image dataset based on a confidence factor related to the class associated to the image.

[0027] The data processing unit as described above, wherein the soft labels are computed for each of the plurality of annotated images in the image dataset based on the following expression:

[0028] Wherein y is the soft label, c(. ) is the soft-max activation function, T is the number of epochs, (x,y)eDtrain is an image-label pair in the training dataset, c is the number of image classes and z(t)(x) is the output logits. This provides the new variable associated with each image in the image dataset to be generated which can then be used for identification of compliant and noncompliant images.

[0029] The data processing unit as described above, wherein the soft label for each of the plurality of annotated images is the same as the associated image label when the following expression is true, argmax (y) = y

[0030] Where y is the soft label and y is the associated image label. This allows for determination of the correct labelling of the images in the image dataset; and the soft label for each of the plurality of annotated images is not the same as the associated image label and the following expression is true, argmax (y) #= y

[0031] Where y is the soft label and y is the associated image label.

[0032] The data processing unit as described above, wherein the one or more processors are further configured to: input the received image dataset to an image classification model for identifying classes present in the images of the image data set; wherein the image classification model is configured to, across one or more epoch, iteratively: generate one or more logits for each image of the image dataset, the logits representing pre-classification un-normalized values across each class; output the one or more logits for each image of the image dataset at each epoch. This allows for logits to be generated that may be used to generate the soft labels used to identify images in the image dataset. The data processing unit as described above, wherein the one or more processors are further configured to: output the modified image data set for training of a weighted processing network for image classification. This provides a means for the data processing unit to distribute the modified image dataset.

[0033] The data processing unit as described above, wherein the one or more processors are further configured to: input the modified image dataset to a weighted processing network for image identification, the weighted processing network comprising parameters; and train the parameters of the weighted processing network using the modified image dataset. This provides the use of the modified image dataset for image classification.

[0034] According to a further aspect of this disclosure there is provided a method of for data pruning in training a weighted processing network; the method comprising: receiving an image dataset comprised of a plurality of annotated images comprising one or more image and one or more associated image labels; computing a soft label for each of the plurality of annotated images in the image dataset, the soft labels representing an object class information within the plurality of annotated images; identifying one or more noncompliant images of the plurality of annotated images in the image data set, based on the computed soft label of each of the plurality annotated images in the received image dataset; modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset comprised of fewer annotated images than the plurality of annotated images in the received image dataset. This allows the removal of easy samples, automatic rectification of mis-labelled samples, and the removal of outlier samples all in an end-to-end framework.

[0035] The method as described above, wherein the method further comprises: in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, computing an entropy value for each of the plurality annotated image based on the computed soft label; and determining that the computed entropy value is equal to or below a predetermined threshold value; wherein the entropy value represents the confidence of assigning an image to one of a plurality of classes; and in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, removing the identified one or more noncompliant image from the received data set when forming the modified data set. This allows uninformative images to be removed from the image dataset.

[0036] The method as described above, wherein the method further comprises: in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, determining whether the soft label for each of the plurality of annotated images is the same as the associated image label; and in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, automatically re-labelling the associated image label to be the soft label when the associated image label does not match the computed soft label. This allows for correction of mislabelled images in the image dataset.

[0037] The method as described above, wherein the method further comprises: in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, determining a maximum class identification for each of the annotated images in the image data set by computing a maximum class identification to a threshold value, where the maximum class identification is based on the computed soft label for the annotated image; determining whether the maximum class identification is equal to or below a predetermined threshold value, where the image is determined to be noncompliant if the maximum class value is equal to or below the predetermined threshold; and in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, removing the identified one or more noncompliant image from the received data set when forming the modified data set. This allows for images which may be assigned to multiple classes and are not determined to be assignable to a single class to be removed from the image dataset.

[0038] The method as described above, wherein the soft labels are computed for each of the plurality of annotated images in the image dataset based on the following expression:

[0039] Wherein y is the soft label, c(. ) is the soft-max activation function, T is the number of epochs, (x,y)e Dtrain is an image label pair in the training dataset, c is the number of image classes and z^(x) is the output logits. This provides the new variable assocaited with each image in the image dataset to be generated which can then be used for identification of compliant and noncompliant images.

[0040] The method as described above, wherein the soft label for each of the plurality of annotated images is the same as the associated image label when the following expression is true, argmax (y) = y

[0041] Where y is the soft label and y is the associated image label. This allows for determination of the correct labelling of the images in the image dataset; and the soft label for each of the plurality of annotated images is not the same as the associated image label and the following expression is true, argmax (y) #= y

[0042] Where y is the soft label and y is the associated image label. The method as described above, wherein the entropy value is determined based on the following expression:

[0043] Where H is the entropy value, y is the soft label, i is the image class, c is the number of image classes and p is probability that an object in the image is in an image class. This provides a means of identifying an entropy of the images in an image dataset based on a confidence factor related to the class associated to the image.

[0044] The method as described above, wherein method further comprises: inputting the received image dataset to an image classification model for identifying classes present in the images of the image data set; wherein the image classification model is configured to, across one or more epoch, iteratively: generating one or more logits for each image of the image dataset, the logits representing preclassification un-normalized values across each class; outputting the one or more logits for each image of the image dataset at each epoch. This allows for logits to be generated that may be used to generate the soft labels used to identify images in the image dataset.

[0045] The method as described above, wherein method further comprises: outputting the modified image data set for training of a weighted processing network for image classification. This provides a means for the data processing unit to distribute the modified image dataset.

[0046] The method as described above, wherein the method further comprises: inputting the modified image dataset to a weighted processing network for image identification, the weighted processing network comprising parameters; and training the parameters of the weighted processing network using the modified image dataset. This provides the use of the modified image dataset for image classification.

[0047] BRIEF DESCRIPTION OF THE FIGURES

[0048] The embodiments of the present disclosure will now be described by way of example with reference to the accompanying drawings. In the drawings:

[0049] Fig. 1 illustrates an example of a first approach for identifying mislabelled samples;

[0050] Fig. 2 illustrates an example of a second approach for identifying image labelling errors;

[0051] Fig. 3 illustrates an example of a third approach for quantifying the importance of a training sample;

[0052] Fig. 4 illustrates an example of a fourth approach for quantifying the importance of atraining sample; Fig. 5 illustrates an example of a general approach for generating a pruned training sample dataset according to this disclosure:

[0053] Fig. 6 illustrates an example of a workflow for generating a training image data set according to this disclosure.

[0054] DETAILED DESCRIPTION

[0055] Large datasets can create bottleneck in the training process of AI / ML systems due to the heightened demand for storage and compute resources. However, a significant portion of the data is redundant and does not contribute effectively to the model’s test accuracy. Furthermore, the presence of incorrect labels and outlier samples negatively impact the model performance.

[0056] This disclosure provides methods and techniques to overcome the above problems discussed above. The idea of using data pruning techniques for efficient training of AI / ML models for image classification represents a key concept described in this disclosure. Specifically, this dsiclosure proposes a apparatus and method to reduce the demand for storage and compute time / cost and enable efficient training by removing redundant / superfluous samples, identifying and correcting mislabeled samples, and removing outlier / ambiguous samples.

[0057] The data processing unit and method of data processing will now be described in relation to Figs. 5 and 6. In a general application case of the data processing unit and method of this disclosure may be seen in Fig. 5.

[0058] In a typical AI / ML model training (e.g., in cloud servers) scenario for image classification, collected images are first annotated and then directly used in training. However, not all annotated images are important for the model. A significant portion of samples can be redundant adding no value to the training. Furthermore, mislabeled and outlier / ambiguous samples can negatively impact the training. The method and data processing unit dsiclosed herein may serve as a pre-processing step in the cloud servers to reduce the number of training samples for training thereby reducing storage and compute costs. The data processing unit and method of this disclosure leverage the complete logit trajectories (i.e., for all samples for all epochs) for efficient data pruning. Similarly, automated label rectification proposed herein is new. The ability of the system to re-use labels by fixing labeling errors removes the need for expensive and time-consuming manual intervention.

[0059] As shown in Fig. 5, one or more devices, in this case a robot fleet 501, may collect one or more unlabelled images that will eventually be used for training. The collected unlabelled images may be annotated 502 with associated image labels to form an image dataset comprised of a plurality of annotated images comprising one or more image and one or more associated image labels. This may be done automatically using an image database, manually by a user, or a combination of both methods. The image dataset that is produced may then be input to the data processing apparatus of this disclosure that may be used for data pruning 503. The data processing unit and method of this disclosure may therefore be comprised of one or more processors that may be configured to receive an image dataset comprised of a plurality of annotated images 502 comprising one or more image and one or more associated image labels. The one or more processors may further be configured to compute a soft label for each of the plurality of annotated images in the image dataset, the soft labels representing an object class information within the plurality of annotated images.

[0060] The data processing unit, comprising one or more processors may also be configured to identify one or more noncompliant images of the plurality of annotated images in the image data set, based on the computed soft label of each of the plurality annotated images in the received image dataset. In other words, noncompliant images may be thought of as images that within the image dataset that are either erroneously labelled and in need of correction, or are considered redundant and should be removed from the image dataset prior to training of the model.

[0061] Once such noncompliant images have been identified the one or more processors may be configured to modify the image dataset based on the identified one or more noncompliant image to form a modified image dataset comprised of fewer annotated images than the plurality of annotated images in the received image dataset. In other words, the one or more processors of the data processing unit may be configured to alter the image dataset to produce a modified data set in dependence on the identified noncompliant images by removing images from the image dataset and / or relablelling other images in the image data set. In this way the image data set can be pruned by the data processing unit and method of this disclosure to produce a training dataset (the modified image dataset) comprised of N samples that is less than the orignal number of M samples present in the image dataset as shown in Fig. 5.

[0062] The modified image dataset may then be output by the data processing unit to be used for training a model 504 to produce a trained model that in some cases can be used for image classification.

[0063] A more detailed version of Fig. 5 can be seen in Fig. 6 that will now be described. In a general sense the system of Fig. 6, that includes the data processing unit (data pruning) according to this disclosure, includes an image annotation module that acts as the entry-point for unlabelled images received from one or more image capturing devices. Upon initiation, this module may receive a set of unlabelled images captured by the robot fleets (or equivalent image capturing devices). For supervised training of an Al model for image classification, these images may be first labelled by the data annotators using the annotation tool. Upon completion, the image-label pairs may be stored in the image database 602.

[0064] Next, this annotated image dataset may be used, by the one or more processors of the data processing unit of this disclosure, to train a data pruning Al model. This may be a dedicated model reserved only for the task of data pruning. On completion of training, this Al model may be used by the one or more processors of the data processing unit to identify and remove redundant samples, automatically rectify incorrect labels, and remove outlier samples from the image dataset. In doing this the Al model may identify noncompliant images within the image dataset for pruning. The term noncompliant as used herein should be understood to mean images that are not desired to be present in the final training image dataset. Such images may be incorrectly or erroneously labelled, or may be outliers or duplicates of other images within the dataset. Conversely, compliant images may be understood to be images that are unique or correctly labelled (annotated) images that would be beneficial in producing a coherent training image dataset (modified image dataset).

[0065] The modified image dataset may then be split into training and validation subsets by a data validator and in some cases may be stored in a dataset database. Finally, in the Al model management module, the training image dataset from the dataset database may be used to train the Al model for image classification for consequent deployment on the robots. This trained model may then be benchmarked using the validation set; if the current model out-performs the existing model on some performance indicator (e.g., classification accuracy), it may be deployed to the robots (or equivalent end-devices) using a model provisioning module.

[0066] Each of the components shown in Fig. 6 will now be described in further detail.

[0067] The initial images may be collected by a robot fleet which is a collection of robots, which may be considered any equivalent image acquisition devices and perform the image classification task. The robots in the robot fleet, may either on request or periodically, send collected images to an image database in an image annotation system. For performing image classification, the robot fleet uses AI / ML models provisioned by an Al model management system.

[0068] The images collected by the robot fleet or equivalent image capture devices once stored in a database may then be annotated with labels. This may be performed by a data annotator which may be a person or group of persons who curates the datasets using an image annotation system by assigning labels (from a pre-defined set of labels) to the images captured and sent by the robots. The image annotation system may produce an image dataset of annotated images for the task of image classification from a set of collected unlabelled images. The image annotation system may comprise: an image database that may initially store the images sent by the robot fleets. After the label assignment, the database may store a set of (image, label) tuples. The label can either be a text or integer. The image annotation system may further comprise an image storage interface 603 that provides methods and routines to manage and access the images stored in the image database. The image annotation system may further include an annotation tool 604 that may enable the data annotators to assign labels to the collected images. This tool 604 can be a user interface (UI) which shows an image to each individual annotator and the list of possible classes (fixed throughout the process) to assign it from. The annotator 605 may select one of the available classes and assigns it to the image. At the end of each annotation cycle, for each image, the tool may assign the most commonly assigned label to an image by the data annotators 605. The APIs 611 may provide a set of endpoints to share data between the different components in the system (image annotation, data management, Al model management). The image dataset that has been annotated may be split into training and validation datasets for AI / ML model training and validation respectively. This may be performed by a data annotator.

[0069] The example configuration of Fig. 6 then demonstrates the image dataset management process 620. Although the one or more processors of the data processing unit may also be configured to perform the functions of the image annotation described above and the Al model management 630 described below, the primary function of the data processing unit of this disclosure may be to perform data pruning of the image dataset as part of the data management process 620 described herein. The image dataset management may be a system that manages the storage, pruning, and splitting of image datasets.

[0070] The data processing unit of this disclosure may comprise a data pruning algorithm 621 module that may perform the removal of uninformative and ambiguous samples and rectifies mislabelled samples from the annotated dataset resulting in a clean dataset. Such uninformative, ambiguous samples and mislabelled samples may be considered noncompliant images within the image dataset. As such the data processing unit may includes one or more processors that are configured to receive an image dataset comprised of a plurality of annotated images comprising one or more image and one or more associated image labels. The one or more processors of the data processing unit may be configured to provide a data pruning model training 622 which may be an Al image classification model that is trained specifically for the task of data pruning and is not used as a model for the robots (or equivalent end devices). The AI / ML model 622 for data pruning may be trained on the image data set comprised of collected and annotated images using any suitable ML technique. During this training, the AI / ML model implemented by the one or more processors may be configured to generate as output, the logits (preclassification un-normalized values across each class), for each sample present in the training for each epoch. More formally, the AI / ML model may generate output logits z^(x) G Rcat epoch t, where c indicates the number of classes. In other words, the one or more processors are further configured to: input the received image dataset to an image classification model for identifying classes present in the images of the image data set. The image classification model may be configured to, across one or more epoch, iteratively: generate one or more logits for each image of the image dataset, the logits representing pre-classification un-normalized values across each class and output the one or more logits for each image of the image dataset at each epoch.

[0071] The data processing unit and method of the present disclosure may be configured to compute a soft label for each of the plurality of annotated images in the image dataset, the soft labels representing an object class information within the plurality of annotated images. The noncompliant images may be identified in a number of ways that will be described below based on some generated soft labels. Once the noncompliant images have been identified the one or more processors may be configured to modify the image dataset based on the identified one or more noncompliant image to form a modified image dataset comprised of fewer annotated images than the plurality of annotated images in the received image dataset.

[0072] The soft labels may be generated in the following manner by a response aggregator 623 that may for part of the data processing unit or be carried out externally. In this example, the one or more processors may be configured to implement a response aggregator that may generate the soft-labels (distribution over all c classes) for each sample in the dataset by (i) accumulating (via summation), (ii) then averaging the logit responses for each sample across the epochs, (iii) and finally applying soft-max to the logits, resulting in a single soft-label y per sample. The soft label for each image sample may be determined based on an expression of the form: where (x, y) G ©train refers to (image, label) pair, z^(x) are the output logits, and c(. ) refers to the soft-max activation function, and c is the number of image categories / classes. In other words, y is the soft label, c(. ) is the soft-max activation function, T is the number of epochs, (x,y)GDtram is an image, label pair, c is the number of image classes and z^(x) is the output logits.

[0073] These generated soft labels may be used by the one or more processors and as part of the method disclosed herein in the identification of noncompliant images. As shown in Fig. 6 there are three main ways 624, 625, 626 that the noncompliant images may be identified. The first of these is an easy sample removal process 624, which acts as a removal module that may compute the entropy H of each sample using the generated soft-label and removes samples with low entropy scores based on a threshold. The entropy value represents the confidence of assigning an image to one of a plurality of classes. The entropy value may be computed on a per-image basis and may represent the confidence of the classifier in assigning an image to one of N classes. An example scenario is described in which there is image 1 and image 2, and two classes cl and c2. Image 1 can be in either cl or c2, and likewise for image 2. If the classifier is confident for image 1, it may assign a probability of 1.0 to cl and 0 (or a very low value) to c2, leading to an entropy of value 0 (i.e., low entropy). If the classifier is not confident for image 2, it may assign an equal probability of 0.5 and 0.5 to cl and c2 respectively, leading to an entropy score of 1 (i.e., high entropy). In the algorithm disclosed herein algorithm, easier samples (i.e., like image 1) are treated as being redundant and may therefore be removed. The entropy value may be determined based on the following expression:

[0074] Where H is the entropy value, y is the soft label, i is the image class, c is the number of image classes and p is probability that an object in the image is in an image class.

[0075] In such a case the one or more processors may be further configured to in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, compute an entropy value for each of the plurality annotated image based on the computed soft label. The one or more processors may then determine that the computed entropy value is equal to or below a predetermined threshold value. The processors may then in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, remove the identified one or more noncompliant image from the received data set when forming the modified data set.

[0076] A further approach that the one or more processors may undertake in order to identify noncompliant images may be label correction 625 which may compare the generated soft-labels with the annotated label. Based on the discrepancy, annotated labels are updated to the most likely class in the soft-label. In other words, the one or more processors are further configured to: in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, determine whether the soft label for each of the plurality of annotated images is the same as the associated image label. In identifying mislabelled images, the soft label for each of the plurality of annotated images is the same as the associated image label when the following expression is true, argmax (y) = y

[0077] Where y is the soft label and y is the associated image label. The soft label for each of the plurality of annotated images is not the same as the associated image label when the following expression is true,

[0078] If argmax (y) y

[0079] In this case, that the soft label for each of the plurality of annotated images is not the same as the associated image label. The one or more processors may then, in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, automatically re-label the associated image label to be the soft label when the associated image label does not match the computed soft label.

[0080] A further approach that may be employed by the one or more processors of the data processing unit and method of this disclosure is that of outlier identification 626. In such an approach the one or more processors and method may be configured to remove samples for which the highest-class score in the soft-label is lower than a specified threshold. In other words, the one or more processors may be further configured to, in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, determine a maximum class identification for each of the annotated images in the image data set by computing a maximum class identification to a threshold value, where the maximum class identification is based on the computed soft label for the annotated image. The one or more processers is then configured to determine whether the maximum class identification is equal to or below a predetermined threshold value, where the image is determined to be noncompliant if the maximum class value is equal to below than the predetermined threshold. The expression of the form below may describe how the maximum class identification may relate to the predetermined threshold in order for the image to be considered noncompliant. max(y)<5

[0081] Where 5 is the predetermined threshold. The one or more processors may then be configured to in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, remove the identified one or more noncompliant image from the received data set when forming the modified data set.

[0082] The system shown in Fig. 6, which may be carried out by the data processing unit, may also include a data validation tool that enables a data validator to specify non-overlapping sub-sets of the modified dataset (referred to as the “clean dataset” in Fig. 4) for the model training and validation respectively. This dataset split can either be done manually by an expert annotator, or automatically via the use of k- Fold Cross Validation. Furthermore, a dataset database 627 may be used to store the datasets that are used for the training and validation of Al image classification model. Both the training and validation dataset may comprise a list of (image, label) tuples. At the end of each annotation cycle, the train and validation datasets may be updated. The data storage interface 628 may provide methods and interfaces for the accessing the datasets from the database.

[0083] Once the modified image dataset has been generated as described above by the one or more processors and method of the present disclosure it may be output for training of a weighted processing network for image classification. Therefore, the one or more processors may further be configured to input the modified image dataset to a weighted processing network for image classification, the weighted processing network comprising parameters. The data processing unit of this disclosure may also in some cases train the parameters of the weighted processing network using the modified image dataset. In other cases, however the training may be performed by another apparatus and the data processing apparatus may be solely used for data pruning by identifying noncompliant images and generating a modified image dataset. The weighted processing network that is to be trained is shown in the example of Fig. 6 as Al model management which may be considered a system that performs the life-cycle management of Al models from training to validation and finally model provisioning 635 to the robots or equivalent end devices. This system, which may be implemented by the data processing unit and method of this disclosure may comprise a model database which may be a database that stores the Al models. The system may also include a model storage interface 633 that may provide methods and routines to access and update the trained Al models stored in the Model Database. The system may further comprise a model training process 634 in which the Al model for deployment is trained for the task of image classification using the available training data set. In addition, there may be included a model quality assessment 632 that may perform validation of the current Al model on the held-out validation image set. If the accuracy of the new model outperforms the existing Al models in database 636, the new model may be stored as the best model. Finally, the system may comprise model provisioning 635 which may be a process which serves the robots, either depending on availability or upon request, with the best currently available Al model.

[0084] The data processing unit and method of this disclosure therefore provide a system to reduce the demand for storage and compute time / cost and enable efficient training by removing redundant / superfluous samples, identifying and correcting mislabelled samples, and removing outlier / ambiguous samples. The apparatus and method further provide efficient model training and deployment for the task of image classification in a cloud infrastructure and analyses the quality of training data, and identifies label ambiguity and label errors. The data processing unit and method identifies and removes uninformative, incorrect, and ambiguous data samples from the training data. Particularly advantageously the unit and method disclosed herein generates, in an end-to-end manner, a 'clean' dataset with fewer samples with respect to the original dataset for the deployment model training while maintaining the test accuracy with little to no loss in performance. The image annotation that may be performed by the one or more processors that may include an annotation tool enables data annotators to label images with their respective category. The number of categories may then stay fixed during the course of the operation. The data pruning model training may train an Al image classification model solely for the task of data pruning. This Al model may then generate logits for each sample over each epoch of training. The Al Model can be from the family of ResNet or Vision Transformers type architectures or other equivalent architectures. The data processing unit and method of this disclosure also provides a response aggregator module that may average the logit responses across epochs to generate logit responses per samples. Then, it may apply a soft-max operation to generate the probability of each sample belonging to a specific class. This generated distribution y may be used as a soft-label for the sample as described above. Particular advantages of the data processing unit and method of the present disclosure are that the solution effectively performs the removal of easy samples, automatic rectification of mis-labelled samples, and the removal of outlier samples all in an end-to-end framework. The negative impact induced by noisy labels on Al models is therefore reduced. Since re-labelling samples via use of manual experts is expensive and time-consuming. The effort required for additional label review is automated hence leading to lower cost.

[0085] The solution described herein also generates a training dataset with fewer but more informative samples thereby enabling efficient training. This removes uninformative, noisy, and ambiguous samples and enables cost and energy efficient training by lowering storage requirements and less use of GPU compute resources. The solution also utilizes the training dynamics by leveraging all logits generated at each epoch of training. Utilization of training dynamics enables efficient pruning in fewer epochs.

[0086] The solution described herein also utilizes a dedicated Al model for performing the dataset pruning. The deployment Al model is kept separate. Data pruning can be performed using a smaller Al model whereas the deployment model, typically a larger model, can benefit from the reduced sample size. A further advantage is that the solution is not limited to image classification and can be extended to other Al applications such as object re-identification (RelD). The solution will have a larger application domain and will not be restricted to standard image classification tasks.

[0087] In addition, the solution guarantees the accuracy of the deployed image classification models by continuously training on newly available annotated data and performing model quality assessment.

[0088] With the arrival of newer images from robots, the deployment model can train and improve its test performance by training on more but informative samples.

[0089] Some potential uses for the above are as follows, datacenters can benefit from data-efficient AI / MU training as (i) it reduces the amount of data that needs to be stored, and (ii) it reduces energy consumption by reducing the amount of data it needs to be processed. Edge Computing can be improved since processing of redundant data at the edge devices increases the computational cost. Companies can deploy data-efficient AI / ME solutions on edge devices to improve performance and efficiency. In the medical industry the above could be used and by reducing redundant and erroneous data, clinicians can have a more complete and accurate view of a patient’s medical history. Furthermore, by removing sensitive data that is no longer useful, organizations can reduce the risk of data breaches and protect the privacy of patients.

[0090] The applicant hereby discloses in isolation each individual feature described herein and any combination of two or more such features, to the extent that such features or combinations are capable of being carried out based on the present specification as a whole in the light of the common general knowledge of a person skilled in the art, irrespective of whether such features or combinations of features solve any problems disclosed herein, and without limitation to the scope of the claims. The applicant indicates that aspects of the present disclosure may consist of any such individual feature or combination of features. In view of the foregoing description, it will be evident to a person skilled in the art that various modifications may be made within the scope of the present disclosure.

Claims

CLAIMS1. A data processing unit for data pruning in training a weighted processing network; the device comprising one or more processors configured to: receive an image dataset comprised of a plurality of annotated images comprising one or more image and one or more associated image labels; compute a soft label for each of the plurality of annotated images in the image dataset, the soft labels representing an object class information within the plurality of annotated images; identify one or more noncompliant images of the plurality of annotated images in the image data set, based on the computed soft label of each of the plurality annotated images in the received image dataset; modify the image dataset based on the identified one or more noncompliant image to form a modified image dataset comprised of fewer annotated images than the plurality of annotated images in the received image dataset.

2. The data processing unit according to claim 1, wherein the one or more processors are further configured to: in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, compute an entropy value for each of the plurality annotated image based on the computed soft label; and determine that the computed entropy value is equal to or below a predetermined threshold value; wherein the entropy value represents the confidence of assigning an image to one of a plurality of classes; and in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, remove the identified one or more noncompliant image from the received data set when forming the modified data set.

3. The data processing unit according to any preceding claim, wherein the one or more processors are further configured to: in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, determine whether the soft label for each of the plurality of annotated images is the same as the associated image label; and in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, automatically re-label the associated image label to be the soft label when the associated image label does not match the computed soft label.

4. The data processing unit according to any preceding claim, wherein the one or more processors are further configured to: in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, determine a maximum class identification for each of the annotated images in the image data set by computing a maximum class identification to a threshold value, where the maximum class identification is based on the computed soft label for the annotated image; determine whether the maximum class identification is equal to or below a predetermined threshold value, where the image is determined to be noncompliant if the maximum class value is equal to or below the predetermined threshold; and in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, remove the identified one or more noncompliant image from the received data set when forming the modified data set.

5. The data processing unit according to any preceding claim, wherein the soft labels are computed for each of the plurality of annotated images in the image dataset based on the following expression:Wherein y is the soft label, c(. ) is the soft-max activation function, T is the number of epochs, (x,y)eDtrain is an image, label pair, c is the number of image classes and z^(x) is the output logits.

6. The data processing unit according to claim 3 or claims 4 to 5 when dependent on claim 3, wherein the soft label for each of the plurality of annotated images is the same as the associated image label when the following expression is true, argmax (y) = yWhere y is the soft label and y is the associated image label; and the soft label for each of the plurality of annotated images is not the same as the associated image label and the following expression is true, argmax (y) #= yWhere y is the soft label and y is the associated image label.

7. The data processing unit according to claim 2 or claims 3 to 5 when dependent on claim 2, wherein the entropy value is determined based on the following expression:Where H is the entropy value, y is the soft label, i is the image class, c is the number of image classes and p is probability that an object in the image is in an image class.

8. The data processing unit according to any preceding claim, wherein the one or more processors are further configured to: input the received image dataset to an image classification model for identifying classes present in the images of the image data set; wherein the image classification model is configured to, across one or more epoch, iteratively: generate one or more logits for each image of the image dataset, the logits representing pre-classification un-normalized values across each class; output the one or more logits for each image of the image dataset at each epoch.

9. The data processing unit according to any preceding claim, wherein the one or more processors are further configured to: output the modified image data set for training of a weighted processing network for image classification.

10. The data processing unit according to any preceding claim, wherein the one or more processors are further configured to: input the modified image dataset to a weighted processing network for image identification, the weighted processing network comprising parameters; and train the parameters of the weighted processing network using the modified image dataset.

11. A method of for data pruning in training a weighted processing network; the method comprising :receiving an image dataset comprised of a plurality of annotated images comprising one or more image and one or more associated image labels; computing a soft label for each of the plurality of annotated images in the image dataset, the soft labels representing an object class information within the plurality of annotated images; identifying one or more noncompliant images of the plurality of annotated images in the image data set, based on the computed soft label of each of the plurality annotated images in the received image dataset; modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset comprised of fewer annotated images than the plurality of annotated images in the received image dataset.

12. The method of claim 11, wherein the method further comprises: in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, computing an entropy value for each of the plurality annotated image based on the computed soft label; and determining that the computed entropy value is equal to or below a predetermined threshold value; wherein the entropy value represents the confidence of assigning an image to one of a plurality of classes; and in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, removing the identified one or more noncompliant image from the received data set when forming the modified data set.

13. The method according to claim 11 or 12, wherein the method further comprises: in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, determining whether the soft label for each of the plurality of annotated images is the same as the associated image label; and in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, automatically re-labelling the associated image label to be the soft label when the associated image label does not match the computed soft label.

14. The method according to any of claims 11 to 13, wherein the method further comprises: in identifying the one or more noncompliant images of the plurality of annotated images in the image data set, determining a maximum class identification for each of the annotated images in theimage data set by computing a maximum class identification to a threshold value, where the maximum class identification is based on the computed soft label for the annotated image; determining whether the maximum class identification is equal to or below a predetermined threshold value, where the image is determined to be noncompliant if the maximum class value is equal to or below the predetermined threshold; and in modifying the image dataset based on the identified one or more noncompliant image to form a modified image dataset, removing the identified one or more noncompliant image from the received data set when forming the modified data set.

15. The method according to any of claims 11 to 13, wherein the soft labels are computed for each of the plurality of annotated images in the image dataset based on the following expression:Wherein y is the soft label, c(. ) is the soft-max activation function, T is the number of epochs, (x,y)eDtrain is an image, label pair, c is the number of image classes and z^(x) is the output logits.

16. The method according to claim 13 or claims 14 to 15 when dependent on claim 3, wherein the soft label for each of the plurality of annotated images is the same as the associated image label when the following expression is true, argmax (y) = yWhere y is the soft label and y is the associated image label; and the soft label for each of the plurality of annotated images is not the same as the associated image label and the following expression is true, argmax (y) #= yWhere y is the soft label and y is the associated image label.

17. The method according to claim 12 or claims 13 to 15 when dependent on claim 12, wherein the entropy value is determined based on the following expression:Where H is the entropy value, y is the soft label, i is the image class, c is the number of image classes and p is probability that an object in the image is in an image class.

18. The method according to any one of claims 11 to 17, wherein method further comprises: inputting the received image dataset to an image classification model for identifying classes present in the images of the image data set; wherein the image classification model is configured to, across one or more epoch, iteratively: generating one or more logits for each image of the image dataset, the logits representing pre-classification un-normalized values across each class; outputting the one or more logits for each image of the image dataset at each epoch.

19. The method according to any one of claims 11 to 18, wherein method further comprises: outputting the modified image data set for training of a weighted processing network for image classification.

20. The method according to any one of claims 11 to 19, wherein the method further comprises: inputting the modified image dataset to a weighted processing network for image identification, the weighted processing network comprising parameters; and training the parameters of the weighted processing network using the modified image dataset.