Systems and methods for training a multi-class object classification model using partially labeled training data

By evaluating the loss function and adjusting parameters, the multi-category object classification model is trained using partially marked training data, which solves the problem of model performance degradation, and realizes efficient training process and low-cost use of labeled data.

CN116018621BActive Publication Date: 2025-07-22GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202080104506.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-10-06
Publication Date
2025-07-22
Estimated Expiration
2040-10-06

AI Technical Summary

Technical Problem

In the prior art, when training multi-category object classification models using partially marked training data, the model performance is prone to degradation, and the process of labeling training data is time-consuming and expensive.

Method used

By evaluating the loss function, adjusting the parameters of the multi-category object classification model, and using partially marked training data to train the model, reducing the assumption impact of unlabeled categories and reducing the negative impact on model performance.

Benefits of technology

The performance of training models under partially marked training data is achieved to maintain stable performance, reduce the time and cost of labeling training data, and improve the training efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116018621B_ABST
    Figure CN116018621B_ABST
Patent Text Reader

Abstract

The systems and methods of the present disclosure relate to a computer-implemented method for training a multi-class object classification model of machine learning using partially labeled training data. The method can include obtaining image data depicting an object and ground truth data including a subset of object class annotations respectively associated with a subset of object classes of a plurality of object classes. The method can include processing the image data using the multi-class object classification model of machine learning to obtain object classification data. The method can include evaluating a loss function that evaluates a multi-class classification loss and adjusting one or more parameters of the multi-class object classification model based on the loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to training machine learning object classification models. More particularly, the present disclosure relates to using partially labeled training data to train a multi-class object classification model of machine learning to detect and identify multiple classes of objects depicted in image data. Background Art

[0002] Training a multi-class object classification model of machine learning to detect and identify multiple classes of objects typically utilizes image training data that is labeled with ground truth bounding boxes for one or more of the multiple classes. This training data is typically not fully labeled. That is, not every class has an explicit label. Instead, some labels may be implicitly inferred. For example, regions of the image data that are not included in the labeled bounding boxes (e.g., unlabeled) are typically assumed to not include any objects belonging to these classes.

[0003] However, these unlabeled regions often include other objects corresponding to the object classes that the model is trained to detect. As an example, a model can be trained to identify cats, and the training images can include cats in the unlabeled regions. If the label for the unlabeled region is implicitly inferred (e.g., that no cats are present in the region), then the model can be incorrectly trained (e.g., trained to not recognize the presence of cats).

[0004] Thus, generating training data for an object classification model typically requires exhaustive annotation of all classes depicted in the images of the training dataset. However, such annotation of a classification training dataset can be very expensive and / or time-consuming. Additionally, current attempts to train a model using image data with partially labeled classes (e.g., only two of three classes are labeled, etc.) typically result in a significant degradation in model quality. Summary of the Invention

[0005] Aspects and advantages of embodiments of the present disclosure will be set forth in part in the following description, or can be learned from the description, or can be learned through practice of the embodiments.

[0006] One example aspect of the present disclosure relates to a computing system that trains a multi-class object classification model using partially labeled training data. The computing system can include one or more processors. The computing system can include a machine learning multi-class object classification model configured to classify multiple object classes. The computing system can include one or more tangible non-transitory computer-readable media storing computer-readable instructions that, when executed by the one or more processors, cause the one or more processors to perform operations. The operations can include obtaining image data depicting one or more objects and ground truth data including a subset of object class annotations respectively associated with a subset of object classes of the multiple object classes. The operations can include processing the image data using the machine learning multi-class object classification model to obtain object classification data. The operations can include evaluating a loss function that evaluates a multi-class classification loss including a difference between the object classification data and the subset of object class annotations, where the loss function includes multiple weighted loss signals respectively associated with the multiple object classes, and where the weight of each of the weighted loss signals is at least partially based on whether the object class associated with the corresponding loss signal is included in the subset of object classes. The operations can include adjusting one or more parameters of the machine learning multi-class object classification model at least partially based on the loss function.

[0007] Another example aspect of the present disclosure relates to a computer-implemented method for training a machine learning multi-class object classification model using partially labeled training data. The method can include obtaining, by a computing system including one or more computing devices, image data depicting one or more objects and ground truth data including a subset of object class annotations respectively associated with a subset of object classes of the multiple object classes. The method can include processing, by the computing system, the image data using the machine learning multi-class object classification model to obtain object classification data. The method can include evaluating, by the computing system, a loss function that evaluates a multi-class classification loss including a difference between the object classification data and the subset of object class annotations, where the loss function includes multiple weighted loss signals respectively associated with the multiple object classes, and where the weight of each of the weighted loss signals is at least partially based on whether the object class associated with the corresponding loss signal is included in the subset of object classes. The method can include adjusting, by the computing system, one or more parameters of the machine learning multi-class object classification model at least partially based on the loss function.

[0008] Another example aspect of the present disclosure relates to one or more tangible non-transitory computer-readable media storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations. The operations can include obtaining image data depicting one or more objects and ground truth data including a subset of object class annotations respectively associated with a subset of object classes of a plurality of object classes. The operations can include processing the image data using a multi-class object classification model of machine learning to obtain object classification data, where the multi-class object classification model of machine learning is configured to classify each of the one or more objects as belonging to an object class of a plurality of object classes. The operations can include modifying a loss function to obtain a modified loss function, where the loss function includes a plurality of loss signals respectively associated with a plurality of object classes, and where the modified loss function includes a subset of the plurality of loss signals respectively associated with the subset of object classes. The operations can include evaluating the modified loss function, where the modified loss function evaluates a multi-class classification loss including a difference between the object classification data and the subset of object class annotations. The operations can include adjusting one or more parameters of the multi-class object classification model of machine learning at least in part based on the loss function.

[0009] Other aspects of the present disclosure relate to various systems, devices, non-transitory computer-readable media, user interfaces, and electronic devices.

[0010] These and other features, aspects, and advantages of the various embodiments of the present disclosure will be better understood with reference to the following description and the appended claims. The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate example embodiments of the present disclosure and, together with the description, serve to explain the relevant principles. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] A detailed discussion of embodiments involving those of ordinary skill in the art is set forth in the specification with reference to the accompanying drawings, in which:

[0012] Figure 1A A block diagram depicts an example computing system that performs multi-class object classification of machine learning in accordance with an example embodiment of the present disclosure.

[0013] Figure 1B A block diagram depicts an example computing device that trains a multi-class object classification model of machine learning in accordance with an example embodiment of the present disclosure.

[0014] Figure 1C A block diagram depicts an example computing device that performs multi-class object classification of machine learning in accordance with an example embodiment of the present disclosure.

[0015] Figure 2 A block diagram depicts an example multi-class object classification model of machine learning in accordance with an example embodiment of the present disclosure.

[0016] Figure 3 A block diagram depicting an example machine learning image analysis model in accordance with an example embodiment of the present disclosure.

[0017] Figure 4 A data flow diagram depicting a method for training a multi-class object classification model of machine learning using partially labeled training data in accordance with an example embodiment of the present disclosure.

[0018] Figure 5 A flowchart depicting an example method 500 of performing training of a multi-class object classification model of machine learning using partially labeled training data.

[0019] Reference numerals that are repeated in the various figures are intended to identify the same features in the various embodiments. Detailed Description

[0020] Overview

[0021] Generally, the present disclosure relates to training a multi-class object classification model of machine learning using partially labeled training data. More particularly, the present disclosure relates to training a multi-class object classification model of machine learning using partially labeled training data to identify multiple class objects depicted in image data. As an example, a multi-class object classification model of machine learning can be configured to classify objects depicted in image data as belonging to multiple classes (e.g., bear class, lion class, tiger class, kangaroo class, etc.). Image data depicting one or more objects can be obtained together with ground truth data that includes a subset of object class annotations (e.g., ground truth class labels) associated with a subset of the multiple object classes (e.g., two out of a total of four classes, etc.). The multi-class object classification model of machine learning can be utilized to process the image data to obtain object classification data for classifying the object(s) depicted in the image data.

[0022] It is able to evaluate a loss function that evaluates a multi-class classification loss, which includes the difference between object classification data and a subset of object class annotations (e.g., the difference between object classification and associated labels, etc.). More particularly, the loss function can include multiple weighted loss signals respectively associated with multiple object classes (e.g., a weighted bear loss signal for the bear class, a weighted lion loss signal for the lion class, etc.). The weight of each of the weighted loss signals can be at least partially based on the object class of the loss signal being included in the subset of object classes. For example, the weighted loss signal from a class (e.g., the kangaroo class) that is not included in the subset of object classes (e.g., the bear, lion, and tiger classes) among multiple object classes can have a lower weight than the weighted loss signal of an object class (e.g., the bear class) that is included in the subset of object classes.

[0023] It is able to adjust the (multiple) parameters of the multi-class object classification model at least partially based on the loss function (e.g., proportional to the weights of the weighted loss signals). In this way, a machine-learned multi-class object classification model can be trained to recognize the labeled classes without degrading the model performance regarding the unlabeled classes. More particularly, by adjusting the relevance of the classes not labeled in the training data (e.g., not included in the subset of classes) or eliminating their loss signals, the model can be trained such that the unlabeled regions of the training images are processed with fewer negative or neutral assumptions regarding the presence of the unlabeled classes. By eliminating or reducing the impact of such assumptions, the model can be trained using partially labeled training data without causing a degradation in the model quality for the classification of unlabeled classes.

[0024] More particularly, image data depicting one or more objects can be obtained together with ground truth data. The ground truth data can include a subset of object class annotations (e.g., ground truth labels) respectively associated with a subset of object classes from multiple object classes. As an example, the image data can be or otherwise include one or more images (e.g., images, multiple video frames, etc.) from a training dataset that is configured to train a machine-learned classification model to recognize and classify multiple classes. An image from the training dataset can include four objects belonging to four independent classes. The subset of object class annotations can be or otherwise include the annotations of the first two of the four independent classes. As another example, the ground truth data of a second image of the training dataset can include the object class annotations of the remaining two of the four independent classes. In this way, in some embodiments, the subset of class annotations included in the ground truth data can include different classes among the multiple classes from different images included in the training dataset.

[0025] In some embodiments, the object category annotation can include a bounding box that defines an image data region and an associated category label. As an example, the image data can depict a lion object. The object category annotation can include a bounding box that defines an image data region that includes the lion and the corresponding "lion" label. Alternatively or additionally, in some embodiments, the object category annotation can label the entire image of the image data as including or not including a category. As an example, the image data may not depict a tiger. The object category annotation can indicate that no tiger is depicted in the image data.

[0026] In some embodiments, the image data can be a portion of the image data from an image. As an example, the image data can be the portion of the image data that is extracted from the image based on a prediction (e.g., extracted by a region proposal network (RPN), etc.) that a portion of the image data depicts an object. As another example, the image data can be a portion of the image data that is not defined by the bounding box of the object category annotation. Alternatively, in some embodiments, the image data can be or otherwise include the entire depiction of the image data. As an example, the image data can be obtained and then iteratively processed by a machine learning multi-class object classification model to identify regions predicted to include an object and then classify the object subsequently or simultaneously.

[0027] The image data can be processed using a machine learning multi-class object classification model to obtain object classification data. The machine learning multi-class object classification model can be or otherwise include any kind of conventional machine learning object classification model. As an example, the machine learning model can be or can otherwise include one or more neural networks (e.g., deep neural networks, recurrent neural networks, graph neural networks, etc.). A neural network (e.g., a deep neural network) can be a feedforward neural network, a convolutional neural network, and / or various other types of neural networks.

[0028] A multi-class object classification model of machine learning can be configured to classify multiple object classes. More particularly, a multi-class object classification model of machine learning can be configured to process image data to classify an object depicted in the image data as belonging to one of multiple classes. Additionally or alternatively, in some embodiments, a multi-class object classification model of machine learning can be part of a machine learning image analysis model that performs object recognition and object classification (e.g., a component of the model, one or more layers of the model, etc.). As an example, a multi-class object classification model of machine learning can be the object classification part of a conventional two-stage object detector model (e.g., Fast R-CNN, etc.). For example, a first stage of a two-stage object detector model (e.g., a Region Proposal Network (RPN), etc.) can obtain image data and output a region of the image data that is predicted to contain an object (e.g., using an anchor-based method, etc.) (e.g., a region not defined by a bounding box, etc.). The region of the image data can be processed using a multi-class object classification model of machine learning to obtain object classification data.

[0029] As another example, a multi-class object classification model of machine learning can be or otherwise incorporated into a one-stage object detector model (e.g., a Single Shot Detector (SSD) model, etc.). For example, a multi-class object classification model of machine learning can be or otherwise include multiple machine learning layers (e.g., (multiple) convolutional layers, (multiple) activation layers, etc.) that predict (multiple) regions that may include a predicted object, and the predicted object can subsequently be processed using an additional classification layer of the multi-class object classification model to generate object classification data. For another example, a multi-class object classification model of machine learning can include multiple layers that predict regions of the image data to include a predicted object and simultaneously classify the predicted object. In this way, a method for training a multi-class object classification model of machine learning can be applied to any conventional machine learning image analysis model (e.g., a Single Shot Detector (SSD) model, a You Only Look Once (YOLO) model, a Fast R-CNN model, etc.).

[0030] Object classification data output by a multi-class object classification model of machine learning can provide a multi-class classification output for an object depicted in image data. As an example, a portion of an image can depict a bear object. The multi-class object classification model of machine learning can be configured to classify the object as belonging to one of four classes (e.g., bear, tiger, kangaroo, lion, etc.). The object classification data can include multiple predicted object class annotations that indicate whether the object belongs to each of the classes (e.g., bear 1, tiger 0, kangaroo 0, lion 0, etc.). Alternatively, in some embodiments, the object classification data can include multiple object class probability predictions that the object depicted in the image data belongs to each of the four classes (e.g., 80% bear, 15% tiger, 15% kangaroo, 10% lion, etc.). In some embodiments, the object classification data can include a single indication that the object depicted in the image data belongs to one class. As an example, the object classification data can be or otherwise include a predicted class annotation (e.g., "bear" class annotation, etc.). Thus, the object classification data can include multiple predicted probability outputs for the respective multiple classes and / or predicted class annotations of the object depicted in the image data or a portion of the image data. Regarding Figure 4 and Figure 5 the object classification data output by the multi-class object classification model of machine learning will be discussed in more detail.

[0031] A loss function can be evaluated that evaluates a multi-class classification loss. The multi-class classification loss can include a difference between the object classification data and a subset of object class annotations (e.g., a difference between a predicted object class annotation and a ground truth object class annotation, etc.). More particularly, the loss function can include multiple weighted loss signals respectively associated with multiple object classes (e.g., a bear weighted loss signal for the bear class, a lion weighted loss signal for the lion, etc.). The weight of each of the weighted loss signals can be at least partially based on whether the object class associated with the respective loss signal is included in the object class subset.

[0032] As an example, the first weighted loss signal can be associated with a first category from multiple object categories, where the first category is not included in the subset of object categories. For example, the first category can be the kangaroo category, and the subset of object categories can include the bear category, the lion category, and the tiger category. Since the first category is not included in the category subset, the weight of the first loss signal can be weighted to reduce the impact of the first loss signal relative to the loss function. In some embodiments, the weight of each of the weighted loss signals can be a normalized value. For example, the weight of the first loss signal can be weight 0, while the weight of the second loss signal associated with the categories included in the category subset can be weighted 1 (e.g., eliminating the loss signal from the loss function, etc.). For another example, the weight of the first loss signal can be weight 0.35, while the weight of the second loss signal associated with the categories included in the category subset can be weighted 1 (e.g., to reduce the correlation of the first loss signal relative to the second loss signal, etc.). In this way, the weights of the weighted loss signals can be configured to reduce or eliminate the impact of the corresponding loss signals that are not associated with the categories included in the category subset (e.g., categories with associated object category annotations (labels), etc.).

[0033] One or more parameters of a multi-class object classification model of machine learning can be adjusted based at least in part on a loss function. More particularly, the (multiple) parameters of the multi-class object classification model of machine learning can be adjusted based on each of the weighted loss signals and their corresponding weights. In some embodiments, the (multiple) parameters of the multi-class object classification model of machine learning can be adjusted proportionally to the weights of each of the weighted loss signals of the loss function. As an example, the final loss value can be calculated using the loss function based on an assessment of the difference between the object classification data and the ground truth data. The final loss value can be proportional to the weights of the corresponding weighted loss signals. For example, a weighted loss signal with a weight of zero may not contribute to the calculation of the final loss value. The loss value can be backpropagated through the multi-class classification model of machine learning, and one or more parameters of the model can be adjusted based on the final loss value (e.g., using a gradient descent algorithm, etc.). In this way, the adjustment of the (multiple) parameters of the model can be proportional to the weights of the weighted loss signals, thus reducing the impact of the loss signals associated with the unlabeled categories during model training.

[0034] In some embodiments, an evaluation loss function can include a subset of weighted loss signals of the evaluation loss function, which subset is respectively associated with a subset of object classes. More particularly, the loss function can include only the weighted loss signals associated with the labeled object classes (e.g., a subset of object class annotations of the ground truth data, etc.). In this way, the loss function can exclude any loss signals associated with object classes that are not annotated by the subset of object class annotations (e.g., not labeled in the training data, etc.).

[0035] In some embodiments, a multi-class object classification model of machine learning can be utilized after training. More particularly, the computing system can obtain additional image data depicting one or more additional objects. The additional image data can be processed using the multi-class object classification model of machine learning to obtain an image classification output. In some embodiments, the image classification output can include one or more labels describing the additional image data. As an example, the image classification output can include one or more object class annotations of one or more additional objects. As another example, the image classification output can include one or more image annotations that annotate the entire image as belonging to one or more classes. For example, the multi-class object classification model of machine learning can label one or more additional objects as being semantically related to a "natural" image (e.g., a rabbit object, a tree object, a boulder object, a grass object, etc.). Based on the (multiple) object classification labels, the image classification data can label the image as belonging to the "natural" image category.

[0036] It should be noted that a multi-class object classification model of machine learning can be utilized to detect the presence of a specific class of objects depicted in the image data. More particularly, a multi-class object classification model of machine learning can be utilized to perform object detection by detecting objects corresponding to the learned object classes. As an example, the object classification model of machine learning can detect a bear object depicted in the image data. Thus, the multi-class object classification model of machine learning according to the embodiments of the present invention can be a machine learning object detection model and / or can be a component of a machine learning object detection model.

[0037] The present disclosure provides a number of technical effects and benefits. As an example technical effect and benefit, the systems and methods of the present disclosure allow for the use of partially labeled training data to train machine learning models. As previously mentioned, the process of labeling training data for multiple classes is a daunting task and is generally considered to be extremely expensive in both time and cost. Additionally, previous attempts to use partially labeled training data to train machine learning models have historically led to a significant degradation in model performance. However, under the proposed method, it is possible to utilize partially labeled training data to train machine learning models with little or no degradation in model performance. Subsequently, this advancement significantly reduces the computational, monetary, and human costs associated with labeling training data for machine learning models. Additionally, this advancement allows for the reuse of previously labeled training data for training multi-class object classification models.

[0038] Reference is now made to the accompanying drawings, and example embodiments of the present disclosure will be discussed in more detail.

[0039] Example Devices and Systems

[0040] Figure 1A A block diagram of an example computing system 100 that performs multi-class object classification for machine learning in accordance with an example embodiment of the present disclosure is depicted. System 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled via a network 180.

[0041] The user computing device 102 can be any type of computing device, such as, for example, a personal computing device (e.g., a laptop computer or a desktop computer), a mobile computing device (e.g., a smart phone or a tablet computer), a gaming console or controller, a wearable computing device, an embedded computing device, or any other type of computing device.

[0042] The user computing device 102 includes one or more processors 112 and a memory 114. The one or more processors 112 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple processors operably connected. The memory 114 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, disks, etc. and combinations thereof. The memory 114 can store data 116 and instructions 118 that are executed by the processor 112 to cause the user computing device 102 to perform operations.

[0043] In some embodiments, the user computing device 102 is capable of storing or including one or more machine-learned multi-class object classification models 120. For example, the machine-learned multi-class object classification model 120 can be or otherwise include various machine learning models, such as neural networks (e.g., deep neural networks) or other types of machine learning models, including non-linear models and / or linear models. The neural network can include a feed-forward neural network, a recurrent neural network (e.g., a long short-term memory recurrent neural network), a convolutional neural network, or other forms of neural networks. Refer to Figures 2-5 Discuss example machine-learned multi-class object classification models 120.

[0044] In some embodiments, one or more machine-learned multi-class object classification models 120 can be received from the server computing system 130 via the network 180, stored in the user computing device memory 114, and then used or otherwise implemented by one or more processors 112. In some embodiments, the user computing device 102 can implement multiple parallel instances of a single machine-learned multi-class object classification model 120 (e.g., to perform parallel multi-class object classification across multiple instances of the machine-learned multi-class object classification model).

[0045] The machine-learned multi-class object classification model 120 can be configured to classify multiple objects. More particularly, the machine-learned multi-class object classification model 120 can be utilized to classify one or more objects depicted in image data as each belonging to one of a plurality of classes. As an example, the user computing device 102 can obtain additional image data depicting one or more additional objects (e.g., via the network 180, data 116, etc.). The machine-learned multi-class object classification model 120 can be used to process the image data to obtain an image classification output. In some embodiments, the image classification output can include one or more labels describing the additional image data. As an example, the image classification output can include one or more object class annotations for one or more additional objects. As another example, the image classification output can include one or more image annotations that annotate the entire image as belonging to one or more classes. For example, the machine-learned multi-class object classification model 120 can label one or more additional objects as semantically related to a "natural" image (e.g., a rabbit object, a tree object, a boulder object, a grass object, etc.). Based on the (multiple) object classification labels, the image classification data can label the image as belonging to the "natural" image class.

[0046] Additionally or alternatively, one or more machine-learned multi-class object classification models 140 can be included in or otherwise stored and implemented by a server computing system 130 that communicates with a user computing device 102 according to a client-server relationship. For example, a machine-learned multi-class object classification model 140 can be implemented by the server computing system 140 as part of a web service (e.g., a multi-class object classification service). Accordingly, one or more models 120 can be stored and implemented at the user computing device 102, and / or one or more models 140 can be stored and implemented at the server computing system 130.

[0047] The user computing device 102 can also include one or more user input components 122 that receive user input. For example, the user input component 122 can be a touch-sensitive component (e.g., a touch-sensitive display screen or a touchpad) that is sensitive to a user input object (e.g., a finger or a stylus). The touch-sensitive component can be used to implement a virtual keyboard. Other example user input components include a microphone, a traditional keyboard, or other components through which a user can provide user input.

[0048] The server computing system 130 includes one or more processors 132 and a memory 134. The one or more processors 132 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.), and can be one processor or multiple processors operably connected. The memory 134 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof. The memory 134 can store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing device 130 to perform operations.

[0049] In some embodiments, the server computing system 130 includes or is otherwise implemented by one or more server computing devices. In cases where the server computing system 130 includes multiple server computing devices, such server computing devices can operate according to a sequential computing architecture, a parallel computing architecture, or some combination thereof.

[0050] As described above, the server computing system 130 can store or otherwise include one or more machine-learned multi-class object classification models 140. For example, a machine-learned multi-class object classification model 140 can be or can otherwise include various machine learning models. Example machine learning models include neural networks or other multi-layer non-linear models. Example neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Reference Figures 2-5Discuss the example model 140.

[0051] The user computing device 102 and / or the server computing system 130 can train the model 120 and / or the model 140 via interaction with a training computing system 150 communicatively coupled via a network 180. The training computing system 150 can be separate from the server computing system 130 or can be part of the server computing system 130.

[0052] The training computing system 150 includes one or more processors 152 and a memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be one processor or multiple processors operably connected. The memory 154 can include one or more non-transitory computer-readable storage media, such as RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc. and combinations thereof. The memory 154 can store data 156 and instructions 158 executed by the processor 152 to cause the training computing system 150 to perform operations. In some embodiments, the training computing system 150 includes one or more server computing devices or is otherwise implemented by one or more server computing devices.

[0053] The training computing system 150 can include a model trainer 160 that uses various training or learning techniques, such as, for example, error backpropagation, to train the machine learning model 120 and / or the model 140 stored at the user computing device 102 and / or the server computing system 130. For example, a loss function can be backpropagated through the (multiple) model to update one or more parameters of the (multiple) model (e.g., based on the gradient of the loss function). Various loss functions can be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent techniques can be used to iteratively update the parameters over multiple training iterations.

[0054] In some embodiments, performing error backpropagation can include performing truncated backpropagation through time. The model trainer 160 can perform various generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the trained model.

[0055] In particular, the model trainer 160 is capable of training the model 120 and / or the model 140 based on the set of training data 162. The training data 162 can include, for example, image data depicting one or more objects, which can be obtained together with the ground truth data. The ground truth data can include a subset of object class annotations respectively associated with a subset of object classes from a plurality of object classes. As an example, the image data can be or otherwise include one or more images from a training data set (e.g., images, multiple video frames, etc.) that is configured to train a classification model of machine learning to identify and classify multiple classes. Images from the training data set can include four objects belonging to four separate classes. The subset of object class annotations can be or otherwise include annotations for the first two of the four separate classes. As another example, the ground truth data for a second image of the training data set can include object class annotations for the remaining two of the four separate classes. In this way, in some embodiments, the ground truth data can include different class annotations for different images included in the training data set.

[0056] In some embodiments, the object class annotations can include bounding boxes that define regions of the image data and associated class labels. As an example, the image data can depict a lion object. The object class annotations can include a bounding box that defines a region of the image data that includes the lion and the corresponding "lion" label. Alternatively or additionally, in some embodiments, the object class annotations can label the entire image of the image data as including or not including the class. As an example, the image data may not depict a tiger. The object class label can indicate that no tiger is depicted in the image data.

[0057] In some embodiments, the image data can be a portion of the image data from an image. As an example, the image data can be the portion of the image data that is extracted from the image based on a prediction (e.g., extracted by a region proposal network (RPN), etc.) that a portion of the image data depicts an object. As another example, the image data can be a portion of the image data that has been defined by the bounding box of the object class annotation. Alternatively, in some embodiments, the image data can be or otherwise include the entire image. As an example, an image can be obtained and subsequently iteratively processed by a multi-class object classification model of machine learning to identify regions predicted to include objects and then or simultaneously classify the objects.

[0058] In some embodiments, if the user has provided consent, training examples can be provided by the user computing device 102. Thus, in such an embodiment, the model 120 provided to the user computing device 102 can be trained by the training computing system 150 based on user-specific data received from the user computing device 102. In some cases, this process can be referred to as personalizing the model.

[0059] The model trainer 160 includes computer logic for providing the required functionality. The model trainer 160 can be implemented in hardware, firmware, and / or software that controls a general-purpose processor. For example, in some embodiments, the model trainer 160 includes program files stored on a storage device, loaded into memory, and executed by one or more processors. In other embodiments, the model trainer 160 includes one or more sets of computer-executable instructions stored in a tangible computer-readable storage medium, such as a RAM hard drive or optical or magnetic media.

[0060] The network 180 can be any type of communication network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or some combination thereof, and can include any number of wired or wireless links. Generally, communication over the network 180 can use various communication protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, secure HTTP, SSL), and be carried over any type of wired and / or wireless connection.

[0061] The method of using partially labeled training data to train a machine learning model according to an embodiment of the present invention can be used for various machine learning models, tasks, applications, and / or use cases.

[0062] In some embodiments, the training data for training the machine learning model(s) of the present disclosure can be partially labeled image data. The machine learning model(s) can process the image data to produce an output. As an example, the machine learning model(s) can process the image data to produce an image recognition output (e.g., recognition of the image data, latent embedding of the image data, encoded representation of the image data, hash of the image data, etc.). As another example, the machine learning model(s) can process the image data to produce an image segmentation output. As another example, the machine learning model(s) can process the image data to produce an image classification output. As another example, the machine learning model(s) can process the image data to produce an image data modification output (e.g., change of the image data, etc.). As another example, the machine learning model(s) can process the image data to produce an encoded image data output (e.g., encoded and / or compressed representation of the image data, etc.). As another example, the machine learning model(s) can process the image data to produce an enlarged image data output. As another example, the machine learning model(s) can process the image data to produce a prediction output.

[0063] In some embodiments, the training data for training the machine learning model(s) of the present disclosure can be partially labeled text or natural language data. The machine learning model(s) can process the text or natural language data to produce an output. As an example, the machine learning model(s) can process the natural language data to produce a language encoding output. As another example, the machine learning model(s) can process the text or natural language data to produce a latent text embedding output. As another example, the machine learning model(s) can process the text or natural language data to produce a translation output. As another example, the machine learning model(s) can process the text or natural language data to produce a classification output. As another example, the machine learning model(s) can process the text or natural language data to produce a text segmentation output. As another example, the machine learning model(s) can process the text or natural language data to produce a semantic intent output. As another example, the machine learning model(s) can process the text or natural language data to produce an enlarged text or natural language output (e.g., text or natural language data of higher quality than the input text or natural language, etc.). As another example, the machine learning model(s) can process the text or natural language data to produce a prediction output.

[0064] In some embodiments, the training data for training the machine learning model(s) of the present disclosure can be partially labeled speech data. The machine learning model(s) can process the speech data to produce an output. As an example, the machine learning model(s) can process the speech data to produce a speech recognition output. As another example, the machine learning model(s) can process the speech data to produce a speech translation output. As another example, the machine learning model(s) can process the speech data to produce a latent embedding output. As another example, the machine learning model(s) can process the speech data to produce an encoded speech output (e.g., an encoded and / or compressed representation of the speech data, etc.). As another example, the machine learning model(s) can process the speech data to produce an enhanced speech output (e.g., speech data of higher quality than the input speech data, etc.). As another example, the machine learning model(s) can process the speech data to produce a text representation output (e.g., a text representation of the input speech data, etc.). As another example, the machine learning model(s) can process the speech data to produce a prediction output.

[0065] In some embodiments, the training data for training the machine learning model of the present disclosure can be partially labeled latent coding data. The machine learning model(s) can process the latent coding data to produce an output. As an example, the machine learning model(s) can process the latent coding data to produce a recognition output. As another example, the machine learning model(s) can process the latent coding data to produce a reconstruction output. As another example, the machine learning model(s) can process the latent coding data to produce a search output. As another example, the machine learning model(s) can process the latent coding data to produce a reclustering output. As another example, the machine learning model(s) can process the latent coding data to produce a prediction output.

[0066] In some embodiments, the input to the machine learning model(s) of the present disclosure can be statistical data. The machine learning model(s) can process the statistical data to produce an output. As an example, the machine learning model(s) can process the statistical data to produce a recognition output. As another example, the machine learning model(s) can process the statistical data to produce a prediction output. As another example, the machine learning model(s) can process the statistical data to produce a classification output. As another example, the machine learning model(s) can process the statistical data to produce a segmentation output. As another example, the machine learning model(s) can process the statistical data to produce a segmentation output. As another example, the machine learning model(s) can process the statistical data to produce a visualization output. As another example, the machine learning model(s) can process the statistical data to produce a diagnostic output.

[0067] In some embodiments, the input to the machine learning model(s) of the present disclosure can be sensor data. The machine learning model(s) can process the sensor data to produce an output. As an example, the machine learning model(s) can process the sensor data to produce an identification output. As another example, the machine learning model(s) can process the sensor data to produce a prediction output. As another example, the machine learning model(s) can process the sensor data to produce a classification output. As another example, the machine learning model(s) can process the sensor data to produce a segmentation output. As another example, the machine learning model(s) can process the sensor data to produce a visualization output. As another example, the machine learning model(s) can process the sensor data to produce a diagnostic output. As another example, the machine learning model(s) can process the sensor data to produce a detection output.

[0068] In some cases, the machine learning model(s) can be configured to perform a task that includes encoding input data for reliable and / or efficient transmission or storage (and / or corresponding decoding). For example, the task can be or otherwise include an audio compression task. The input can include audio data, and the output can include compressed audio data. In another example, the input includes visual data (e.g., one or more images or videos), the output includes compressed visual data, and the task is a visual data compression task. In another example, the task can include generating an embedding for the input data (e.g., input audio or visual data).

[0069] In some cases, the training data can include partially labeled visual data, and the task is a computer vision task. In some cases, the input includes pixel data for one or more images, and the task is an image processing task. For example, the image processing task can be image classification, where the output is a set of scores, each score corresponding to a different object class and representing the likelihood that one or more images depict an object belonging to that object class. The image processing task can be object detection, where the image processing output identifies one or more regions in one or more images, and for each region, identifies the likelihood that the region depicts an object of interest. As another example, the image processing task can be image segmentation, where the image processing output defines, for each pixel in one or more images, the respective likelihoods of each category in a predefined set of categories. For example, the set of categories can be foreground and background. As another example, the set of categories can be object classes. As another example, the image processing task can be depth estimation, where the image processing output defines a respective depth value for each pixel in one or more images. As another example, the image processing task can be motion estimation, where the network input includes multiple images, and the image processing output defines, for each pixel in one of the input images, the motion of the scene depicted at the pixel between the images in the network input.

[0070] In some cases, the training data can include partially labeled audio data representing spoken utterances, and the task is a speech recognition task. The output can include a text output mapped to the spoken utterance. In some cases, the task includes encrypting or decrypting input data. In some cases, the task includes microprocessor performance tasks, such as branch prediction or memory address translation.

[0071] Figure 1A The illustration can be used to implement an example computing system of the present disclosure. Other computing systems can also be used. For example, in some embodiments, the user computing device 102 can include a model trainer 160 and a training data set 162. In such an embodiment, the model 120 can be trained and used locally on the user computing device 102. In some such embodiments, the user computing device 102 can implement the model trainer 160 to personalize the model 120 based on user-specific data.

[0072] Figure 1B A block diagram of an example computing device 10 performing training of a multi-class object classification model for machine learning according to an example embodiment of the present disclosure is depicted. The computing device 10 can be a user computing device or a server computing device.

[0073] Computing device 10 includes a number of applications (e.g., Application 1 to Application N). Each application contains its own machine learning library and (multiple) machine learning models. For example, each application can include a machine learning model. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like.

[0074] As Figure 1B Illustrated, each application can communicate with a number of other components of the computing device, such as one or more sensors, a context manager, a device status component, and / or additional components. In some embodiments, each application can communicate with each device component using an API (e.g., a common API). In some embodiments, the API used by each application is specific to that application.

[0075] Figure 1C A block diagram of an example computing device 50 that performs multi-class object classification of machine learning according to an example embodiment of the present disclosure is depicted. Computing device 50 can be a user computing device or a server computing device.

[0076] Computing device 50 includes a number of applications (e.g., Application 1 to N). Each application communicates with a central intelligence layer. Example applications include a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application, and the like. In some embodiments, each application can communicate with the central intelligence layer (and the (multiple) models stored therein) using an API (e.g., a common API across all applications).

[0077] The central intelligence layer includes a number of machine learning models. For example, as Figure 1C Illustrated, a corresponding machine learning model (e.g., a model) can be provided for each application and managed by the central intelligence layer. In other embodiments, two or more applications can share a single machine learning model. For example, in some embodiments, the central intelligence layer can provide a single model (e.g., a single model) for all applications. In some embodiments, the central intelligence layer is included within the operating system of computing device 50 or otherwise implemented by the operating system of computing device 50.

[0078] The central intelligence layer can communicate with a central device data layer. The central device data layer can be a centralized data repository of computing device 50. As Figure 1C Illustrated, the central device data layer can communicate with a number of other components of the computing device, such as for example one or more sensors, a context manager, a device status component, and / or additional components. In some embodiments, the central device data layer can communicate with each device component using an API (e.g., a private API).

[0079] Example Model Arrangement

[0080] Figure 2 A block diagram of an example machine - learned multi - class object classification model 200 according to an example embodiment of the present disclosure is depicted. In some embodiments, the machine - learned multi - class object classification model 200 is trained to receive a set of input data 204 that describes an image, and as a result of receiving the input data 204, provide output data 206 that classifies one or more objects depicted in the image.

[0081] More particularly, the machine - learned multi - class object classification model 200 is capable of being configured to classify multiple object classes. The machine - learned multi - class object classification model 200 is capable of obtaining input data 204 that includes image data depicting one or more objects. The machine - learned multi - class object classification model 200 is capable of processing the image data to obtain output data 206. The output data 206 can include object classification data that classifies one or more objects as belonging to one or more corresponding object classes. As an example, the output data 204 can include one or more object class annotations for one or more objects depicted in the input data 204. As another example, the output data 204 can include one or more image annotations that annotate the entire image depicted by the input data 204 as belonging to one or more classes. For example, the machine - learned multi - class object classification model 200 can label one or more additional objects as semantically related to a "natural" image (e.g., a rabbit object, a tree object, a boulder object, a grass object, etc.). Based on the (multiple) object classification labels, the image classification data can label the image depicted by the image data 204 as belonging to the "natural" image category.

[0082] Figure 3 A block diagram of an example machine - learned image analysis model 300 according to an example embodiment of the present disclosure is depicted. The machine - learned image analysis model 300 is similar to Figure 2 the machine - learned multi - class object classification model 200, except that the machine - learned image analysis model 300 further includes a machine - learned object recognition model 302. The machine - learned object recognition model 302 is operable to predict the presence of one or more objects in a portion of the input data 204.

[0083] More particularly, the multi-class object classification model 304 of machine learning can be included as the object classification part of the image analysis model 300 of machine learning (e.g., a conventional two-stage object detector model, etc.). In addition, the object recognition model 302 of machine learning can be included as the object recognition part of the image analysis model 300 of machine learning. For example, the object recognition model 302 of machine learning (e.g., a region proposal network (RPN), etc.) can obtain the input data 204 including the image data and output a part of the image data 304 predicted to contain an object (e.g., defined as a bounding box) (e.g., using an anchor-based method, etc.). The part of the image data 304 can be processed using the multi-class object classification model 202 of machine learning to obtain the output data 206 including the object classification data.

[0084] Figure 4 A data flow diagram of a method 400 for training a multi-class object classification model of machine learning using partially labeled training data according to an example embodiment of the present disclosure is depicted. Image data 402 (e.g., partially labeled training data, etc.) depicting three objects 402A - 402C can be obtained, with each object classified as belonging to a corresponding object category from a plurality of object categories. As depicted, example objects 402A - 402C can belong to the lion object category (e.g., 402A), the tiger object category (e.g., 402B), and the bear object category (e.g., 402C), respectively. Ground truth data 404 can be obtained together with the image data 402, and the image data 402 includes a subset of object category annotations 404A / 404B respectively associated with a subset of the object categories from the plurality of object categories. More particularly, the ground truth data 404 can include the ground truth lion category label 404A (e.g., object category annotation) for the object 402A and the ground truth tiger label 404B (e.g., object category annotation) for the tiger object 402B, but can lack the corresponding object category annotation for the bear object 402C (e.g., the label for the region of the image depicting the bear).

[0085] As depicted, the object category annotations 404A - 404B can be or otherwise include bounding boxes defining regions of the image data 402 and associated class labels (e.g., 404A / 404B). Alternatively or additionally, in some embodiments, the ground truth data 404 can include (a) object category annotation(s) that label the entire image depicted in the image data 402 as including or not including an object of a particular object category. As an example, the ground truth kangaroo label of the ground truth data 404 can indicate whether the image data 402 depicts or does not depict a kangaroo.

[0086] In some embodiments, the machine - learned multi - class object classification model 406 is capable of processing image data 402. Alternatively, in some embodiments, the machine - learned multi - class object classification model 406 is capable of processing a portion of the image data 402 that is predicted to include an object. For example, the image data can be a portion of the image data 402 that includes a bear - class object 402C (e.g., a portion of the image data 402 not defined by a bounding box associated with the labels 404A - 404B, etc.). The image data 402 can be processed to obtain object classification data 408. The object classification data 408 output by the machine - learned multi - class object classification model 406 can include multi - class classification outputs 408A - 408C of the (multiple) objects (e.g., 402A - 402C) that the machine - learned multi - class object classification model is configured to classify.

[0087] The object classification data 408 can include a class output 408A - 408C for each of the classes (e.g., lion class, tiger class, bear class, etc.) that the machine - learned multi - class object classification model 406 is configured to classify. As an example, each of the class outputs 408A - 408C can include a prediction as to whether the object depicted in the image data 402 belongs to the corresponding class. For example, the class 1 output 408A can indicate that the object belongs to the lion class 402A, and the class outputs 408B / 408C can indicate that the object does not belong to their respective classes 402B / 402C. As another example, each of the class outputs 408A - 408C can indicate the probability that the object depicted in the image data (or a portion of the image data) belongs to each of the four classes. For example, the class 1 output 408A can indicate a 15% probability that the depicted object is a tiger, while the class 2 output 408B and the class 3 output 408C can indicate other probabilities that the object is a certain class.

[0088] A loss function 410 can be evaluated, which evaluates the difference between the object classification data 408 and a subset of the ground - truth data 404's object class annotations (e.g., the difference between the predicted object class annotations 408A - 408C and the ground - truth object class annotations 404A - 404B, etc.). More particularly, the loss function 410 can include multiple weighted loss signals 410A - 410C (e.g., a weighted lion loss signal 410A for the lion class 402A, a weighted tiger loss signal 410B for the tiger class 402B, a weighted bear loss signal 410C for the bear class 402C, etc.) respectively associated with multiple object classes 402A - 402C. The weight of each of the weighted loss signals 410A - 410C can be at least partially based on whether the object class 402A - 402C associated with the corresponding loss signal 410A - 410C is included within the subset of object classes 402A - 402C.

[0089] The weighted lion loss signal 410A and the weighted tiger loss signal 410B are respectively associated with classes 402A and 402B included in the subsets of classes 402A / 402B of the multiple classes 402A - 402C (e.g., marked with object class annotations 404A / 404B in the ground truth data 402). The weighted bear loss signal 410C is associated with class 402C that is not included in the subsets of classes 402A / 402B of the multiple classes 402A - 402C (e.g., not marked with the corresponding object class annotation in the ground truth data 402). Since there is no object class annotation in the ground truth data 404 of class 402C, the accuracy (e.g., loss) of the class output 408C of the bear class 402C may not be correctly evaluated. Thus, due to the unknown classification accuracy of the output 408C, the corresponding weight of the weighted bear loss signal 410C can be lower than the weights of the loss signals (e.g., 410A and 410B) that provide labels, thereby reducing the overall impact of the weighted bear loss signal 410C on the final loss value 412.

[0090] The final loss value 412 can be determined based at least in part on the weighted loss signals 410A, 410B, and 410C. In some embodiments, the weighted loss signals of classes not included in the class subset (e.g., the weighted loss signals of classes without corresponding object class annotations) can be weighted such that the loss signals are excluded from the determination of the final loss value 412. As an example, the weight of the weighted loss signal 410C can be a zero weight, and the final loss value 412 can be based on each weighted loss signal (e.g., 410A and 410B) with a weight greater than zero. In some embodiments, the weight of each of the weighted loss signals can be a normalized value. For example, the weight of the weighted bear loss signal 410C can be 0.35, while the weights of the weighted lion and tiger loss signals 410A / 410B can each be weighted to 1 (e.g., to reduce the correlation of the weighted bear loss signal 410C with the calculation of the final loss value, etc.). The final loss value 412 can be determined based on the respective weights of each of the loss signals 410A - 410C.

[0091] One or more parameter adjustments 414 can be generated based at least in part on the loss function 410 and / or the final loss value 412, and the parameter adjustments can be applied to a multi-class object classification model of machine learning (e.g., using a gradient descent algorithm, etc.). In this way, for model outputs corresponding to classes without corresponding class annotations, the penalty on the multi-class object classification model of machine learning derived from the loss value 412 can be reduced, thus helping to train the multi-class object classification model 406 of machine learning using only partially labeled image data 402 (e.g., image data including objects without corresponding labels, etc.).

[0092] Example Method

[0093] Figure 5 FIG. 500 is a flow chart depicting an example method 500 of performing training of a multi-class object classification model of machine learning using partially labeled training data in accordance with an example embodiment of the present disclosure. Although, for purposes of illustration and discussion, Figure 5 steps are depicted as being performed in a particular order, the methods of the present disclosure are not limited to the particular order or arrangement shown. The various steps of method 500 can be omitted, rearranged, combined, and / or adapted in various ways without departing from the scope of the present disclosure.

[0094] At 502, the computing system can obtain image data depicting an object and ground truth data including object class annotations. More particularly, the image data depicting one or more objects can be obtained by the computing system together with the ground truth data. The ground truth data can include a subset of object class annotations (e.g., ground truth labels) respectively associated with a subset of object classes from a plurality of object classes. As an example, the image data can be or otherwise include one or more images (e.g., images, multiple video frames, etc.) from a training data set that is configured to train a classification model of machine learning to identify and classify multiple classes. An image from the training data set can include four objects belonging to four separate classes. The subset of object class annotations can be or otherwise include annotations for the first two of the four separate classes. As another example, the ground truth data for a second image of the training data set can include object class annotations for the remaining two of the four separate classes. In this way, in some embodiments, the subset of class annotations included in the ground truth data can include different classes from multiple classes in different images included in the training data set.

[0095] In some embodiments, an object class annotation can include a bounding box that defines a region of the image data and an associated class label. As an example, the image data can depict a lion object. The object class annotation can include a bounding box that defines a region of the image data that includes the lion and the corresponding "lion" label. Alternatively or additionally, in some embodiments, the object class annotation can label the entire image of the image data as including or not including the class. As an example, the image data may not depict a tiger. The object class annotation can indicate that no tiger is depicted in the image data.

[0096] In some embodiments, the image data can be a portion of the image data from an image. As an example, the image data can be the portion of the image data that is extracted from the image based on a prediction (e.g., extracted by a region proposal network (RPN), etc.) that depicts an object in a portion of the image data. As another example, the image data can be a portion of the image data that is not defined by the bounding box of the object class annotation. Alternatively, in some embodiments, the image data can be or otherwise include the entire depiction of the image data. As an example, the image data can be obtained and then iteratively processed by a machine learning multi-class object classification model to identify regions predicted to include an object and then classify the object subsequently or simultaneously.

[0097] At 504, the computing system can utilize a machine learning multi-class object classification model to process the image data. More particularly, the computing system can utilize a machine learning multi-class object classification model to process the image data to obtain object classification data. The machine learning multi-class object classification model can be or otherwise include any kind of conventional machine learning object classification model. As an example, the machine learning model can be or can otherwise include one or more neural networks (e.g., deep neural networks, recurrent neural networks, graph neural networks, etc.). The neural network (e.g., deep neural network) can be a feedforward neural network, a convolutional neural network, and / or various other types of neural networks.

[0098] A multi-class object classification model of machine learning can be configured to classify multiple object classes. More particularly, a multi-class object classification model of machine learning can be configured to process image data to classify an object depicted in the image data as belonging to one of multiple classes. Additionally or alternatively, in some embodiments, a multi-class object classification model of machine learning can be part of a machine learning image analysis model that performs object recognition and object classification (e.g., a component of the model, one or more layers of the model, etc.). As an example, a multi-class object classification model of machine learning can be the object classification part of a conventional two-stage object detector model (e.g., Fast R-CNN, etc.). For example, the first stage of a two-stage object detector model (e.g., a Region Proposal Network (RPN), etc.) can obtain image data and output regions of the image data that are predicted to contain an object (e.g., a region not defined by a bounding box, etc.) (e.g., using an anchor-based method, etc.). The regions of the image data can be processed using a multi-class object classification model of machine learning to obtain object classification data.

[0099] As another example, a multi-class object classification model of machine learning can be or otherwise incorporated into a one-stage object detector model (e.g., a Single Shot Detector (SSD) model, etc.). For example, a multi-class object classification model of machine learning can be or otherwise include multiple machine learning layers (e.g., (multiple) convolutional layers, (multiple) activation layers, etc.) that predict (multiple) regions that may include a predicted object, and the predicted object can subsequently be processed using an additional classification layer of the multi-class object classification model to produce object classification data. For another example, a multi-class object classification model of machine learning can include multiple layers that predict regions of the image data to include a predicted object and simultaneously classify the predicted object. In this way, a method of training a multi-class object classification model of machine learning can be applied to any conventional machine learning image analysis model (e.g., a Single Shot Detector (SSD) model, a You Only Look Once (YOLO) model, a Fast R-CNN model, etc.).

[0100] Object classification data output by a multi-class object classification model of machine learning can provide a multi-class classification output for an object depicted in image data. As an example, a portion of an image can depict a bear object. The multi-class object classification model of machine learning can be configured to classify the object as belonging to one of four classes (e.g., bear, tiger, kangaroo, lion, etc.). The object classification data can include a plurality of predicted object class annotations that indicate whether the object belongs to each of the classes (e.g., bear 1, tiger 0, kangaroo 0, lion 0, etc.). Alternatively, in some embodiments, the object classification data can include a plurality of object class probability predictions that the object depicted in the image data belongs to each of the four classes (e.g., 80% bear, 15% tiger, 15% kangaroo, 10% lion, etc.). In some embodiments, the object classification data can include a single indication that the object depicted in the image data belongs to one class. As an example, the object classification data can be or otherwise include a predicted class annotation (e.g., a "bear" class annotation, etc.). Thus, the object classification data can include a plurality of predicted probability outputs for the respective plurality of classes and / or a predicted class annotation of the object depicted in the image data or a portion of the image data.

[0101] At 506, the computing system can evaluate a loss function. More particularly, the computing system can evaluate a loss function that evaluates a multi-class classification loss. The multi-class classification loss can include a difference between the object classification data and a subset of object class annotations (e.g., a difference between a predicted object class annotation and a ground truth object class annotation, etc.). More particularly, the loss function can include a plurality of weighted loss signals associated with the respective plurality of object classes (e.g., a bear weighted loss signal for the bear class, a lion weighted loss signal for the lion, etc.). The weight of each of the weighted loss signals can be at least partially based on whether the object class associated with the respective loss signal is included in the object class subset.

[0102] As an example, the first weighted loss signal can be associated with a first category from multiple object categories, where the first category is not included in the subset of object categories. For example, the first category can be the kangaroo category, and the subset of object categories can include the bear category, the lion category, and the tiger category. Since the first category is not included in the category subset, the weight of the first loss signal can be weighted to reduce the impact of the first loss signal relative to the loss function. In some embodiments, the weight of each of the weighted loss signals can be a normalized value. For example, the weight of the first loss signal can be weight 0, while the weight of the second loss signal associated with a category included in the category subset can be weighted 1 (e.g., to eliminate the loss signal from the loss function, etc.). As another example, the weight of the first loss signal can be weight 0.35, while the weight of the second loss signal associated with a category included in the category subset can be weighted 1 (e.g., to reduce the correlation of the first loss signal relative to the second loss signal, etc.). In this way, the weights of the weighted loss signals can be configured to reduce or eliminate the impact of the corresponding loss signals that are not associated with the categories included in the category subset (e.g., categories with associated object category annotations (labels), etc.).

[0103] At 508, the computing system can adjust the parameters of a multi-class object classification model of machine learning. More particularly, the computing system can adjust one or more parameters of the multi-class object classification model of machine learning at least in part based on the loss function. More particularly, the (multiple) parameters of the multi-class object classification model of machine learning can be adjusted based on each of the weighted loss signals and their respective weights. In some embodiments, the (multiple) parameters of the multi-class object classification model of machine learning can be adjusted proportionally to the weights of each weighted loss signal of the loss function. As an example, the loss function can be used to calculate a final loss value based on an evaluation of the difference between the object classification data and the ground truth data. The final loss value can be proportional to the weights of the corresponding weighted loss signals. For example, a weighted loss signal with a weight of zero may not contribute to the calculation of the final loss value. The loss value can be backpropagated through the multi-class classification model of machine learning, and one or more parameters of the model can be adjusted based on the final loss value (e.g., using a gradient descent algorithm, etc.). In this way, the adjustment of the (multiple) parameters of the model can be proportional to the weights of the weighted loss signals, thus reducing the impact of the loss signals associated with the unlabeled categories during model training.

[0104] In some embodiments, an evaluation loss function can include a subset of weighted loss signals of the evaluation loss function, the subset being associated with a subset of object classes respectively. More particularly, the loss function can include only weighted loss signals associated with labeled object classes (e.g., a subset of object class annotations of ground truth data, etc.). In this way, the loss function can exclude any loss signals associated with object classes not annotated by a subset of object class annotations (e.g., not labeled in the training data, etc.).

[0105] In some embodiments, a multi-class object classification model of machine learning can be utilized after training. More particularly, a computing system can obtain additional image data depicting one or more additional objects. The machine learning multi-class object classification model can be used to process the image data to obtain an image classification output. In some embodiments, the image classification output can include one or more labels describing the additional image data. As an example, the image classification output can include one or more object class annotations of one or more additional objects. As another example, the image classification output can include one or more image annotations that annotate the entire image as belonging to one or more classes. For example, the machine learning multi-class object classification model can label one or more additional objects as being semantically related to a "natural" image (e.g., a rabbit object, a tree object, a boulder object, a grass object, etc.). Based on the (multiple) object classification labels, the image classification data can label the image as belonging to the "natural" image category.

[0106] Additional Disclosure

[0107] The techniques discussed herein relate to servers, databases, software applications, and other computer-based systems, as well as the actions taken and the information sent to and received from such systems. The inherent flexibility of computer-based systems allows for many possible configurations, combinations, and divisions of tasks and functions among components. For example, the processes discussed herein can be implemented using a single device or component or multiple devices or components working in combination. The database and application can be implemented on a single system or can be distributed across multiple systems. The distributed components can operate sequentially or in parallel.

[0108] While the subject matter has been described in detail with respect to various specific example embodiments, each example has been provided by way of explanation and not as a limitation of the disclosure. Those skilled in the art will be able to readily generate alterations, variations, and equivalents of such embodiments after understanding the foregoing. Accordingly, the subject matter disclosure does not exclude including such modifications, variations, and / or additions to the subject matter that would be apparent to a person of ordinary skill in the art. For example, features that are illustrated or described as part of one embodiment can be used with another embodiment to yield yet another embodiment. Accordingly, the disclosure is intended to cover such alterations, variations, and equivalents.

Claims

1. A computer-implemented method for training a multi-class object classification model of machine learning using partially labeled training data, comprising: obtaining, by a computing system including one or more computing devices, image data depicting one or more objects and ground truth data including a subset of object class annotations respectively associated with a subset of object classes of a plurality of object classes; processing, by the computing system, the image data using the multi-class object classification model of the machine learning to obtain object classification data; evaluating, by the computing system, a loss function that evaluates a multi-class classification loss including a difference between the object classification data and the subset of object class annotations, wherein the loss function includes a plurality of weighted loss signals respectively associated with the plurality of object classes, and wherein a weight of each of the weighted loss signals is at least partially based on whether the object class associated with the corresponding loss signal is included in the subset of object classes; adjusting, by the computing system, one or more parameters of the multi-class object classification model of the machine learning at least partially based on the loss function; wherein: the weight of each of the weighted loss signals is a normalized value; and the weight of the weighted loss signal associated with the object class included in the subset of object classes is greater than the weight of the weighted loss signal associated with the object class excluded from the subset of object classes; and wherein the subset of object class annotations associated with the subset of object classes includes a label indicating that an object of the object class in the subset of object classes is not depicted in the image data.

2. The computer-implemented method according to claim 1, wherein: the weight of the weighted loss signal associated with the object class included in the subset of object classes is one; and the weight of the weighted loss signal associated with the object class excluded from the subset of object classes is zero.

3. The computer-implemented method according to claim 1, wherein, Adjusting the one or more parameters of the multi-class object classification model of the machine learning includes adjusting, by the computing system, the one or more parameters of the multi-class object classification model of the machine learning at least partially based on each of the plurality of weighted loss signals having a weight greater than zero.

4. The computer-implemented method according to claim 1, wherein, The one or more parameters of the multi-class object classification model of the machine learning are adjusted proportionally to the weight of each of the weighted loss signals of the loss function.

5. The computer-implemented method according to claim 1, wherein: the annotation of the subset of object class annotations includes a bounding box and a label of the corresponding object of the first object class depicted in the image data.

6. The computer-implemented method according to claim 1, wherein, The object classification data includes one or more predicted object class annotations predicting one or more objects to be depicted in the image data.

7. The computer-implemented method according to claim 1, wherein Evaluating the loss function includes evaluating, by the computing system, a subset of the weighted loss signals of the loss function respectively associated with the subset of object classes.

8. The computer-implemented method according to any one of claims 1-7, wherein the method further comprises: Obtain additional image data depicting one or more additional objects by the computing system; and Process the additional image data by the computing system using the multi-class object classification model of the machine learning to obtain an image classification output, the image classification output including one or more labels describing the additional image data.

9. One or more tangible non-transitory computer-readable media storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 1-8.

10. A computing system for training a multi-class object classification model using partially labeled training data, comprising: One or more processors; A multi-class object classification model of machine learning, the multi-class object classification model of machine learning being configured to classify multiple object classes; and One or more tangible non-transitory computer-readable media storing computer-readable instructions that, when executed by the one or more processors, cause the one or more processors to perform the method according to any one of claims 1-8.