Non-forgetting dynamic class incremental learning

By using a forgetting-free dynamic class incremental learning method, new categories are dynamically added without retraining the network, solving the catastrophic forgetting problem of deep neural networks when faced with new data, and improving the adaptability and efficiency of neural networks.

CN120917459APending Publication Date: 2025-11-07QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380096007.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-03-24
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing deep neural networks are prone to catastrophic forgetting when faced with new data, and cannot effectively learn new categories without affecting performance on old categories.

Method used

We employ a forgetting-free dynamic class incremental learning method, which receives input through a first artificial neural network, extracts features to generate representations, and generates embeddings based on multiple similarity measures, dynamically adding new categories without retraining the network.

Benefits of technology

This enables dynamic learning of new categories without affecting the performance of old categories, improving the adaptability and efficiency of neural networks while reducing computational resources and storage requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120917459A_ABST
    Figure CN120917459A_ABST
Patent Text Reader

Abstract

A processor-implemented method for forgetting-free dynamic class incremental learning includes receiving an input through an artificial neural network (ANN). The ANN extracts features of the input to generate a representation of the input. An embedding is generated by the ANN based on the representation and a multi-similarity metric. The ANN generates a new class without retraining the ANN. The new class is generated based on a comparison of the embedding and a set of previous embedding.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Aspects of the disclosure generally relate to artificial neural networks, and more particularly, to dynamic class-incremental learning without forgetting. BACKGROUND

[0002] An artificial neural network can include an interconnected set of artificial neurons (e.g., neuron models). An artificial neural network can be a computing device or a method represented as being performed by a computing device.

[0003] Neural networks are composed of operations that consume and produce tensors. Neural networks can be used to solve complex problems; however, the time it takes for the network to complete the task can be long due to the size of the network and the amount of operations that can be performed to produce a solution can be vast. Moreover, since these tasks can be performed on mobile devices that can have limited computational capabilities, the computational cost of deep neural networks can be problematic.

[0004] A convolutional neural network is a type of feed-forward artificial neural network. A convolutional neural network can include a collection of neurons, where each neuron has a receptive field and collectively tiles the input space. Convolutional neural networks (CNNs), such as deep convolutional neural networks (DCNs), have numerous applications. In particular, these neural network architectures are used in various technologies, such as image recognition, pattern recognition, speech recognition, autonomous driving, and other classification tasks.

[0005] With the rapid development of deep learning, current deep models can learn a fixed number of classes with high performance. However, data often comes from open environments and can be in a streaming format or can only be temporarily available, e.g., due to privacy concerns. Thus, learning new classes of data without restarting the training process is challenging. SUMMARY

[0006] The disclosure is set forth in independent claims. Some aspects of the disclosure are described in dependent claims.

[0007] In one aspect of the disclosure, a processor-implemented method includes receiving an input by a first artificial neural network (ANN). The method also includes extracting, by the first ANN, features of the input to generate a representation of the input. The method yet further includes generating, by the first ANN, an embedding based on the representation and a plurality of similarity measures. The method further includes generating, by the first ANN, a new class without retraining the first ANN, the new class being generated based on a comparison of the embedding and a set of previous embeddings.

[0008] Another aspect of the disclosure relates to an apparatus having a memory and one or more processors coupled to the memory. The processor is configured to receive an input through a first artificial neural network (ANN). The processor is further configured to extract features of the input through the first ANN to generate a representation of the input. The processor is yet further configured to generate an embedding through the first ANN based on the representation and a plurality of similarity measures. The processor is also configured to generate a new class through the first ANN without retraining the first ANN, the new class being generated based on a comparison of the embedding and a set of previous embeddings.

[0009] In another aspect of the disclosure, a processor-implemented method includes receiving an input at a user device through an artificial neural network (ANN). The method also includes extracting features of the input through the ANN to generate a representation of the input. The method yet further includes generating an input embedding through the ANN based on the representation and a plurality of learned similarities. The method also includes comparing the input embedding to a set of embeddings stored at the user device. The method further includes generating an inference for the input based on the comparison of the input embedding and the set of embeddings.

[0010] Another aspect of the disclosure relates to an apparatus having a memory and one or more processors coupled to the memory. The processor is configured to receive an input at a user device through an artificial neural network (ANN). The processor is further configured to extract features of the input through the ANN to generate a representation of the input. The processor is yet further configured to generate an input embedding through the ANN based on the representation and a plurality of learned similarities. The processor is also configured to compare the input embedding to a set of embeddings stored at the user device. The processor is further configured to generate an inference for the input based on the comparison of the input embedding and the set of embeddings.

[0011] Additional features and advantages of the disclosure will be described in the following detailed description. It will be appreciated that the disclosure can be readily adapted for use as a basis for modifying or designing other structures for carrying out the same purposes as the disclosure. It will also be appreciated that such equivalent constructions do not depart from the spirit and scope of the disclosure as set forth in the appended claims. Novel features which are believed to be characteristic of the disclosure, both as to its organization and method of operation, together with further objects and advantages will be better understood from the following description when considered in connection with the accompanying drawings. It is to be expressly understood, however, that the drawings are included merely for purposes of illustration and description and are not intended to limit the disclosure, as defined by the appended claims. BRIEF DESCRIPTION OF DRAWINGS

[0012] The features, nature, and advantages of the present disclosure will become more apparent from the detailed description set forth below when taken in conjunction with the drawings in which like reference characters identify correspondingly throughout and wherein:

[0013] Figure 1 An example implementation of a neural network using a system on a chip (SoC) including a general purpose processor according to certain aspects of the present disclosure is illustrated.

[0014] Figure 2A Figure 2B and Figure 2C are diagrams illustrating a neural network according to aspects of the present disclosure.

[0015] Figure 2D is a diagram illustrating an example deep convolutional network (DCN) according to aspects of the present disclosure.

[0016] Figure 3 is a block diagram illustrating an example deep convolutional network (DCN) according to aspects of the present disclosure.

[0017] Figure 4 is a block diagram illustrating an example software architecture that can modularize artificial intelligence (AI) functionality.

[0018] Figure 5 is a block diagram illustrating an example architecture for dynamic class-incremental learning without forgetting according to aspects of the present disclosure.

[0019] Figure 6A is a diagram illustrating an example multi-similarity loss according to aspects of the present disclosure.

[0020] Figure 6B shows an example plot illustrating weighting positive and negative pairs shown in Figure 6A based on multiple similarities according to aspects of the present disclosure.

[0021] Figure 7 is a visual flow diagram illustrating an example process for computing a multi-similarity loss according to aspects of the present disclosure.

[0022] Figure 8 and Figure 9 are flow diagrams illustrating a processor-implemented method for dynamic class-incremental learning without forgetting according to aspects of the present disclosure. DETAILED DESCRIPTION

[0023] ​The detailed description set forth below, in connection with the appended drawings and description of specific configurations, is intended as a description of various configurations and is not intended to represent the only configurations in which the concepts described herein can be practiced. The detailed description includes specific details for the purpose of providing a thorough understanding of various concepts. However, it will be apparent to those skilled in the art that these concepts can be practiced without these specific details. In some instances, well-known structures and components are shown in block diagram form in order to avoid obscuring such concepts.

[0024] Based on the teachings herein one of ordinary skill in the art will appreciate that the scope of the disclosure is intended to cover any aspect of the disclosure, whether implemented independently of, or combined with, any other aspect of the disclosure. For example, an apparatus can be implemented or a method can be practiced using any number of the aspects set forth. In addition, the scope of the disclosure is intended to cover such an apparatus or method which is practiced using, as substitute for, or in combination with, some other aspect of the disclosure. It is understood that any aspect of the disclosure disclosed can be embodied by one or more elements of a claim.

[0025] The word "exemplary" is used herein to mean "serving as an example, instance, or illustration." Any aspect described as "exemplary" is not necessarily to be construed as preferred or advantageous over other aspects.

[0026] While specific aspects are described, numerous variations and permutations of the aspects are possible. While some aspects are described with respect to particular examples, other implementations are possible. While features, benefits, and advantages of the preferred aspects are discussed herein, additional features, benefits, and advantages can be realized. Although specific configurations have been illustrated and described herein, it will be appreciated by those of ordinary skill in the art that any arrangement which is calculated to achieve the same purpose can be substituted for the specific embodiments shown. This disclosure is intended to cover all adaptations or variations of preferred aspects. Therefore, it is intended that the description contained herein be considered in all respects as illustrative only and not as limiting the scope of the disclosure, which is set out in the appended claims and their equivalents.

[0027] Incremental learning aims to develop artificial intelligence systems that can continually learn to solve new tasks from new data while retaining knowledge learned from previously learned tasks. Incremental learning can enable efficient resource usage by eliminating retraining from scratch as new data arrives. In addition, incremental learning can reduce memory usage by limiting the amount of data to store, which can be important when privacy restrictions are imposed. Furthermore, incremental learning can enable artificial neural networks to learn in a way that is more highly analogous to human learning.

[0028] When new data different from the data included in the initial training dataset for the artificial neural network is presented to the artificial neural network, attempts to classify the new data can result in erroneous classifications. One conventional approach to address this issue involves retraining the deep neural network to modify the class head (e.g., the last fully connected (FC) layer) and retrain / fine-tune the state-of-the-art (SOTA) model deep neural network (DNN) (e.g., convolutional neural network such as convolutional neural network (CNN) / transformer) with the incoming new data (e.g., unseen classes). This approach can be applied, for example, when training a robot in the open world (when the robot observes new objects or in e-commerce platforms when new types of products appear). Additionally, examples can include accident data in autonomous driving scenarios and expensive data annotation (e.g., disease data for medical diagnosis). However, retraining / fine-tuning can suffer from catastrophic forgetting. Catastrophic forgetting refers to a neural network model losing its generalization capability for a task after being trained on a new task. The new task can override the weights learned by the neural network and thus can degrade performance (e.g., classifier accuracy) for past tasks. That is, performance for inference on previous classes can sharply decrease due to the lack of previous data, especially on-device data.

[0029] Accordingly, to address these and other challenges, aspects of the present disclosure relate to no-forgetting dynamic class-incremental learning.

[0030] Figure 1 An example implementation of a system on chip (SoC) 100 is illustrated, which can include a central processing unit (CPU) 102 or multi-core CPU configured for no-forgetting dynamic class-incremental learning (neuro end-to-end network). Variables (e.g., neural signals and synaptic weights), system parameters associated with the computing device (e.g., neural network with weights), delays, frequency bin information, and task information can be stored in a memory block associated with a neural processing unit (NPU) 108, a memory block associated with the CPU 102, a memory block associated with a graphics processing unit (GPU) 104, a memory block associated with a digital signal processor (DSP) 106, a memory block 118, or can be distributed across multiple blocks. Instructions executed at the CPU 102 can be loaded from a program memory associated with the CPU 102 or can be loaded from the memory block 118.

[0031] SoC 100 can also include additional processing blocks tailored for specific functions, such as GPU 104, DSP 106, connectivity block 110 (which can include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and a multimedia processor 112 that may, for example, detect and recognize gestures. In one implementation, NPU 108 is implemented in CPU 102, DSP 106, and / or GPU 104. SoC 100 can also include a sensor processor 114, an image signal processor (ISP) 116, and / or a navigation module 120 (which can include a global positioning system).

[0032] SoC 100 can be based on an ARM instruction set. In an aspect of the disclosure, instructions loaded into general purpose processor 102 can include code to receive an input by a first artificial neural network (ANN). General purpose processor 102 can also include code to extract features of the input by the first ANN to generate a representation of the input. General purpose processor 102 can additionally include code to generate an embedding by the first ANN based on the representation and a multi-similarity measure. General purpose processor 102 can also include code to generate a new class by the first ANN without retraining the first ANN. The new class is generated based on a comparison of the embedding and a set of previous embeddings.

[0033] In another aspect of the disclosure, instructions loaded into general purpose processor 102 can include code to receive an input by an artificial neural network (ANN) at a user device. General purpose processor 102 can also include code to extract features of the input by the ANN to generate a representation of the input. General purpose processor 102 can additionally include code to generate an input embedding by the ANN based on the representation and a learned multi-similarity. General purpose processor 102 can also include code to compare the input embedding to a set of embeddings stored at the user device. General purpose processor 102 can also include code to generate an inference for the input based on the comparison of the input embedding and the set of embeddings.

[0034] Deep learning architectures can perform object recognition tasks by learning to represent inputs at successively higher levels of abstraction in each layer, building a useful feature representation of the input data. In this way, deep learning addresses a major bottleneck of traditional machine learning. Prior to the advent of deep learning, machine learning approaches to object recognition problems can have relied heavily on human-designed feature objects, possibly in conjunction with a shallow classifier. A shallow classifier can be a two-class linear classifier, for example, in which a weighted sum of the components of a feature vector can be compared to a threshold to predict which class the input belongs to. Human-designed feature objects can be templates or kernels customized for a particular problem domain by engineers with domain expertise. In contrast, while deep learning architectures can learn to represent features similar to those that a human engineer might design, this occurs through training. Moreover, deep networks can learn to represent and recognize new types of features that a human might not have considered.

[0035] Deep learning architectures can learn a hierarchy of features. For example, if presented with visual data, a first layer can learn to recognize relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, a first layer can learn to recognize spectral power in particular frequencies. A second layer, taking the output of the first layer as input, can learn to recognize combinations of features, such as simple shapes in visual data or sound combinations in auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Still higher layers can learn to recognize common visual objects or spoken phrases.

[0036] Deep learning architectures can perform particularly well when applied to problems with a natural hierarchical structure. For example, classification of motorized vehicles can benefit from first learning to recognize wheels, windshields, and other features. These features can be combined in different ways at higher layers to recognize cars, trucks, and airplanes.

[0037] Neural networks can be designed with a variety of connectivity patterns. In feedforward networks, information passes from lower to higher layers, with each neuron in a given layer communicating with neurons in higher layers. As described above, a hierarchical representation can be built in successive layers of a feedforward network. Neural networks can also have recurrent or feedback (also known as top-down) connections. In recurrent connections, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can be helpful in recognizing patterns that span more than one block of input data delivered to the neural network in sequence. Connections from a neuron in a given layer to a neuron in a lower layer are known as feedback (or top-down) connections. Networks with many feedback connections can be helpful when recognition of high-level concepts can aid in discriminating particular low-level features of an input.

[0038] The connections between the layers of a neural network can be fully connected, or locally connected. Figure 2A An example of a fully connected neural network 202 is illustrated. In a fully connected neural network 202, a neuron in a first layer can communicate its output to every neuron in a second layer, so that every neuron in the second layer will receive input from every neuron in the first layer. Figure 2B An example of a locally connected neural network 204 is illustrated. In a locally connected neural network 204, a neuron in a first layer can be connected to a limited number of neurons in a second layer. More generally, a locally connected layer of a locally connected neural network 204 can be configured so that every neuron in the layer will have the same or similar connectivity pattern, but the connection strengths can have different values (e.g., 210, 212, 214, and 216). The locally connected connectivity pattern can result in spatially distinct receptive fields in higher layers, as higher layer neurons in a given region can receive input that is tuned through training to properties of a restricted portion of the total input to the network.

[0039] One example of a locally connected neural network is a convolutional neural network. Figure 2C An example of a convolutional neural network 206 is illustrated. A convolutional neural network 206 can be configured so that the connection strengths associated with input to each neuron in a second layer are shared (e.g., 208). Convolutional neural networks can be well suited for problems in which the spatial location of input is meaningful.

[0040] One type of convolutional neural network is a deep convolutional network (DCN). Figure 2D A detailed example of a DCN 200 designed to recognize visual features from images 226 input by an image capture device 230, such as a vehicle-mounted camera, is illustrated. The DCN 200 of the current example can be trained to identify traffic signs and the numbers provided on traffic signs. Of course, the DCN 200 can be trained for other tasks, such as identifying lane markings or identifying traffic signal lights.

[0041] Supervised learning can be used to train the DCN 200. During training, an image (such as image 226 of a speed limit sign) can be presented to the DCN 200, and forward passes can then be computed to produce output 222. The DCN 200 may include a feature extraction part and a classification part. Upon receiving image 226, convolutional layer 232 can apply a convolutional kernel (not shown) to image 226 to generate a first set 218 of feature maps. As an example, the convolutional kernel used for convolutional layer 232 may be a 5x5 kernel that generates 28x28 feature maps. In this example, because four different feature maps are generated in the first set of feature maps 218, four different convolutional kernels are applied to image 226 at convolutional layer 232. Convolutional kernels may also be referred to as filters or convolutional filters.

[0042] The first set of feature maps 218 can be subsampled by a max-pooling layer (not shown) to generate a second set of feature maps 220. The max-pooling layer reduces the size of the first set of feature maps 218. That is, the size of the second set of feature maps 220 (e.g., 14x14) is smaller than the size of the first set of feature maps 218 (e.g., 28x28). The reduced size provides similar information to subsequent layers while reducing memory consumption. The second set of feature maps 220 can be further convolved via one or more subsequent convolutional layers (not shown) to generate one or more subsequent sets of feature maps (not shown).

[0043] exist Figure 2D In the example, the second set of feature maps 220 is convolved to generate a first feature vector 224. Furthermore, the first feature vector 224 is further convolved to generate a second feature vector 228. Each feature of the second feature vector 228 may include a number corresponding to a possible feature of image 226, such as "sign", "60", and "100". A softmax function (not shown) can convert the numbers in the second feature vector 228 into probabilities. Therefore, the output 222 of DCN 200 is the probability that image 226 includes one or more features.

[0044] In this example, the probabilities for "sign" and "60" in output 222 are higher than the probabilities for other numbers in output 222 (such as "30", "40", "50", "70", "80", "90", and "100"). Before training, output 222 generated by DCN 200 may be incorrect. Therefore, the error between output 222 and the target output can be calculated. The target output is the ground truth value of image 226 (e.g., "sign" and "60"). The weights of DCN 200 can then be adjusted so that output 222 of DCN 200 is more closely aligned with the target output.

[0045] To adjust the weights, a learning algorithm can compute a gradient vector for the weights. The gradient can indicate the amount by which the error will increase or decrease if the weights are adjusted. At the top level, the gradient can directly correspond to the value of the weights connecting the activation neurons in the second-to-last layer and the neurons in the output layer. In lower levels, the gradient can depend on the value of the weights and the computed error gradient of the higher level. The weights can then be adjusted to reduce the error. This way of adjusting the weights can be referred to as "backpropagation" because it involves a "backward pass" through the neural network.

[0046] In practice, the error gradient for the weights can be computed in a small number of examples, causing the computed gradient to approximate the true error gradient. This approximation method can be referred to as stochastic gradient descent. Stochastic gradient descent can be repeated until the achievable error rate for the entire system stops decreasing or until the error rate reaches a target level. After learning, new images can be presented to the DCN, and a forward pass through the network can produce an output 222 that can be considered an inference or prediction of the DCN.

[0047] A deep belief network (DBN) is a probabilistic model that includes multiple layers of hidden nodes. A DBN can be used to extract a hierarchical representation of a training dataset. A DBN can be obtained by stacking layers of restricted Boltzmann machines (RBMs). An RBM is a type of artificial neural network that can learn a probability distribution over a set of inputs. Because an RBM can learn a probability distribution without information about the class to which each input should be classified, RBMs are often used for unsupervised learning. Using a mixed paradigm of supervised and unsupervised learning, the bottom RBMs of a DBN can be trained in an unsupervised manner and can be used as feature extractors, while the top RBMs can be trained in a supervised manner (on the joint distribution of inputs from the previous layer and target classes) and can be used as classifiers.

[0048] A deep convolutional network (DCN) is a network of convolutional networks configured with additional pooling and normalization layers. DCNs have achieved state-of-the-art performance on many tasks. A DCN can be trained using supervised learning, in which both input and output targets are known for many examples and are used to modify the weights of the network by using a gradient descent method.

[0049] A DCN can be a feedforward network. Furthermore, as described above, connections from a neuron in a first layer of the DCN to a group of neurons in a next higher layer are shared across the neurons in the first layer. The feedforward and shared connections of a DCN can be used for fast processing. For example, the computational burden of a DCN can be much less than that of a similarly sized neural network that includes recurrent or feedback connections.

[0050] The processing of each layer of a convolutional network can be thought of as a spatially invariant template or basis projection. If the input is first decomposed into multiple channels, such as the red, green, and blue channels of a color image, then a convolutional network trained on that input can be thought of as three-dimensional, with two spatial dimensions along the axes of the image, and a third dimension capturing color information. The output of a convolutional connection can be viewed as forming a feature map in the next layer, where each element in the feature map (e.g., 220) receives input from a range of neurons in the previous layer (e.g., feature map 218) and from each of the multiple channels. The values in the feature map can be further processed with a nonlinearity, such as a rectification, max(0, x). Values from neighboring neurons can be further pooled, which corresponds to down-sampling, and can provide additional local invariance and dimensionality reduction. Normalization corresponding to whitening can also be applied through lateral inhibition between neurons in the feature map.

[0051] The performance of deep learning architectures can increase as more labeled data points become available or as computing power increases. Modern deep neural networks are typically trained with computational resources that are thousands of times the computational resources available to a typical researcher only fifteen years ago. New architectures and training paradigms can further boost the performance of deep learning. Rectified linear units can reduce a training problem known as vanishing gradients. New training techniques can reduce overfitting, and thus enable larger models to achieve better generalization. Encapsulation techniques can extract data in a given receptive field and further improve overall performance.

[0052] Figure 3 is a block diagram illustrating a deep convolutional network 350. Based on connectivity and weight sharing, the deep convolutional network 350 can include multiple different types of layers. As shown, the deep convolutional network 350 includes convolutional blocks 354A, 354B. Each of the convolutional blocks 354A, 354B can be configured with a convolutional layer (CONV) 356, a normalization layer (LNorm) 358, and a max-pooling layer (MAX POOL) 360. Figure 3

[0053] The convolutional layer 356 can include one or more convolutional filters that can be applied to input data to generate a feature map. Although only two convolutional blocks 354A, 354B are shown, the present disclosure is not so limited, but instead any number of convolutional blocks 354A, 354B can be included in the deep convolutional network 350 according to design preference. The normalization layer 358 can normalize the output of the convolutional filters. For example, the normalization layer 358 can provide whitening or lateral inhibition. The max-pooling layer 360 can provide a down-sampling aggregation over space to achieve local invariance and dimensionality reduction.

[0054] ​For example, a parallel filter bank of a deep convolutional network can be loaded onto CPU 102 or GPU 104 of SoC 100 to achieve high performance and low power consumption. In alternative embodiments, the parallel filter bank can be loaded onto DSP 106 or ISP 116 of SoC 100. Further, deep convolutional network 350 can access other processing blocks that can be present on SoC 100, such as sensor processor 114 and navigation module 120 that are specialized for sensors and navigation, respectively.

[0055] Deep convolutional network 350 can also include one or more fully connected layers 362 (FC1 and FC2). Deep convolutional network 350 can also include a logistic regression (LR) layer 364. There are weights (not shown) to be updated between each layer 356, 358, 360, 362, 364 of deep convolutional network 350. The output of each of these layers (e.g., 356, 358, 360, 362, 364) can be used as input to the next one of these layers (e.g., 356, 358, 360, 362, 364) in deep convolutional network 350 to learn hierarchical feature representations from input data 352 (e.g., images, audio, video, sensor data, and / or other input data) supplied at first convolutional block 354A. The output of deep convolutional network 350 is a classification score 366 of input data 352. Classification score 366 can be a set of probabilities, where each probability is a probability that input data includes a feature in a set of features.

[0056] Figure 4 is a block diagram illustrating an example software architecture 400 that can modularize artificial intelligence (AI) functionality. Using this architecture, various processing blocks (e.g., CPU 422, DSP 424, GPU 426, and / or NPU 428) of a system on a chip (SoC) 420 can be designed to support dynamic class-incremental learning without forgetting for AI applications 402, in accordance with aspects of the present disclosure.

[0057] AI applications 402 can be configured to invoke functionality defined in user space 404, which can, for example, provide detection and recognition of a scene that indicates a current operating location of a device. For example, AI applications 402 can configure microphones and cameras differently depending on whether the recognized scene is an office, a lecture hall, a restaurant, or an outdoor environment such as a lake. AI applications 402 can make a request for compiled program code associated with a library defined in AI function application programming interface (API) 406. This request can ultimately rely on an output of a deep neural network that is configured to provide an inference response based on, for example, video and positioning data.

[0058] The runtime engine 408, which can be compiled code of a runtime framework, can further be accessible by the AI application 402. For example, the AI application 402 can cause the runtime engine to request an inference at specific time intervals or triggered by events detected by a user interface of the application. When causing the runtime engine to provide an inference response, the runtime engine can in turn transmit a signal to an operating system, such as the kernel 412, in an operating system (OS) space running on the SoC 420. The operating system can in turn cause continuous quantization relaxation to be performed on the CPU 422, the DSP 424, the GPU 426, the NPU 428, or some combination thereof. The CPU 422 can be directly accessible by the operating system, while the other processing blocks can be accessed through drivers, such as the drivers 414, 416, or 418 for the DSP 424, the GPU 426, or the NPU 428, respectively. In an example instance, a deep neural network can be configured to run on a combination of processing blocks, such as the CPU 422, the DSP 424, and the GPU 426, or can run on the NPU 428.

[0059] The application 402, e.g., an AI application, can be configured to call functions defined in the user space 404, e.g., that can provide detection and recognition of a scene indicative of a current operating location of a device. For example, the application 402 can configure microphones and cameras differently depending on whether the recognized scene is an office, a lecture hall, a restaurant, or an outdoor environment such as a lake. The application 402 can make a request for compiled program code associated with a library defined in a SceneDetect application programming interface (API) 406 to provide an estimate of a current scene. The request can ultimately rely on an output of a differential neural network configured to provide a scene estimate based on, e.g., video and positioning data.

[0060] The runtime engine 408, which can be compiled code of a runtime framework, can further be accessible by the application 402. For example, the application 402 can cause the runtime engine to request a scene estimate at specific time intervals or triggered by events detected by a user interface of the application. When causing the runtime engine to estimate a scene, the runtime engine can in turn transmit a signal to the operating system 410, such as the kernel 412, in an operating system (OS) space running on the SoC 420. The operating system 410 can in turn cause computations to be performed on the CPU 422, the DSP 424, the GPU 426, the NPU 428, or some combination thereof. The CPU 422 can be directly accessible by the operating system, and the other processing blocks can be accessed through drivers, such as the drivers 414-418 for the DSP 424, for the GPU 426, or for the NPU 428. In an example instance, a differential neural network can be configured to run on a combination of processing blocks, such as the CPU 422 and the GPU 426, or can run on the NPU 428.

[0061] As described, aspects of the present disclosure relate to dynamic class-incremental learning without forgetting.

[0062] Figure 5 is a block diagram illustrating an example architecture 500 for dynamic class-incremental learning without forgetting, in accordance with aspects of the present disclosure. The example architecture can be, for example, an artificial neural network (ANN) model, in accordance with aspects of the present disclosure. The architecture 500 can be included on a server 510. The server 510 can be, for example, a remote computing device such as a cloud server.

[0063] Referring to Figure 5 , the example architecture 500 can include a main chain 504 and a projector 506. The main chain 504 can include, for example, a CNN model that can learn representations for downstream tasks (e.g., 350 shown in Figure 3 In some aspects, the main chain 504 can also utilize and include an autoencoder, a self-attention mechanism, and a transformer to effectively leverage the correlation between support sets in order to learn highly discriminative global features. As shown in Figure 5 The example architecture 500 can receive an input 502, as shown. The input 502 can include an image, a video, sequence data, or other types of data. For example, the input 502 can be supplied directly from one or more user devices, or can be streamed via the Internet. The input 502 can include, for example, a pair of images or a batch of samples. The main chain 504 can process the input 502, extracting features of the input 502 to generate a representation of the input 502. The representation can be supplied to the projector 506.

[0064] The projector 506 can be, for example, a multi-layer perceptron (MLP) model. The projector 506 can receive the representation from the main chain 504 and can project the representation to an embedding that can be contrasted with other representations, for example, using a multi-similarity loss.

[0065] For example, the example architecture 500 can be self-supervised and trained using a multi-similarity loss on an uncurated dataset of images. The input 502 can be mined for pairs of samples (also referred to as “examples”) and weighted against a pair-based loss function. The example architecture 500 can learn to minimize the distance (e.g., cosine distance) between similar samples and maximize the distance between dissimilar samples. Positive pairs of samples with the lowest (or lower) similarity score and negative pairs with the highest (or higher) similarity score can be assigned the highest (or higher) weight. Thus, the example architecture 500 can train an artificial neural network (ANN) model on the server 510. The ANN model can in turn be distributed to one or more user devices 520a, 520b. For example, the user devices (e.g., 520a or 520b) can be, for example, mobile electronic devices (e.g., smartphones, wearable computing devices (e.g., smart glasses)) or autonomous vehicles.

[0066] A user device (e.g., 520a or 520b) can train an ANN model on-device. Once the ANN model is trained on-device at the user device (e.g., 520a or 520b), an index (e.g., 522a or 522b) can be generated, and the index can include, for example, embeddings of various items (e.g., classes in the training dataset, no forgetting). In some aspects, the index (e.g., 522a or 522b) can be searchable. For example, a user device 520b can observe a vehicle 524. The on-device ANN model can process the observation (e.g., an image or a frame of a video of the vehicle 524 including the vehicle 524). The on-device ANN model can operate in a manner similar to that described for the example architecture 500. That is, the on-device ANN model can extract features of the observation of the vehicle 524 using the main chain 526 to generate a representation of the observation of the vehicle 524. The representation can be supplied to the projector 528, which can project the representation to an embedding. The embedding of an item can include a sample mean (e.g., a prototype feature) of the same class.

[0067] Additionally, an index (e.g., 522a or 522b) can be generated for new classes. New classes can be added dynamically. Advantageously, new classes can be added on-device without the need to retrain the on-device ANN (e.g., the main chain 526 or the projector 528) expensively. For example, the on-device ANN can generate an embedding (e.g., using the main chain 526 and the projector 528) for a representative sample of a new class. The embedding can then be saved and added to the index (e.g., 522a or 522b). Classification can be done by the nearest sample mean rule. Thus, at inference (e.g., query) time, a fast approximate nearest neighbor search (e.g., using an ANN) can be utilized to retrieve the closest matching item from the index (e.g., 522a or 522b) in sublinear time.

[0068] Figure 6A is a plot illustrating a multi-similarity loss of an example dataset 600 in accordance with aspects of the present disclosure. Reference is made to Figure 6AThe example dataset can include a collection of data samples (e.g., 602a-602d, 604a-604d, 606a-606d). The dataset 600 can be, for example, a training dataset. As described, the dataset 600 (e.g., training dataset) can be mined for pairs of samples (which can also be referred to as “examples”). Each shape can represent a class of each of the data samples. The rectangular shapes can represent positive pairs of samples (e.g., 602a-602d). The circular shapes can represent negative pairs of samples (e.g., 604a-604d). The triangular shapes can represent other pairs of samples (e.g., 606a-606d). Although four samples of each of the three classes are shown, this is merely for ease of illustration and understanding and not a limitation. Rather, any number of samples and classes can be included in the dataset.

[0069] A data sample from the dataset 600 can be randomly selected as an anchor sample. For example, as shown in the example of FIG. 6B, the sample 604a is selected as an anchor. The anchor sample 604a can be used to evaluate the similarity of pairs of samples in the dataset 600. Two other data samples in the dataset 600 can be selected and used to determine a plurality of similarities. A data sample from the same class as the anchor sample can be referred to as a “positive” sample (e.g., 606d). Another data sample from a different class than the anchor sample 604a can be referred to as a “negative” sample (e.g., 602a). Figure 6A

[0070] According to aspects of the disclosure, the degree of similarity of pairs of samples can be determined based on a plurality of similarity metrics. The plurality of similarity metrics can include, for example, similarity-S, similarity-P, and similarity-N. However, this is merely an example and not a limitation. Rather, other similarity metrics can also be used. Similarity-P refers to a relative similarity metric compared to positive pairs. Similarity-N refers to a relative similarity metric compared to negative pairs.

[0071] Similarity-S can refer to a self-similarity of negative pairs to positive pairs of a given anchor. Specifically, both negative pairs with large similarity-S and positive pairs with small similarity-S can be considered hard pairs that are more helpful for learning the embedding space compared to pairs that are easier to classify. As shown in the example of FIG. 6C, the self-similarity can be computed using the sample 604d and the anchor 604a as a positive pair and the sample 602a and the anchor 604a as a negative pair. Figure 6A

[0072] ​​Further, relative similarity measures, such as similarity-P and similarity-N, can be computed. For example, each of the similarity measures can be computed using a cosine similarity distance. A similarity score can be determined based on the plurality of similarity measures (e.g., similarity-S, similarity-P, or similarity-N). In turn, the sample pairs can be weighted according to the degree of similarity, which can be indicated by the similarity score.

[0073] Figure 6B Example plot 650a and 650b illustrating weighting positive pairs and negative pairs based on a plurality of similarities, respectively, in accordance with aspects of the disclosure are shown. As shown in example plot 650a, the positive sample pairs with the lowest similarity scores can be assigned the highest weights. As the similarity scores of the positive pairs increase, lower weights can be assigned to such positive pairs in descending order accordingly. On the other hand, as shown in example plot 650b, the negative sample pairs with the lowest similarity scores can be assigned the lowest weights. The weights can be increased according to the similarity scores, where the negative pairs with the highest similarity scores are assigned the highest weights. Figure 6A

[0074] For example, using a multi-similarity loss function can enable an ANN model to learn directly from a random set of images on the Internet. Further, the ANN model can learn from a random set of images without the need for careful and time-consuming data governance and labeling used in conventional computer vision training. The ANN can then output image embeddings. Accordingly, aspects of the disclosure can be beneficially applied to generate more powerful, more reasonable, and more robust computer vision models that can discover salient information in images by considering relationships between different objects observed. In some aspects, the ANN model (e.g., on-device ANN model) can learn from an ungoverned dataset of images via a greedy search or grid search of hyperparameters.

[0075] Figure 7 is a visual flowchart 700 illustrating an example process for computing a multi-similarity loss in accordance with aspects of the disclosure. With reference to Figure 7 , the flowchart 700 can retrieve index information from an index (e.g., in 522a or 522b). Each of the indices in the index can represent a same class of sample mean (e.g., a prototype of a feature). At 702, a miner can mine or search the index information for potential pairs. The potential pairs can be included in a distance matrix 704. And mining can be performed for similar sample pairs at 706. As Figure 7 shown, the multi-similarity can be determined according to a batched multi-similarity loss function. However, the disclosure is not limited thereto. Rather, other measures and techniques for determining the multi-similarity can also be employed.

[0076] In Figure 7 ​In the example, a greedy search can be implemented to search for hyperparameters, such as α, β, and λ, used to compute similarity scores (e.g., pair weights at 708). Hyperparameter α can refer to the weights applied to positive pairs. Hyperparameter β can refer to the weights applied to negative pairs. Hyperparameter λ can refer to the offset applied to the exponent in the batch pair multi-similarity loss function, which can be given by:

[0077]

[0078] Where m is the batch size (x1, x2, x3, ..., x...). m ), where i is the anchor sample (x) i The index of P) i Given an anchor point (x) i Positive samples, N i Given an anchor point (x) i The negative samples of P, k is the negative sample of P. i or N i The item index in S, and S ik It is the anchor point (x) i ) and its pair (x) k Distance metrics between (e.g., Euclidean distance or cosine similarity).

[0079] Based on similarity scores, it can be similar to, for example, relative to Figure 6B The described methods are weighted. For example, such as... Figure 7 As shown, given the example dataset 730 with anchor x1, the similarity of pairs can be determined. Samples x2 and x3 are in the same class (class A) and can each be paired with anchor x1 to form pairs (x1, x2) and (x1, x3) respectively. The similarity of pairs can be determined. In the example dataset 730, given an offset λ of 0.5, the similarity of pairs (x1, x2) can be e. -α(0.7-0.5) And the similarity with respect to (x1, x3) can be e -α(0.4-0.5) Weighting can be performed on the pairs based on similarity in weighting step 708. Because e -α(0.7-0.5) Less than e -α(0.4-0.5) Therefore, compared to the pair (x1, x3), the pair (x1, x2) can be assigned a greater weight.

[0080] On the other hand, in example dataset 750, given an offset λ of 0.5, the similarity of negative pairs can be determined. Samples x2 and x3 are in a different class (class B) from anchor x1, and can each be paired with anchor x1 to form negative pairs (x1, x2) and (x1, x3), respectively. In the example of dataset 750, given an offset λ of 0.5, the similarity of the negative pair (x1, x2) can be e. -β(0.3-0.5)And the similarity of the negative pair (x1, x3) can be e -β(0.1-0.5) The pair can be weighted based on the similarity in the weighting step 708. Because e -β(0.3-0.5) is greater than e -β(0.1-0.5) so the negative pair (x1, x2) can be assigned a greater weight than the negative pair (x1, x3).

[0081] Accordingly, aspects of the present disclosure can beneficially combine embedding representation learning plus indexing and searching incremental learning without forgetting learning on-device. Accordingly, aspects of the present disclosure can enable efficient resource usage by reducing, and in some aspects eliminating, retraining from scratch when new data arrives. In some aspects, memory usage can also be reduced by limiting the amount of data to store (e.g., when privacy restrictions are imposed). Additional aspects can include learning that is more highly analogous to human learning and on-device one-shot “main chain and protector” for low-power optimization (e.g., quantization, etc.).

[0082] Further benefits can be realized because, as described, an embedding of a representative item of a new class can be computed and added to the index. Accordingly, an unlimited new number of classes can be added to the index without having to retrain the on-device ANN model. This ability to dynamically add new classes is particularly useful when dealing with problems where the number of different items is not known in advance, is constantly changing, or is very large.

[0083] Self-supervised training on diverse, real, and unfiltered internet data can also lead to interesting properties to emerge, such as geo-localization, fairness, multilingual topic label embeddings, and artistic information and improved semantic information. That is, in addition to capturing semantic information, aspects of the present disclosure can also capture information about artistic style and learn salient information such as geo-location and multilingual word embeddings based on visual content only, for example.

[0084] In some aspects, the on-device ANN model can be fine-tuned quickly for a particular task. For example, if it is desired to train a visual model on semantic segmentation for autonomous driving, a camera can be installed in a car and driving through a city for an hour can collect a much larger amount of data. In contrast, in the case of a conventional approach using supervised learning, where all images would have to be manually labeled. Manual labeling is extremely expensive and manually labeling the same amount of data is very time consuming (e.g., can take months).

[0085] Further, aspects of the present disclosure can learn a metric embedding space in which distances between embedding points are a function of an effective distance metric. These distance metrics can satisfy the triangle inequality such that the space obeys approximate nearest neighbor search and thus can result in high retrieval accuracy. By applying pair mining and pair weighting using a multi-similarity loss function, many types of relative similarity between pairs can be determined. Additionally, pairs assigned a higher weight can be considered to be pairs that are more informative and are retained. On the other hand, pairs assigned a lower weight can be considered to be pairs that are less informative and can be discarded. The weight of a given pair can depend on the self-similarity pair as well as the relative comparison / similarity to other pairs in the training set.

[0086] Figure 8 is a flowchart illustrating a processor-implemented method 800 for dynamic class-incremental learning without forgetting in accordance with aspects of the present disclosure. For example, the processor-implemented method 800 can be performed by a processor such as the CPU 102 or the NPU 108. As shown at block 802, the processor receives an input through a first artificial neural network (ANN). The first ANN can be trained based on a multi-similarity loss function. In some aspects, the ANN can be included at, for example, a server such as a cloud server. The input can include image data, video data, sensor data, sequence data, or other types of data. For example, the input can be provided directly from a user device or can be streamed via the Internet. Figure 8

[0087] At block 804, the processor extracts features of the input through the first ANN to generate a representation of the input. For example, as discussed with respect to Figure 5 the main chain 504 can process the input 502, extract features of the input 502 to generate a representation of the input 502.

[0088] At block 806, the processor generates an embedding through the first ANN based on the representation and a multi-similarity metric. As described with respect to Figure 5 the projector 506 can receive the representation from the main chain 504 and can project the representation to an embedding that can be contrasted with other representations, for example, using a multi-similarity loss.

[0089] At block 808, the processor generates a new class through the first ANN without retraining the first ANN. The new class is generated based on a comparison of the embedding and a set of previous embeddings. As described with respect to Figure 5 the new class can be dynamically added on a device without time-consuming retraining. The ANN can generate an embedding for a representative sample of the new class (e.g., using the main chain 504 and the projector 506).

[0090] Figure 9 ​is a flow diagram illustrating a processor-implemented method 900 for dynamic class-incremental learning without forgetting, in accordance with aspects of the present disclosure. For example, the processor-implemented method 900 can be performed by a processor such as the CPU 102 or the NPU 108. The processor-implemented method 900 can be implemented by an artificial neural network (ANN) model. The ANN model can be included at a user device. The user device can include, for example, a mobile computing device such as a smartphone, an autonomous vehicle, or other mobile computing device.

[0091] As Figure 9 shown, at block 902, the processor receives, at a user device, an input by an artificial neural network (ANN). For example, the input can include image data, video data, sensor data, sequence data, or other types of data.

[0092] At block 904, the processor extracts, by the ANN, features of the input to generate a representation of the input. For example, as discussed with respect to Figure 5 the on-device ANN model can extract features of the observation of the vehicle 524 using the main chain 526 to generate a representation of the observation of the vehicle 524.

[0093] At block 906, the processor generates, by the ANN, an input embedding based on the representation and a learned plurality of similarities. As described with respect to Figure 5 the representation can be supplied to a projector 528, which can project the representation to an embedding. In some aspects, the embedding of an item can include a same class sample mean (e.g., a prototype feature).

[0094] At block 908, the processor compares the input embedding to a set of embeddings stored at the user device. As described with respect to Figure 5 the embedding can be saved and added to an index (e.g., 522a or 522b) at the user device.

[0095] At block 910, the processor generates an inference for the input based on the comparison of the input embedding and the set of embeddings. As described with respect to Figure 5 for example, classification of the input can be made by a nearest sample mean rule. At inference (e.g., query) time, a fast approximate nearest neighbor search (e.g., using an ANN) can be utilized to retrieve the closest matching item from the index (e.g., 522a or 522b) in sublinear time.

[0096] Implementation examples are provided in the following numbered clauses:

[0097] 1. A processor-implemented method comprising:

[0098] receiving, by a first artificial neural network (ANN), an input;

[0099] extracting features of the input through the first ANN to generate a representation of the input;

[0100] generating, through the first ANN, an embedding based on the representation and a multi- similarity measure; and

[0101] generating, through the first ANN, a new class without retraining the first ANN, the new class being generated based on a comparison of the embedding and a set of previous embeddings.

[0102] 2. The processor-implemented method of clause 1, wherein the first ANN is trained based on a multi-similarity loss function.

[0103] 3. The processor-implemented method of clause 1 or 2, wherein the first ANN is implemented at a server, and the method further comprises:

[0104] training a second ANN based on the multi-similarity loss function; and distributing, by the server, the second ANN model to a user device.

[0105] 4. The processor-implemented method of any of clauses 1-3, further comprising determining a score for a pair of samples included in the input based on the multi-similarity measure for the pair of samples, the embedding being generated based on the score.

[0106] 5. The processor-implemented method of any of clauses 1-4, further comprising:

[0107] mining the input and training samples to determine similar pairs of samples, the similar pairs of samples being determined based on a self-similarity and a relative similarity; and

[0108] assigning a weight to the similar pairs of samples based on the self-similarity and the relative similarity.

[0109] 6. A processor-implemented method comprising:

[0110] receiving, at a user device, an input through an artificial neural network (ANN);

[0111] extracting features of the input through the ANN to generate a representation of the input;

[0112] generating, through the ANN, an input embedding based on the representation and a learned multi-similarity;

[0113] comparing the input embedding to a set of embeddings stored at the user device; and

[0114] generating an inference for the input based on the comparison of the input embedding and the set of embeddings.

[0115] 7. The processor-implemented method of clause 6, further comprising:

[0116] determining that the input comprises never-seen data;

[0117] generating an index of the embedding corresponding to the never-seen data; and

[0118] storing the embedding with the set of embeddings at the user device.

[0119] 8. The processor-implemented method of clause 6 or 7, further comprising generating, by the ANN, a new class, the embedding corresponding to the never-seen data is a prototype of the new class, and the new class is added without retraining the ANN.

[0120] 9. An apparatus comprising:

[0121] a memory; and

[0122] at least one processor coupled to the memory, the at least one processor configured to:

[0123] receive an input by a first artificial neural network (ANN), the first ANN trained based on a multi-similarity loss function;

[0124] extract, by the first ANN, features of the input to generate a representation of the input;

[0125] generate, by the first ANN, an embedding based on the representation and a multi-similarity metric; and

[0126] generate, by the first ANN without retraining the first ANN, a new class, the new class generated based on a comparison of the embedding and a set of previous embeddings.

[0127] 10. The apparatus of clause 9, wherein the first ANN is trained based on a multi-similarity loss function.

[0128] 11. The apparatus of clause 9 or 10, wherein the first ANN is implemented at a server, and the at least one processor is further configured to:

[0129] train a second ANN based on the multi-similarity loss function; and

[0130] distributing, by the server, the second ANN model to a user device.

[0131] 12. The apparatus of any of clauses 9-11, wherein the at least one processor is further configured to determine a score for a pair of samples included in the input based on the plurality of similarity measures for the pair of samples, the embedding being generated based on the score.

[0132] 13. The apparatus of any of clauses 9-12, wherein the at least one processor is further configured to:

[0133] mining the input and training samples to determine similar pairs of samples, the similar pairs of samples being determined based on a self-similarity and a relative similarity; and

[0134] assigning a weight to the similar pairs of samples based on the self-similarity and the relative similarity.

[0135] 14. An apparatus comprising:

[0136] a memory; and

[0137] at least one processor coupled to the memory, the at least one processor configured to:

[0138] receive, at a user device, an input by an artificial neural network (ANN);

[0139] extract, by the ANN, features of the input to generate a representation of the input;

[0140] generate, by the ANN, an input embedding based on the representation and a learned plurality of similarity;

[0141] compare the input embedding to a set of embeddings stored at the user device; and

[0142] generate an inference for the input based on the comparison of the input embedding and the set of embeddings.

[0143] 15. The apparatus of clause 14, wherein the at least one processor is further configured to:

[0144] determine that the input includes never seen data;

[0145] generate an index of the embedding corresponding to the never seen data; and store the embedding with the set of embeddings at the user device.

[0146] 16. The apparatus of clause 14 or 15, wherein the at least one processor is further configured to generate, by the ANN, a new class, the embedding corresponding to the never- seen data is a prototype of the new class, and the new class is added without retraining the ANN.

[0147] In one aspect, the receiving component, the extracting, comparing, and / or generating component can be the CPU 102, a program memory associated with the CPU 102, the dedicated memory block 118, the fully connected layer 362, the NPU 428, and / or the routing and connection processing unit 216 configured to perform the recited functions. In another configuration, the aforementioned components can be any module or any means for performing the functions recited by the aforementioned components.

[0148] The various operations of methods described above can be performed by any suitable means depending on the functionality of the means. These means can include various hardware and / or software component(s) and / or module(s), including, but not limited to a circuit, an application specific integrated circuit (ASIC), or processor. Generally, where there are operations illustrated in figures, those operations can have corresponding counterpart means-plus-function components with similar numbering.

[0149] As used herein, the term “determining” encompasses a wide variety of actions. For example, “determining” can include calculating, computing, processing, deriving, investigating, looking up (e.g., looking up in a table, a database or another data structure), ascertaining and the like. Additionally, “determining” can include receiving (e.g., receiving information), accessing (e.g., accessing data in a memory) and the like. Furthermore, “determining” can include resolving, selecting, choosing, establishing and the like.

[0150] As used herein, a phrase referring to “at least one of’ a list of items refers to any combination of those items, including single members. As an example, “at least one of a, b, or c” is intended to cover a, b, c, a-b, a-c, b-c, and a-b-c.

[0151] The various illustrative logical blocks, modules, and circuits described in connection with the disclosure can be implemented or performed with a general purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic device (PLD), discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor can be a microprocessor, but in the alternative, the processor can be any commercially available processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0152] The steps of a method or algorithm described in connection with the present disclosure can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in any form of storage medium that is known in the art. Some examples of storage media that can be used include random access memory (RAM), read only memory (ROM), flash memory, erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), registers, a hard disk, a removable disk, a CD-ROM, and so forth. A software module can comprise a single instruction, or many instructions, and can be distributed over several different code segments, among different programs, and across multiple storage media. A storage medium can be coupled to a processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor.

[0153] The methods disclosed herein comprise one or more steps or actions for achieving the described method. The method steps and / or actions can be interchanged with one another without departing from the scope of the claims. In other words, unless a specific order of steps or actions is specified, the order and / or use of specific steps and / or actions can be modified without departing from the scope of the claims.

[0154] The described functions can be implemented in hardware, software, firmware, or any combination thereof. If implemented in hardware, an example hardware configuration can include a processing system in a device. The processing system can be implemented with a bus architecture. The bus can include any number of interconnecting buses and bridges depending on the specific application of the processing system and the overall design constraints. The bus can link together various circuits including processors, machine-readable media, and bus interface circuits. A bus interface circuit can be used to connect a network adapter to the processing system via the bus. The network adapter can be used to implement signal processing functionality. For certain aspects, a user interface (e.g., keypad, display, mouse input, joystick, etc.) can also be connected to the bus. The bus can also link various other circuits such as timing sources, peripherals, voltage regulators, power management circuits, and the like, which are well known in the art, and therefore, will not be described any further.

[0155] The processor can be responsible for managing a bus and general processing, including the execution of software stored on the machine-readable media. The processor can be implemented with one or more general-purpose processors and / or a special-purpose processor. Examples include microprocessors, microcontrollers, DSP processors and other circuitry that can execute software. Software shall be construed broadly to mean any instructions, code, data, or any combination thereof, whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise. Machine-readable media can include, for example, random access memory (RAM), flash memory, read only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, magnetic disks, optical disks, hard drives, floppy disks, tape, compact disks, other appropriate storage media, or any combination thereof. The machine-readable media can be embodied in a computer- program product. The computer-program product can comprise packaging materials.

[0156] In software implementations, the machine-readable media can include, for example, RAM memory, flash memory, ROM memory, PROM memory, EPROM memory, EEPROM memory, registers, magnetic disks, optical disks, hard drives, floppy disks, tape, compact disks, other appropriate storage media, or any combination thereof. The machine-readable media can be embodied in a computer-program product. The computer-program product can comprise packaging materials.

[0157] The processing system can be configured as a general-purpose processing system with one or more microprocessors and external memory providing the processor functionality and at least a portion of the machine-readable media, all linked together with other supporting circuitry through an external bus architecture. Alternatively, the processing system can include one or more neuromorphic processors for implementing the described neuron models and nervous system models. As another alternative, the processing system can be implemented with an application-specific integrated circuit (ASIC) having the processor, bus interface, user interface, supporting circuitry, and at least a portion of the machine-readable media integrated into a single chip, or with one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gated logic, discrete hardware components, or any other suitable circuitry, or any combination of circuitry that can perform the various functionality described throughout this disclosure. Those skilled in the art will recognize how best to implement the described functionality for the processing system, depending on the particular application and general design constraints imposed on the overall system.

[0158] A machine-readable medium can include a plurality of software modules. These software modules include instructions that, when executed by a processor, cause the processing system to perform various functions. The software modules can include a transmission module and a receiving module. Each software module can reside in a single storage device or be distributed across multiple storage devices. By way of example, a software module can be loaded into RAM from a hard drive when a triggering event occurs. During execution of the software module, the processor can load some of the instructions into cache to increase access speed. One or more cache lines can then be loaded into a general register file for execution by the processor. When referring to the functionality of a software module below, it will be understood that such functionality is implemented by the processor when executing instructions from that software module. Furthermore, it should be appreciated that aspects of the present disclosure are capable of operation with other systems and are not limited to the system depicted.

[0159] If implemented in software, the functions can be stored or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media include both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. A storage medium can be any available medium that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Additionally, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. Disk and disc, as used herein, include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray® disc where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Thus, in some aspects computer readable medium can comprise non-transitory computer readable medium (e.g., tangible media). In addition, for other aspects, computer readable medium can comprise transitory computer read medium (e.g., a signal). Combinations of the above should also be included within the scope of computer readable media.

[0160] Thus, certain aspects can comprise a computer program product for performing the operations presented herein. For example, such a computer program product can comprise a computer-readable medium having instructions stored thereon, the instructions being executable by one or more processors to perform the operations described. For certain aspects, the computer program product can include packaging material.

[0161] For ease of presentation, the description has been split into multiple sections. The technical features of the various aspects are presented in Section I, the example systems and techniques are presented in Section II, and the example use cases are presented in Section III. The description begins in Section I with a general overview of aspects of the present disclosure.

[0162] It should be understood that the claim is not limited to the precise arrangements and components exemplified above. Various modifications, changes and variations can be made to the arrangements, operations and details of the methods and apparatuses described above without departing from the scope of the claims.

Claims

1. A processor-implemented method comprising: receiving, by a first artificial neural network (ANN), an input; extracting, by the first ANN, features of the input to generate a representation of the input; generating, by the first ANN, an embedding based on the representation and a multi- similarity measure; and generating, by the first ANN, a new class without retraining the first ANN, the new class generated based on a comparison of the embedding and a set of previous embeddings.

2. The processor-implemented method of claim 1, wherein the first ANN is trained based on a multi-similarity loss function.

3. The processor-implemented method of claim 2, wherein the first ANN is implemented at a server, and the method further comprises: training a second ANN based on the multi-similarity loss function; and distributing, by the server, the second ANN model to a user device.

4. The processor- implemented method of claim 1, further comprising: determining a score for a pair of samples included in the input based on the multi- similarity measure for the pair of samples, the embedding generated based on the score.

5. The processor-implemented method of claim 1, the processor-implemented method further comprising: mining the input and training samples to determine pairs of similar samples, the pairs of similar samples determined based on a self-similarity and a relative similarity; and assigning a weight to the pairs of similar samples based on the self-similarity and the relative similarity.

6. A processor-implemented method comprising: receiving, by an artificial neural network (ANN) at a user device, an input; extracting, by the ANN, features of the input to generate a representation of the input; generating, by the ANN, an input embedding based on the representation and a learned multi-similarity; comparing the input embedding to a set of embeddings stored at the user device; and generating an inference for the input based on the comparison of the input embedding and the set of embeddings.

7. The processor-implemented method of claim 6, the processor-implemented method further comprising: determining that the input includes never seen data; generating an index of the embedding corresponding to the never seen data; and storing, at the user device, the embedding with the set of embeddings.

8. The processor- implemented method of claim 7, further comprising: generating, by the ANN, a new class, the embedding corresponding to the never seen data a prototype of the new class, and the new class added without retraining the ANN.

9. An apparatus comprising: a memory; and at least one processor coupled to the memory, the at least one processor configured to: receive, by a first artificial neural network (ANN), an input, the first ANN trained based on a multi-similarity loss function; extract, by the first ANN, features of the input to generate a representation of the input; generate, by the first ANN, an embedding based on the representation and a multi- similarity measure; and generating, by the first ANN without retraining the first ANN, a new class, the new class being generated based on a comparison of the embedding and a set of previous embeddings.

10. The apparatus of claim 9, wherein the first ANN is trained based on a multi similarity loss function.

11. The apparatus of claim 10, wherein the first ANN is implemented at a server, and the at least one processor is further configured to: train a second ANN based on the multi similarity loss function; and distribute, by the server, the second ANN model to a user device.

12. The apparatus of claim 9, wherein the at least one processor is further configured to determine a score for a pair of samples included in the input based on the multi similarity metric for the pair of samples, the embedding being generated based on the score.

13. The apparatus of claim 9, wherein the at least one processor is further configured to: mine the input and training samples to determine pairs of similar samples, the pairs of similar samples being determined based on a self similarity and a relative similarity; and assign a weight to the pairs of similar samples based on the self similarity and the relative similarity.

14. An apparatus, the apparatus comprising: a memory; and at least one processor coupled to the memory, the at least one processor configured to: receive, at a user device, an input by an artificial neural network (ANN); extract, by the ANN, features of the input to generate a representation of the input; generate, by the ANN, an input embedding based on the representation and a learned multi similarity; compare the input embedding to a set of embeddings stored at the user device; and generate an inference for the input based on the comparison of the input embedding and the set of embeddings.

15. The apparatus of claim 14, wherein the at least one processor is further configured to: determine that the input includes an unseen data; generate an index of the embedding corresponding to the unseen data; and store, at the user device, the embedding with the set of embeddings.

16. The apparatus of claim 15, wherein the at least one processor is further configured to generate, by the ANN, a new class, the embedding corresponding to the unseen data being a prototype of the new class, and the new class being added without retraining the ANN.