Methods and apparatus for parallel training of adaptive resonance theory neural networks

L-DNN addresses the limitations of conventional DNNs by integrating a DNN-based subsystem with a fast learning ART module for parallel training and knowledge merging, enabling efficient real-time learning and adaptation on edge devices.

WO2025155707A1PCT designated stage expired Publication Date: 2025-07-24NEURALA INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/011850
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-16
Filing Date
2025-01-16
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Conventional Deep Neural Networks (DNNs) face challenges in real-time learning and knowledge updates due to high computational requirements, catastrophic forgetting, and inefficient data transfer, making them impractical for edge devices and multi-edge systems.

Method used

The Lifelong Deep Neural Network (L-DNN) combines a representation-rich DNN-based subsystem with a fast learning subsystem using Adaptive Resonance Theory (ART) to enable continuous, online learning, allowing parallel training of nodes and merging knowledge across devices without data sharing, thus reducing training time and computational load.

Benefits of technology

L-DNN enables real-time learning and knowledge sharing on edge devices with reduced training time, memory requirements, and improved accuracy, facilitating seamless updates and adaptations in real-world applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025011850_24072025_PF_FP_ABST
    Figure US2025011850_24072025_PF_FP_ABST
Patent Text Reader

Abstract

An apparatus may separate inputs to the ART classifier into a first input group including inputs to be learned by existing nodes, and into a second input group including inputs to be learned by new nodes, wherein the first node group includes the existing nodes, and the second node group includes new nodes. An apparatus may train nodes in a first node group of the ART classifier in parallel with each other. An apparatus may train nodes in a second node group of the ART classifier in parallel with each other, wherein training the nodes in the second node group includes consolidating weights of all of the new nodes, and training the nodes in the first node group is in parallel with training the second node group. An apparatus may merge weights of the nodes in the first node group with weights of the nodes in the second node group.
Need to check novelty before this filing date? Find Prior Art

Description

METHODS AND APPARATUS FOR PARALLEL TRAINING OF ADAPTIVE RESONANCE THEORY NEURAL NETWORKSCROSS-REFERENCE TO RELATED APPLICATION(S)

[0001] This application claims the priority benefit, under 35 U.S.C. § 119(e), of U.S. Application No. 63 / 621,541, filed January 16, 2024, which is incorporated herein by reference in its entirety.BACKGROUND

[0002] Traditional Neural Networks, including Deep Neural Networks (DNNs) that include many layers of neurons interposed between the input and output layers, require thousands or millions of iteration cycles over a particular dataset to train. These cycles are frequently performed in a high-performance computing server. In fact, some traditional DNNs may take days or even weeks to be trained, depending on the size of the training dataset.

[0003] One technique for training a DNN involves backpropagation. Backpropagation computes changes of all the weights in the DNN in proportion to the error gradient from a labeled dataset, via application of the chain rule in order to backpropagate error gradients. Backpropagation makes small changes to the weights for each datum and runs over all data in the set for many epochs.

[0004] The larger the learning rate taken per iteration cycle, the more likely that the gradient of the loss function will settle to a local minimum instead of a global minimum, which could lead to poor performance. To increase the likelihood that the loss function will settle to a global minimum, DNNs decrease the learning rate, which leads to small changes to their weights every training epoch. This increases the number of training cycles and the total learning time.

[0005] Advancements of Graphic Processing Unit (GPU) technology has led to massive improvement in compute capability for the highly parallel operations, with training jobs that used to take weeks or months now taking hours or days. This is still not fast enough for a real-time knowledge update. Furthermore, utilizing a high-performance computational server for updating a DNN brings up the cost in terms of server prices and energy consumption.This makes it extremely difficult to update the knowledge of DNN based systems on-the-fly, which is desired for many cases of real-time operations.

[0006] Furthermore, since the gradient of the loss function computed for any single training sample can affect all the weights in the network (due to the typically distributed representations), standard DNNs are vulnerable to forgetting previous knowledge when they learn new objects. This loss of the ability to correctly classify old examples is called “catastrophic forgetting”. Repetitive presentations of the same inputs over multiple epochs mitigate this issue, with the drawback of making it extremely difficult to quickly add new knowledge to the system. This is one reason why conventional learning is impractical or altogether impossible on a computationally limited edge device (e.g., a cell phone, a tablet, or a small form factor processor). Even if the problem of forgetting is solved, conventional learning on edge devices would still be impractical due to the high computational load of the training, the small training steps, and the repetitive presentation of all inputs.

[0007] These limitations are true for a single compute edge device (e.g., a smart phone, smart camera, drone, self-driving vehicle, and the like) across its deployment lifespan, where the edge device may need to update its knowledge, and becomes even more pronounced due to the need of data sharing for distributed, multi-edge systems (e.g., smart phones connected in a network, networked smart cameras, a fleet of drones or self-driving vehicles, and the like), where quick sharing of newly acquired knowledge is a desirable for an intelligent agent across its deployment life cycle.

[0008] In order to learn knowledge, a real-time operating machine that uses a traditional DNN may have to accumulate a large amount of data to retrain the DNN. The accumulated data is transferred from the “edge” of the real-time operating machine (i.e., the device itself, for example, a self-driving car, a drone, a robot, etc.) to a central server (e.g., a cloud-based server) in order to get the labels from the operator and then retrain the DNN executed on the edge. The more accumulated data there is, the more expensive the transfer process in terms of time and network bandwidth. In addition, an interleaved training on the central server has to combine the new data with the original data that is stored for the whole life cycle of the system. This creates severe transmission bandwidth and data storage limitations.

[0009] In summary, applying conventional backpropagation-based DNN training to a realtime operating system suffers from the following drawbacks: a. Updating the system with new knowledge on-the-fly is impractical, if not impossible; b. Learning throughout the deployment cycle of an edge device is impossible without regular communications with the servers and a significant wait time for knowledge update; c. Learning new information requires server space, energy consumption, and disk space consumption to store all input data indefinitely for further training; d. Inability to learn on a small form factor computing edge devices; and e. Inability to merge knowledge across multiple edge devices without engaging slow and expensive data transfer, server-side retraining, and redeployment.SUMMARY

[0010] A Lifelong Deep Neural Network (L-DNN) enables continuous, online, lifelong learning in Artificial Neural Networks (ANNs) and Deep Neural Networks (DNNs) in a lightweight compute device (edge) with learning that is much less time consuming and computationally intensive. An L-DNN enables real-time learning from continuous data streams, bypassing the need to store input data for multiple iterations of backpropagation learning.

[0011] L-DNN technology combines a representation-rich, DNN-based subsystem (Module A) with a fast learning subsystem (Module B) to achieve fast, yet stable learning of features that represent entities or events of interest. These feature sets can be pre-trained by slow learning methodologies, such as backpropagation. In the DNN-based case, described in detail in this disclosure (other feature descriptions are possible by employing non-DNN methodologies for Module A), the outputs of the high-level feature extraction layers of the DNN serve as inputs into the fast learning system in Module B, which classifies familiar entities and events and adds knowledge of unfamiliar entities and events on the fly. Module B is able to learn important information and capture descriptive and highly predictive features of the environment without the drawback of slow learning.

[0012] L-DNN techniques can be applied to visual, structured light, LIDAR, SONAR, RADAR, or audio data, among other modalities. For visual or similar data, L-DNN techniques can be applied to visual processing, such as enabling whole-image classification (e.g., scene detection), bounding box-based object detection, pixel-wise segmentation, and other visual recognition tasks. They can also perform non-visual recognition tasks, such as classification of non-visual signals, and other tasks, such as updating Simultaneous Localization and Mapping (SLAM) generated maps by incrementally adding knowledge as the robot, self-driving car, drone, or other device is navigating the environment.

[0013] Memory consolidation in an L-DNN keeps memory requirements under control in Module B as the L-DNN learns more entities or events (in visual terms, ‘objects’ or ‘categories’). Additionally, the L-DNN methodology enables multiple edge computing devices to merge their knowledge (or ability to classify input data). The merging can occur on a peer-to-peer basis, by direct exchange of neural network representations between two Modules B, or via an intermediary server that merges representations of multiple Modules B from several edges. In either case only neural representations must be shared, but sharing of training data is not necessary. Finally, L-DDN does not rely on backpropagation, thereby dramatically decreasing training time, power requirements, and compute resources to update L-DNN knowledge using new input data.

[0014] Example implementation of Module B can utilize Adaptive Resonance Theory (ART) that by design avoids catastrophic forgetting, allows one-shot learning, and provides an ability to share knowledge across devices without sharing the original data. Conventional ART has a downside that inputs have to be learned sequentially one after another, which is a problem in itself as well as a cause of notorious order dependency of ART networks where same inputs in different order can lead to drastic accuracy changes.

[0015] The inventors have appreciated that an ART -based fast learning system (e.g., Module B) and other ART-based systems can be further improved by learning more than one input at a time, such as in parallel learning. Conventionally, entirely parallel learning has for decades been considered impossible for ART networks, even among prominent experts in the ART field. For example, some experts have stated that if inputs are considered in parallel, it would not be known which inputs will create the nodes that will then be further updated by other inputs during training, versus the inputs that will just update already existing nodes.

[0016] However, the inventors have recognized and appreciated that parallel learning for ART networks is possible and advantageous. Regarding possibility, the inventors recognized that when the information about the inputs is available, a system can compare the inputs across themselves (in addition to comparison with the weights as in the conventional ART) and know which of inputs will go into the same node and which of them will go into different node(s). As for example advantages, in addition to faster training, ART parallel learning reduces dependency on the order of inputs since large chunks (or groups) of inputs are considered simultaneously and individual inputs within a chunk have no advantage over each other. Given enough memory to fit the data, the inventors recognized that ART parallel learning can learn an entire dataset in one go, with no order dependency. That is, the inventors have recognized that if a system has enough memory (which may be CPU memory, not necessarily or even typically GPU memory) to fit the data, it is advantageous for the system to learn all inputs at once, in terms of accuracy, speed, and / or power consumption.

[0017] In some embodiments, the techniques described herein relate to a method of implementing an Adaptive Resonance Theory (ART) classifier operating in real-time, the method including: training nodes in a first node group of the ART classifier in parallel with each other; training nodes in a second node group of the ART classifier in parallel with each other; and merging weights of the nodes in the first node group with weights of the nodes in the second node group.

[0018] In some embodiments, the techniques described herein relate to a method, wherein the method includes separating inputs to the ART classifier into a first input group including inputs to be learned by existing nodes, and into a second input group including inputs to be learned by new nodes, the nodes in the first node group include the existing nodes, and the nodes in the second node group include new nodes.

[0019] In some embodiments, the techniques described herein relate to a method, wherein: training the nodes in the first node group of the ART classifier in parallel with each other includes training the existing nodes using the first input group, and training the nodes in the second node group of the ART classifier in parallel with each other includes training the new nodes using the second input group.

[0020] In some embodiments, the techniques described herein relate to a method, wherein: the method includes training the existing nodes in parallel with the new nodes.

[0021] In some embodiments, the techniques described herein relate to a method, wherein: training the nodes in the second node group includes consolidating weights of all of the new nodes.

[0022] In some embodiments, the techniques described herein relate to a method, wherein: the method includes: computing a similarity matrix between inputs to the ART classifier and weights of nodes of the ART classifier; and for each input to the ART classifier: determining a maximal similarity value in the similarity matrix; comparing the maximal similarity value to a training vigilance; in response to determining the maximal similarity value is below the training vigilance, assigning the input to the second input group; and in response to determining the maximal similarity value is not below the training vigilance, assigning a node corresponding to the maximal similarity value as a node to train for the input and assigning, to the first input group, at least one input with a corresponding node to train for the at least one input.

[0023] In some embodiments, the techniques described herein relate to an apparatus for realtime operation on continuous sensory input data to learn objects on-the-fly, the apparatus including: at least one processor configured to: provide an Adaptive Resonance Theory (ART) classifier operating in real-time to classify objects represented by the continuous sensory input data based on features extracted by a convolutional deep neural network; train nodes in a first node group of the ART classifier in parallel with each other; train nodes in a second node group of the ART classifier in parallel with each other; and merge weights of the nodes in the first node group with weights of the nodes in the second node group.

[0024] In some embodiments, the techniques described herein relate to an apparatus, wherein: the at least one processor is configured to separate inputs to the ART classifier into a first input group including inputs to be learned by existing nodes, and into a second input group including inputs to be learned by new nodes, the nodes in the first node group include the existing nodes, and the nodes in the second node group include new nodes.

[0025] In some embodiments, the techniques described herein relate to an apparatus, wherein the at least one processor is configured to: train the nodes in the first node group of the ARTclassifier in parallel with each other by training the existing nodes using the first input group, and train the nodes in the second node group of the ART classifier in parallel with each other by training the new nodes using the second input group.

[0026] In some embodiments, the techniques described herein relate to an apparatus, wherein the at least one processor is configured to: train the existing nodes in parallel with the new nodes.

[0027] In some embodiments, the techniques described herein relate to an apparatus, wherein: to train the second node group includes to consolidate weights of all of the new nodes.

[0028] In some embodiments, the techniques described herein relate to an apparatus, wherein the at least one processor is configured to: compute a similarity matrix between inputs to the ART classifier and weights of nodes of the ART classifier; and for each input to the ART classifier: determine a maximal similarity value in the similarity matrix; compare the maximal similarity value to a training vigilance; in response to determining the maximal similarity value is below the training vigilance, assign the input to the second input group; and in response to determining the maximal similarity value is not below the training vigilance, assign a node corresponding to the maximal similarity value as a node to train for the input and assign, to the first input group, at least one input with a corresponding node to train for the at least one input.

[0029] In some embodiments, the techniques described herein relate to a method of implementing a Lifelong Learning Deep Neural Network (L-DNN) including an Adaptive Resonance Theory (ART) classifier operating in real-time, the method including: separating inputs to the ART classifier into a first input group including inputs to be learned by existing nodes, and into a second input group including inputs to be learned by new nodes, wherein the first node group includes the existing nodes, and the second node group includes new nodes; training nodes in a first node group of the ART classifier in parallel with each other; training nodes in a second node group of the ART classifier in parallel with each other, wherein training the nodes in the second node group includes consolidating weights of all of the new nodes, and training the nodes in the first node group is in parallel with training thesecond node group; and merging weights of the nodes in the first node group with weights of the nodes in the second node group.

[0030] In some embodiments, the techniques described herein relate to a method, wherein: training the nodes in the first node group of the ART classifier in parallel with each other includes training the existing nodes using the first input group, and training the nodes in the second node group of the ART classifier in parallel with each other includes training the new nodes using the second input group.

[0031] In some embodiments, the techniques described herein relate to a method, wherein: the method includes: computing a similarity matrix between inputs to the ART classifier and weights of nodes of the ART classifier; and for each input to the ART classifier: determining a maximal similarity value in the similarity matrix; comparing the maximal similarity value to a training vigilance; in response to determining the maximal similarity value is below the training vigilance, assigning the input to the second input group; and in response to determining the maximal similarity value is not below the training vigilance, assigning a node corresponding to the maximal similarity value as a node to train for the input and assigning, to the first input group, at least one input with a corresponding node to train for the at least one input.

[0032] It should be appreciated that all combinations of the foregoing concepts and additional concepts discussed in greater detail below (provided such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of claimed subject matter appearing at the end of this disclosure are contemplated as being part of the inventive subject matter disclosed herein. It should also be appreciated that terminology explicitly employed herein that also may appear in any disclosure incorporated by reference should be accorded a meaning most consistent with the particular concepts disclosed herein.

[0033] Other systems, processes, and features will become apparent to those skilled in the art upon examination of the following drawings and detailed description. It is intended that all such additional systems, processes, and features be included within this description, be within the scope of the present invention, and be protected by the accompanying claims.BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The skilled artisan will understand that the drawings primarily are for illustrative purposes and are not intended to limit the scope of the inventive subject matter described herein. The drawings are not necessarily to scale; in some instances, various aspects of the inventive subject matter disclosed herein may be shown exaggerated or enlarged in the drawings to facilitate an understanding of different features. In the drawings, like reference characters generally refer to like features (e.g., functionally similar and / or structurally similar elements).

[0035] FIG. 1 illustrates an overview of a Lifelong Deep Neural Network (L-DNN) as it relates to multiple compute edges, either acting individually over a data stream, or connected peer-to-peer or via an intermediary compute server.

[0036] FIG. 2 illustrates an example L-DNN architecture.

[0037] FIG. 3 A illustrates an implementation of the concept of unknown in neural networks.

[0038] FIG. 3B illustrates an additional implementation of the concept of unknown in neural networks.

[0039] FIG. 3C illustrates an example of calibration.

[0040] FIG. 3D illustrates an example of conventional sequential Adaptive Resonance Theory (ART) network training in the unsupervised case with suboptimal order of inputs and limitation on cluster size.

[0041] FIG. 3E illustrates an example of conventional sequential Adaptive Resonance Theory (ART) network training in the unsupervised case with improved order of inputs and limitation on cluster size as well as an example of parallel ART network training.

[0042] FIG. 3F illustrates a flowchart of an example method of implementing an ART classifier operating in real-time.

[0043] FIG. 3G illustrates an additional flowchart of an example method of implementing an ART classifier operating in real-time.

[0044] FIG. 4A illustrates consolidation and melding using an Adaptive Resonance Theory (ART) neural network.

[0045] FIG. 4B shows computing the similarity matrix for consolidation in parallel implementation.

[0046] FIG. 5 illustrates selection of inputs in two groups using masks.

[0047] FIG. 6 illustrates computing the similarity matrix for input separation.

[0048] FIG. 7 illustrates a VGG- 16-based L-DNN classifier as one example implementation.

[0049] FIG. 8 illustrates non-uniform multiscale object detection.

[0050] FIG. 9 illustrates a Mask R-CNN-based L-DNN for object segmentation.DETAILED DESCRIPTION

[0051] Continual Learning in Real-Time Operating Machines

[0052] A Lifelong Learning Deep Neural Network or Lifelong Deep Neural Network (L- DNN) enables a real-time operating machine to learn on-the fly at the edge without the necessity of learning on a central server or cloud. This reduces or eliminates network latency, increases real-time performance, and ensures privacy when desired. In some instances, realtime operating machines can be updated for specific tasks in the field using an L-DNN. For example, with L-DNNs, inspection drones can learn how to identify problems at the top of cell towers or solar panel arrays, smart toys can be personalized based on user preferences without the worry about privacy issues since data is not shared outside the local device, smart phones can share knowledge learned at the edge (peer to peer or globally with all devices) without shipping information to a central server for lengthy learning, or self-driving cars can learn and share knowledge as they operate.

[0053] An L-DNN also enables learning new knowledge without forgetting old knowledge, thereby mitigating or eliminating catastrophic forgetting. In other words, L-DNN technology enables real-time operating machines to continually and optimally adjust behavior at the edge based on user input without a) needing to send or store input images, b) time-consuming training, or c) large computing resources. Learning after deployment with an L-DNN allows areal-time operating machine to adapt to changes in its environment and to user interactions, handle imperfections in the original data set, and provide customized experience for a user.

[0054] The disclosed technology can also merge knowledge from multiple edge devices. This merging includes a “crowd collection” and labeling of knowledge and sharing this collected knowledge among edge devices, eliminating hours of tedious centralized labeling. In other words, the brain (e.g., the weight matrix in Module B representing learned knowledge about objects and their features, storing compressed representations of objects as generalized weight patterns) from one or more of the edge devices can be merged with other edge device brains either one onto another (peer-to-peer) or into a shared brain that is pushed back to some or all of the devices at the edge. L-DNN ensures that the merging / melding / sharing / combining of knowledge results in a growth in memory footprint that is no faster than linear in the number of objects, happens in real-time, and results in small amount of information exchanged between devices. These features make L-DNNs practical for real-world applications.

[0055] An L-DNN implements a heterogeneous Neural Network architecture characterized by two modules:1) Slow learning Module A, which includes a neural network (e.g., a Deep Neural Network) that is either factory pre-trained and fixed or configured to learn via backpropagation or other learning processes based on sequences of data inputs; and2) Module B, which provides an incremental classifier able to change synaptic weights and representations instantaneously, with very few training samples. Example instantiations of this incremental classifier include, for example, an Adaptive Resonance Theory (ART) network or Restricted Boltzmann Machine (RBM) with contrastive divergence training neural networks, as well as non-neural methods, such as Support Vector Machines (SVMs) or other fast learning supervised classification processes.

[0056] Typical application examples of L-DNN are exemplified by, but not limited to, an Internet of Things (loT) device that learns a pattern of usage based on the user’s habits; a self-driving vehicle that can adapt its driving ‘style’ from the user, quickly learn a new skill on-the-fly, or park in a new driveway; a drone that is able to learn, on-the-fly, a new class of damage to an infrastructure and can spot this damage after a brief period of learning while inoperation; a home robot, such as a toy or companion robot, which is able to learn (almost) instantaneously and without pinging the cloud for its owner’s identity; a robot that can learn to recognize and react to objects it has never seen before, avoid new obstacles, or locate new objects in a world map; an industrial robot that is able to learn a new part and how to manipulate it on-the-fly; and a security camera that can learn a new individual or object and quickly find it in imagery provided by other cameras connected to a network. The applications above are only examples of a class of problems that are unlocked and enabled by the innovation(s) described herein, where learning can occur directly in the computing device embedded in a particular application, without being required to undertake costly and lengthy iterative learning on the server.

[0057] L-DNN technology can be applied to several input modalities, including but not limited to video streams, data from active sensors (e.g., infrared (IR) imagery, LIDAR data, SONAR data, and the like), acoustic data, other time series data (e.g., sensor data, real-time data streams, including factory -generated data, loT device data, financial data, and the like), and any multimodal linear / nonlinear combination of the such data streams.

[0058] Overview of L-DNN

[0059] As disclosed above, L-DNN implements a heterogeneous neural network architecture to combine a fast learning mode and a slow learning mode. In the fast learning mode, a realtime operating machine implementing L-DNN learns new knowledge and new experiences quickly so that it can respond to the new knowledge almost immediately. In this mode, the learning rate in the fast learning subsystem is high to favor new knowledge and the corresponding new experiences, while the learning rate in the slow learning subsystem is set to a low value or zero to preserve the consistent feature space on which the fast learning system operates.

[0060] FIG. 1 provides an overview of an L-DNN architecture where multiple devices, including a master edge / central server and several compute edge devices (e.g., drones, robots, smartphones, or other loT devices), running L-DNNs operate in concert. Each device receives sensory input 100 and feeds it to a corresponding L-DNN 106 comprising a slow learning Module A 102 and a fast learning Module B 104. Each Module A 102 is based on a pre-learned (fixed weight) DNN and serves as feature extractor. It receives the input 100,extracts the relevant features into compressed representations of objects and feeds these representations to the corresponding Module B 104. Module B 104 is capable of fast learning of these object representations. Through the interactions with the user, it receives correct labels for the unfamiliar objects, quickly learns the association between each feature vector and corresponding label, and as a result can recognize these new objects immediately. As multiple L-DNNs 106 learn different inputs, the L-DNNs can connect peer-to-peer (dashed line) or to the central server (dotted line) to meld (fuse, merge, or combine) the newly acquired knowledge and share it with other L-DNNs 106 as disclosed below.

[0061] An example object detection L-DNN implementation presented below produced the following pilot results comparing to the traditional object detection DNN “You only look once” (YOLO). The same small (600-image) custom dataset with one object was used to train and validate both networks. 200 of these images were used as validation set. Four training sets of different sizes (100, 200, 300, and 400 images) were created from the remaining 400 images. For L-DNN training, each image in the training set was presented once. For the traditional DNN YOLO, batches were created by randomly shuffling the training set and training proceeded over multiple iterations through these batches. After training, the validation was run on both networks and produced the following mean average precision (mAP) results:

[0062] Furthermore, the training time for L-DNN using 400 images training set was 1.1 seconds; the training time for YOLO was 21.5 hours. This is a shockingly large performance improvement. The memory footprint of the L-DNN was 320 MB, whereas the YOLO footprint was 500 MB. These results clearly show that an L-DNN can achieve better precision than a traditional DNN YOLO and do this with smaller data sets, much faster training time, and smaller memory requirements.

[0063] Example L-DNN Architecture

[0064] FIG. 2 illustrates an example L-DNN architecture used by a real-time operating machine, such as a robot, drone, smartphone, or loT device. The L-DNN 106 uses twosubsystems, slow learning Module A 102 and fast learning Module B 104. In one implementation, Module A includes a pre-trained DNN, and Module B is based on a fast learning Adaptive Resonance Theory (ART) paradigm, where the DNN feeds to the ART the output of one of the latter feature layers (typically, the last or the penultimate layer before the DNN’s own fully connected layers). Other configurations are possible, where multiple DNN layers can provide inputs to one or more Modules B (e.g., in a multiscale, voting, or hierarchical form).

[0065] An input source 100, such as a digital camera, detector array, or microphone, acquires information / data from the environment (e.g., video data, structured light data, audio data, a combination thereof, and / or the like). If the input source 100 includes a camera system, it can acquire a video stream of the environment surrounding the real-time operating machine. The input data from the input source 100 is processed in real-time by Module A 102, which provides a compressed feature signal as input to Module B 104. In this example, the video stream can be processed as a series of image frames in real-time by Modules A and B. Module A and Module B can be implemented in suitable computer processors, such as graphics processor units, field-programmable gate arrays, or application-specific integrated circuits, with appropriate volatile and non-volatile memory and appropriate input / output interfaces.

[0066] In one implementation, the input data is fed to a pre-trained Deep Neural Network (DNN) 200 in Module A. The DNN 200 includes a stack 202 of convolutional layers 204 used to extract features that can be employed to represent an input information / data as detailed in the example implementation section. The DNN 200 can be factory pre-trained before deployment to achieve the desired level of data representation. It can be completely defined by a configuration file that determines its architecture and by a corresponding set of weights that represents the knowledge acquired during training.

[0067] The L-DNN system 106 takes advantage of the fact that weights in the DNN are excellent feature extractors. In order to connect Module B 104, which includes one or more fast learning neural network classifiers, to the DNN 200 in Module A 102, some of the DNN’s upper layers engaged in only classification by the original DNN (e.g., layers 206 and 208 in FIG. 2) are ignored or even stripped from the system altogether. A desired raw convolutional output of high-level feature extraction layer 204 is accessed to serve as input toModule B 104. For instance, the original DNN 200 usually includes a number of fully connected, averaging, and pooling layers 206 plus a cost layer 208 that is used to enable the gradient descent technique to optimize its weights during training. These layers are used during DNN training or for getting direct predictions from the DNN 200, but are not necessary for generating an input for the Module B 104 (the shading in FIG. 2 indicates that layers 206 and 208 are unnecessary). Instead, the input for the neural network classifier in Module B 104 is taken from a subset of the convolutional layers of the DNN 204. Different layers, or multiple layers can be used to provide input to Module B 104.

[0068] Each convolutional layer on the DNN 200 contains filters that use local receptive fields to gather information from a small region in the previous layer. These filters maintain spatial information through the convolutional layers in the DNN. The output from one or more late stage convolutional layers 204 in the feature extractor (represented pictorially as a tensor 210) are fed to input neural layers 212 of a neural network classifier (e.g., an ART classifier) in Module B 104. There can be one-to-one or one-to-many correspondence between each late stage convolutional layer 204 in Module A 102 and a respective fast learning neural network classifier in Module B 104 depending on whether the L-DNN 106 is designed for whole image classification or object detection as described in detail in the example implementation section.

[0069] The tensor 210 transmitted to the Module B system 104 from the DNN 200 can be seen as an n-layer stack of representations from the original input data (e.g., an original image from the sensor 100). In this example, each element in the stack is represented as a grid with the same spatial topography as the input images from the camera. Each grid element, across n stacks, is the actual input to the Module B neural networks.

[0070] The initial Module B neural network classifier can be pre-trained with arbitrary initial knowledge or with a trained classification of Module A 102 to facilitate learning on-the-fly after deployment. The neural network classifier continuously processes data (e.g., tensor 210) from the DNN 200 as the input source 100 provides data relating to the environment to the L- DNN 106. The Module B neural network classifier uses fast, preferably one-shot learning. An ART classifier uses bottom -up (input) and top-down (feedback) associative projections between neuron-like elements to implement match-based pattern learning as well as horizontal projections to implement competition between categories.

[0071] In the fast learning mode, when a novel set of features is presented as input from Module A 102, ART -based Module B 104 puts the features as an input vector in Fl layer 212 and computes a distance operation between this input vector and existing weight vectors 214 to determine the activations of all category nodes in F2 layer 216. The distance is computed either as a fuzzy AND (in the default version of ART), dot product, or Euclidean distance between vector ends. The category nodes are then sorted from highest activation to lowest to implement competition between them and considered in this order as winning candidates. These winning candidates are filtered by the vigilance parameter so that the candidates with the distance that exceeds the maximal distance defined by vigilance are removed from further consideration. Further filtering is done in the case of supervised learning based on the ground truth. If the label of the winning candidate matches the label provided by the ground truth, then the corresponding weight vector is updated to generalize and cover the new input through a learning process that in the simplest implementation takes a weighted average between the new input and the existing weight vector for the winning node. If none of the winners has a correct label in supervised case or none of the winners has distance allowed by vigilance, then a new category node is introduced in category layer F2 216 with a weight vector that is a copy of the input. In either case, Module B 104 is now familiar with this input and can recognize it on the next presentation.

[0072] The result of Module B 104 serves as an output of L-DNN 106 either by itself or as a combination with an output from a specific DNN layer from Module A 102, depending on the task that the L-DNN 106 is solving. For whole scene object recognition, the Module B output may be sufficient as it classifies the whole image. For object detection, Module B 104 provides class labels that are superimposed on bounding boxes determined from Module A activity, so that each object is located correctly by Module A 102 and labeled correctly by Module B 104. For object segmentation, the L-DNN with multiple module Bs attached to different spatial resolutions in Module A can be utilized. In this case each Module B will learn and predict the class label for individual pixels at their respective resolutions. The final output segmentation mask is determined by combining individual pixel masks from all Module Bs by a combination of averaging and thresholding. In an ideal case, even a superresolution of the final mask can be achieved as shown in FIG. 8, but usually the inner layers of DNNs are limited to the spatial resolutions that are r / 2", where a is 7 or 8 and n is gradually reduced from input layer to the final layers of the DNN, so the maximal resolutionof the final map will be the resolution of the Module B attached to the earliest DNN layer.More details about Module A 102 and Module B 104 are provided below.

[0073] Real-time operation and a concept of unknown in neural networks

[0074] Since L-DNN in general and Module B in particular are designed to operate in realtime on continuous sensory input, a neural network in Module B should be implemented so that it is not confused by when no familiar objects are presented to it. A conventional neural network targets datasets that usually contain a labeled object in the input; as a result, it does not need to handle inputs without familiar objects present. Thus, to use such a network in Module B of an L-DNN an additional special category of “Nothing I know” should be added to the network to alleviate Module B’s attempts to erroneously classify unfamiliar objects as familiar (false positives).

[0075] This concept of “Nothing I know” may be implemented through a dynamic threshold that adapts based on the activations derived from the object being classified. Because these activations are based on the object itself, their average — and hence the dynamic threshold — varies from object to object. This approach provides significant advantages over static thresholds or conventional Softmax implementations.

[0076] For example, conventional metrics based on an entire dataset (rather than the average of activations for a single object) may work for a conventional classifier that learns on the whole dataset and never changes afterwards. However, such conventional metrics may not work for a classifier designed for incremental learning because every time the classifier learns something new, the classifier would need to recompute the metrics by rerunning validation on an ever-larger dataset. In contrast, a dynamic threshold is based on activations derived from the object itself, so a classifier can correctly classify objects even as it learns new objects.

[0077] As another example, consider two input activation vectors: [0.1, 0.1, 0.3, 0.1, 0.1] and [0.8, 0.8, 1.0, 0.8, 0.8], While Softmax would produce identical outputs [0.191519, 0.191519, 0.233922, 0.191519, 0.191519] for both vectors, the dynamic threshold can distinguish between them. The first vector’s average (0.14) and second vector's average (0.84) produce different thresholds when multiplied by a scaling factor. For a factor greater than 1 (e.g., 1.25), the resulting thresholds (0.175 and 1.05) effectively differentiate between highconfidence classification (0.3 > 0.175) and low confidence or potentially “Nothing I know” cases (1.0 < 1.05).

[0078] This concept of “Nothing I know” is useful when processing a live sensory stream that can contain previously unseen and unlabeled objects. It allows Module B and the L-DNN to identify an unfamiliar object as “Nothing I know” or “not previously seen” instead of potentially identifying the unfamiliar object as incorrectly being a familiar object. Extending the conventional design with the implementation of “Nothing I know” concept can be as simple as adding a bias node to the network. The “Nothing I know” concept can also be implemented in a version that automatically scales its influence depending on the number of known classes of objects and their corresponding activations.

[0079] One possible implementation of “Nothing I know” concept works as an implicitly dynamic threshold that favors predictions in which the internal knowledge distribution is clearly focused on a common category as opposed to flatly distributed over several categories. In other words, when the neural network classifier in Module B indicates that there is a clear winner among known object classes for an object, then it recognizes the object as belonging to the winning class. But when multiple different objects have similar activations (i.e., there is no clear winner), the system reports the object as unknown. Since learning process explicitly uses a label, the “Nothing I know” implementation may only affect the recognition mode and may not interfere with the learning mode.

[0080] An example implementation of the “Nothing I know” concept using an ART network is presented in FIG. 3A and FIG. 3B. During presentation of an input, the category layer F2 216 responds with an activation pattern across its nodes. The inputs that do contain a familiar object are likely to have prominent winners like 300A and 300B in FIG. 3A and FIG. 3B. The inputs that do not contain familiar objects are likely to have a flatter distribution of activity in the F2 layer as shown in the lower case in FIG. 3A and the top case in FIG. 3B. Calculating the average of all activations and using it as a threshold is insufficient to distinguish these two cases, because even in the second case there may be nodes 302 A with activity higher than a threshold (dotted lines in FIG. 3 A and FIG. 3B). Multiplying the average by a parameter (which may be called dominance) that is no less than 1 raises the threshold (dashed lines in FIG. 3 A and FIG. 3B, which may be called the dominance threshold) so that only clearwinners like 300A and 300B remain above it as in the upper case in FIG. 3A and the middle case in FIG. 3B.

[0081] The exact value of this parameter depends on multiple factors and can be calculated automatically based on the number of categories the network has learned and the total number of category nodes in the network. An example calculation iswhere 0 is the threshold, C is the number of known categories, N is the number of category nodes, and scaling factor ,s is set based on the type of DNN used in Module A and is finetuned during L-DNN preparation. Setting it too high may increase the false negative rate of the neural network, and setting it too low may increase the false positive rate.

[0082] The bottom case of FIG. 3B is an edge case that the inventors recognized can create some difficulties. The dotted line in FIG. 3B is the threshold that is the average of all activations in nodes at this time, and the dashed line is the same threshold multiplied by the dominance, which as above may be called the dominance threshold. For example, occasionally when more than one activation is higher than the dominance threshold, the system may select the best winner. In this example, classification is split between two classes: one with the probability of 50.5%, and another with probability of 49.5%. This is an undesirable situation, at least because of the lack of a clear winner.

[0083] Training a standalone Module B utilizing the “Nothing I know” concept produced the following results. 50 objects from the 100 objects in the Columbia Object Image Library 100 (COIL- 100) dataset were used as a training set. All 100 objects from the COIL- 100 dataset were used as testing set so that 50 novel objects would be recognized as “Nothing I know” by the standalone Module B. During training, an ART classifier in the standalone Module B was fed objects one by one without any shuffling to simulate real-time operation. After training, the ART classifier demonstrated a 95.5% correct recognition rate (combined objects and “Nothing”). For comparison, feeding an unshuffled dataset of all 100 objects in the COIL- 100 dataset to a conventional ART only produced 55% correct recognition rate. This could be due to the order dependency of conventional ART discussed below.

[0084] If an input is not recognized by the ART classifier in Module B, it is up to the user to introduce corrections and label the desired input. If the unrecognized input is of no importance, the user can ignore it and the ART classifier will continue to identify it as “Nothing I know”. If the object is important to the user, she can label it, and the fast learning Module B network will add the features of the object and the corresponding label to its knowledge. Module B can engage a tracker system to continue watching this new object and add more views of it to enrich the feature set associated with this object.

[0085] Example Implementation of Module A

[0086] In operation, Module A extracts features and creates compressed representations of objects. Convolutional deep neural networks are well suited for this task as outlined below.

[0087] Convolutional neural networks (CNNs) are DNNs that use convolutional units, where the receptive field of the unit's filter (weight vector) is shifted stepwise across the height and width dimensions of the input. When applied to visual input, the input to the initial layer in the CNN is an image, with height (h), width (w), and one to three channel (c) dimensions (e.g., red, green, and blue pixel components), while the inputs to later layers in the CNN have dimensions of height (h), width (w), and the number of filters (c) from the preceding layers. Since each filter is small, the number of parameters is greatly reduced compared to fully- connected layers, where there is a unique weight projecting from each of (h, w, c) to each unit on the next layer. For convolutional layers, each unit has number of weights equal to (f, f, c) where f is the spatial filter size (typically 3), which is much smaller than either h or w. The application of each filter at different spatial locations in the input provides the appealing property of translation invariance in the following sense: if an object can be classified when it is at one spatial location, it can be classified at all spatial locations, as the features that comprise the object are independent of its spatial locations.

[0088] Convolutional layers are usually followed by subsampling (downsampling) layers. These reduce the height (h) and width (w) of their input by reducing small spatial windows (e.g., 2 x 2) of the input to single values. Reductions have used averaging (average pooling) or taking the maximum value (max pooling). Responses of sub sampling layers are invariant to small shifts in the image, and this effect is accumulated over the multiple layers of a typical CNN. In inference, when several layers of convolution and subsampling are appliedto an image, the output exhibits impressive stability with respect to various deformations of the input, such as translation, rotation, scaling, and even warping, such as a network trained on unbroken (written without lifting up the pen) handwritten digits having similar responses for a digit "3" from the training set and a digit "3" written by putting small circles together.

[0089] These invariances provide a feature space in which the encoding of the input has enhanced stability to visual variations, meaning as the input changes (e.g., an object slightly translates and rotates in the image frame), the output values change much less than the input values. This enables learning - it can be difficult to learn on top of another method in which, for example, the encoding of two frames with an object translated by a few pixels have little to no similarity.

[0090] Further, with the recent use of GPU-accelerated gradient descent techniques for learning the filters from massive datasets, CNNs are able to reach impressive generalization performance for well-trained object classes. Generalization means that the network is able to produce similar outputs for test images that are not identical to the trained images, within a trained class. It takes a large quantity of data to learn the key regularities that define a class. If the network is trained on many classes, lower layers, whose filters are shared among all classes, provide a good set of regularities for all natural inputs. Thus, a DNN trained on one task can provide excellent results when used as an initialization for other tasks, or when lower layers are used as preprocessors for new higher-level representations. Natural images share a common set of statistical properties. The learned features at low-layers are fairly classindependent, while higher and higher layers become more class-dependent, as shown by recent work in visualizing the internals of well-trained neural networks.

[0091] An L-DNN exploits these capabilities of CNNs in Module A so that Module B gets the high quality compressed and generalized representations of object features for classification. To increase or maximize this advantage, a DNN used for L-DNN may be pretrained on as many different objects as possible, so that object specificity of the high-level feature layers does not interfere with the fast learning capability of L-DNN.

[0092] Example Implementations of Module B

[0093] In operation, Module B learns new objects quickly and without catastrophic forgetting.

[0094] Adaptive Resonance Theory (ART)

[0095] One example implementation of Module B is an ART network. ART avoids catastrophic forgetting by utilizing competition among category nodes to determine a winning node for each object presentation. In the case of supervised learning, if and only if this winning node is associated with the correct label for the object, the ART classifier updates its weights. Since each node is associated with only one object, and the ART classifier updates weights only for the winning node, any learning episode in ART affects one and only one object. Therefore, there is no interference with previous knowledge when new objects are added to the system; rather, ART simply creates new category nodes and updates the corresponding weights.

[0096] The inventors have appreciated that conventional ART as described in the literature has several disadvantages that limit its successful use as L-DNN Module B without improvements. The list of conventional ART-specific problems and solutions to these problems are disclosed below.

[0097] Conventional classical fuzzy ART does not handle sparse inputs well due to complement coding, which is an integral part of its design. When a sparse input is complement coded, the complement part has high activations in most of the components since complements of zeroes abundantly present in sparse inputs are ones. With all these ones in the complement part of the inputs, it becomes very hard to separate different inputs from each other during distance computation, so the system becomes confused. On the other hand, powerful feature extractors like DNN tend to provide exclusively sparse signals on the high levels of feature extraction. Keeping the ART paradigm but stepping away from the classical fuzzy design and complement coding becomes useful for using ART in Module B of an L-DNN. One of the solutions is to remove complement coding and replace the fuzzy AND distance metric used by fuzzy ART with a dot product based metric. This dot product based metric has the advantage that every input is interpreted as an angle in a multidimensional feature space, and this helps alleviate the change of conditions in which the model operates as follows.

[0098] Frequently, the network or model is trained under one set of conditions, for example lighting, and is used for inference under another set of conditions. Normally, to maintain theaccuracy of the network or model, a new set of training data must be collected under new conditions, and the new network or model must be trained using these data. The inventors appreciated that the relationships between features stay consistent under environment changes, and the entire change of environmental conditions can be represented as a single rotation of all the feature representations in multidimensional space. This allows to quickly calibrate the model to new conditions without retraining by simply rotating all weights by a given angle. This process is sketched in FIG. 3C. The angle of rotation can be determined by acquiring just a few images under new conditions and comparing their statistics in the feature space to the statistics of existing weights.

[0099] The conventional ART family of neural networks is very sensitive to the order of presentation of inputs. In other words, conventional ART lacks the property of consistency; a different order of inputs leads to a different representation of corresponding objects in the ART network as shown in FIG. 3D and FIG. 3E where inputs are provided in the sequence as marked, and each node represents a spatial average of its inputs. Nodes 305B and 305A average over inputs 1, 2, and 3, while nodes 306 A and 306B average over inputs 4 and 5. Clearly, representation of FIG. 3E provides much better clustering than the one in FIG. 3D.

[0100] The inventors recognized that real-time operating systems like L-DNN generally cannot shuffle their training data to provide consistency because they consume their training data as they receive it from the sensors. Frequently during real-time operation, the sensors provide most or all samples of a first object before all samples of subsequent objects, so the system learns one object representation at a time. This may lead to the situation where only a few nodes represent the first object, as without competition from other objects the system may not make mistakes and thus may refine the object representation properly. On the other hand, the subsequent objects may be overrepresented as the system would squeeze its representation into a hyperspace that is already mostly occupied by the representation of the first object. Enforcing a limit on the number of inputs each node can learn introduces competition at the early stage and ensures fine grain representation of the first object. For example, the cartoonish representation of FIG. 3D and FIG. 3E uses the hard limit of 3 inputs per node. In actual implementation, the limit of inputs per node is usually a soft limit based on gradually increasing class-specific vigilance, which increases with the number of inputslearned by the node. The consolidation described below reduces or eliminates overrepresentation for subsequent objects.

[0101] Consolidation also reduces the memory footprint of the object representations, which is especially beneficial for edge devices with limited memory. Creating a new category node for every view of an object that the system cannot classify otherwise leads to a constant increase of memory footprint for ART systems as new objects are added as inputs. During real-time operation and sequential presentation of objects as described above, the system creates a superlinearly increasing number of nodes for each successive object. Thus, the memory footprint of Module B using conventional ART may grow at a rate faster than a linear increase with the number of objects. Consolidation significantly reduces the memory growth within ART, allows reduced order dependency, and forms more consistent object representations.

[0102] Example Implementation of Brain Consolidation and Melding Process Using ART

[0103] FIG. 4A shows an example brain consolidation and melding process using ART. It extends the default ART process as follows. Each of ART’s category nodes in layer F2 216 represents a certain object, and each object has one or more category nodes representing it. The left side of FIG. 4 A shows a weight pattern 402 for a category node in layer F2 216 activated by layer Fl 212. Each weight pattern represents the generalized input pattern that was learned from multiple real feature inputs 210 provided to layer Fl 212. Note, that learning in ART only happens when the category node and the corresponding object are winning the competition and correctly identifying the object in question.

[0104] The middle of FIG. 4 A shows weight patterns 402 for different category nodes in layer F2 216 after multiple inputs 210 for different objects were presented to the ART network. Each weight pattern 402 represents a generalized version of inputs that the corresponding node learned to associate with a corresponding object label. The weight patterns 402 in the middle of FIG. 4 A become the consolidation inputs 404 shown at right in FIG. 4A.

[0105] The original inputs 210 are generally not available to the system at the time of consolidation or melding. On the other hand, the collection of weight patterns 402 is a generalization over all the inputs 210 that the ART network was exposed to during training.As such, the weight patterns 402 represent the important features of the inputs 210 as well or better than the original inputs 210 and can serve as a substitute for real inputs during training process.

[0106] Consolidation uses the weight patterns as substitutes for real inputs. During consolidation the following steps happen:• The weight vectors 402 in the existing weight matrix (e.g., the matrix of weight vectors 214 in FIG. 2) are added to a consolidation input set (right side of FIG. 4 A) a = Wi where a are input vectors, w are weight vectors, and i is an index from 1 to the number of existing category nodes in the network. If the ART network uses complement coding, the complement half of the weight vector is decomplemented and averaged with the initial half of the vector (di = (wi + (1 - Wic)) / 2). Each vector in the consolidation input set receives a corresponding label extracted from its respective category node.• All existing F2 nodes and corresponding weights are removed from the ART network, so the ART network is in its blank initial state.• The consolidation input set is randomly shuffled, and the ART network learns this set the same way it learned original inputs. Random shuffling reduces the effect of order dependency in a conventional ART network and allows the ART network to build more compact (fewer category nodes created) and more optimal (better generalization) representation.

[0107] Using weights for the consolidation input set has a further advantage that a single vector replaces many original input vectors, so that the consolidation processes have reduced complexity and faster computational time than the original learning process.

[0108] The consolidation process can happen at any time during L-DNN-based system operation. It reduces the memory footprint of ART -based implementation of Module B and reduces the order dependency of ART -based system. Reducing the order dependency is beneficial for any real-time operating machine based on L-DNN since during such operation there is no way to change the order of sensory inputs that are coming into the system as it operates. Consolidation can be triggered by user action, or automatically when the memoryfootprint becomes too large (e.g., reaches or exceeds a threshold size), or regularly based on the duration of operation.

[0109] Example consolidation was done for the COIL dataset intentionally presented to the ART network as if it was operating in real-time and saw objects one after another. Initial training took 4.5 times longer than consolidation training. Consolidation reduced the memory footprint by 25% and improved object recognition performance from 50% correct to 75% correct. For the case where the training dataset was initially shuffled to reduce order artifacts, consolidation still showed performance improvement. There was no significant memory footprint reduction since the system was already well compressed after initial training, but the percent correct for object recognition went up from 87% to 98% on average. These experimental results represent unexpectedly large performance improvements.

[0110] Melding is an extension of consolidation where the consolidation training set is combined from weight matrices of more than one ART network. It inherits all the advantages of consolidation and capitalizes on the generalization property of ART networks. As a result, when multiple melded ART networks have knowledge of the same object, all similar representations of this object across multiple ART networks are combined together naturally by the ART learning process, while all distinct representations are preserved. This leads to smart compression (e.g., significant compression while maintaining or improving classification accuracy) of object representations and further reduction of memory footprint of the melded system.[OHl] For example, learning 50 objects from the COIL dataset with one ART instance, and learning 33 objects (17 objects being the same for two sets) with another ART instance leads to 92.9% correct for the first instance and 90.5% correct for the second instance. Melding them together creates a network that is 97% correct on all 66 unique objects learned by both ART instances. In addition, the melded version has memory footprint 83% of what the brute force combination of the two networks would have. Furthermore, the memory footprint of the melded version is 3% smaller than the combination of the first network with only new objects of the second network (excluding the overlapping 17 objects). Thus, melding indeed does the smart compression and refines the object representations to increase accuracy. If the inputs are not randomly shuffled, the results of melding are even more prominent in terms of correctness: 85.3 and 77.6% correct networks are melded into a 96.6% correct network thathas 84.6% of memory footprint of the two networks combined. These melding experimental results represent unexpectedly large performance improvements.

[0112] Consolidation process described above can be implemented in matrix form as shown in FIG. 4B, where F is the number of features and N is the number of nodes. By multiplying the combined weight matrix with its own transpose, the similarity matrix S is produced. Each value in this matrix shows the closeness of each pair of nodes in the feature space. By comparing these values with consolidation vigilances, a decision can be made if the two nodes should be consolidated into one, and then a new node can be created with the weight equal to a weighted average of the original node’s weights. While working on matrix consolidation implementation, the inventors recognized that the same process could allow to overcome the inherent sequential nature of original ART algorithm and its order dependency.

[0113] Examples of Parallel Training for ART

[0114] The inventors recognized and appreciated that parallel training for ART can be used in many different ways, not limited to L-DNN or any other specific implementation. For example, an ART system can work independently or as part of any suitable system.

[0115] FIG. 3D shows an example of conventional unsupervised sequential ART training with suboptimal order of inputs. FIG. 3E shows an example the same training with more optimal order of inputs as well as a target for some embodiments of parallel ART training. As shown, the inventors have recognized and appreciated that conventional sequential ART training (in which inputs 1-5 come one-by-one and the order will change the resulting nodes) produces nodes (e.g., nodes 305A and 306A) that are less consistent with existing classes than nodes (e.g., nodes 305B and 306B) with the optimal order or in parallel ART training (in which inputs 1-5 can come in any order and provide the same result).

[0116] The inventors have recognized that some features of conventional sequential ART can be omitted in parallel ART without loss of performance.

[0117] In some embodiments, the techniques described herein relate to an apparatus for realtime operation. For example, the apparatus may be similar to Module B 104 or L-DNN 106 as described herein, or any other suitable apparatus. The apparatus may include at least one processor, including but not limited to a CPU or GPU, a Field-programmable gate array(FPGA), application-specific integrated circuits (ASIC), volatile and non-volatile memory, input / output interfaces, and / or any other suitable components and structure.

[0118] In some embodiments, the apparatus may operate on continuous sensory input data to learn objects on-the-fly, for example, by providing an ART classifier operating in real-time to classify objects represented by the continuous sensory input data based on features extracted by a convolutional deep neural network. In some embodiments, the apparatus may (e.g., via the at least one processor) perform any or all of the suitable functions described herein.

[0119] FIG. 3F is a flowchart of an example method of implementing an Adaptive Resonance Theory (ART) classifier operating in parallel, which may be implemented via an apparatus as described herein. In some embodiments, the method may begin at step 320.

[0120] At step 320, the apparatus may train nodes in a first node group of the ART classifier in parallel with each other. The method may then proceed to step 330.

[0121] At step 330, the apparatus may train nodes in a second node group of the ART classifier in parallel with each other. The method may then proceed to step 340.

[0122] At step 340, the apparatus may merge weights of the nodes in the first node group with weights of the nodes in the second node group. The method may then end or repeat, as appropriate.

[0123] In conventional sequential ART, the first step is to determine for each input whether there exists a node in the system that is close enough to the input so it can be updated without creating a new node. This is done by comparing the distance to the vigilance parameter. The difficulty of parallelizing this operation for multiple inputs is that each input will change the weights of some existing node or create a new node, which appears to require to rerun the comparison for every input.

[0124] Inventors appreciated the following facts to enable proper parallelization: a. If the input from a batch of inputs is far enough from existing nodes to warrant the creation of a new node, it is highly unlikely that training any other or even all other inputs in a batch will allow to train this input without creating a new node.b. If the input from a batch of inputs is close enough to some existing node, it is likely that it will remain close to it even if all other inputs in a batch are trained before this input. c. Neural networks are statistical by nature, so following the statistically likely procedures rather than exact computations should not impair the outcome and in many cases will actually improve it.

[0125] Thus, in order to implement parallel ART, the inputs should be separated into two groups: those that will lead to creation of new nodes and those that would update existing nodes. When training happens from scratch, the system does not initially have any existing nodes, so all inputs will go in the group that creates new nodes.

[0126] In FIG. 3G this process of separation of inputs is shown in step 310, as discussed below.

[0127] FIG. 3G is a flowchart of an example method of implementing a Lifelong Learning Deep Neural Network (L-DNN) including an Adaptive Resonance Theory (ART) classifier operating in real-time. In some embodiments, the method may begin at step 310.

[0128] At step 310, the apparatus may separate inputs to the ART classifier into a first input group including inputs to be learned by existing nodes, and into a second input group including inputs to be learned by new nodes. Although any node can learn any input, one consideration is which input should be learned by existing nodes so that the accuracy of the resulting system is increased, and which inputs should be learned only by new nodes to avoid disrupting the previously learned knowledge represented by the existing nodes. The inventors recognized from sequential ART training that every input can take one of two paths: either the input is sufficiently different from previous inputs and should create a new node, or the input is sufficiently similar and should update an existing node.

[0129] In some embodiments, the apparatus may use index masks, slicing, and / or any other technique that can split an array of inputs into two arrays (which may be called groups or chunks) without copying the data. For example, Eigen masks may be used, such as two Eigen masks: one created for inputs for new nodes, and another created for inputs for existingnodes. FIG. 5 illustrates how masks can provide different views of the same input matrix to separate it in two groups.

[0130] In some embodiments, a similarity matrix computed by multiplying input matrix I and weight matrix W as shown in FIG. 6, where F is the number of features and B is the number of inputs in a batch. Elements of this matrix are compared with the training vigilance parameter set by the user. Usually, for supervised learning, training vigilance is set around 0 to allow maximum generalization, while for unsupervised learning, training vigilance is set higher (between 0.5 and 0.99) to allow more detailed clustering. Inputs that have no “winners” (e.g., no nodes higher than training vigilance) are put in the set for new node creation (i.e., a new node should be used for these inputs) by adding these inputs to the first mask. For those inputs that have “winners,” the apparatus may pick the best winner. In the case of supervised training, the additional condition of assigning the “winner” to be trained by particular input is matching the “winner’s” class label to the input’s ground truth label.

[0131] In some embodiments, step 310 may optionally include step 312. At step 312, the apparatus may assign each input based on a similarity matrix between inputs. In some embodiments, step 312 may include step 313.

[0132] At step 313, the apparatus may compare a max similarity value to a training vigilance. For example, the apparatus may compute a similarity matrix between inputs to the ART classifier and weights of nodes of the ART classifier. For each input to the ART classifier, the apparatus may: determine a maximal similarity value in the similarity matrix; and compare the maximal similarity value to a training vigilance. If the maximal similarity value is below the training vigilance, the method may proceed to step 314. If the maximal similarity value is not below the training vigilance, the method may proceed to step 315.

[0133] At step 314, the apparatus may assign the input to the second input group.

[0134] At step 315, the apparatus may assign the input to a node corresponding to the maximal similarity value (e.g., assign this node as a node to train for the input).

[0135] In some embodiments, if training is supervised, the apparatus may determine if the ground truth for the input matches the class label for the node that produced this maximal similarity value. If the ground truth for the input does not match the class label for the node,the apparatus may determine if there is another node that has matching class label and produces a similarity value within match tracking parameter of ART from the original similarity value. If the determination is false, the apparatus may assign the input to the second input group and stop processing this input. If the determination is true, the apparatus may select this node as a node to train for this input instead of the original “winner” node.

[0136] The method may proceed from step 315 to step 316.

[0137] At step 316, the apparatus may assign, to the first input group, at least one input with a corresponding node to train for the at least one input. In some embodiments, it is acceptable if all inputs are assigned to the second input group. Alternatively or additionally, it is acceptable if all inputs are assigned to the first input group.

[0138] For example, in some embodiments, the apparatus may compute a similarity matrix between inputs to the ART classifier and weights of nodes of the ART classifier; and for each input to the ART classifier: determine a maximal similarity value in the similarity matrix; compare the maximal similarity value to a training vigilance; in response to determining the maximal similarity value is below the training vigilance, assign the input to the second input group; and in response to determining the maximal similarity value is not below the training vigilance, assign a node corresponding to the maximal similarity value as a node to train for the input and assign, to the first input group, at least one input with a corresponding node to train for the at least one input.

[0139] As another example, if the network or model has existing nodes, the apparatus may compute a similarity matrix between inputs to the ART classifier and weights of nodes of the ART classifier; and for each input to the ART classifier: determine a maximal similarity value in the similarity matrix; and compare the maximal similarity value to a training vigilance. In response to determining the maximal similarity value is below the training vigilance, the apparatus may assign the input to the second input group. In response to determining the maximal similarity value is not below the training vigilance, the apparatus may assign a node corresponding to the maximal similarity value as a node to train for the input and assign, to the first input group, at least one input with a corresponding node to train for the at least one input.

[0140] In some embodiments, the apparatus may compute the matrix multiplication of the input feature vector array represented as a IxF matrix (where I is the number of inputs and F is the number of features in each input) and network or model weight vector array represented as a FxN matrix (where F is the number of features in each input and N is the number of existing nodes in the network or model) to produce the IxN similarity matrix between inputs and node weights. For each input vector, the apparatus may find the maximal similarity value in the similarity matrix. The apparatus may also determine if the maximal similarity value is below the global training vigilance. If the maximal similarity value is below the global training vigilance, the apparatus may assign the input to the second input group and stop processing this input. If the maximal similarity value is not below the global training vigilance, the apparatus may assign the node corresponding to the maximal similarity value as a node to train for this input.

[0141] Following step 310, the method may then proceed to both step 320A, step 330A, or both, which may be similar to the step 320 and step 330 described above, respectively. As shown in FIG. 3G, step 320A and step 330A may proceed in parallel or sequentially.

[0142] For example, at step 320A the apparatus may train nodes in a first node group of the ART classifier in parallel with each other using the first input group, and in parallel, at step 330A the apparatus may train nodes in a second node group of the ART classifier in parallel with each other using the second input group. For example, the apparatus may train the existing nodes in parallel with the new nodes. Similarly, in some embodiments, training from the first input group is done in parallel with training from the second input group. The inventors recognized these parallel training processes are possible because training from each input group is independent.

[0143] In some embodiments, the first node group includes the existing nodes, and the second node group includes new nodes.

[0144] In some embodiments, for all inputs in the first input group and corresponding nodes to train, the apparatus may apply the appropriate learning rule (e.g., ART and its variations have several different rules depending on the data representation, etc.; some processes described herein can be agnostic to which flavor is used). Since there are no interdependencies between different nodes, these processes can be done in parallel. For caseswhere multiple inputs update the same node, the average or weighted average of these inputs can be used so that the learning rule is only applied once per node.

[0145] In step 320 A, for each node that needs an update, all inputs that will update this node are selected, and the new weight is computed based on a weighted average of original weight and selected input vectors.

[0146] In step 330A, to ensure that no extra nodes are created for the similar inputs, the process similar to consolidation can be implemented, where the similarity matrix is computed for all inputs in the group, and all subgroups of similar inputs are combined into corresponding nodes.

[0147] In some embodiments, step 332 is included in step 330A. For example, the first part of step 330A may be computing the similarity matrix across all inputs in this group and deriving the weights from it. This may be an exact mirror of the consolidation procedure as depicted in FIG. 4B, except happening in step 332 before / during training rather than after training. In some embodiments, consolidation can be used as a separate call to consolidate across existing nodes. If an additional consolidation is desired, such additional consolidation can be applied after step 340 in FIG. 3G.

[0148] If the training is supervised, the apparatus may use ground truth to compute classspecific vigilance for the training based on how many inputs belong to each class. The apparatus may select the maximum of computed class-specific vigilance and the global training vigilance as a final value for class-specific vigilance for each class. If the training is unsupervised, the apparatus may use the global training vigilance as class-specific vigilance for all classes. This step is independent from computing the matrix multiplication described above and can be done in parallel.

[0149] In some embodiments, the apparatus may split the inputs in the second input group into subgroups so that in each subgroup: (a) If using supervised training, all the inputs belong to the same class; and (b) the similarity scores between each pair of inputs is above the classspecific vigilance for the corresponding class.

[0150] In some embodiments, for each of these subgroups, the apparatus may create a new node with weight equal to the average or weighted average of all input vectors that belong tothis subgroup. The inventors have recognized and appreciated that such subgroups are independent, so this process can also be done in parallel.

[0151] In some embodiments, step 320A and / or step 330A may proceed to step 340, which may be similar to step 340 described above.

[0152] At step 340, the apparatus may merge weights of the nodes in the first node group with weights of the nodes in the second node group. For example, the apparatus may merge the previously existing nodes used to train from the first input group, and newly created nodes from training from the second input group, into one combined weight matrix. The method may then end or repeat, as appropriate.

[0153] The inventors recognized that this process may lose the check in conventional ART if each input produces the correct output with updated weights. The inventors have recognized this does not seem to affect the results significantly, based on performed tests.

[0154] Examples of Parallel Inference for ART

[0155] The inventors have recognized that inference in ART networks may also or alternatively be done in parallel using a similar procedure. In some embodiments, parallel inference includes computing the similarity matrix and finding the maximal similarity value in the similarity matrix, followed by selection of the maximal similarity for each input and looking up the corresponding class label. In some embodiments, postprocessing may be used to receive a confidence score from the maximal similarity value.

[0156] Examples of Complete L-DNN Implementations

[0157] L-DNN Classifier

[0158] FIG. 7 represents example L-DNN implementation for whole image classification using a modified VGG-16 DNN as the core of Module A. Softmax and the last two fully connected layers are removed from the original VGG-16 DNN, and an ART-based Module B is connected to the first fully connected layer of the VGG-16 DNN. A similar but much simpler L-DNN can be created using Al exnet instead of VGG-16. This is a very simple and computationally cheap system that runs on any modern smartphone, does not require a GPUor any other specialized processor, and can learn any set of objects from a few frames of input provided by the smartphone camera.

[0159] L-DNN Grid-Based Detector

[0160] One way to detect objects of interest in an image is to divide the image into a grid and run classification on each grid cell. In this implementation of an L-DNN, the following features of CNNs are especially useful.

[0161] In addition to the longitudinal hierarchical organization across layers described above, each layer processes data maintaining a topographic organization. This means that irrespective of how deep in the network or kernel, stride, or pad sizes, features corresponding to a particular area of interest on an image can be found on every layer at various resolutions in the similar area of the layer. For example, when an object is in the upper left corner of an image, the corresponding features will be located in the upper left comer of each layer along the hierarchy of layers. Therefore, attaching a Module B to each of the locations in the layer allows the Module B to run classification on a particular location of an image and determine whether any familiar objects are present in this location.

[0162] Furthermore, only one Module B must be created per each DNN layer (or scale) used as input because the same feature vector represents the same object irrespective of the position in the image. Learning one object in the upper right corner thus allows Module B to recognize it anywhere in the image. Using multiple DNN layers of different sizes (scales) as inputs to separate Modules B allows detection on multiple scales. This can be used to fine tune the position of the object in the image without processing the whole image at finer scale as in the following process.

[0163] In this process, Module A provides the coarsest scale (for example, 7 * 7 in the publicly available Extract! onNet) image to Module B for classification. If Module B says that an object is located in the cell that is second from the left edge and fourth from top edge, only the corresponding part of the finer DNN input (for example, 14 x 14 in the same Extract! onNet) should be analyzed to further refine the location of the object.

[0164] Another application of multiscale detection can use a DNN design where the layer sizes are not multiples of each other. For example, if a DNN has a 30 x 30 layer it can bereduced to layers that are 2 * 2 (compression factor of 15), 3 * 3 (compression factor of 10), and 5 x 5 (compression factor of 6). As shown in FIG. 8, attaching Modules B to each of these compressed DNNs gives coarse locations of an object (indicated as 802, 804, 806). But if the output of these Modules B is combined (indicated as 808), then the spatial resolution becomes a nonuniform 8 x 8 grid with higher resolution in the center and lower resolution towards the edges.

[0165] Note that to achieve this resolution, the system runs the Module B computation only (2 x 2) + (3 x 3) + (5 x 5) = 38 times, while to compute a uniform 8 x 8 grid it does 64 Module B computations. In addition to being calculated with fewer computations, the resolution in the multiscale grid in FIG. 8 for the central 36 locations is equal to or finer than the resolution in the uniform 8 x 8 grid. Thus, with multiscale detection, the system is able to pinpoint the location of an object (810) more precisely using only 60% of the computational resources of a comparable uniform grid. This performance difference increases for larger layers because the square of the sum (representing the number of computations for a uniform grid) grows faster than sum of squares (representing the number of computations for a non- uniform grid).

[0166] Non-uniform (multiscale) detection can be especially beneficial for moving robots as the objects in the center of view are most likely to be in the path of the robot and benefit from more accurate detection than objects in the periphery that do not present a collision threat.

[0167] L-DNN for Image Segmentation

[0168] For images, object detection is commonly defined as the task of placing a bounding box around an object and labeling it with an associated class (e.g., “dog”). In addition to the grid-based method of the previous section, object detection techniques are commonly implemented by selecting one or more regions of an image with a bounding box, and then classifying the features within that box as a particular class, while simultaneously regressing the bounding box location offsets. Region-based CNN (R-CNN), Fast R-CNN, and Faster R- CNN implement this method of object detection, although any method that does not make the localization depend directly on classification information may be substituted as the detection module.

[0169] Image segmentation is the task of determining a class label for all or a subset of pixels in an image. Segmentation may be split into semantic segmentation, where individual pixels from two separate objects of the same class are not disambiguated, and instance segmentation, where individual pixels from two separate objects of the same class are uniquely identified or instanced. Image segmentation is commonly implemented by taking the bounding box output of an object detection method (such as R-CNN, Fast R-CNN, or Faster R-CNN) and segmenting the most prominent object in that box. The class label that is associated with the bounding box is then associated with segmented object. If no class label can be attributed to the bounding box, the segmentation result is discarded. The resulting segmented object may or may not have instance information. Mask R-CNN implements this method of segmentation is.

[0170] An L-DNN design for image detection or segmentation based on the R-CNN family of networks is presented in FIG. 9. Consider an image segmentation process that uses a static classification module, such as Mask R-CNN. In this scenario, the static classification module 900 may be replaced with an L-DNN Module B 104. That is, the segmentation pathway of the network remains unchanged; region proposals are made as usual, and subsequently segmented. Just as the case with a static classification module, when the L-DNN Module B 104 returns no positive class predictions that pass threshold (e.g., as would happen when the network is untrained or recognizes a segmented area as “Nothing I know” as described above), the segmentation results are discarded. Similarly, when the L-DNN Module B 104 returns an acceptable class prediction, the segmentation results are kept, just as with the static classification module. Unlike the static classification module 900, the L-DNN Module B 104 offers continual adaptation to change state from the former to the latter via user feedback.

[0171] User feedback may be provided directly through bounding box and class labels, such as is the case when the user selects and tags an object on a social media profile, or through indirect feedback, such as is the case when the user selects an object in a video, which may then be tracked throughout the video to provide continuous feedback to the L-DNN on the new object class. This feedback is used to train the L-DNN how to classify novel class networks over time. This process does not affect the segmentation component of the network.

[0172] The placement of Module B 104 in this paradigm also has some flexibility. The input to Module B 104 should be directly linked to the output of Module A convolutional layers 202, so that class labels may be combined with the segmentation output to produce a segmented, labeled output 902. This constraint may be fulfilled by having both Modules A and B take the output of a region proposal stage. Module A should not depend on any dynamic portion of Module B. That is, because Module B is adapting its network’s weights, but Module A is static, if Module B were to change its weights and then pass its output to Module A, Module A would likely see a performance drop due to the inability of most static neural networks to handle a sudden change in the input representation of its network.

[0173] Brain Consolidation and Brain Melding

[0174] Multiple real-time operating machines implementing L-DNN can individually learn new information on-the-fly through L-DNN. In some situations, it may be advantageous to share knowledge between real-time operating machines as outlined in several use cases described in the next sections. Since the real-time operating machines learn new knowledge on the edge, in order to share the new knowledge, each real-time operating machine sends a compressed and generalized representation of new information (represented in the network in terms of a synaptic weight matrix in Module B) from the edge to a central server or to other real-time operating machine(s). By implementing the following steps, knowledge acquired by each real-time operating machine can be extracted, appended, and consolidated either in a central server or directly on the edge device and shared with other real-time operating machines through centralized or peer-to-peer communication.• Learning new information in field deployment - As discussed above, a real-time operating machine can learn new information on-the-fly through L-DNN. When a user sees that a real-time operating machine is encountering a new object and / or new knowledge, she can provide a label for the new object and trigger the fast learning mode so that the real-time operating machine can learn the new object and / or new knowledge on-the-fly. In this manner, the real-time operating machine can modify its behavior and adapt to the new object and / or new knowledge quickly.• Consolidation of new knowledge - after one or more objects are learned on the fly, the system runs a consolidation process in the fast learning Module B. This process compresses the representation(s) of the new object(s), integrates it with therepresentations of previously known objects, improves network generalization abilities, and reduces the memory footprint of Module B. An example implementation based on ART network is detailed below.• Communicating consolidated individual brains to other devices - At any point during operation or after completing a mission, a real-time operating machine can transmit the consolidated weight matrix of its fast learning module (Module B) to a central server (e.g., a cloud-based server) through a wired or wireless communication channel. In some instances, the weight matrix of the fast learning module of each realtime operating machine can be downloaded to an external storage device and can be physically coupled to the central server. When the central server is not available or not desirable, the communication can happen in peer-to-peer fashion among real-time operating machines (edge devices).• Brain melding (or fusing, merging, combining) - After the weight matrices from several real-time operating machines are collected at the central server or one of the edge devices, the central server or edge device can run a melding utility that combines, compresses, and consolidates the knowledge that is newly acquired from each real-time operating machine into a single weight matrix. The melding utility reduces the memory footprint of the resulting matrix and removes redundancy while preserving accuracy of the entire system. An example implementation based on ART network is detailed below.• Updating individual brains after melding - The resulting weight matrix that is created during brain melding is then downloaded to one or more real-time operating machines through wired or wireless communication channels or by downloading it to a physical external storage / memory device and physically transferring the storage / memory device to the real-time operating machine.

[0175] In this manner, the knowledge from multiple real-time operating machines can be consolidated and new knowledge learned by each of these machines can be shared with other real-time operating machines.

[0176] Example Use Cases of L-DNN and / or Parallel ART

[0177] The following use cases are non-limiting examples of how an L-DNN and / or parallel ART can address technical problems in a variety of fields.

[0178] Automate Inspections: Single or Multiple Sources of Imagery

[0179] Consider a drone service provider wanting to automate the inspection process for industrial infrastructure, for example, power lines, cell towers, or wind turbines. Existing solutions require an inspector to watch hours of drone videos to find frames that include key components that need to be inspected. The inspector must manually identify these key components in each of the frames.

[0180] In contrast, an L-DNN based assistant can be introduced to the identification tool. Data that includes labels for objects or anomalies of interest can be provided to the L-DNN based assistant as a pre-trained set during conventional slow DNN factory training. Additions can be made by the user to this set during fast learning mode as described below.

[0181] An L-DNN based assistant can be included on a “smart” drone or on a computer used to review videos acquired by a “dumb” drone. A drone inspects a structure, such as a telecommunication tower, a solar panel array, a wind turbine farm, or a power line distribution (these are only example structures, others can be envisioned). A drone operator may be using manual control of the drone, or supervising a drone functioning automatically. A human analyst, such as an analyst in a control room, can provide labels as the drone is flying, or post flight, to Module B 104 in an L-DNN system 106 that processes sensory input (e.g., video, LIDAR, etc.) 100 from the drone.

[0182] Initially, the drone receives a copy of the L-DNN 106 as its personal local classifier. When the drone acquires video frames 100 while inspecting these power lines, cell towers, and wind turbines, Module A 102 of the L-DNN 106 extracts image features from the video frames 100 based on pre-trained data. Module B 104 then provides a probable label for each object based on these features. This information is passed to the user. If the user finds the labels unsatisfactory, she can engage a fast learning mode to update the Module B network with correct labels. In this manner, user-provided information may correct the current label. Thus, the fast learning subsystem can utilize one-trial learning to determine the positions and features of already learned objects, such as power lines, cell towers, and wind turbines, as early as in the first frame after update. In the case of analyzing the video taken earlier, that means immediately after the user introduced correction. The system 106 thus becomes more knowledgeable over time and with user’s help provides better identification over time.

[0183] L-DNN technology described herein can be applied as a generalized case of that above, where multiple drones collect data synchronously or asynchronously. The information learned by the L-DNN associated with each drone can be merged (combined or melded) and pushed back to the other drones, shared peer-to-peer among drones, or shared with a central server that contains a Module B 104. The central server merges the individual L-DNN learned information and pushes the merged information back to all drones, including a drone that has not been exposed to the information about telecommunication towers, solar panel arrays, wind turbine farm, or power line distribution derived from data acquired by other drones, but is now able to understand and classify these items thanks to the merging process.

[0184] Parallel ART can be applied by training an ART network (e.g., as Module B or as an independent system) in parallel and / or using parallel inference. As discussed herein, this can reduce or eliminate input order dependency, among other advantages, which could make drone pathing and / or image sequences more flexible.

[0185] Automate Warehouse Operations: Consolidating and Melding Knowledge from Multiple Sources

[0186] The system described above can be extended for multiple machines or cameras (fixed, drone bound, etc.) that operate in concert. Consider a company with large warehouses in multiple varied geographic locations. Taking manual inventory in large warehouses can take many man-hours and usually requires the warehouses to be shut during this time. Existing automated solutions have difficulty identifying stacked objects that may be hidden. In addition, in existing automated solutions, information that is learned at one geographic location is not transferred to other locations. In some instances, because of the vast amount of data that is collected at different geographic locations, these automated solutions can take weeks to learn new data and act upon the new data.

[0187] In contrast, the L-DNN technology described herein can be applied to a warehouse, industrial facility, or distribution center environment, where sensors (e.g., fixed cameras or moving cameras mounted on robots or drones) can learn on-the-fly new items in inventory via various L-DNN modules connected to the sensors. Additionally, operators could teach new information to various L-DNN modules in a decentralized fashion. This new knowledgecan be integrated centrally or communicated peer-to-peer and pushed back after melding to each individual device (e.g., a camera).

[0188] For instance, consider an example of fixed cameras. Each camera acquires corresponding video imagery 100 of objects on a conveyor belt and provides that imagery 100 to a corresponding L-DNN 106a-106c (collectively, L-DNNs 106). The L-DNNs 106 recognize known objects in the imagery 100, e.g., for inspection, sorting, or other distribution center functions.

[0189] Each L-DNN 106 tags unknown objects for evaluation by human operators or as “nothing I know.” For instance, when presented with an unknown object, L-DNN 106a flags the unknown object for classification by human operator. Similarly, L-DNN 106c flags unknown object for classification by human operator. When L-DNN 106b is presented with an unknown object, it simply tags the unknown object as “nothing I know.” A standalone Module B 104d coupled to the L-DNNs 106 merges the knowledge acquired by L-DNNs 106a and 106c from the human operators and pushes it to the Module B 104b in L-DNN 106b so that Module B 104b can recognize future instances of objects.

[0190] The L-DNN 106 for each device can be pre-trained to recognize existing landmarks in the warehouses, such as pole markings, features like EXIT signs, a combination thereof, and / or the like. This enables the system to triangulate a position of the unmanned vehicle equipped with a sensor or appearing in an image acquired by a sensor (e.g., a camera). The L- DNN in each vehicle operate in exact same way as in the use case described above. In this manner, knowledge from multiple unmanned vehicles can be consolidated, melded, and redistributed back to each unmanned vehicle. The consolidation and melding of knowledge from all locations can be carried out by a central server as described in the consolidation and melding section above; additionally, peer-to-peer melding can be applied as well. Thus, inventory can be taken at multiple warehouses and the knowledge consolidated with minimal disruption to warehouse operations.

[0191] Parallel ART can be applied by training an ART network (e.g., as one of the Module Bs, the standalone Module B, or as an independent system) in parallel and / or using parallel inference. As discussed herein, this can reduce or eliminate input order dependency, amongother advantages, which could make improve the flexibility of how inputs are processed to or from camera-equipped robots.

[0192] In a Fleet of Mobile Devices

[0193] Consider distributed networks of consumer mobile devices, such as consumer smart phones and tablets, or professional devices, such as the mobile cameras, body -worn cameras, and LTE handheld devices used by first responders and public safety personnel for public safety. Consumer devices can be used to understand the consumer’s surroundings, such as when taking pictures. In these cases, L-DNN technology described herein can be applied to smart phone or tablet devices. Individuals (e.g., users) could teach knowledge to the L-DNN modules 106 in devices and merge this information peer-to-peer or on a server containing a Module B 104. The server pushes the merged knowledge back to some or all connected devices, possibly including devices that did not participate in the original training.

[0194] The L-DNN modules can learn, for example, to apply image processing techniques to pictures taken by users, where users teach each L-DNN some customized actions associated with aspects of the picture (e.g., apply a filter or image distortion to these classes of objects, or areas). The combined learned actions could be shared, merged, or combined peer-to-peer or collectively across devices. Additionally, L-DNN technology can be applied to the generalized use case of smart phone usage, where input variables can be sensory or non- sensory (any pattern of usage of the smart phone). These patterns of usages, which can be arbitrary combinations of input variables and output variables, can be learned at the smart phone level, and pushed to a central L-DNN module 104, merged, and pushed back to individual devices.

[0195] In another example, a policeman can be looking for a lost child, a suspect, or a suspicious object using a professional device running an L-DNN. In such a situation, officers and / or first responders cannot afford to waste time. Existing solutions provided to officers and / or first responders require video feeds from the cameras to be manually analyzed and coordinated. Such solutions take too long since they require using a central server to analyze and identify objects. That is, such solutions have major latency issues since the video data needs to be analyzed in the cloud / central server. This could be a serious hurdle for first responders / officers who often need to act immediately as data is received. In addition,sending video data continuously to the central server can put a strain on communication channels. Accuracy is also of great importance to avoid incorrect identification and speedy discovery, which parallel ART can provide (e.g., as one of the Module Bs, the server Module B, or as an independent system) in parallel and / or using parallel inference.

[0196] Instead, by using L-DNN in mobile phones, body-worn cameras, and LTE handheld devices, data can be learned and analyzed on the edge itself. Consumers can learn to customize their device in-situ, and officers / first responders can look for and provide a location of the person / object as well as search and identify the person / object of interest in places that the officers might not be actively looking at. The L-DNN can utilize a fast learning mode to learn from an officer on the device in the field instead of learning from an operator on a remote server, reducing or eliminating latency issues associated with centralized learning.

[0197] FIG. 1 can illustrate the operation of an L-DNN in a mobile phone for labeling components in an image as a consumer points the phone to a scene and labels the whole scene or sections of the scene (objects, section of scene, such as sky, water, etc.). Additionally, a police officer can access video frames and identifying suspicious person / object. When the mobile phone acquires video frames 100, a Module A 102 can extract image features from these frames based on pre-trained data. A Module B 104 can then provide a probable label for each object using these features. For instance, if person A lives in a neighborhood B and has been observed in neighborhood B in the past, then person A may be labeled as a “resident” of neighborhood B. Thus, the fast learning subsystem can utilize one-trial learning to determine the relative positions and features of already learned objects, such as, houses, trees, as early as immediately after the first frame of learning. More importantly, a dispatcher on the central server 110 can introduce a new object to find to the server-side L-DNN, and it will be melded and distributed to local first responders as needed.

[0198] This use case is very similar to the previous use cases, but takes greater advantage of L-DNN’ s ability to learn new objects quickly without forgetting old. While during inspections and inventory collections there is usually little time pressure and memory consolidation can be done in the slow learning mode, in the case of first responders it can be important to consolidate and meld knowledge from multiple devices as fast as possible so that all devices in the area can start searching for the suspect or missing child. Thus, the ability ofL-DNN to quickly learn a new object guided by one first responder, almost instantaneously consolidate it on the server and distribute to all first responders in the area becomes a tremendous advantage for this use case.

[0199] Replacing conventional DNNs with L-DNN s and / or Parallel ART in data centers

[0200] The L-DNN and / or parallel ART technology described herein can be applied as a tool to decrease computational time for DNN processes in individual compute nodes or servers in large data centers. L-DNN technology speeds up learning in DNN by several orders of magnitude. This feature can be used to dramatically decrease the need in, or cut the consumption of, computational resources on the servers, where information can be learned over a few seconds across massive datasets 100 that usually requires hours / days / weeks of training time. The use of L-DNN also results in reduction of power consumption, and overall better utilization of server resources in data centers. The use of parallel ART can result in improved accuracy and further reduction in power consumption and improved resource utilization.

[0201] Conclusion

[0202] As described above, an L-DNN can provide on-the-fly (one-shot) learning for neural network systems. Conversely, traditional DNNs often require thousands or millions of iteration cycles to learn a new object. The larger the step size taken per iteration cycle, the less likely that the gradient of the loss function can lead to actual performance gains. Hence, these traditional DNN make small changes to their weights per training sample. This makes it extremely difficult to add new knowledge on-the-fly. In contrast, an L-DNN with fast learning neural networks can learn stable object representations with very few training examples. In some instances, just one training example can suffice for an L-DNN.

[0203] ART is a good choice for fast learning neural network in L-DNN for multiple reasons. For example, ART is resistant to the “catastrophic forgetting” that plagues traditional DNNs. Moreover, ART allows one-shot learning, ART’s computational complexity only mildly grows with the amount of knowledge, and the knowledge from several ARTs can be easily merged together.

[0204] Moreover, parallel ART can improve the accuracy, speed, and efficiency of fast training portions of an L-DNN or as an independent system, as described herein.

[0205] While various inventive embodiments have been described and illustrated herein, those of ordinary skill in the art will readily envision a variety of other means and / or structures for performing the function and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is deemed to be within the scope of the inventive embodiments described herein. More generally, those skilled in the art will readily appreciate that all parameters, dimensions, materials, and configurations described herein are meant to be exemplary and that the actual parameters, dimensions, materials, and / or configurations will depend upon the specific application or applications for which the inventive teachings is / are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific inventive embodiments described herein. It is, therefore, to be understood that the foregoing embodiments are presented by way of example only and that, within the scope of the appended claims and equivalents thereto, inventive embodiments may be practiced otherwise than as specifically described and claimed. Inventive embodiments of the present disclosure are directed to each individual feature, system, article, material, kit, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits, and / or methods, if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent, is included within the inventive scope of the present disclosure.

[0206] The above-described embodiments can be implemented in any of numerous ways. For example, embodiments may be implemented using hardware, software or a combination thereof. When implemented in software, the software code can be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers.

[0207] Further, it should be appreciated that a computer may be embodied in any of a number of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, a computer may be embedded in a device not generally regarded as a computer but with suitable processing capabilities, including a Personal Digital Assistant (PDA), a smart phone or any other suitable portable or fixed electronic device.

[0208] Also, a computer may have one or more input and output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that can be used to provide a user interface include printers or display screens for visual presentation of output and speakers or other sound generating devices for audible presentation of output. Examples of input devices that can be used for a user interface include keyboards, and pointing devices, such as mice, touch pads, and digitizing tablets. As another example, a computer may receive input information through speech recognition or in other audible format.

[0209] Such computers may be interconnected by one or more networks in any suitable form, including a local area network or a wide area network, such as an enterprise network, and intelligent network (IN) or the Internet. Such networks may be based on any suitable technology and may operate according to any suitable protocol and may include wireless networks, wired networks or fiber optic networks.

[0210] The various methods or processes outlined herein may be coded as software that is executable on one or more processors that employ any one of a variety of operating systems or platforms. Additionally, such software may be written using any of a number of suitable programming languages and / or programming or scripting tools, and also may be compiled as executable machine language code or intermediate code that is executed on a framework or virtual machine.

[0211] Also, various inventive concepts may be embodied as one or more methods, of which an example has been provided. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed in which acts are performed in an order different than illustrated, which may include performing some acts simultaneously, even though shown as sequential acts in illustrative embodiments.

[0212] All publications, patent applications, patents, and other references mentioned herein are incorporated by reference in their entirety.

[0213] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0214] The indefinite articles “a” and “an,” as used herein in the specification and in the claims, unless clearly indicated to the contrary, should be understood to mean “at least one.”

[0215] The phrase “and / or,” as used herein in the specification and in the claims, should be understood to mean “either or both” of the elements so conjoined, i.e., elements that are conjunctively present in some cases and disjunctively present in other cases. Multiple elements listed with “and / or” should be construed in the same fashion, i.e., “one or more” of the elements so conjoined. Other elements may optionally be present other than the elements specifically identified by the “and / or” clause, whether related or unrelated to those elements specifically identified. Thus, as a non-limiting example, a reference to “A and / or B”, when used in conjunction with open-ended language such as “comprising” can refer, in one embodiment, to A only (optionally including elements other than B); in another embodiment, to B only (optionally including elements other than A); in yet another embodiment, to both A and B (optionally including other elements); etc.

[0216] As used herein in the specification and in the claims, “or” should be understood to have the same meaning as “and / or” as defined above. For example, when separating items in a list, “or” or “and / or” shall be interpreted as being inclusive, i.e., the inclusion of at least one, but also including more than one, of a number or list of elements, and, optionally, additional unlisted items. Only terms clearly indicated to the contrary, such as “only one of’ or “exactly one of,” or, when used in the claims, “consisting of,” will refer to the inclusion of exactly one element of a number or list of elements. In general, the term “or” as used herein shall only be interpreted as indicating exclusive alternatives (i.e. “one or the other but not both”) when preceded by terms of exclusivity, such as “either,” “one of,” “only one of,” or “exactly one of.” “Consisting essentially of,” when used in the claims, shall have its ordinary meaning as used in the field of patent law.

[0217] As used herein in the specification and in the claims, the phrase “at least one,” in reference to a list of one or more elements, should be understood to mean at least one element selected from any one or more of the elements in the list of elements, but not necessarily including at least one of each and every element specifically listed within the list of elements and not excluding any combinations of elements in the list of elements. This definition also allows that elements may optionally be present other than the elements specifically identified within the list of elements to which the phrase “at least one” refers, whether related orunrelated to those elements specifically identified. Thus, as a non-limiting example, “at least one of A and B” (or, equivalently, “at least one of A or B,” or, equivalently “at least one of A and / or B”) can refer, in one embodiment, to at least one, optionally including more than one, A, with no B present (and optionally including elements other than B); in another embodiment, to at least one, optionally including more than one, B, with no A present (and optionally including elements other than A); in yet another embodiment, to at least one, optionally including more than one, A, and at least one, optionally including more than one, B (and optionally including other elements); etc.

[0218] In the claims, as well as in the specification above, all transitional phrases such as “comprising,” “including,” “carrying,” “having,” “containing,” “involving,” “holding,” “composed of,” and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases “consisting of’ and “consisting essentially of’ shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03.

Claims

CLAIMS1. A method of implementing an Adaptive Resonance Theory (ART) classifier operating in real-time, the method comprising: training nodes in a first node group of the ART classifier in parallel with each other; training nodes in a second node group of the ART classifier in parallel with each other; and merging weights of the nodes in the first node group with weights of the nodes in the second node group.

2. The method of claim 1, wherein: the method comprises separating inputs to the ART classifier into a first input group including inputs to be learned by existing nodes, and into a second input group including inputs to be learned by new nodes, the nodes in the first node group include the existing nodes, and the nodes in the second node group include new nodes.

3. The method of claim 2, wherein: training the nodes in the first node group of the ART classifier in parallel with each other comprises training the existing nodes using the first input group, and training the nodes in the second node group of the ART classifier in parallel with each other comprises training the new nodes using the second input group.

4. The method of claim 3, wherein: the method comprises training the existing nodes in parallel with the new nodes.

5. The method of claim 3, wherein: training the nodes in the second node group includes consolidating weights of all of the new nodes.

6. The method of claim 2, wherein: the method comprises:computing a similarity matrix between inputs to the ART classifier and weights of nodes of the ART classifier; and for each input to the ART classifier: determining a maximal similarity value in the similarity matrix; comparing the maximal similarity value to a training vigilance; in response to determining the maximal similarity value is below the training vigilance, assigning the input to the second input group; and in response to determining the maximal similarity value is not below the training vigilance, assigning a node corresponding to the maximal similarity value as a node to train for the input and assigning, to the first input group, at least one input with a corresponding node to train for the at least one input.

7. An apparatus for real-time operation on continuous sensory input data to learn objects on-the-fly, the apparatus comprising: at least one processor configured to: provide an Adaptive Resonance Theory (ART) classifier operating in real-time to classify objects represented by the continuous sensory input data based on features extracted by a convolutional deep neural network; train nodes in a first node group of the ART classifier in parallel with each other; train nodes in a second node group of the ART classifier in parallel with each other; and merge weights of the nodes in the first node group with weights of the nodes in the second node group.

8. The apparatus of claim 7, wherein: the at least one processor is configured to separate inputs to the ART classifier into a first input group including inputs to be learned by existing nodes, and into a second input group including inputs to be learned by new nodes, the nodes in the first node group include the existing nodes, and the nodes in the second node group include new nodes.

9. The apparatus of claim 8, wherein the at least one processor is configured to: train the nodes in the first node group of the ART classifier in parallel with each other by training the existing nodes using the first input group, and train the nodes in the second node group of the ART classifier in parallel with each other by training the new nodes using the second input group.

10. The apparatus of claim 9, wherein the at least one processor is configured to: train the existing nodes in parallel with the new nodes.

11. The apparatus of claim 9, wherein: to train the second node group includes to consolidate weights of all of the new nodes.

12. The apparatus of claim 8, wherein the at least one processor is configured to: compute a similarity matrix between inputs to the ART classifier and weights of nodes of the ART classifier; and for each input to the ART classifier: determine a maximal similarity value in the similarity matrix; compare the maximal similarity value to a training vigilance; in response to determining the maximal similarity value is below the training vigilance, assign the input to the second input group; and in response to determining the maximal similarity value is not below the training vigilance, assign a node corresponding to the maximal similarity value as a node to train for the input and assign, to the first input group, at least one input with a corresponding node to train for the at least one input.

13. A method of implementing a Lifelong Learning Deep Neural Network (L-DNN) including an Adaptive Resonance Theory (ART) classifier operating in real-time, the method comprising: separating inputs to the ART classifier into a first input group including inputs to be learned by existing nodes, and into a second input group including inputs to be learned by new nodes, wherein the first node group includes the existing nodes, and the second node group includes new nodes; training nodes in a first node group of the ART classifier in parallel with each other;training nodes in a second node group of the ART classifier in parallel with each other, wherein training the nodes in the second node group includes consolidating weights of all of the new nodes, and training the nodes in the first node group is in parallel with training the second node group; and merging weights of the nodes in the first node group with weights of the nodes in the second node group.

14. The method of claim 13, wherein: training the nodes in the first node group of the ART classifier in parallel with each other comprises training the existing nodes using the first input group, and training the nodes in the second node group of the ART classifier in parallel with each other comprises training the new nodes using the second input group.

15. The method of claim 14, wherein: the method comprises: computing a similarity matrix between inputs to the ART classifier and weights of nodes of the ART classifier; and for each input to the ART classifier: determining a maximal similarity value in the similarity matrix; comparing the maximal similarity value to a training vigilance; in response to determining the maximal similarity value is below the training vigilance, assigning the input to the second input group; and in response to determining the maximal similarity value is not below the training vigilance, assigning a node corresponding to the maximal similarity value as a node to train for the input and assigning, to the first input group, at least one input with a corresponding node to train for the at least one input.

Citation Information

Patent Citations

  • Systems and methods to enable continual, memory-bounded learning in artificial intelligence and deep learning continuously operating applications across networked compute edges

    US20180330238A1

  • Integrated Circuit Designs for Reservoir Computing and Machine Learning

    US20210406648A1

  • Decentralized artificial intelligence (AI) / machine learning training system

    US20220344049A1