Systems and methods for enabling memory-bounded continuous learning in artificial intelligence and deep learning for continuously operating applications across the network computational edge
Lifetime Deep Neural Networks (L-DNNs) address the inefficiencies of traditional DNNs in real-time learning by employing a heterogeneous neural network architecture, enabling continuous online learning and efficient knowledge updates on edge devices.
Patent Information
- Application Number
- JP2023117276
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2017-12-31
- Filing Date
- 2023-07-19
- Publication Date
- 2025-05-22
- Estimated Expiration
- 2038-05-09
AI Technical Summary
Traditional deep neural networks (DNNs) face challenges in real-time learning and knowledge updates due to high computational requirements, memory constraints, and the need for extensive training, making them inefficient for edge devices and real-time applications.
The implementation of Lifetime Deep Neural Networks (L-DNNs) which utilize a heterogeneous neural network architecture comprising a slow learning module A and a fast learning module B, enabling continuous online learning and knowledge updates without the need for extensive training or storage of input data.
L-DNNs facilitate rapid learning and adaptation in real-time environments, reducing computational resources and memory requirements, allowing for efficient knowledge updates and merging across edge devices, and enabling learning on resource-constrained devices.
Smart Images

Figure 0007681650000003 
Figure 0007681650000004 
Figure 0007681650000005
Abstract
Description
[Technical field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a joint venture of U.S. Patent No. 5,399,414, filed on December 31, 2017 under 35 U.S.C. § 119(e). Application No. 62 / 612,529, filed May 9, 2017, and U.S. patent application Ser. No. 503,639, each of which is incorporated by reference in its entirety. and is hereby incorporated by reference. [Background technology]
[0002] Deep neural networks that contain many layers of neurons placed between the input and output layers. Conventional neural networks, including deep learning neural networks (DNNs), are limited to a specific data set. For example, a scalar requires thousands or millions of iteration cycles for training. These cycles occur frequently in high performance computing servers. Depending on the size of the dataset, some traditional DNNs can take days or weeks to train. This may also be necessary.
[0003] One technique for training a DNN involves the backpropagation algorithm. Rhythm uses the chain rule to backpropagate the error gradient to find the label Calculate the change in all weights of the DNN proportional to the error gradient from the filtered dataset. changes the weights for each data point by a small amount and spreads it across all data points in the set for many epochs. .
[0004] The larger the learning rate in each iteration, the more likely it is that the gradient of the loss function will move toward a local minimum instead of a minimum. loss function, which may lead to poorer performance. To increase the likelihood that the numbers settle to a minimum, the DNN decreases the learning rate, thereby causing the weights to change slightly over all training epochs. This increases the number of training cycles and the total learning time.
[0005] Advances in graphics processing unit (GPU) technology have led to a significant improvement in the computing power for highly parallel operations that were once used to achieve training jobs that took weeks or months. These jobs can now be completed on a GPU in hours or days, but this is still not fast enough to update knowledge in real time. Furthermore, using a high- performance computing server to update the DNN is criticized in terms of server cost and energy consumption. This makes it very difficult to update the knowledge of an on-the- fly DNN-based system, which is desirable in many cases of real-time operation.
[0006] Furthermore, the gradient of the loss function computed for any single training sample can affect all the weights in the network (by the normal distributive representation), so a standard DNN tends to forget previous knowledge when learning new objects. Repeated representation of the same input over multiple epochs mitigates this problem, but has the drawback that it is very difficult to quickly add new knowledge to the system. This is one reason why learning is difficult or completely impossible to implement on resource- constrained edge devices (e.g., mobile phones, tablets, or small form factor processors). Even if the forgetting problem is solved, learning on edge devices has the high This remains infeasible due to the small training steps and the iterative representation of the entire input. .
[0007] These limitations are due to the fact that over the life of a deployment, the edge may need to update its knowledge Rapidly apply newly acquired knowledge not only at the compute edge but across the deployment lifecycle In distributed multi-edge systems ( For example, smartphones connected in a network, networked smart phones This also applies to cameras, drones or fleets of autonomous vehicles, and similar.
[0008] The processor that runs the backpropagation algorithm calculates the contribution of each neuron to the error at the output. The weights of all neurons are then calculated by the loss function The gradient of the training set is then calculated, so that the network can correctly predict old examples. To avoid losing the ability to classify, new training examples are added without retraining old examples. However, it cannot be added to a pre-trained network to correctly classify old examples. The loss of ability is called "catastrophic forgetting." This forgetting problem occurs when In real time, where there is often a need to rapidly learn and incorporate new information on the fly This is particularly relevant when considered in connection with a working machine.
[0009] Real-time machines that use traditional DNNs to learn knowledge are It may be necessary to accumulate large amounts of data in order to retrain the NN. The data is then fed to the operator, who gets the labels and then re-runs the DNN that runs on the edge. To train them, we use machines that operate in real time on the “edge” (i.e. For example, devices themselves (e.g., autonomous vehicles, drones, robots, etc.) can transmit information to a central server (e.g., The more data that is stored, the more time it takes to The transfer process is more expensive in terms of bandwidth and network usage. Interleaving training on the system reduces the time it takes to update new data. It must be combined with the original data that is stored for the entire cycle. This makes it difficult Transmission bandwidth and data storage limitations are created.
[0010] In summary, we propose a method to improve the performance of traditional backpropagation-based DNN training in real-time operating environments. When applied to a computing system, it suffers from the following drawbacks: a. It is not possible to update the system with new knowledge on the fly. b. Edge,without regular communication with the server and significant latency for,knowledge updates. Learning throughout the deployment cycle is not possible. c. Learning new information requires storing all input data indefinitely for further training. This requires server space, energy consumption, and disk space consumption to store the data. d. Learning cannot be done on small form factor computing edge devices. e. Multiple edge deployments are not possible without slow and expensive server-side retraining and redeployment. You cannot merge knowledge across pages. Summary of the Invention
[0011] Lifetime Deep Neural Networks (L-DNN) reduce time-consuming computations It is possible to develop artificial neural networks on lightweight computing devices (edge) without the need for intensive training by human. Neural Networks (ANN) and Deep Neural Networks (DNN) L-DNN enables continuous online lifelong learning by learning from continuous data streams. This allows real-time training from the input data for multiple iterations of backpropagation training. This avoids the need to store the data.
[0012] L-DNN technology involves the fast yet robust learning of features that represent entities or events of interest. To achieve this, we developed a subsystem (Module A) based on a representation-rich DNN for rapid learning. Combine with the subsystem (Module B). These feature sets are based on low level It can be pre-trained using fast learning methodologies, as described in detail in this disclosure (among other features). The description of is possible by adopting a non-DNN methodology in module A)DN For N-based examples, the high-level feature extraction layer of the DNN classifies known entities and events. and add knowledge of unknown entities and events on the fly. It serves as the input to a fast learning system. Module B can learn to: It can learn important information and capture descriptive and highly predictive environmental features.
[0013] The L-DNN technique works by analyzing visual data, structured light data, and Lidar (LI), among other modalities. DAR data, SONAR data, RADAR data, or It can be applied to audio data. For visual or similar data, the L-DNN technique: Full-image classification (e.g., scene detection), bounding box-based object recognition, pixel-wise segmentation The method can be applied to visual processing, enabling cognition, retrieval, and other visual recognition tasks. L-DNN techniques are also useful for non-visual recognition tasks such as classification of non-visual signals, and for robots, autonomous vehicles, and other As a self-driving car, drone, or other device navigates the environment, it gradually learns By adding It can also perform other tasks, such as updating the generated map.
[0014] L-DNN can learn to recognize more entities or events (in visual terms, “objects” or “categories”). When learning the “real-time” feature, we reduce memory requirements by aggregating memory within the L-DNN. In addition, the L-DNN methodology is used to keep multiple edge nodes under the control of module B. Computing devices traverse the edge and apply their knowledge (or classify input data) The merge is the process of changing the new module B between the two modules. by direct exchange of the neural network representation or by multiple modules from some edges. This can happen peer-to-peer, via an intermediate server, which merges the representations of role B and role B. In addition, L-DDN does not rely on backpropagation, thereby reducing training time, power requirements, and It is possible to update the L-DNN knowledge using new input data with dramatically reduced computational resources. do.
[0015] Of course, all of the above concepts and the additional concepts discussed in more detail below Combinations (provided that such concepts are not mutually exclusive) are not intended to be limiting unless otherwise disclosed herein. In particular, the claims appearing at the end of this disclosure are considered to be part of the subject matter of the present invention. All combinations of subject matter recited in the ranges are part of the inventive subject matter disclosed herein. It is also understood that any disclosure incorporated by reference is expressly construed as including the The terms used as shown are to be given the meaning that most closely matches the specific concepts disclosed in this specification. It is necessary.
[0016] Upon consideration of the following charts and detailed description, other systems, processes, and features will become apparent to those skilled in the art. All such additional systems, processes, and features are included within this description, are within the scope of the present invention, and are intended to be protected by the appended claims. **Brief Description of the Drawings**
[0017] Those skilled in the art will understand that the drawings are primarily for illustrative purposes and are not intended to limit the scope of the subject matter of the present invention described herein. The drawings are not necessarily to scale, and in some instances, various aspects of the subject matter of the present invention disclosed herein may be shown exaggerated or enlarged in the drawings to facilitate understanding of different features. In the drawings, like reference characters generally mean like features (e.g., functionally similar and / or structurally similar elements).
[0018] [Figure 1] Figure 1 shows an overview of a lifetime deep neural network (L-DNN) when associated with a plurality of computing edges, each of which either acts on a data stream individually, or is connected peer-to-peer or via an intermediate computing server. [Diagram 2] Figure 2 shows an example of an L-DNN architecture. [Diagram 3] Figure 3 shows an implementation of the concept of the unknown in a neural network. [Figure 4] Figure 4 shows a VGG-16 based L-DNN classifier as an exemplary implementation. [Diagram 5]FIG. 5 shows heterogeneous multi-scale object detection. [Figure 6] Figure 6 shows L-DNN based on Mask R-CNN for object segmentation. [Figure 7A] FIG. 7A shows aggregation and blending using an adaptive resonance theory (ART) neural network. [Figure 7B] Figure 7B shows how locally ambiguous information, e.g., a pixelated image of a camel (in the desert in the first scene) or a dog (in the suburbs in the second scene), can be found using global information about the scene and past associations learned between objects. [Figure 8] Figure 8 shows the application of L-DNN to the drone-based industrial survey use case. [Figure 9] In Figure 9, we extend the drone-based industrial survey use case from Figure 8 to a situation where multiple drones with L-DNNs are operating in a concert. [Figure 10] Figure 10 shows the application of L-DNN to a warehouse inventory use case. [Figure 11] FIG. 11 shows a large number of smart devices using L-DNN to collectively acquire and share knowledge. [Figure 12] Figure 12 shows an example where L-DNN replaces traditional DNN in a data center-based application. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0019] Real-time machine learning Lifelong Learning Deep Neural Networks or Lifelong Deep Neural Networks (L-DNN) enables machines to operate in real time on a central server or on the cloud. This allows for on-the-fly learning at the edge, without the need for training at the to eliminate network latency, increase real-time performance, and prioritize when needed. In some cases, real-time machines use L-DNNs. For example, L-DNN can be used to map survey dropouts. A robot can learn how to identify problems at the top of a cell tower or solar panel, No need to worry about privacy issues as data is not shared outside of the local device Instead of relying on the user's preferences, the smart toy can be personalized and the smartphone can Knowledge learned at the edge, without having to send information to a central server for constant, lengthy learning Share your knowledge (peer-to-peer or globally with all devices) or use self-driving cars When it works, knowledge can be learned and shared.
[0020] L-DNN also makes it possible to learn new knowledge without forgetting old knowledge. In other words, this technology: A machine that operates in real time a) does not need to transmit or store input images, and b) c) without spending time on training, and b) without large computing resources. This allows the edge to continually and optimally adjust its behavior based on user input. By learning using L-DNN, a machine that operates in real time can learn to recognize its environment and users. It adapts to changes in user interactions, addresses deficiencies in the original dataset, and provides customized It becomes possible to provide a user with an experience.
[0021] The disclosed technology also allows for merging knowledge from multiple edge devices. The company is developing a cloud collection, as well as labeling and disseminating knowledge from this collection. Including sharing among edge devices and eliminating the time of monotonous and boring centralized labeling . In other words, the intelligence from one or more of the edge devices can be merged (peer-to-peer) into one another, or pushed to part or all of the devices at the edge and merged into the intelligence shared therein. The L-DNN enables the merging / mixing / sharing / combining of knowledge to result in an increase in the memory footprint at a rate that does not change with the linear increase in the number of objects, occurring in real time, and as a result, it is guaranteed that only a small amount of information is exchanged between devices. These features make the L-DNN practical for real-world applications.
[0022] The L-DNN implements a heterogeneous neural network architecture characterized by the following two modules. 1) A slow learning module A that includes a neural network (e.g., a deep neural network) that is either pre-trained and fixed in the factory or configured to learn by backpropagation or other learning algorithms based on a sequence of data inputs. 2) A module B that provides an incremental classifier that can instantaneously change synaptic weights and representations with few training samples. Examples of instantiations of this incremental classifier include non-neural methods such as, for example, Adaptive Resonance Theory (ART) networks, or Restricted Boltzmann Machines (RBMs) that train neural networks with contrastive divergence, and Support Vector Machines (SVMs) or other supervised classification processes for fast learning.
[0023] Typical applications of L-DNN include, but are not limited to, using Internet of Things (IoT) devices that learn patterns of behavior and driving style Can it adapt its "user" style to the user's? Can it quickly learn new skills on the fly? Autonomous vehicles that can park in new driveways or new roads, and new infrastructure damage detection on the fly A drone that can learn a new class of damage and can distinguish this damage after a short period of learning during operation. It can learn (almost) instantly without pinging the cloud to identify the owner. Domestic robots, such as toys or companion robots, and never-before-seen It can recognize and react to new objects, avoid new obstacles, and discover new objects in the world map. A robot that can learn new parts and how to manipulate them on the fly. A learning-capable industrial robot, able to learn and network with new individuals or objects A security camera that can quickly detect it among the images provided by other cameras The above application is lifted and made possible by the technical innovations described herein. The training is done on a server using expensive and lengthy iterative training. Embedded computing in specific applications without the need to This can occur directly on the device.
[0024] The techniques disclosed herein can be used to analyze video streams, data from active sensors (e.g., Infrared (IR) imagery, lidar data, sonar data, and similar), acoustic data , and other time series data (e.g., sensor data, real-time data including data generated in factories) data streams, IoT device data, financial data, and similar), and Any multi-modal linear / non-linear combination of such data streams may be used. The present invention is applicable to several input formats, but is not limited to these.
[0025] Overview of L-DNN As disclosed above, L-DNN combines fast and slow training modes. We implement a heterogeneous neural network architecture to allow for fast learning mode. The paper proposes that real-time machines implementing L-DNN can respond to new knowledge almost instantly. In this mode, you learn new knowledge and new experiences so that you can answer questions quickly. The learning rate of the brain system is high enough to favor new knowledge and corresponding new experiences. On the other hand, the learning rate of the slow learning subsystem is set to preserve old knowledge and corresponding old experiences. otherwise, it is set to a low value or to zero.
[0026] Figure 1 shows a master edge / central server and several compute edges (e.g., drones). , robots, smartphones, or other IoT devices) to run the L-DNN We provide an overview of the L-DNN architecture, where multiple devices work in concert. The device receives sensory input 100 and transmits it to a slow learning module A102 and a fast learning module A103. Each module A supplies the L-DNN 106 with a learning module B 104. 102 acts as a feature extractor based on a pre-trained (fixed weights) DNN. Module A 102 receives input 100 and extracts relevant features into a compressed representation of an object. and supplies these representations to the corresponding module B104. These object representations can be learned quickly through user interaction. receives the correct labels for unknown objects and maps each feature vector to the corresponding label. They quickly learn the associations between the objects and, as a result, can instantly recognize these new objects. The L-DNN106 mixes newly acquired knowledge when learning different inputs. , merge, or combine it with other L-DNNs, as disclosed below. 06 can be connected peer-to-peer (dashed line) or to a central server (dotted line) for sharing .
[0027] The following example of an L-DNN implementation for object detection is based on the conventional object detection DNN "You o The next test will compare it to "You only look once" (YOLO). Produced results. The same small (600 images) custom dataset with one object. We trained and validated both networks using the Four training sets of different sizes (100, 200, 300, and 400 images) were generated from the remaining 400 images. In training, each image in the training set was presented once. creates batches by randomly shuffling the training set, Training was then repeated multiple times on these batches. After training, both networks We perform validation on the network and calculate the average of the following average precision (mAP: mean average precision): This resulted in a positive outcome. [Table 1]
[0028] Furthermore, the training of L-DNN using a training set of 400 images took 1.1 seconds, while the training time for YOLO was 21.5 hours. This is a surprisingly large performance improvement. The memory footprint of L-DNN is 320MB while, on the other hand, the footprint of YOLO was 500MB. These results show that L- DNN can achieve better accuracy than the conventional DNN YOLO, and can do so with a smaller dataset , a much faster training time, and fewer memory requirements.
[0029] Example of L-DNN Architecture Figure 2 shows an example of an L-DNN architecture used by a machine operating in a real-time image, such as a robot, drone, smartphone, or IoT device. L-D NN106 uses two subsystems, a slow learning module A102 and a fast learning module B104. In one implementation, module A includes a pre-trained DNN and module B is based on the Adaptive Resonance Theory (ART) paradigm for fast learning, where the DNN feeds the output of one of the latter's feature layers (usually the layer immediately before the layer that the DNN itself classifies as a fully connected layer, or the second layer from the end) to ART. Multiple DNN layers can provide input to more than one module B (e.g., in a multi-scale, voting, or hierarchical format ), and other configurations are possible.
[0030] An input source 100, such as a digital camera, detector array, or microphone, obtains information / data (e.g., video data, structured light data, audio data, combinations thereof, from the environment and / or similar things). When the input source 100 includes a camera system, it can obtain a video stream of the environment surrounding the machine operating in real time. The input data from the input source 100 is processed in real time by module A102, and module A102 provides a compressed feature signal as an input to module B104. In this example, the video stream can be processed as a series of image frames in real time by modules A and B. Modules A and module B are accompanied by appropriate volatile and non- volatile memories, as well as appropriate input / output interfaces, and can be implemented on a graphic process ing unit, a field programmable gate array, or an application specific integrated circuit, or any appropriate computer processor.
[0031] In one implementation, the input data is supplied to the pre-trained deep neural network (DNN) 200 of module A. The DNN 200 includes a stack 202 of convolutional layers 204 that are used to extract features that can be used to represent the input information / data, as shown in detail in the exemplary implementation section. The DNN 200 can be pre-trained at the factory before deployment to achieve the desired level of data representation. The DNN 200 can be completely defined by a configuration file that determines its architecture and a corresponding set of weights that represent the knowledge obtained during training.
[0032] The L-DNN system 106 utilizes the fact that the weights of the DNN are excellent feature extractors. Module B104 includes one or more high-speed learning neural network classifiers. ,To connect to DNN200 in module A102, the first DNN is used for classification. Either ignore only some of the upper layers of the DNN (e.g., layers 206 and 208 in Figure 2) or even removed from the system altogether. The convolution output of the process is accessed to serve as an input to module B 104. For example, the original DNN200 was mostly able to optimize weights during training using gradient descent techniques. In addition to the cost layer 208, a number of fully connected averaging These layers are used during DNN training or 200 to obtain predictions directly, but generates input for module B 104. (The shading in FIG. 2 indicates that layers 206 and 208 are unnecessary). Instead, the input to the neural network classifier of module B104 is a DNN The layers are taken from a subset of the 204 convolutional layers. Different layers or multiple layers can be used to refine the model. The input to module B104 can be provided by the NI 8110.
[0033] Each convolutional layer on the DNN200 uses local receptive fields to extract the local image from a small region in the previous layer. These filters are then fed into the convolutional layers of the DNN to The output from one or more later convolutional layers 204 of the feature extractor (pictorially The tensor 210 is a neural network classifier in module B 104. The input neurons are fed to the input neuron layer 212 of a classifier (e.g., an ART classifier). As described in detail in this section, L-DNN106 is used for global image classification or object detection. Depending on how the module A102 is designed, each later convolutional layer 204 of the module A102 and the module B104 with each fast learning neural network classifier, one-to-one or one-to-one There can be a one-to-many correspondence.
[0034] The tensor 210 transmitted from the DNN 200 to the module B system 104 is the original input It can be seen as a representation of an n-layer stack from the force data (e.g., the raw image from the sensor 100). In this example, each element of the stack represents the same spatial topography as the input image from the camera. Each grid element across n stacks is represented as a grid with is the actual input to the neural network.
[0035] The neural network classifier in the initial module B is trained on the fly after deployment. To facilitate learning, any initial knowledge or training in module A102 The input source 100 may be pre-trained with the classifications that have been created using the L-DN algorithm. When provided to N106, the neural network classifier uses the data from DNN200 ( For example, tensor 210) is processed continuously. Neural network of module B The classifier uses fast, preferably one-shot learning. The ART classifier is To implement pattern learning based on bottom-up (input-based) interactions between neuron-like elements, ) and top-down (feedback) associative projections, as well as competition between categories. Use a horizontal projection to do this.
[0036] In the fast learning mode, a new set of features is presented as input from module A102. When the feature vector is generated, the ART-based module B 104 outputs the feature vector to the F1 layer 212 as an input vector Then, calculate the distance between this input vector and the existing weight vector 214 and ,F2 layer 216 determines the activation of all category nodes. The distance is calculated by fuzzy AND (AR T's default version), the dot product, or the Euclidean distance between vector ends. The category nodes are then sorted from most activated to least activated. The nodes are then sorted into the most popular order, and a competition is run between the category nodes, with the nodes in this order considered as potential winners. If the label of a potential winner matches the label provided by the user, the corresponding In the simplest implementation, the weight vector is updated to match the new input and the existing Through a learning process that takes a weighted average with the weight vector of If there is no winner with the correct label, the new category node is created with a copy of the input. The weight vectors are then introduced into the categorical layer F2 at 216. Module B 104 now knows this input and can recognize it on the next presentation.
[0037] The result of module B104 depends on the task that L-DNN106 is solving. As the output of L-DNN106 in the field, or from a specific DNN layer from module A102 It works either as a combination of the outputs from the 3D model or as a full-scale model for object recognition in the whole scene. When classifying body images, the output of module B may be sufficient. The activity of module B 104 is superimposed on the bounding box determined by the activity of module A. A class label is provided so that each object is correctly identified by module A102. The object is found and correctly labeled by module B 104. For this purpose, the bounding box from module A102 is placed on the pixel-wise mask. The masks may be replaced and module B 104 may provide labels for these masks. Further details of module A 102 and module B 104 are provided below.
[0038] Real-time behavior and the concept of unknowns in neural networks. The normal L-DNN and specific module B operate in real time in response to continuous sensory input. Module B's neural network is designed to operate on what it knows. It should be implemented so that it is not confusing when no body is presented at all. The neural network targets datasets that typically contain labeled objects in the input. As a result, if a known object is not present, there is no need to act on the input. Therefore, to use such a network in module B of L-DNN, A further special category called "Nothing I know" was added to the network. Module B attempts to add to the queue and mistakenly classify unknown objects as known (false positives). The effectiveness of the measures should be improved.
[0039] This concept of "knowing nothing" encompasses previously unseen and unlabeled objects exclusively. This is useful for processing live sensory streams. And L-DNN can potentially mistakenly identify unknown objects as known ones instead. In contrast, people describe unfamiliar objects as "unfamiliar" or "not seen before." The concept of "knowing nothing" can be distinguished from "usually seen." Extending a conventional design with a bias node is equivalent to adding a bias node to the network. The notion of "knowing nothing" can also be expressed as the number of known object classes and their These can be implemented in a version that automatically increases or decreases the effect according to the corresponding activation.
[0040] One possible implementation of the concept of "knowing nothing" is the uniform distribution of knowledge into several categories. In contrast to the hypothesis that knowledge variance within a group is clearly concentrated in common categories, the hypothesis that knowledge variance within a group is clearly concentrated in common categories. In other words, the neural network of module B A classifier is a classifier that is used to classify objects into groups of known objects. However, if multiple different objects have similar activations, the object is recognized as belonging to the winning class. When this occurs (i.e., there is no clear winner), the system reports that the object is unknown. Because we explicitly use labels in the training process, our “no-knowledge” implementation allows us to It may only affect the mode and may not interfere with the learn mode.
[0041] An exemplary implementation of the “know nothing” concept using an ART network is presented in Figure 3. During the presentation of the input, the categorical layer F2 216 exhibits activation patterns across its nodes. For inputs that contain known objects, such as the top example in Figure 3, the winner is the winner. 300. For input that does not contain known objects, the bottom example in Figure 3 shows As shown in Fig. 1, the activity of the F2 layer is likely to be more uniformly distributed. Even in the second case, the threshold (Fig. 3) there may be nodes 302 with higher activity, so we calculate the average of all activations, Using the threshold as a threshold is insufficient to distinguish between these two cases. As in the example, the average number of clear winners 300 is chosen so that only the clear winners 300 remain above the threshold (dashed line in Figure 3). The threshold is increased by multiplying the average by a parameter of 1 or more.
[0042] The exact value of this parameter depends on several factors and is determined by the categories that the network has been trained on. It can be calculated automatically based on the number of categories and the total number of category nodes in the network. An example of the calculation is as follows:
number
[0043] By training independent module B using the concept of “knowing nothing”, the following results are obtained: Columbia Object Image Library10 50 objects from the 100 objects in the COIL-100 dataset were traced. The 50 novel objects were presented to the participants by independent module B as a training set. We used all 100 objects from the COIL-100 dataset to recognize them as “unfamiliar.” , was used as the test set. During training, we simulated real-time behavior. In this way, no shuffling is performed, and each one is input to the ART classifier of the independent module B. After training, the ART classifier achieved a recognition rate of 95.5% (object and “non-object”). For comparison, we used the COIL-100 dataset. If we feed the unshuffled dataset of all 100 objects to a conventional ART, The recognition rate was only 55%. This is due to the order dependency of ART, which is discussed below. Ugh.
[0044] If the input is not recognized by the ART classifier in module B, a correction is introduced to It is up to the user to label the correct input. If unrecognized input is not important , the user can ignore the input and the ART classifier will continue to identify it as "I don't know anything." If an object is important to the user, the user can label it, and the learning process can be accelerated. The features of objects and their corresponding labels are added to the knowledge by the network of learning module B. Module B will be added to the feature set associated with this new object. So you can activate a tracker system that will keep watching this object and add more perspectives. Cut.
[0045] Example implementation of module A In operation, module A extracts features and creates a compressed representation of the object. Neural networks are well suited to this task, as outlined below.
[0046] A convolutional neural network (CNN) is a DNN that uses convolutional units. The receptive field of the unit's filter (weight vector) is scaled by the height and width of the input. When applied to a visual input, the input to the early layers of the CNN is The length (h), width (w), and dimensions of one to three channels (c) (e.g., red, green, and an image with the components of the called blue pixels), while the input to the subsequent layers of the CNN is has dimensions of height (h), width (w), and the number of filters (c) from the previous layer. Since each filter is small, the number of parameters is significantly reduced compared to a fully connected layer and there are unique weights projecting from each of (h, w, c) to each unit on the next layer. For the convolutional layer, each unit has a number of weights equal to (f, f, c), where f is the spatial filter size (usually 3), which is much smaller than either h or w. By applying each filter at different spatial positions in the input, features that can be classified when the object is at one spatial position are independent of the object's spatial position, so the object can be classified at all spatial positions, providing the desirable property of translational invariance. After the convolutional layer, generally, subsampling (downsampling) layers follow. These reduce the height (h) and width (w) of the input by reducing a small spatial window of the input (e.g., 2×2) to a single value. For the reduction, averaging (average pooling) or taking the maximum value (max pooling) is used. The response of the subsampling layer is
[0047] invariant to small changes in the image, and this effect accumulates over multiple layers of a normal CNN. In inference, when applying several convolutional layers and subsampling layers to an image, the output exhibits excellent stability to various deformations of the input, such as translation, rotation, scaling, and warping. For example, a network trained on (handwritten digits written without lifting the pen) subsequent reduces the height (h) and width (w) of the input by reducing a small spatial window of the input (e.g., 2×2) to a single value. For the reduction, averaging (average pooling) or taking the maximum value (max pooling) is used. The response of the subsampling layer is invariant to small changes in the image, and this effect accumulates over multiple layers of a normal CNN. In inference, when applying several convolutional layers and subsampling layers to an image, the output exhibits excellent stability to various deformations of the input, such as translation, rotation, scaling, and warping. For example, a network trained on (handwritten digits written without lifting the pen) subsequent is invariant to small changes in the image, and this effect accumulates over multiple layers of a normal CNN. In inference, when applying several convolutional layers and subsampling layers to an image, the output exhibits excellent stability to various deformations of the input, such as translation, rotation, scaling, and warping. For example, a network trained on (handwritten digits written without lifting the pen) subsequent applying several convolutional layers and subsampling layers to an image, the output exhibits excellent stability to various deformations of the input, such as translation, rotation, scaling, and warping. For example, a network trained on (handwritten digits written without lifting the pen) subsequent is invariant to small changes in the image, and this effect accumulates over multiple layers of a normal CNN. In inference, when applying several convolutional layers and subsampling layers to an image, the output exhibits excellent stability to various deformations of the input, such as translation, rotation, scaling, and warping. For example, a network trained on (handwritten digits written without lifting the pen) subsequent exhibits wonderful stability to various deformations of the input, such as translation, rotation, scaling, and warping. For example, a network trained on (handwritten digits written without lifting the pen) subsequent handwritten digits is invariant to small changes in the image, and this effect accumulates over multiple layers of a normal CNN. In inference, when applying several convolutional layers and subsampling layers to an image, the output exhibits excellent stability to various deformations of the input, such as translation, rotation, scaling, and warping. For example, a network trained on (handwritten digits written without lifting the pen) subsequent er than the training set digit '3' and smallerIt responds in a similar way to the number "3" written with connected small circles.
[0048] These invariances allow the coding of the input to be more robust to visual perturbations. A feature space is provided, i.e., as the input changes (e.g., the object becomes smaller in the image frame), When the object is moved (slightly translated and rotated), the change in the output value is much smaller than the change in the input value. For example, two frames with objects that are translated by a few pixels It is difficult to learn from another method that has little or no similarity to the coding of the Possible.
[0049] Furthermore, recent advances in GPU-accelerated algorithms have been developed to learn filters from large datasets. Using gradient descent techniques, CNNs can train well-trained object classes. , we can achieve excellent generalized performance. Generalization means that the performance of the trained classes is For test images that are not identical to the training images, the network This means that it can produce output. Learning the key regularities that define the classes requires a large amount of If the network is trained on many classes, In this case, the lower layer, in which the filters are shared among all classes, can provide a good regularity set for all natural inputs. Thus, a DNN trained on one task can be used to train a When used as task initialization or to pre-write lower layers to a new high-level representation When used as a processor, it can provide excellent results. Natural images share common statistical properties. Share a common set. Visualize the inside of a well-trained neural network. Recent studies have shown that the features learned in the lower layers are largely class independent. On the other hand, the higher the layer, the more class dependent it becomes.
[0050] In L-DNN, module B generates high-quality compressed and generalized object features for classification. We exploit the capabilities of these CNNs in module A to obtain a representation of the features. In order to maximize the number of distinct objects as possible, the DNN used for L-DNN is Therefore, the object specificity of the high-level feature layer can be pre-trained on L- Does not interfere with the DNN's ability to learn quickly.
[0051] Example implementation of module B In operation, module B learns new objects quickly and without catastrophic forgetting.
[0052] Adaptive Resonance Theory (ART) An exemplary implementation of module B is the ART neural network. In ART, By utilizing competition between category nodes to determine the winner node for each object presentation, This avoids catastrophic forgetting. The winning node is the one associated with the correct label of the object. The learning algorithm updates the weights of that node if and only if the Each node is associated with only one object, and the learning algorithm weights only the winning node. Because the learning process updates the memory, any learning episode in ART affects only one object. Thus, when a new object is added to the system, there is no interference from previous knowledge. Rather, ART only creates new category nodes and updates the corresponding weights. .
[0053] Unfortunately, the ARTs described in the literature do not have a successful implementation as module B of L-DNN. There are several disadvantages that prevent the use of The lack of a concept of “not knowing anything” is not unique to ART and is discussed above. A list of ART-specific problems and solutions to these problems is disclosed below. do.
[0054] Classical fuzzy ARTs, because of the complement coding that is an integral part of their design, Does not handle sparse inputs well. When sparse inputs are complement coded, The complement part is high for most of the components because the complement of zero is 1, which is abundant in the input. Since all of these 1s are in the complement of the input, they are different during the distance calculation. The system becomes confused because it becomes very difficult to separate the inputs from each other. Powerful feature extractors such as NNs provide mostly sparse signals when extracting high-level features. The tendency is to retain the ART paradigm but to use classical fuzzy design and complement codes. Moving away from the dummy tagging is useful for using ART in module B of L-DNN. One solution is to remove the complement coding and use the fuzzy ART. The aim of this study is to replace the fuzzy AND distance metric used in the previous study with a metric based on the dot product. For the dot product based metric, the results remain normalized, and other methods to Fuzzy ART The advantage is that no modifications are required.
[0055] The ART family of neural networks is highly sensitive to the order of input presentation. In other words, ART lacks the consistency property, and different input orders can lead to the same results in the ART network. Unfortunately, in real-time operations like L-DNN, The rating system consumes training data as it receives it from the sensors. Therefore, the training data cannot be shuffled to provide consistency. The sensor frequently samples most or all of the first object during real-time operation. It then provides all subsequent samples of the object, so that the system learns the object representation one at a time. This allows the system to learn without competing with other objects and therefore without making mistakes in object representation. This leads to a situation where only some nodes present the first object. On the other hand, subsequent objects may be captured by the system in a representation that is already largely occupied by the representation of the first object. It can be exaggerated because it would push the expression into the hyperspace in which it is expressed. The "no idea" mechanism described introduces competition early on and The aggregation described above results in an exaggerated representation of the subsequent objects, Reduce or eliminate it.
[0056] Aggregation also reduces the memory footprint of object representations, making it ideal for memory-limited tasks. This is particularly useful for edge devices, as the system can detect all objects that it cannot classify otherwise. For every viewpoint, creating a new category node means that a new object is tracked as input. Note that when more memory is added, the memory footprint of the ART system increases at a constant rate. During the real-time operation described above and the subsequent presentation of the object, the system , then for each subsequent object, we create a superlinearly increasing number of nodes. ,The system experiences an exponential growth in the number of nodes with the number of objects.,Thus, Using conventional ART, the memory footprint of module B is linear with the number of objects. In the worst case, this growth can be exponential. The memory growth is suppressed to a rate that is linear with the number of objects, and each object that the L-DNN learns is It becomes possible to create a fixed-size, near-optimal representation of the body.
[0057] A complete L-DNN implementation example
[0058] L-DNN classifier Figure 4 shows the overall performance of the modified VGG-16 DNN as the core of module A. An example L-DNN implementation for image classification using softmax and fully connected The last two layers were removed from the original VGG-16 DNN, and the ART-based module B was added. , which is connected to the fully connected first layer of the VGG-16 DNN. A similar but much simpler L -DNNs can be created using Alexnet instead of VGG-16. This is a very simple and computationally inexpensive system that can be run on any modern smartphone. It works without the need for a GPU or any other dedicated processor, and works just like a smartphone camera. We can learn any set of objects from a few input frames provided by
[0059] L-DNN Grid-Based Detector One approach to detecting objects of interest in an image is to divide the image into grids and This implementation of L-DNN takes advantage of the following features of CNN: It is useful for.
[0060] In addition to the vertical hierarchical organization of the layers described above, each layer contains data that maintains a topographical organization. This can be done by specifying the network or kernel, stride, or pad size. Features that correspond to specific areas of interest in the image, regardless of how deep they are, This means that similar areas of a layer can be found on all layers at different resolutions. For example, When an object is in the upper left corner of an image, the corresponding feature is found in the upper left corner of each layer along the layer hierarchy. Thus, by attaching module B to each of the layer locations, Module B performs classification on a specific location in the image and detects if any known object is in this location. It allows one to determine whether a
[0061] Furthermore, the same feature vector represents the same object regardless of its location in the image, For each DNN layer (or scale) used as input, create only one module B. Therefore, by studying one object in the upper right corner, This allows module B to recognize the object anywhere in the image. By using multiple DNN layers of different sizes (scales) as input to module B, This allows detection at multiple scales, which is the basis for the next process. It is used to fine-tune the position of an object in an image without processing the entire image at a large scale. Can be used.
[0062] In this process, module A selects the coarsest scale (e.g., publicly available) for classification. Provide a (7×7) image of possible ExtractionNet to module B. If rule B indicates that it found the object in the second cell from the left and the fourth cell from the top, The corresponding finer DNN input (e.g., 14 × 14 for the same ExtractionNet) Only a portion of the image should be analyzed to further refine the object location.
[0063] Another application of multiscale detection is to design DNNs where the layer sizes are not multiples of each other. For example, if a DNN has 30x30 layers, it can be compressed to 2x2 (compression factor 15). , 3×3 (compression factor 10), and 5×5 (compression factor 6) layers. As shown in Fig. 1, by attaching module B to each of these compressed DNNs, the size of the object is Rough locations (shown as 502, 504, and 506) are obtained. However, these modules When the outputs of filter B are combined (shown as 508), the spatial resolution is The result is a non-uniform 8x8 grid with higher resolution towards the edges and lower resolution towards the edges.
[0064] To achieve this resolution, the system must multiply the calculations of module B by (2×2)+(3×3 )+(5×5)=38 times, whereas to calculate a uniform 8×8 grid, it takes 64 Note that we need to calculate Joule B. In addition to requiring fewer calculations, The resolution of the multi-scale grid in Fig. 5 for the central 36 locations is a uniform 8× 8 grid resolution or finer. Therefore, multi-scale Through detection, the system achieves better performance using only 60% of the computational resources of an equivalent uniform grid. This performance difference is due to the sum of the squares (uniform gradients). The sum of squares (representing the number of calculations for a non-uniform grid) is increases for larger layers because it grows faster than the
[0065] Heterogeneous (multi-scale) detection detects that an object in the center of the field of view is likely to be in the robot's path. are more likely to be detected accurately than surrounding objects that do not present any signs of a dangerous collision. This can be particularly beneficial for moving robots, as they can benefit from
[0066] L-DNN for Image Segmentation For images, object detection typically involves placing a bounding box around the object and detecting the associated class ( For example, the task of labeling a character (e.g., "dog") with a In addition to methods based on object detection, object detection techniques typically identify one or more regions of an image with a bounding box. Select a box and then classify the features within that box as a particular class while simultaneously filtering the boundaries. This is implemented by regressing the offset of the box location. The algorithm used to implement the method is Region-based CNN (R-CNN), Fa Includes st R-CNN and Faster R-CNN, but does not perform localization. Both methods rely directly on classification information, which may be substituted as a detection module. .
[0067] Image segmentation involves computing a set of pixels for all or a subset of the pixels in an image. Segmentation is the task of determining the class label of two distinct images of the same class. Semantic segmentation, which disambiguates individual pixels from objects, ,Individual pixels from two distinct objects of the same class can be uniquely,identified or instantiated. Image segmentation can be divided into image segmentation and instance segmentation. The method is usually based on object detection methods (Fast R-CNN, Fast R-CNN, or Faster R - Take the bounding box output of a neural network (CNN, etc.) and segment the most prominent object in that box. The bounding box is then associated with a class label. After that, it is associated with the segmented object. If there is no segmentation result, the segmentation result is discarded. The object may or may not have instance information. One algorithm that implements this method is Mask R-CNN.
[0068] Based on the R-CNN family of networks, for image detection or segmentation The L-DNN design is presented in Figure 6. We use a static classification module, such as Mask R-CNN. Consider the image segmentation process you want to use. In this scenario, we use static classification The module 600 may be replaced with the module B 104 of the L-DNN. The network segmentation pathway remains unchanged and the area is as usual. As with the static classification module, L- Module B104 of the DNN does not return any positive class predictions that pass the threshold (e.g. For example, areas where the network is not trained or segmented are When a person recognizes that he or she "knows nothing," as described in Similarly, the L-DNN module B104 selects the acceptable classes. When returning predictions, the segmentation results are preserved, just like in the static classification module. Unlike the static classification module 600, the L-DNN module B 104 Provide continuous adaptation so that the state changes from the former to the latter through user feedback. .
[0069] User feedback is based on the user's selection of an object on their social media profile. directly through bounding boxes and class labels, or through user-defined tags, such as The user selects an object in the video, and then the object is tracked through the video and a new object is Through indirect feedback, such as when providing continuous feedback on body classes. This feedback may be used to classify new class networks over time. This process is used to train the L-DNN on how to It does not affect the segmentation components of the network.
[0070] There is also some flexibility in the placement of module B 104 in this paradigm. The input to module B104 is the class labels combined with the segmentation output. The modules may be processed to produce a segmented and labeled output 602. The output of the convolutional layer 202 of module A should be directly linked to the output of the convolutional layer 202 of module A. This constraint is This can be satisfied by having modules A and B take the output of the region proposal stage. A should not depend on any dynamic parts of module B. B adapts to the weights of the network, but since module A is static, module If B should change the weights and then pass its output to module A, module A should Most static neural networks cannot cope with sudden changes in the network's input representation. will likely see a degradation in performance because
[0071] Brain aggregation and brain blending Multiple real-time machines implementing L-DNN generate new In some cases, the following information can be learned on the fly: As outlined in several use cases, knowledge sharing between machines operating in real time is It can be advantageous for machines operating in real time to learn new knowledge at the edge. So each machine operates in real time to share new knowledge from the edge to the center. Compressed and generalized transfer of new information to a server or other real-time machine. representation (represented in the network in terms of the synaptic weight matrix of module B) By carrying out the following steps, each real-time operating machine transmits Extract the knowledge obtained by the,process, either at a central server or directly on the edge device, In addition, it can aggregate and communicate with other real-time systems through centralized or peer-to-peer communication. It can be shared with other machines.
[0072] Learn new information while deployed in the field – as discussed above, it operates in real time. The machine can learn new information on the fly through L-DNN. Real-time machines can be expected to encounter new objects and / or new knowledge. and a machine operating in real time can create new objects and / or new knowledge on the fly. To allow the system to learn from the object, the system can provide labels for new objects and trigger a fast learning mode. In this way, a machine operating in real time can modify its behavior and rapidly adapt to new objects and and / or be able to adapt to new knowledge.
[0073] New knowledge aggregation – After learning on the fly about one or more objects, the system The fast learning module B runs the aggregation process. This process creates a representation of the new object. The network compresses the representation of the object and integrates it with previously known object representations to improve the network's generalization ability. This reduces the memory footprint of module B. The practical implementation is detailed below.
[0074] Transferring the collective intelligence to other devices – at any point during operation, or After completing the mission, the machine operates in real time and its fast learning module The aggregated weight matrix of each node (node B) is transmitted to a central server (e.g., In some cases, the data can be transmitted to a cloud-based server (e.g., a cloud-based server). The weight matrix of the machine's high-speed learning module can be downloaded to an external storage device and Can be physically linked to a central server when a central server is not available or desirable Communication can occur in a peer-to-peer fashion between machines (edge devices) operating in real time. .
[0075] Brain blending (or merging, combining, combining) - some real-time After the weight matrices from the machines are collected by a central server or one of the edge devices, The central server or edge device receives newly acquired data from each real-time operating machine. Run a hybrid utility that combines the knowledge collected, compresses it into a single weight matrix, and aggregates it. The mixed utility allows you to mix results while maintaining system-wide accuracy. This reduces the memory footprint of the resulting matrix and removes redundancies. An exemplary implementation based on the metric is detailed below.
[0076] Individual brain updates after blending - the resulting weight rows produced during the blending The sequence is then transmitted through a wired or wireless communication channel or to a physical external storage / memory. Download to the device and physically transfer the storage / memory device to a machine that operates in real time and then downloaded to one or more real-time machines via electronic delivery. do.
[0077] In this way, knowledge from multiple machines operating in real time can be aggregated, The new knowledge learned by each of these machines is shared with other machines operating in real time. Yes, you can.
[0078] An Example Implementation of the Brain Aggregation and Blending Process Using ART FIG. 7A illustrates an exemplary brain aggregation and blending process using ART. Extend the default ART process as follows: Each node represents an object, and each object has one or more category nodes that represent it. The left side of FIG. 7A shows the category in layer F2 at 216 that is activated by layer F1 at 212. 7 shows weight patterns 702 for the nodes. Each weight pattern is provided to 212 of layer F1. 2 represents a generalized input pattern learned from multiple real feature inputs 210. Learning in ART is based on the fact that category nodes and corresponding objects win the competition and solve the problem. Note that this only occurs if you correctly identify the object.
[0079] In the center of FIG. 7A, multiple inputs 210 for different objects are presented to the ART network. 7 shows the weight patterns 702 of the different category nodes in layer F2 216 after the weights have been calculated. Each weight pattern 702 is set so that the corresponding node is assumed to be the same as the corresponding object label. The weight pattern 702 in the center of Figure 7A represents a generalized version of the input learned in this way. , resulting in aggregate input 704 shown on the right in FIG. 7A.
[0080] The original input 210 is generally not available to the system at the time of aggregation or mixing. The collection of weight patterns 702 represents the inputs that the ART network touched during training. 210. Thus, the weight pattern 702 is a generalization of all the important The features also represent important features that are better than the original input 210 during the training process. , which can act as a substitute for real-world input.
[0081] The aggregation uses weight patterns as a proxy for the real input. During aggregation, the next step A pop occurs.
[0082] A weight vector 702 of an existing weight matrix (e.g., the matrix of weight vector 214 in FIG. 2) ,Aggregated input set (right side of Fig. 7A) a i =w i where a is the input vector, w is a weight vector, and i ranges from 1 to the number of existing category nodes in the network. If the ART network uses complement coding, half of the weight vectors Decomplement the minus sign and average it over the first half of the vector. (a i =(w i +(1-w ic )) / 2). Each vector in the aggregate input set is Receive the corresponding labels extracted from the category nodes.
[0083] All existing F2 nodes and corresponding weights are removed from the ART network. Therefore, the ART network is in a blank initial state.
[0084] The aggregate input set is randomly shuffled and the ART network learns the original inputs. By randomly shuffling the set, we train it in the same way as we did in the previous example. The effect of order dependency in the ART network is reduced, making the ART network more Compact (fewer category nodes are created) and more optimal (better generalization) ) representations can be constructed.
[0085] Using weights on an aggregate input set involves converting a single vector into many original input vectors. has the added benefit of replacing , thus reducing the complexity of the aggregation process. and the computation time is faster than the original learning process.
[0086] The aggregation process can occur at any time during the operation of the L-DNN-based system. The process reduces the memory footprint of the ART-based implementation of module B, and Reduces the order dependency of RT-based systems. During such operations, the system There is no way to change the order of sensory inputs entering the system as the The reduction in dependency will be beneficial for any real-time machine based on L-DNN. Aggregation is disabled by user action or when the memory footprint becomes too large. automatically upon reaching or exceeding a threshold size, or It can be triggered periodically based on time.
[0087] The example aggregation works in real time, as if one object after another was seen, and the ART The first trial was conducted on the COIL dataset, which was presented to the network intentionally. Training took 4.5 times longer than aggregate training. Aggregation reduces memory The footprint has been reduced by 25%, and object recognition performance has improved from 50% accuracy to 75% accuracy. The training dataset was first shuffled to reduce ordering artifacts. In the cases where the number of samples was reduced, aggregation still showed an improvement in performance. There was no significant reduction in memory footprint since the system was already well compressed, but The average accuracy rate for object recognition increased from 87% to 98%. This represents an unexpectedly large performance improvement.
[0088] Fusion is a method to assemble an aggregate training set from the weight matrices of two or more ART networks. Fusion is an extension of aggregation, which is combined with ART network. As a result, multiple fused ART networks can be created. ,When ,having ,knowledge ,of ,the ,same ,representation ,of ,this ,object ,across ,multiple ,ART ,networks, , All are naturally combined together by the ART learning process, while all the distinctive This is due to the intelligent compression of object representations and the memory of the fused system. This will lead to further reduction in the footprint.
[0089] For example, one ART instance trains 50 objects from the COIL dataset. Then, we trained 33 objects (17 objects were the same in both sets) on another ART instance. When the first instance is correct, the result is 92.9% correct, and when the second instance is correct, the result is 92.9% correct. By fusing them together, both ART instances It is trained on all 66 unique objects, producing a network that is 97% accurate. In addition, the fusion version will have a brute force combination of the two networks It has 83% of the memory footprint of the fused version. The toplining is a combination of the first network with only new objects in the second network. (excluding the 17 overlapping objects). Therefore, the mixture actually Intelligently compress and refine object representations to improve accuracy. Randomly shuffle inputs. Without it, the mixed results were even more outstanding in terms of accuracy, 85.3% and 77.6%. Mixing the correct network reduces the memory footprint of the combined network. The result of these mixed experiments is a 96.6% correct network with 84.6% of the inputs. The results show an unexpectedly large performance improvement.
[0090] Using Context Information to Improve Performance L-DNN-based systems further combine context information with current object information. By using context L-DNN, it is possible to improve the performance accuracy. It can learn things that are likely to co-occur in a stream, e.g. camels, palm trees, and sand dunes. , and off-road vehicles are typical objects in desert scenes (see Fig. 7B), while houses Sports cars, oak trees, and dogs are context-typical objects in suburban scenes. The locally ambiguous pixel-level information obtained as drone input is Depending on the text, it can be mapped to two object classes (e.g. camel or dog). In both cases, the focus of the object has some ambiguity, which is typical of low-resolution images. So, the pixelated image of a camel is the probability that “camel” can be inferred by L-DNN. The fourth highest class, and the most likely, based on local pixel information alone, is "horse". However, it is possible to use global information about the scene and past associations learned between objects. The objects in the context (sand dunes, off-road vehicles, palm trees) are Therefore, the context classifier recognizes "horse" as a " class and select the "Camel" class. In an urban scene that includes a car, an oak tree, and a dog, the same set of pixels can be mapped to
[0091] As a side note, when an object is identified as ambiguous or abnormal, like the camel in the example above, the L- The DNN system may prompt the human analyst / user to look more closely at the object. This anomaly detection and alerting subsystem uses context to distinguish between normal and By resolving the ambiguity of the identification of objects of interest that do not belong to the scene, This can be achieved through
[0092] Infinite regress problem, i.e., before the context module can produce object classes, The need for object classification requires that the label with the highest probability is chosen as the input to the context classifier. Thus, at each fixation of the object, the context classifier It is possible to iteratively refine guesses about object labels.
[0093] L-DNN can leverage huge amounts of unlabeled data The sheer volume of unstructured content means that even without labels, valuable training data is not possible. The data is provided to module A of L-DNN. Greedy layer-wise pre-training (gre layer-wise pre-training) By training each layer in turn, the DNN can achieve bottom-up unsupervised learning. The layer-by-layer training mechanism is a contrastive learning mechanism. Convergence, i.e., denoising autoencoder and convolutional autoencoder An autoencoder takes an input and encodes it using weights and transfer functions. After training a layer, its output is The pre-trained network is the input of the next layer. The network also benefits from the learning of edges on layer 1, corners on layer 2, and useful features such as edge groups and other edge groups, as well as features specific to higher-level data in later layers. Moreover, the convolutional variants often capture the hierarchical feature relationships of convolutional networks. It enjoys the inherent translational invariance of the
[0094] This process tends to precede later supervised learning ("fine-tuning"), so In many cases, the performance of a pre-trained network is Outperforms networks without pre-training, when there is a large amount of labeled data ,Because the labels put some burden on the analyst, the ,network that is pre-trained is , However, it cannot compete with networks without pre-training. The "real" network achieves the recognition performance of the L-DNN system while keeping the labeling burden low. will improve performance over other pre-trained networks. In other words in addition to limited labels resulting from analyst reports, training DNNs on unlabeled data will lead to improved performance over DNNs trained from only a relatively small number of analyst reports as well as another large labeled dataset
[0095] Finally, ART as an implementation of Module B has another advantage in that it can also perform unsupervised learning ART does not require labels for learning, but when available it can be considered "semi-supervised", meaning that it can utilize labels. ART stores search information for the most matching observation frames and image regions for each node and, while operating in unsupervised learning mode, assists in organizing unlabeled data By doing so, analysts may be able to access and examine many similar observations for each ART node
[0096] Exemplary Use Cases of L-DNN The following use cases are non-limiting examples of how L-DNN can address technical problems in various fields
[0097] Automating Surveys Using L-DNN: Single or Multiple Image Resources For example, consider a drone service provider who wants to automate the inspection process of industrial infrastructure such as power transmission lines, base station towers, or wind turbines With existing solutions, an inspector has to watch hours of drone video to find frames containing the main components that need to be inspected In each frame, the inspector has to look for these main components must be identified manually.
[0098] In contrast, L-DNN-based assistants can be deployed in identification tools. Data containing object or anomaly labels are pre-trained during traditional slow DNN factory training. As a pre-trained set, it can be provided to the L-DNN-based assistant. The user can add to this set during the quick learn mode, as described below.
[0099] Figure 8 shows the results of a drone attack on a "smart" drone or a "dumb" drone. L-, which may be included on a computer used to review the captured video. The DNN-based assistant in action. The drone 800 is connected to a communication tower 820, a solar panel 830, and a A structure such as a panel 830, a wind turbine farm 840, or a power distribution line 850 (which (These are merely exemplary configurations, other configurations may be envisioned.) Drone operators 810 may use manual control of the drone or may control the drone to function automatically. A human analyst 805, such as an analyst in a control room, may monitor the drone 800. The L-DNN system has 100 sensory inputs (e.g., video, lidar, etc.) from Module B104 of module 106 to provide labels while the drone is flying, or Flights can be posted.
[0100] First, drone 800 receives a copy of L-DNN106 as a separate local classifier. The drone 800 is used to detect these power lines 850, base station towers 820, and wind turbines. While investigating the 840th,video frame 100 is acquired, and the L-DNN106 Module A102 is configured to generate a video frame 10 based on the pre-trained data. Module B104 then extracts image features from the image based on these features. It provides a likely label for each object. This information is passed to the user 805. 05: If the label is not satisfactory, the fast learning mode is activated to learn the network of module B. This allows the network to be updated with the correct labels. Therefore, the fast learning subsystem may correct the label of the first frame after the update. It can learn to spot objects it has already learned to spot, such as power lines, cell towers, and wind turbines, as quickly as a real-time image. One-trial learning can be used to determine the location and characteristics of the When analyzing the system 10, it means immediately after the user introduces the correction. 6 becomes more knowledgeable over time and provides better discrimination over time with the user's help. do.
[0101] FIG. 9 illustrates the L-DNN techniques described herein in a multi-drone (e.g., data processing The drones (non-capable 800 and 900, as well as the smart drone 910) can be synchronized How can this be applied as a generalized example of Figure 8, where data is collected periodically or asynchronously? The information learned by the L-DNN associated with each drone is merged. They are then pushed back (combined or mixed) to other drones or drawn. The data is shared peer-to-peer between the users or in collaboration with a central server that contains the module B104. The central server merges the information learned by each L-DNN and distributes it to the drones. The merged information is pushed back to all drones, including drone 910, and drone 910 The communication tower 820, the solar panel, and the like are obtained from the data acquired by the satellites 800 and 900. For more information about wind turbines 830, wind turbine farms 840, or power distribution 850, But thanks to the merge process, these items can be understood and categorized here.
[0102] Automating Warehouse Operations Using L-DNN: Aggregating and Blending Knowledge from Multiple Sources The system described above is a system for recording multiple machines or cameras (fixed or can be expanded to multiple different geographic locations (e.g., mounted on drones, etc.) Consider a company that has a large warehouse. If you were to inventory a large warehouse manually, you would end up with many This can take up a lot of man-hours and often requires the warehouse to be closed during that time. The solution makes it difficult to identify stacked objects that may be hidden. In addition, with existing automation solutions, information learned in one geographic location is not shared across other In some cases, vast amounts of data are collected in different geographic locations. Therefore, these automated solutions have the ability to learn and act on new data. It can take weeks to do so.
[0103] In contrast, the L-DNN technique described herein relies on sensors (e.g., fixed cameras 101 0a-1010c, or a moving camera mounted on a robot or drone) New items in the inventory are automatically added on-the-fly via various L-DNN modules that connect to the sensors. The environment in which the system can be trained is a warehouse, industrial facility, or distribution center, as shown in Figure 10. In addition, operators at 805 and 1005 can New information can be taught to various L-DNN modules. This new knowledge is then centrally integrated. The data is then sent to each individual device (e.g. , camera 1010).
[0104] For example, the fixed cameras 1010a to 1010c (collectively, cameras 1010) in FIG. Each of these cameras 1010 captures a corresponding view of an object on the conveyor belt. A video image 100 is acquired, and the image 100 is input to the corresponding L-DNNs 106a to 106c (or The L-DNN 106 then performs tasks such as investigation, classification, or to recognize known objects in the image 100 for other distribution center functions.
[0105] Each L-DNN 106 is then subjected to evaluation by human operators 805 and 1005. or tagging the unknown object as “nothing known.” For example, unknown object 104 When presented with 0, the L-DNN 106a outperforms the classification by the human operator 805. In order to detect the unknown object 1040, the L-DNN 106c flags it. Flag the unknown object 1060 for classification by the operator 805. L-DN N106b is when an unknown object 1050 is presented, and the subject simply sees the unknown object 1050 as “nothing.” The independent module B104d connected to the L-DNN106 is , by L-DNN 106a and 106c from human operators 805 and 1005 The acquired knowledge is merged and the module B 104b calculates the current To recognize the later instances, the module B 104b of the L-DNN 106b is informed Push awareness.
[0106] The L-DNN106 for each device can be equipped with features such as post markings, exit signs, or combinations thereof. The system is pre-programmed to recognize existing landmarks in the warehouse, such as This allows the system to be trained to equip sensors or to use sensors (e.g. For example, the position of the unmanned vehicle appearing in the images acquired by the camera 1010 is triangulated. The L-DNN in each vehicle operates exactly the same as the use cases described above. In this way, knowledge from multiple unmanned vehicles is aggregated, blended, and fed back to each unmanned vehicle. The aggregation and mixing of knowledge from all places can be achieved by the aggregation and mixing center above. This can be done by a central server, as described in the previous section, and also on a peer-to-peer basis. A mix of stocks can also be applied. Thus, stocks can be distributed across multiple warehouses with minimal disruption to warehouse operations. You can take inventory and consolidate knowledge.
[0107] Using L-DNN on a fleet of mobile devices Consumers’ mobile devices, such as their smartphones and tablets, or mobile Mobile cameras, body-worn cameras, and public safety first responders and public safety officials A distributed network of specialized devices, such as handheld LTE devices, used by Think about your device. Your device is using sensors to understand your surroundings, for example when taking a photo. In these cases, the L-DNN technique described herein can be used to Applicable to smartphone or tablet devices 1110, 1120 and 1130 Individuals (e.g., users 1105 and 1106) can use devices 1110 and 113 Each L-DNN module 106 is taught knowledge and this information is shared between peers. The server 1190 may include a module B 104. ,The merged knowledge may be applied to devices that did not ,participate in the original training 1020. Push back to some or all connected devices, including
[0108] The L-DNN module can, for example, apply image processing techniques to photos taken by the user. Each L-DNN can learn a set of customizable features associated with the image. The effects of applying filters or image distortions to these object classes or areas The combined learned actions can be taught across devices. They may be shared, merged, or combined, either peer-to-peer or collectively. -DNN technology allows input variables to be sensory or non-sensory (all of the The input variables and These usage patterns can be any combination of the number of users and output variables. The L-DNN modules 104 are trained at the individual levels and pushed to the central L-DNN module 104, where they are merged and Push back to individual devices.
[0109] In another example, police officers could use specialized devices powered by L-DNN to identify lost, suspected, or suspicious objects. In such a situation, the police and / or the first Responders cannot afford to waste time. Provided to police officers and / or first responders Existing solutions require manual analysis and organization of video feeds from cameras. Such solutions use a central server to analyze and identify objects. This takes too much time. There is a big latency problem because the analysis needs to be done on the cloud / central server. This poses a significant obstacle for first responders / officers who often need to act immediately upon receiving the data. In addition, continuous transmission of video data to a central server can be used to may place a burden on
[0110] Instead, we use L-DNNs in mobile phones, body-worn cameras, and handheld LTE devices. By using the edge, data can be learned and analyzed at the edge itself. can learn to customize their devices so officers / first responders can easily find the location of a person / object Find and provide, as well as identify, people / things of interest that officers may not be actively looking at. It can also search and identify bodies in various locations. L-DNN learns from operators on a remote server. Instead, it utilizes a fast learning mode to learn from officers on the device in the field. It reduces or eliminates latency issues associated with centralized learning.
[0111] FIG. 1 illustrates how a consumer can point a phone at a scene and capture the entire scene or a portion of the scene (an object, e.g., When labeling a scene (such as sky, water, or other parts of a scene), we label the components in the image. In addition, police officers can use video cameras to track the movements of the L-DNN on their mobile phones. The mobile phone can access 10 video frames to identify suspicious people / objects. When it gets 0, module A102 will, based on the pre-trained data, Features of the image can be extracted from these frames. Module B104 can then use these features to provide likely labels for each object. For example, if person A lives in neighborhood B and has been observed in neighborhood B in the past, person A may be labeled as a "resident" of neighborhood B. Thus, the fast learning subsystem can instantly determine the relative positions and features of already learned objects, such as houses and trees, using one-shot learning as soon as after the first frame of learning. More importantly, the dispatcher on the central server 110 can introduce new objects to be found by the server-side L-DNN, and the new objects will be
[0112] mixed as needed and distributed to the local first responders. This use case is very similar to the previous one, but it makes more use of the ability of the L-DNN to quickly learn new objects without forgetting old ones. Surveys and inventory collections generally have little time constraint and can aggregate memory in low-speed learning mode, while for first responders, it is important to quickly aggregate and mix knowledge from multiple devices so that all devices in the area can start looking for suspects or missing children. Thus, the ability of the L-DNN to quickly learn new objects introduced by a single first
[0113] responder, aggregate them almost instantaneously on the server, and distribute them to all first responders in the area is an extremely great advantage for this Reducing the computation time of the DNN process on a large data center server 1200 L-DNN technology can be used as a tool to accelerate learning in DNNs by orders of magnitude. Using this feature, you can dramatically reduce the need for computing resources on the server, or Reduces the consumption of computing resources and information is often acquired without the need for hours / days / weeks of training time. It can be trained on the required large dataset of 100 in just a few seconds. The use of L-DNN also reduces power consumption and frees up server resources in data centers. It also leads to better utilization overall.
[0114] conclusion As described above, L-DNN allows for on-the-fly (one-shot) learning of new Conversely, traditional DNNs often It takes thousands or even millions of iteration cycles to learn a The larger the step size per loop, the more the gradient of the loss function will change, which will lead to actual performance improvement. Therefore, these conventional DNNs require a large number of samples per training sample. , introducing small changes to the weights, which allows us to add new knowledge on the fly. In contrast, L-DNN, which involves fast learning neural networks, can learn stable object representations with few training examples. ,Just one training example can be sufficient for L-DNN.
[0115] L-DNN uses fast training neural networks in addition to traditional DNNs. Because it uses a new algorithm, it is resistant to the "catastrophic forgetting" that plagues conventional DNNs. When new input is provided to the DNN, all weights of the DNN are adjusted for each sample presented. When learning new inputs, we make the DNN "forget" how to classify old inputs. can be avoided by simply retraining the complete set of inputs, including the new input. However, retraining would take too long to be practical. Some existing approaches Selectively limit weights based on their importance or train sub-networks of DNNs or use a modular approach to avoid catastrophic forgetting? However, such an approach is not only slow, but also It also requires multiple iterative cycles to train the NN. In contrast, L-DNN This provides a means to achieve fast and stable learning capabilities without retraining. L-DNN also allows for the generation of object representations with a single example and / or in a single iteration cycle. It also promotes ongoing, stable learning.
[0116] Although various embodiments of the invention have been described and illustrated herein, those skilled in the art will appreciate that the present disclosure to perform the functions and / or obtain the results and one or more advantages described in the document, Various other means and / or structures are readily envisioned, and each such variation and / or modification is incorporated herein by reference. are considered within the scope of the inventive embodiments described herein. More generally It will be understood by those skilled in the art that all parameters, dimensions, materials, and configurations described herein are illustrative. The actual parameters, dimensions, materials, and / or configurations may vary depending on the teachings of the present invention. As those of ordinary skill in the art will readily appreciate, the nature of the invention will depend on the particular application or applications to be implemented. Those skilled in the art will recognize and appreciate that there are many equivalents to the specific inventive embodiments described herein, or that are merely familiar to those skilled in the art. Any changes or modifications can be confirmed using routine experimentation. Accordingly, the above-described embodiments are intended to be illustrative only. As set forth herein, within the scope of the appended claims and equivalents thereto, embodiments of the invention are It is to be understood that the invention may be practiced otherwise than as specifically described and claimed. Embodiments of the present invention relate to each individual feature, system, article, material, kit, and / or methods. In addition, any two or more of such features, systems, articles, Any combination of such features, systems, materials, kits, and / or methods may be used. Unless mutually inconsistent, the articles, materials, kits, and / or methods of the present disclosure Included in the range.
[0117] The above described embodiments can be implemented in any of numerous ways. For example, the embodiments may include: It may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code is provided to a single computer. The present invention may be implemented in any suitable processor or processors, whether integrated or distributed across multiple computers. It may be executed on a collection of processors.
[0118] Furthermore, the computer may be a rack mounted computer, a desktop computer, It can be used in a variety of forms, including laptop computers and tablet computers. It should be understood that a computer may be embodied in any of the following ways: A personal digital assistant (PDA) rather than a device considered a computer. , a smartphone, or any other suitable portable or fixed electronic device; It may be embedded within a device with suitable processing capabilities.
[0119] A computer may also have one or more input and output devices. The data can be used, among other things, to present a user interface. Examples of output devices that can be used to provide an interface include a printer or a display screen for visual presentation, and a speaker or other device for audible presentation of the output Examples of input devices that can be used in a user interface include voice generation devices. This includes keyboards, mice, touchpads, and digitizer tablets. As another example, a computer can be used to recognize The input information may be received in an audible or other audible format.
[0120] Such computers can be part of a local area network or an enterprise network. Wide area networks such as the ICT network, and Intelligent Networks (IN) are interconnected by one or more networks of any suitable form, including the Internet. Such a network may be based on any suitable technology and may include any It may operate according to any suitable protocol and may be used over a wireless network, a wired network, or may include a fiber optic network.
[0121] The various methods or processes outlined herein may be implemented on a variety of operating systems or or platform, and can run on one or more processors. In addition, such software may be coded as software capable of executing a variety of A number of appropriate programming languages and / or programming or scripting tools It may be written using either the The program may be compiled into executable machine code or intermediate code that is executed.
[0122] Also, various inventive concepts may be embodied in one or more methods, examples of which are provided herein. The acts performed as part of the method may be ordered in any suitable manner. Thus, embodiments may be constructed in which the acts occur in an order different from that illustrated. This means that some acts may be performed simultaneously even though the exemplary embodiments show sequential acts. This may include:
[0123] All publications, patent applications, patents, and other references mentioned herein are hereby incorporated by reference. The whole of them is incorporated.
[0124] All definitions and definitions used herein are taken from dictionary definitions, texts incorporated by reference. The definitions in this document and / or the ordinary meaning of the defined terms should be understood as governing the It is.
[0125] As used in this specification and claims, the indefinite articles "a" and "an" are used interchangeably. Unless expressly indicated otherwise, "at least one" should be understood to mean "at least one." do.
[0126] As used in this specification and claims, the term "and / or" means means "either or both" of the elements connected together, i.e., "and" should be understood to mean elements that are present in some cases and disjunctively present in others. The elements listed with "or" are of the same form, i.e., among the coordinated elements. The other elements are to be interpreted as relating to the element specifically identified. Specifically identified by the "and / or" clause, whether relevant or unrelated. In addition to the elements listed above, the following may be optionally present: References to "or B" when used in conjunction with open-ended phrases such as "have" are In one embodiment, only A (optionally including elements other than B), in another embodiment, In yet another embodiment, both A and B ( Optionally including other elements), etc.
[0127] As used in this specification and claims, "or" means any combination of "or" as defined above. and / or." For example, items in a list When separating terms, "or" or "and / or" is used inclusively, i.e. or list of elements, and optionally additional items not listed. "one" or "two or more" shall be construed as including one, but also including two or more unless expressly stated to the contrary. Only the indicated term, e.g., "only one of" or "exactly one of", or "Consisting of," when used in the claims, means a number or list of elements. In general, as used herein, "or" refers to the inclusion of exactly one element of a The term "any," "one of," "only one of," or "the When preceded by a term of exclusivity, such as "exactly one of," it indicates exclusive alternatives (i.e. "The term "essentially" refers to the use of "either / or" in a manner that is consistent with the "constitutional" principle. "Consisting of," when used in the claims, shall have its ordinary meaning as used in the field of patent law. It shall have the following.
[0128] As used herein and in the claims, "list" refers to a list of one or more elements. The phrase "at least one" refers to a selection of one or more of the elements in a list of elements. means at least one element, not just any element specifically listed in the list of elements. Any element in a list of elements, but not necessarily including at least one of every element It should be understood that this definition does not exclude combinations of elements. The phrase "at least one" refers to any element other than those specifically identified in the list of elements. Anything that may be present, whether or not related to the specifically identified elements Thus, as a non-limiting example, "at least one of A and B" (or, equivalently, "at least one of A or B" or, equivalently, "A and / or at least one of B) in one embodiment, B is absent. , including at least one, and optionally two or more, A (optionally including elements other than B) In another embodiment, A is absent and has at least one, and optionally two. In another embodiment, B (optionally including elements other than A) includes at least At least one, optionally including two or more, A, and at least one, optionally including two or more, can refer to two or more B (optionally including other elements), etc.
[0129] In the claims, as well as in the above specification, all transitional phrases, such as "comprising," "including," "carrying," "having," "containing," "involving," "holding," "composed of," and the like, are to be understood as open-ended, i.e., meaning including but not limited to. Only the transitional phrases "consisting of" and "consisting essentially of" shall be closed or semi-closed transitional phrases, respectively, as set forth in the United States Patent Office Manual of Patent Examining Procedures, Section 2111.03. The claims of the parent application as originally filed were as follows: Claim 1: 1. A method for analyzing objects in an environment, comprising: collecting, by a sensor, a data stream representative of the objects in the environment; extracting, by a neural network running on a processor operatively coupled to the sensor, a convolution output from the data stream, the convolution output representing a feature of the object; and and classifying the object based on the convolution output by a classifier operatively coupled to the neural network. Claim 2: The method of claim 1 , wherein the sensor is an image sensor and the data stream includes an image. Claim 3: Extracting the set of features comprises: generating a plurality of segmented sub-areas of the first image; and coding each of the plurality of segmented subareas by the neural network. Claim 4: Extracting the feature set comprises: allowing a user to select portions of said data stream that are of interest; enabling the user to divide the portion of interest into a plurality of segments; and encoding each of the plurality of segments by the neural network. Claim 5: 2. The method of claim 1, wherein the sensor is at least one of a lidar, radar, or acoustic sensor, and the data stream includes a corresponding one of lidar data, radar data, or acoustic data. Claim 6: a sensor that collects a data stream of an environment, the data stream representing objects in the environment; 1. An apparatus comprising: at least one processor operatively coupled to the image sensor, the at least one processor executing: (i) a neural network to extract convolution outputs from the data stream representative of the features of the object; and (ii) a classifier to classify the object based on the convolution outputs. Claim 7: The apparatus of claim 6 , wherein the sensor comprises at least one of an image sensor, a lidar, a radar, or an acoustic sensor. Claim 8: The apparatus of claim 6 , wherein the neural network comprises a deep neural network (DNN). Claim 9: The apparatus of claim 6 , wherein the neural network comprises an Adaptive Resonance Theory (ART) network. Claim 10: A method for implementing a lifelong learning deep neural network (L-DNN) in a machine operating in real time, comprising: predicting, by the L-DNN, a first action for the real-time operating machine based on (i) observing an environment of the real-time operating machine using sensors, and (ii) predetermined weights of the L-DNN; determining, by the L-DNN, a discrepancy between an expectation and a perception regarding the machine operating in real time based on the observation; and triggering a fast learning mode by the L-DNN in response to the discrepancy, the fast learning mode generating predictions that are revised based on the observations without changing the pre-determined weights of the L-DNN. Claim 11: determining that the real-time operating machine is offline; and 11. The method of claim 10, further comprising: in response to determining that the real-time operating machine is offline, triggering a slow learning mode, the slow learning mode modifying the pre-determined weights of the L-DNN based on the observations. Claim 12: 1. A method of extracting, aggregating and sharing knowledge among a plurality of real-time operating machines, comprising: each real-time operating machine among said plurality of real-time operating machines implementing a neural network with a respective copy of a weight matrix; learning at least one new object with a rapid learning subsystem of a first real-time machine among the plurality of real-time machines; transmitting a representation of the at least one new object from the first real-time machine to a server over a communication channel; and forming, at the central server, an updated weight matrix based at least in part on the representation of the at least one new object from the first real-time operating machine; transmitting a copy of the updated weighting matrix from the server to at least a second real-time machine of the plurality of real-time machines. Claim 13: The learning of the new object includes: acquiring an image of the at least one new object with an image sensor operatively coupled to the rapid learning subsystem of the first real-time machine; and processing the image of the at least one new object with the rapid learning subsystem of the first real-time machine. Claim 14: the neural network comprises an adaptive resonance theory (ART) neural network; 13. The method of claim 12, further comprising generating the representation of the at least one new object with the ART neural network. Claim 15: the representation of the at least one new object includes a weight vector; 15. The method of claim 14, further comprising aggregating the copy of the weight matrix and the weight vector used by the first real-time operating machine to reduce a memory footprint of the representation of the at least one new object. Claim 16: The method of claim 12 , further comprising aggregating the representation of the new object with a representation of at least one previously known object. Claim 17: Transmitting the representation of the at least one new object comprises: 13. The method of claim 12, further comprising transmitting the representation of the at least one new object to the server by a second real-time machine of the plurality of real-time machines. Claim 18: Forming the updated weight matrix comprises: 13. The method of claim 12, comprising blending the representation of the at least one new object with a representation of at least one other new object from at least one other real-time operating machine among the plurality of real-time operating machines. Claim 19: 1. A method of classifying objects with a neural network trained to recognize objects in a plurality of categories, comprising: presenting an object to the neural network; determining, by the neural network, a plurality of confidence levels, each confidence level among the plurality of confidence levels representing a likelihood that the object falls into a corresponding category among the plurality of categories; comparing the plurality of confidence levels to a threshold; and determining, based on the comparison, that the object does not fall into any of the plurality of categories. Claim 20: Conducting the comparison comprises: 20. The method of claim 19, comprising determining that no confidence level in the plurality of confidence levels exceeds the threshold. Claim 21: The method of claim 19 , further comprising setting the threshold greater than an average of the plurality of confidence levels.
Claims
1. 1. A method for implementing a lifelong learning deep neural network (L-DNN) in a machine operating in real time, the machine including a fast learning subsystem and a slow learning subsystem with pre-determined weights, the method comprising: predicting, by the L-DNN, a first action for the real-time operating machine based on (i) observing an environment of the real-time operating machine by a sensor, and (ii) the predetermined weights of the L-DNN; determining, by the L-DNN, a discrepancy between an expectation and a perception regarding the real-time operating machine based on the observations; triggering a fast learning mode by the L-DNN in response to the mismatch, where the fast learning mode includes updating a weight vector of a corresponding category node of the fast learning subsystem based on the observation, or adding a category node to the fast learning subsystem with a weight vector based on the observation without changing the pre-determined weights of the slow learning subsystem; The method includes:
2. determining that the real-time operating machine is offline; and triggering a slow learn mode in response to determining that the real-time operating machine is offline, the slow learn mode modifying the pre-determined weights based on the observations; The method of claim 1 further comprising:
3. The method of claim 1 , wherein predicting the first action of the real-time operating machine is further based on (iii) features extracted from the observations by the slow learning subsystem.
4. 2. The method of claim 1, further comprising aggregating the weights of the fast learning subsystem to limit the memory footprint of the fast learning subsystem and limit the growth of the L-DNN's memory to grow no faster than a linear growth in the number of objects the L-DNN is trained to recognize.
5. 5. The method of claim 4, further comprising applying the aggregated weights of the fast learning subsystem to predict a second action of the real-time operating machine based on subsequent observations of the environment of the real-time operating machine by the sensors.
6. 10. The method of claim 1, wherein the real-time operating machine is one of a robot, a drone, a self-driving car, a smartphone, and an Internet of Things (IoT) device.
7. A sensor that observes the environment of a machine operating in real time; a processor operatively coupled to the sensors and implementing a lifelong learning deep neural network (L-DNN) including a slow learning subsystem and a fast learning subsystem having predetermined weights, the L-DNN being configured to: predict a first action on the real-time operating machine based on (i) the observations and (ii) the pre-determined weights of the slow learning subsystem; determine a discrepancy between an expectation and a perception regarding the real-time operating machine based on the observations; and, in response to the discrepancy, trigger a fast learning mode, which updates a weight vector of a corresponding category node of the fast learning subsystem based on the observations or adds a category node to the fast learning subsystem with a weight vector based on the observations without changing the pre-determined weights of the slow learning subsystem; Machines that operate in real time, including:
8. 8. The real-time operating machine of claim 7, wherein the real-time operating machine is configured to modify the pre-determined weights based on the observations in a slow learning mode when the real-time operating machine is offline.
9. The machine operating in real time according to claim 7 , wherein the L-DNN is further configured to (iii) make predictions based on features extracted from the observations by the slow learning subsystem.
10. 8. The real-time machine of claim 7, wherein the processor is further configured to aggregate the weights of the fast learning subsystem, limit the memory footprint of the fast learning subsystem, and limit memory growth of the L-DNN to no faster than a linear growth in the number of objects the L-DNN is trained to recognize.
11. 11. The real-time operating machine of claim 10, further configured to apply the aggregated weights of the fast learning subsystem to predict a second action of the real-time operating machine based on subsequent observations of the environment of the real-time operating machine by the sensors.
12. The machine that operates in real time according to claim 7, wherein the machine that operates in real time is one of a robot, a drone, an autonomous vehicle, a smartphone, and an Internet of Things (IoT) device.
Citation Information
Patent Citations
Device and method for information processing, and device and method for pattern recognition
JP2005352900A