Systems and Methods for Incremental Learning for Object Detection
The two-stage neural network approach with a teacher-student model and combined loss function allows neural networks to learn new object classes without forgetting previous classes, maintaining detection accuracy and reducing resource requirements.
Patent Information
- Application Number
- CN202080022272.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-03-21
- Filing Date
- 2020-03-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2040-03-13
AI Technical Summary
Existing neural network object detectors are prone to catastrophic forgetting when incrementally learning new object categories, and cannot effectively maintain the detection capabilities of previous learning object categories, and have high computing resources requirements.
A two-level neural network object detector is used to balance the ability to detect new class objects and maintain old class objects through the total loss function. It uses the iterative update method of the teacher network and student network, and combines the distillation loss of regional proposals and object predictions, and gradually adjusts the model to minimize the total loss.
It realizes that when learning new object categories incrementally, maintains detection capabilities for previous learning object categories, reduces computing resource requirements, and provides an end-to-end learning mechanism, which can temporarily customize and update models when requested by users.
Smart Images

Figure CN113614748B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to detecting objects in images and / or videos. Background Art
[0002] Various real-world applications, such as automatic analysis of medical images and security detection based on image recognition, detect objects (i.e., instances of known object classes) in images and / or videos. Recently, object detection in images and / or videos has significantly benefited from the development of neural networks. As used herein, the term “one or more object detectors” is a shorter form of “one or more neural network object detectors”. A neural network object detector is essentially a model characterized by a set of parameters. Object detection includes determining the location and type or category of an object in an image and / or video. Conventional object detectors utilize one-stage or two-stage object detectors. One-stage object detectors (e.g., You-Only-Look-Once, YOLO, and Single Shot Detection, SSD) simultaneously determine where and what objects are present by making a fixed number of grid-based object predictions. Two-stage object detectors separately propose regions where known objects may be located and then associate specific object predictions (e.g., probabilities) with the proposed regions.
[0003] However, neural networks suffer from catastrophic forgetting, i.e., the loss of the ability to detect objects of a first class when using incremental learning to retrain a given object detector to detect objects of a second class, which makes incremental learning incompatible with lifelong detection. Attempts to bypass catastrophic forgetting include running a new object detector dedicated to the new class in parallel with the object detector used to detect objects of the previously learned class, which doubles the necessary computational resources (memory and execution capabilities). Another approach uses images that include a combination of objects of both the new class and the previously learned class to train the object detector, but such a combined dataset may not be available. Training the object detector with a large dataset that includes both images of objects of the new class and images of objects of the previously learned class also significantly increases the necessary computational resources.
[0004] K. Schmelkov proposed in “Incremental Learning of Object Detectors without Catastrophic Forgetting” at the International Conference on Computer Vision (2017) a fast region-based convolutional neural network (RCNN) architecture that uses pre-computed object proposals and thus runs slowly. The article discusses adding new classes to the region classifier using knowledge distillation, but this is problematic because the above method cannot be directly applied to more advanced architectures, such as faster RCNN.
[0005] Neural network-based object detection models are routinely trained on object detection datasets such as PASCAL VOC and COCO which have 20 and 80 classes respectively. The object classes that can be detected are determined by the dataset used to train the model. Once the model has been trained, the object classes cannot be changed without retraining the entire network with labeled data for all classes of interest.
[0006] If the training data (images) of an existing model is not available to the user (e.g., the corresponding website is slow or unavailable) or the user lacks the privilege to access the original data, the user cannot update the model by adding new object classes that have become relevant to the user.
[0007] Accordingly, it is desirable for a neural network object detector (i.e., method and device) to be able to avoid catastrophic forgetting of previously learned object classes when being trained to detect objects belonging to new object classes. There is a need for a more advanced object detection system that can be temporarily customized in the case where a user wants to add new object classes and can update the object detection model upon request. SUMMARY OF THE INVENTION
[0008] Exemplary embodiments relate to a method for incrementally learning new object classes (i.e., training an object detector to detect objects belonging to new classes) without catastrophically forgetting objects belonging to previously learned classes. The object detector is updated using a total loss that enables balancing the ability to detect objects of new classes and the ability to detect objects of previously learned classes, without the need to access the images previously used to train the object detector. The trained and updated object detector retains the ability to detect objects belonging to previously learned object classes while extending its detection ability to objects belonging to new object classes.
[0009] Other exemplary embodiments are systems that detect known object classes, interact with the user (e.g., via natural language), receive a request from the user to add new object classes, autonomously collect the necessary training data, and retrain the object detector to locate and identify objects belonging to the new classes without significantly losing the ability to detect objects belonging to previously learned object classes. In one embodiment, the system provides the user with recommendations and requests for new object classes.
[0010] An exemplary embodiment relates to a method for incrementally learning object detection in an image without catastrophically forgetting previously learned object classes. The method includes training an existing two-stage neural network object detector to locate and identify objects belonging to at least one additional object class in an image by iteratively updating the two-stage neural network object detector until an overall detection accuracy criterion is met. The update is performed to balance minimizing the loss of the initial ability to locate and identify objects belonging to one or more previously learned object classes and maximizing the ability to locate and identify objects belonging to the additional object class.
[0011] Another exemplary embodiment relates to a method for incrementally learning object detection without catastrophically forgetting previously learned object classes, the method including receiving a training image that includes an object belonging to an object class unknown to an initial version of a two-stage neural network object detector that is capable of detecting objects belonging to at least one previously learned object class. The method further includes using a plurality of images one by one in the training image to train the two-stage neural network object detector to detect objects of the initially unknown object class until a predetermined condition is met. The training includes: (1) inputting one image of the plurality of images into an initial version of the two-stage neural network object detector to obtain a first region proposal and a first object prediction for at least one previously learned object class, (2) inputting the one image into a current version of the two-stage neural network object detector to obtain a second region proposal and a second object prediction for the at least one previously learned object class and the initially unknown object class, (3) comparing the first region proposal with the second region proposal to estimate a region proposal distillation loss that quantifies a reduction in the ability to locate objects of the at least one previously learned object class, (4) comparing the first object prediction with the second object prediction to estimate an object recognition distillation loss that quantifies a reduction in the ability to identify objects of the at least one previously learned object class, (5) comparing the second region proposal with a ground truth label of the one image to estimate a region proposal network loss for the initially unknown object class, (6) comparing the second object prediction with the ground truth label to estimate an object recognition loss for the initially unknown object class, (7) calculating a total loss that combines the region proposal distillation loss, the object recognition distillation loss, the region proposal network loss, and the object recognition loss, and (8) updating the current version of the two-stage neural network object detector to minimize the total loss. The predetermined condition is met when the number of training iterations reaches a predetermined number or when the total loss reduction rate is lower than a predetermined threshold.
[0012] An exemplary embodiment relates to a computer-readable medium containing computer-readable code that, when read by a computer, causes the computer to execute a method for incremental learning for object detection in images without catastrophically forgetting one or more previously learned object classes. The method includes training a two-stage neural network object detector to localize and identify objects belonging to additional object classes by iteratively updating the two-stage neural network object detector until an overall detection accuracy criterion is met. The update is performed to balance minimizing the loss of the initial ability to localize and identify objects belonging to one or more previously learned object classes and maximizing the ability to localize and identify objects belonging to additional object classes.
[0013] Other exemplary embodiments relate to systems for incremental learning for object detection without catastrophically forgetting previously learned object classes. The systems each include a computing device configured to receive an image, use a two-stage neural network object detector to localize and identify objects, and communicate with a cloud server, and the cloud server. The cloud server is configured to collect training images, store an initial version of the two-stage neural network object detector that localizes and identifies one or more object classes, and communicate with the computing device and the Internet. At least one of the computing device and the cloud server is configured to train the two-stage neural network object detector to localize and identify objects belonging to additional object classes by iteratively updating the two-stage neural network object detector, which is initially capable of localizing and identifying objects belonging to one or more object classes, until an overall detection accuracy criterion is met, and perform the update so as to balance minimizing the loss of the initial ability of the two-stage neural network detector to localize and identify objects belonging to the one or more object classes and maximizing the ability of the two-stage neural network detector to localize and identify objects belonging to the additional object classes. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Embodiments of the present invention will now be described by way of example only with reference to the accompanying drawings, in which:
[0015] Figure 1 is a schematic diagram of incremental learning object detection without catastrophic forgetting according to an embodiment;
[0016] Figure 2 is a flowchart of a method for incremental learning object detection in an image without catastrophically forgetting one or more previously learned object classes according to an embodiment;
[0017] Figure 3 shows the outputs of an initial version of an object detector and a trained version of the object detector using the same image as the input;
[0018] Figure 4 shows an object detection system according to an embodiment;
[0019] Figure 5 depicts a cloud computing environment according to an embodiment of the present invention;
[0020] Figure 6 depicts an abstract model layer according to an embodiment of the present invention;
[0021] Figure 7 is a flowchart showing object detection according to an embodiment;
[0022] Figure 8 is a flowchart of a method for adding a new class according to an embodiment;
[0023] Figure 9 is a flowchart showing accuracy-aware model improvement according to an embodiment; and
[0024] Figure 10 is a flowchart showing adding a new object class to an object detection model upon receiving a user request according to an embodiment. Detailed Description
[0025] The flowcharts and block diagrams in the accompanying drawings used in the following description illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to different embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or part of an instruction, which includes one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions labeled in the blocks may occur in an order different from that labeled in the figures. For example, depending on the functions involved, two consecutive blocks shown may actually be executed substantially simultaneously, or these blocks may sometimes be executed in the reverse order. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a system based on dedicated hardware that performs the specified functions or actions or executes a combination of dedicated hardware and computer instructions.
[0026] The exemplary embodiments described below relate to incremental learning to detect objects related to new classes without catastrophically forgetting to detect objects related to previously learned object classes. The training of an object detector aims to: (A) enable the detector to detect (i.e., localize and identify) objects belonging to new classes, regardless of where such objects are located and what other objects appear in the image; (B) maintain good performance for detecting objects belonging to previously learned classes; (C) employ a reasonable number of model parameters and have an acceptable memory requirement for using the model; and (D) provide an end-to-end learning mechanism that jointly updates the classifier to identify the new object classes and the feature representation to localize objects belonging to the new object classes.
[0027] Figure 1 Schematically shows incremental learning of object detection in an image according to an embodiment without catastrophic forgetting of previously learned object categories. An initial version of a two-stage object detector named "Teacher Network (T)" is capable of detecting (i.e., localizing and identifying) objects belonging to previously learned object categories. A current version of a two-stage neural network object detector named "Student Network (S)" is trained to detect objects belonging to a new object category. The following description relates to a single new object category, but multiple object categories can be added successively or simultaneously. The student network starts with parameters from the teacher network, but the output from the student network is extended to detect objects belonging to the new object category.
[0028] Training images 100 include objects belonging to the new object category, but may also include objects belonging to previously learned object categories. These training images are input one by one or in batches into the teacher network 110 and the student network 120 (i.e., in parallel to the two networks and serially through the training images). For each input image, the teacher network 110 generates teacher region proposals (RPN(T)) 111 and teacher object predictions 113, where the teacher region proposals 111 are used to extract feature vectors 112, and the teacher object predictions 113 identify one of the previously learned object categories as corresponding to the object in each region of the teacher RPN. Additionally, for each input image, the student network 120 generates student region proposals (RPN(S)) 121 and student object predictions 123, where the student region proposals 121 are used to extract feature vectors 122, and the student object predictions 123 identify one of the previously learned object categories or the new object category as corresponding to the object in each region of the RPN(S).
[0029] The detection accuracy can be estimated based on the outputs of the teacher network and the student network. Based on comparing the teacher region proposals 111 with the student region proposals 121, the region proposal network distillation loss 131 quantifies the reduction in the ability of the student network to localize objects of previously learned object categories. Based on comparing the teacher object predictions 112 with the student object predictions 122, the object recognition distillation loss 132 quantifies the reduction in the ability of the student network to identify objects of previously learned object categories.
[0030] Further, comparing the output of the student network with the ground truth labels associated with the training images enables the estimation of the region proposal loss 141 of the student network and the object recognition loss 142 of the student network, where the region proposal loss 141 of the student network is related to localizing objects belonging to the new object category, and the object recognition loss 142 of the student network is related to identifying objects belonging to the new object category. In machine learning, the term "ground truth" refers to the classification labels of the training set or more simply the true location and identity of the objects in the training images.
[0031] In one embodiment, the total loss is defined by combining the region proposal network distillation loss, the object recognition distillation loss, the region proposal network loss, and the object recognition loss. The student network can then be updated to minimize the total loss. This combination can be a simple summation of the relative measurements of these losses, or a weighted summation thereof. It is represented in the following mathematical form:
[0032]
[0033] where L total is the total loss, is the region proposal network distillation loss, is the object recognition distillation loss, is the region proposal network loss and is the object recognition loss, and λ1...λ4 are the corresponding weights.
[0034] Using the total loss when updating the object detector is a way to balance minimizing the loss of its ability to detect objects belonging to previously learned object classes and maximizing its ability to detect objects belonging to newly added object classes. However, there are other ways to achieve this balance; for example, there may be a predefined limit on the loss of the ability to detect objects belonging to any individual previously learned object class, the average loss of the ability, and / or the ability to detect objects belonging to the added object classes (in this case, the predefined limit is a threshold that must be exceeded), and so on.
[0035] Figure 2 is a flowchart of method 200 for incremental learning of object detection without catastrophic forgetting according to an embodiment. A training image 210 is received, which includes objects belonging to object classes unknown to an initial version of a two-stage neural network object detector that is capable of detecting objects belonging to at least one previously learned object class. The training image has associated ground truth labels indicating the locations (e.g., regions that can be boxes) and object classes of the objects therein.
[0036] At 220, the two-stage neural network object detector is trained to detect objects of the initially unknown object classes one by one or in batches using the training image until a predetermined condition is met. The predetermined condition can be completing a predetermined number of iterations, achieving a threshold accuracy, or being unable to further reduce the loss. The training of the two-stage neural network object detector includes performing the following operations for each of a plurality of images:
[0037] (1) Input the image into an initial version of the two-stage neural network object detector (e.g., the teacher network) to obtain a first region proposal and a first object prediction for the previously learned object classes,
[0038] (2) Input the image into the current version of a two-stage neural network object detector (e.g., the student network) to obtain second region proposals and second object predictions for the previously learned object classes and the new object class,
[0039] (3) Compare the first region proposals with the second region proposals to estimate a region proposal distillation loss that quantifies the reduction in the ability to localize objects belonging to the previously learned object classes,
[0040] (4) Compare the first object predictions with the second object predictions to estimate an object recognition distillation loss that quantifies the reduction in the ability to recognize objects of the previously learned object classes,
[0041] (5) Compare the second region proposals with the ground truth labels of the image to estimate a region proposal network loss for the new object class,
[0042] (6) Compare the second object predictions with the ground truth labels to estimate an object recognition loss for the new object class,
[0043] (7) Calculate a total loss by combining the region proposal distillation loss, the object recognition distillation loss, the region proposal network loss, and the object recognition loss, and
[0044] (8) Update the current version of the two-stage neural network object detector to minimize the total loss. The update can use any of various known methods (e.g., by applying a stochastic gradient descent method to the total loss) to adjust the parameters of the neural network.
[0045] More generally, a method for incremental learning of object detection in images without catastrophically forgetting previously learned object classes performs training of a two-stage neural network object detector to localize and identify objects belonging to additional object classes until an overall detection accuracy criterion is met. The overall detection accuracy describes the ability of the trained object detector to detect previously learned object classes and additional object classes. For example, the sum of the predictions expressed as probabilities must exceed an accuracy threshold, or the predicted probability for each object must exceed a low threshold accuracy limit.
[0046] Iteratively update a two-stage neural network object detector to balance minimizing the loss of its initial ability to localize and identify objects belonging to previously learned object classes and maximizing its ability to additionally localize and identify objects belonging to additional object classes. The ability can be measured by, for example, the percentage of correct detections in all detections, the average of the prediction probabilities, a combination with the number of false positive detections or false negative detections, etc. The loss of the initial ability to localize and identify objects belonging to previously learned object classes can be the sum (or weighted sum) of the distillation losses. This loss is determined by comparing the output of the initial (untrained) version of the object detector with the output of the trained (current) version of the object detector.
[0047] The overall detection accuracy criterion can be evaluated by determining a region proposal distillation loss and an object recognition distillation loss. The "distillation loss" indicates the reduction in the corresponding ability due to the addition of new object classes. The region proposal distillation loss is based on comparing the region proposals output by the initial version of the two-stage neural network object detector with the current region proposals output by the current version of the two-stage neural network object detector for the same input training image.
[0048] The object recognition distillation loss is based on comparing the previously learned object predictions output by the initial version of the two-stage neural network object detector with the object predictions output by the current version of the two-stage neural network object detector. The comparison for determining the distillation loss is made from the perspective of the previously learned object classes. The two-stage neural network object detector is updated to minimize the region proposal distillation loss and the object recognition distillation loss.
[0049] Figure 3 Exemplarily shown is the object detection (i.e., the position in the box and object recognition) of two objects (chair and banana) in the left image 310 by the initial version of the object detector trained for 19 object classes, and the object detection of two objects and an additional object (i.e., laptop) in the right image 320 by the version of the two-stage neural network object detector trained for 20 object classes (i.e., also for detecting laptops). The region proposals include rectangular boxes 312, 314 and 322 - 326 respectively. The object predictions can be represented as specific object probabilities. Taking "chair" 0.98 and "banana" 0.95 as old classes, and "laptop" 0.99 as a new class (afterwards: chair 0.96, banana 0.95), and updating the text accordingly.
[0050] For the left image, the probability of the chair in box 312 is 0.98 (1 is certain), and the probability of the banana in box 314 is 0.99. For the right image, the probability of the chair in box 322 is 0.96, the probability of the banana in box 324 is 0.95, and the probability of the laptop in box 328 is 0.99.
[0051] Figure 4 FIG. 2 shows an object detection system 400 according to an embodiment. The object detection system includes a computing device 410 and a cloud device 420 (e.g., a server). The computing device is capable of receiving image / video input from a camera 402 and transmitting an image / video output to a screen 408. The computing device is also capable of receiving audio input from a microphone 404 and transmitting an audio output to a speaker 406.
[0052] The camera 402 that provides the image / video input to the computing device may be integrated in a smart phone or smart glasses, or may be a stand-alone camera. The microphone 404 enables a user request expressed by a user to be provided as audio input to the computing device.
[0053] The computing device 410 is configured to process audio and video input and output for object detection and communicate with the cloud / server 420. The cloud / server 420 is configured to perform object detection, process user requests received from the computing device, collect training images from the web 430, and store an initial version of an object detector.
[0054] At least one of the computing device 410 and the cloud / server 420 is configured to execute a method for performing incremental learning object detection in an image without catastrophically forgetting previously learned object categories (e.g., as previously described). When the incremental learning is performed by the cloud / server, an object detector trained to detect objects belonging to an additional object category is sent to the computing device. When the method is performed by the computing device, the cloud / server provides the computing device with training images and an initial version of the object detector.
[0055] It should be understood that although a detailed description of cloud computing is provided, the implementation of the teachings provided herein is not limited to a cloud computing environment. Instead, embodiments of the present invention are capable of being implemented in conjunction with any other type of computing environment now known or later developed. Cloud computing is a model for service delivery that enables convenient, on-demand network access to a shared pool of configurable computing resources (e.g., networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly configured and released with minimal management effort or interaction with a service provider.
[0056] The cloud model can include at least five characteristics, at least three service models, and at least four deployment models. These five characteristics are on-demand self-service, broad network access, resource pooling, rapid elasticity, and measured service. Regarding on-demand self-service, cloud consumers can automatically and unilaterally provision computing capabilities such as server time and network storage on demand without human interaction with the service provider. Broad network access refers to the ability to be accessed over a network and through standard mechanisms that facilitate the use of heterogeneous thin or thick client platforms (e.g., mobile phones, laptops, and PDAs). For resource pooling, the provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, where different physical and virtual resources are dynamically assigned and reassigned as needed. There is a sense of location independence as consumers typically have no control or knowledge of the exact location of the provided resources but may be able to specify a location at a higher level of abstraction (e.g., country, state, or data center). Rapid elasticity refers to the ability to provision (automatically in some cases) quickly and elastically, to scale out rapidly, and to release quickly to scale in rapidly. For consumers, the capabilities available for provisioning generally appear unlimited and can be purchased in any quantity at any time. For measured service, the cloud system automatically controls and optimizes resource use by leveraging a metering capability at some level of abstraction appropriate to the service type (e.g., storage, processing, bandwidth, and active user accounts). Resource use can be monitored, controlled, and reported, providing transparency for both the provider and the consumer of the utilized service.
[0057] These three service models are Software as a Service (SaaS), Platform as a Service (PaaS), and Infrastructure as a Service (IaaS). Software as a Service provides consumers with the ability to use the provider's applications running on the cloud infrastructure. The applications can be accessed from different client devices through a thin client interface such as a web browser (e.g., web-based email). Except for limited user-specific application configuration settings, the consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, storage, or even the individual application capabilities. Platform as a Service provides consumers with the ability to deploy onto the cloud infrastructure, and the applications created or acquired by the consumer are created using programming languages and tools supported by the provider. The consumer does not manage or control the underlying cloud infrastructure, including the network, servers, operating systems, or storage, but has control over the deployed applications and possibly the application hosting environment configuration. Infrastructure as a Service provides consumers with the ability to provide processing, storage, network, and other basic computing resources in which the consumer can deploy and run any software that may include operating systems and applications. The consumer does not manage or control the underlying cloud infrastructure but has control over the operating systems, storage, deployed applications, and possibly limited control over select networking components such as the host firewall.
[0058] The deployment models are private cloud, community cloud, public cloud, and hybrid cloud. A private cloud infrastructure is for an organization's exclusive use. It can be managed by the organization or a third party and can be located on-premises or off-premises. A community cloud infrastructure is shared by several organizations and supports a specific community with shared concerns (e.g., missions, security requirements, policies, and compliance considerations). It can be managed by the organization or a third party and can be located on-premises or off-premises. A public cloud infrastructure is available to the public or large industry groups and is owned by an organization selling cloud services. A hybrid cloud infrastructure is a composition of two or more clouds (private, community, or public) that remain distinct entities but are bound together by standardized or proprietary technologies that enable data and application portability, such as cloud bursting for load balancing between clouds.
[0059] The cloud computing environment is service-oriented, focusing on statelessness, low coupling, modularity, and semantic interoperability. The core of cloud computing is an infrastructure that includes a network of interconnected nodes. Now refer to Figure 5 , which depicts an illustrative cloud computing environment 500. As shown, the cloud computing environment 500 includes one or more cloud computing nodes 510 with which local computing devices used by cloud consumers (such as, for example, a personal digital assistant (PDA) or cellular phone 520, desktop computer 530, laptop computer 540, and / or in-vehicle computer system 550) can communicate. The nodes 510 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as the private cloud, community cloud, public cloud, or hybrid cloud or a combination thereof described above. This allows the cloud computing environment 500 to provide infrastructure, platform, and / or software as services without the cloud consumer having to maintain resources on a local computing device. It should be understood that Figure 5 the types of computing devices 520 - 550 shown are only illustrative, and the computing nodes 510 and the cloud computing environment 500 can communicate with any type of computerized device via any type of network and / or network addressable connection, such as using a web browser.
[0060] Now refer to Figure 6 , which shows a set of functional abstraction layers provided by the cloud computing environment 500. It should be understood in advance that Figure 1 、 2The components, layers, and functions shown in FIGS. 4 are intended to be illustrative only, and embodiments of the present invention are not limited thereto. As depicted, the following layers and corresponding functions are provided. The hardware and software layer 660 is made up of hardware and software components. Examples of hardware components include: mainframes 661, servers 662 based on RISC (Reduced Instruction Set Computer) architecture, servers 663, blade servers 664, storage devices 665, and network and networking components 666. In some embodiments, software components include web application server software 667 and database software 668.
[0061] The virtualization layer 670 is an abstraction layer that includes one or more virtual entities such as: virtual servers 671, virtual storage 672, virtual networks 673 (which may be virtual private networks), virtual applications and operating systems 674, and virtual clients 675.
[0062] In one example, the management layer 680 may provide the functions described below. Resource provisioning 681 provides for the dynamic acquisition of computing resources and other resources used to perform tasks within the cloud computing environment. Metering and pricing 682 provides cost tracking when resources are utilized within the cloud computing environment and bills or invoices for the consumption of these resources. In one example, these resources may include application software licenses. Security provides authentication for cloud consumers and tasks, as well as protection for data and other resources. The user portal 683 provides access to the cloud computing environment for consumers and system administrators. Service level management 684 provides cloud computing resource allocation and management such that the required service levels are met. Service level agreement (SLA) planning and fulfillment 685 provides pre-arrangement and procurement of cloud computing resources for anticipated future requirements in accordance with the SLA.
[0063] The workload layer 690 provides examples of functions that can utilize the cloud computing environment. Examples of workloads and functions that may be provided from this layer include: maps and navigation 691, software development and lifecycle management 692, virtual classroom education delivery 693, data analysis processing 694, transaction processing 695, and incremental learning object detection 696.
[0064] Figures 7 - 10 is a flowchart of a method that can be performed by Figure 4 the system 400 shown in FIGS.. These flowcharts illustrate steps performed by the computing device 410 or the cloud / server 420. The object detector may be run by the computing device 410 or by the cloud / server 420. Figure 7It is a flowchart showing object detection including local or cloud object detection. At 710, the computing device receives an image from a camera (e.g., 402). If object detection is performed locally (i.e., the "yes" branch of box 720), then at 730, the object detector runs on the local computing device. If object detection is performed in the cloud (i.e., the "no" branch of box 720), then at 760 the computing device sends the received image(s) to the cloud / server, enabling the object detector to run on the cloud / server at 770. The detection results are then output to the user: shown at 740 (e.g., on screen 408) and / or notified at 750 (e.g., using speaker 406). If object detection continues (i.e., the "no" branch of box 780), then the computing device is ready to process other image(s) at 710. For example, when a video stream is received as a batch of images, the detection can continue.
[0065] Figure 8 It is a flowchart of a method for adding a new object category designated as category A to an existing object detector (such as the detector used in Figure 7 ). The method shows obtaining a training data set, which can be used for incremental training and is based on the content provided in the input (image or name), the ability to obtain the existing training data set, or the need to create a new training data set.
[0066] At 810, the computing device receives a user request to add category A as an audio input to the previously learned object categories of the object detector. At 820, the computing device extracts the name of category A from the user request received as natural language from the microphone, or obtains a representative image of object category A provided by the camera. In the former case, at 830, the computing device sends the name of category A to the cloud / server. Otherwise, the computing device sends the representative image to the cloud / server.
[0067] Once the name of category A is received, the cloud / server searches for existing images on the website including the requested category A at 840. If it is determined at 850 that such images are not available, then at 860 the user is requested to provide a sample image of category A object. If such images are available, then at 870, they are collected to form a training image set for updating the object detector to be able to detect objects belonging to category A via incremental learning. The updated object detector is provided to the computing device at 880, which can then provide audio or video feedback to the user regarding the category A detection ability at 890.
[0068] Now returning to 820, if the computing device obtains a representative image of object category A, the computing device prepares a photo of category A objects at 825 and sends it / them to the cloud / server at 835. When receiving a photo or a sample image from the user, at 845, the cloud / server searches websites for images using visual features or category names similar to the received photo or sample image. At 855, the cloud / server creates a training image set for category A, and at 870, the training images are used to update the object detector.
[0069] Figure 9 is a flowchart showing the training of a precision-aware model. The method shows options that can be provided when adding a new category does not go smoothly. Starting from a detection model D_n+m that can detect objects belonging to an original set of n object categories and an added set of m categories at 910, at 920, the computer device receives a user request to add a new category c. At 930, the computer device sends the name of category c of its representative image to the cloud / server. At 940, the cloud / server updates the detection model to add object category c to the m+n object categories that the object detector can already detect. Then, at 950, the computer device checks for a decrease in detection accuracy for the m+n object categories (also called classes) due to training (i.e., the difference between the detection accuracy of D_n+m and the detection accuracy of D_m+n+c). If the detection accuracy of any class decreases significantly (the "yes" branch in box 960), the system reverts to the D_n+m model at 970, creates a new detection model D_n+c at 980, and uses D_n+m and Dn+c in parallel at 990. Otherwise (the "no" branch in box 960), at 965, the computing device uses the D_n+m+c model for object detection.
[0070] Figure 10 is a flowchart showing adding a new object to an object detection model when receiving a user request. At 1010, the computing device receives an audio input including a user request to add a new object to an existing object detection model. At 1020, the computing device receives an image of the new object (e.g., a photo of this new object viewed from different angles). At 1030, the computing device sends the image (photo) to the cloud / server. At 1040, the server updates the existing detection model to include the new object category and sends the updated model to the computing device at 1050.
[0071] The methods and systems according to exemplary embodiments of the present invention may take the form of a full hardware embodiment, a full software embodiment, or an embodiment containing both hardware and software elements. In a preferred embodiment, the present invention is implemented in software, which includes but is not limited to firmware, resident software, and microcode. In addition, the exemplary methods and systems may take the form of a computer program product accessible from a computer-usable or computer-readable medium that provides program code for use by or in conjunction with a computer, a logic processing unit, or any instruction execution system. For the purposes of this description, a computer-usable or computer-readable medium can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. Suitable computer-usable or computer-readable media include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems (or apparatus or devices) or propagation media. Examples of computer-readable media include semiconductor or solid state memories, magnetic tape, removable computer disks, random access memory (RAM), read-only memory (ROM), rigid disks, and optical disks. Current examples of optical disks include compact disk-read only memory (CD-ROM), compact disk-read / write (CD-R / W), and DVD.
[0072] Suitable data processing systems for storing and / or executing program code include, but are not limited to, at least one processor directly or indirectly coupled to memory elements through a system bus. The memory elements include local memory employed during actual execution of the program code, mass storage devices, and cache memories that provide temporary storage of at least some program code to reduce the number of times code must be retrieved from the mass storage device during execution. Input / output or I / O devices (including but not limited to keyboards, displays, and pointing devices) can be coupled to the system directly or through intervening I / O controllers. Exemplary embodiments of the methods and systems according to the present invention also include a network adapter coupled to the system to enable the data processing system to be coupled to other data processing systems or remote printers or storage devices through an intervening private or public network. Suitable currently available types of network adapters include, but are not limited to, modems, cable modems, DSL modems, Ethernet cards, and combinations thereof.
[0073] In one embodiment, the present invention is directed to a machine-readable or computer-readable medium that includes machine-executable or computer-executable code that, when read by a machine or computer, causes the machine or computer to perform a method for incrementally learning object detection in an image without catastrophically forgetting previously learned object classes. The machine-readable or computer-readable code can be any type of code or language that can be read and executed by a machine or computer and can be expressed in any suitable language or syntax known and available in the prior art, including machine language, assembly language, high-level languages, object-oriented languages, and scripting languages. The computer-executable code can be stored on any suitable storage medium or database, including databases disposed therein, communicatively coupled to and accessible by a computer network utilized by a system according to the present invention, and can be executed on any suitable hardware platform known and available in the art, including a control system for controlling the presentation of the present invention.
[0074] While it is apparent that the illustrative embodiments of the invention disclosed herein achieve the objectives of the invention, it should be understood that those skilled in the art can design many modifications and other embodiments. Additionally, features and / or elements from any embodiment can be used alone or in combination with other embodiments, and steps or elements from the methods according to the present invention can be performed or executed in any suitable order. Accordingly, it should be understood that the appended claims are intended to cover all such modifications and embodiments that fall within the spirit and scope of the invention.
Claims
1. A method for incrementally learning object detection in an image without catastrophic forgetting of one or more previously learned object classes, the method comprising: Training the two-stage neural network object detector to locate and identify objects in the image belonging to an additional object class by iteratively updating the two-stage neural network object detector until a global detection accuracy criterion is met, wherein the update is performed to balance minimizing the loss of the initial ability to locate and identify objects belonging to the one or more previously learned object classes and maximizing the ability to additionally locate and identify objects belonging to the additional object class, and evaluating whether the global detection accuracy criterion is met includes: Determining a region proposal distillation loss based on a comparison of region proposals output by an initial version of the two-stage neural network object detector with current region proposals output by a current version of the two-stage neural network object detector, wherein the region proposal distillation loss quantifies the reduction in the ability to locate objects belonging to previously learned object classes; and Determining a previously learned object recognition distillation loss based on a comparison of previously learned object predictions obtained by the initial version of the two-stage neural network object detector with object predictions output by the current version of the two-stage neural network object detector.
2. The method according to claim 1, wherein the two-stage neural network object detector is updated to minimize the region proposal distillation loss and the previously learned object recognition distillation loss.
3. The method according to claim 2, wherein the two-stage neural network object detector is updated by applying an optimization method to a total loss, the total loss being a combination of: the region proposal distillation loss, the previously learned object recognition distillation loss, a region proposal network loss of the current version of the two-stage neural network object detector, and an object recognition loss of the current version of the two-stage neural network object detector, and the region proposal network loss and the object recognition loss respectively quantify the ability to locate and identify the objects belonging to the additional object class.
4. The method according to claim 3, wherein, if the total loss exceeds a first predetermined threshold or the ability to detect an object belonging to one of the one or more object classes decreases below a second predetermined threshold, another optimization method is applied, or the training is abandoned.
5. The method according to claim 3, wherein the optimization method is the stochastic gradient descent method applied to the total loss.
6. The method according to claim 1, wherein the training uses training images, and validation images different from the training images are used to evaluate the global detection accuracy criterion.
7. The method according to claim 1, further comprising: Additionaltraining the two-stage neural network object detector to locate and identify objects belonging to another additional object class by iteratively additionally updating the two-stage neural network object detector until another global detection accuracy criterion is met, wherein the additional updating is performed to minimize a loss in the ability to localize and identify the objects belonging to the one or more previously learned object categories and the other additional object category, while maximizing an additional ability to localize and identify the objects belonging to the other additional object category.
8. The method according to claim 1, wherein the two-stage neural network object detector is trained concurrently to localize and identify the objects belonging to the additional object category and the objects belonging to another additional object category.
9. The method according to claim 1, further comprising at least one of the following: receiving a user request for training the two-stage neural network object detector to localize and identify the objects belonging to the additional object category, and retrieving from the Internet training images including the objects belonging to the additional object category.
10. A method for incremental learning of object detection without catastrophic forgetting of previously learned object categories, the method comprising: receiving training images including objects belonging to an object category unknown to an initial version of a two-stage neural network object detector capable of detecting objects belonging to at least one previously learned object category; and using a plurality of images one by one in the training images to train the two-stage neural network object detector to detect the objects of the initially unknown object category until a predetermined condition is satisfied by: inputting one image of the plurality of images into the initial version of the two-stage neural network object detector to obtain a first region proposal and a first object prediction for the at least one previously learned object category, inputting the one image into a current version of the two-stage neural network object detector to obtain a second region proposal and a second object prediction for the at least one previously learned object category and the initially unknown object category, comparing the first region proposal with the second region proposal to estimate a region proposal distillation loss quantifying a reduction in the ability to localize the objects of the at least one previously learned object category, comparing the first object prediction with the second object prediction to estimate an object recognition distillation loss quantifying a reduction in the ability to identify the objects of the at least one previously learned object category, comparing the second region proposal with a ground truth label of the one image to estimate a region proposal network loss for the initially unknown object category, comparing the second object prediction with the ground truth label to estimate an object recognition loss for the initially unknown object category, calculating a total loss combining the region proposal distillation loss, the object recognition distillation loss, the region proposal network loss, and the object recognition loss, and updating the current version of the two-stage neural network object detector to minimize the total loss, wherein the predetermined condition is satisfied when the number of training iterations reaches a predetermined number or when a total loss reduction rate is lower than a predetermined threshold.
11. A computer-readable medium comprising computer-readable code that, when read by a computer, causes the computer to perform a method for incremental learning for object detection in an image without catastrophically forgetting one or more previously learned object classes, the method comprising: Training the two-stage neural network object detector to locate and identify objects in the image belonging to an additional object class by iteratively updating the two-stage neural network object detector until an overall detection accuracy criterion is met, wherein Performing the update to balance minimizing the loss of the initial ability to locate and identify objects belonging to the one or more previously learned object classes and maximizing the ability to additionally locate and identify objects belonging to the additional object class, and Evaluating whether the overall detection accuracy criterion is met includes: Determining a region proposal distillation loss based on a comparison of region proposals output by an initial version of the two-stage neural network object detector with current region proposals output by a current version of the two-stage neural network object detector, wherein the region proposal distillation loss quantifies a reduction in the ability to locate objects belonging to previously learned object classes; and Determining a previously learned object recognition distillation loss based on a comparison of previously learned object predictions obtained by the initial version of the two-stage neural network object detector with object predictions output by the current version of the two-stage neural network object detector.
12. The computer-readable medium of claim 11, wherein The two-stage neural network object detector is updated to minimize the region proposal distillation loss and the previously learned object recognition distillation loss.
13. The computer-readable medium of claim 12, wherein The two-stage neural network object detector is updated by applying an optimization method to a total loss that is a combination of: The region proposal distillation loss, The previously learned object recognition distillation loss, The region proposal network loss of the current version of the two-stage neural network object detector, and The object recognition loss of the current version of the two-stage neural network object detector, and The region proposal network loss and the region convolutional neural network loss respectively quantify the ability to locate and identify the objects belonging to the additional object class.
14. The computer-readable medium of claim 13, wherein If the total loss exceeds a first predetermined threshold or the ability to detect an object belonging to one of the one or more object classes decreases below a second predetermined threshold, another optimization method is applied, or the training is abandoned.
15. The computer-readable medium of claim 13, wherein The optimization method is the stochastic gradient descent method applied to the total loss.
16. The computer-readable medium of claim 11, wherein The training uses training images and uses validation images different from the training images to evaluate the overall detection accuracy criterion.
17. The computer-readable medium according to claim 11, wherein, The method further includes: Additional training of the two-stage neural network object detector to locate and identify objects belonging to another additional object class by iteratively and additionally updating the two-stage neural network object detector until another overall detection accuracy criterion is met. Wherein the additional updating is performed to minimize the loss of the ability to locate and identify the objects belonging to the one or more previously learned object classes and the other additional object class, while maximizing the additional ability to locate and identify the objects belonging to the other additional object class.
18. The computer-readable medium of claim 11, wherein The two-stage neural network object detector is simultaneously trained to locate and identify the objects belonging to the additional object class and the objects belonging to another additional object class.
19. The computer-readable medium of claim 11, wherein the method further comprises at least one of the following: Receiving a user request to train the two-stage neural network object detector to locate and identify the objects belonging to the additional object class, and Retrieving from the Internet training images including objects belonging to the additional object class.
20. A computer-readable medium comprising computer-readable code which, when read by a computer, causes the computer to perform a method for incremental learning in object detection in images without catastrophic forgetting of previously learned object classes, the method comprising: Using a plurality of images one by one in the training images to train a two-stage neural network object detector to detect the objects of the initially unknown object class until a predetermined condition is met by: Inputting one image of the plurality of images into an initial version of the two-stage neural network object detector to obtain a first region proposal and a first object prediction for at least one previously learned object class, Inputting the one image into a current version of the two-stage neural network object detector to obtain a second region proposal and a second object prediction for the at least one previously learned object class and the initially unknown object class, Comparing the first region proposal with the second region proposal to estimate a region proposal distillation loss that quantifies the reduction in the ability to locate the objects of the at least one previously learned object class, Comparing the first object prediction with the second object prediction to estimate an object recognition distillation loss that quantifies the reduction in the ability to recognize the objects of the at least one previously learned object class, Comparing the second region proposal with the ground truth label of the one image to estimate a region proposal network loss for the initially unknown object class, Comparing the second object prediction with the ground truth label to estimate an object recognition loss for the initially unknown object class, Calculating a total loss that combines the region proposal distillation loss, the object recognition distillation loss, the region proposal network loss, and the object recognition loss, and Updating the current version of the two-stage neural network object detector to minimize the total loss. Among them, when the number of training iterations reaches a predetermined number, or when the total loss reduction rate is lower than a predetermined threshold, the predetermined condition is satisfied.
21. A system for incremental learning of object detection without catastrophic forgetting of previously learned object classes, the system comprising: A computing device having a user interface and configured to use a two-stage neural network object detector to locate and identify objects and communicate with a cloud server; And The cloud server is configured to collect training images, store an initial version of the two-stage neural network object detector that locates and identifies previously learned object classes, and communicate with the computing device and the Internet, Wherein At least one of the computing device and the cloud server is configured to train the two-stage neural network object detector to locate and identify objects belonging to additional object classes by iteratively updating the two-stage neural network object detector that can locate and identify objects belonging to previously learned object classes until a general detection accuracy standard is met, Performing the update to balance minimizing the loss of the initial ability of the two-stage neural network detector to locate and identify objects belonging to the one or more object classes and maximizing the ability of the two-stage neural network detector to additionally locate and identify objects belonging to the additional object classes, and Evaluating whether the general detection accuracy standard is met includes: Determining a distillation loss of region proposals based on a comparison between region proposals output by an initial version of the two-stage neural network object detector and current region proposals output by a current version of the two-stage neural network object detector, wherein the distillation loss of the region proposals quantifies a reduction in the ability to locate objects belonging to previously learned object classes; And Determining a distillation loss of previously learned object recognition based on a comparison between previously learned object predictions obtained from the initial version of the two-stage neural network object detector and object predictions output by the current version of the two-stage neural network object detector.
22. The system according to claim 21, wherein The two-stage neural network object detector is updated by applying an optimization method to a total loss, the total loss being a combination of: The distillation loss of the region proposals, The distillation loss of previously learned object recognition, A region proposal network loss of the current version of the two-stage neural network object detector, and An object recognition loss of the current version of the two-stage neural network object detector, and The region proposal network loss and the object recognition loss respectively quantify the ability to locate and identify the objects belonging to the additional object classes.
23. The system according to claim 22, wherein If the total loss exceeds a first predetermined threshold or the ability to detect an object belonging to one of the one or more object classes decreases to below a second predetermined threshold, another optimization method is applied, or the training is abandoned.
24. The system according to claim 21, wherein The two-stage neural network object detector is trained to locate and identify objects belonging to the other additional object class simultaneously or sequentially with training the two-stage neural network object detector to locate and identify objects belonging to the other additional object class.
Citation Information
Patent Citations
Augmented neural network configuration, training method therefor, and computer readable storage medium
CN107256423A
Group behavior identification method and a device based on a semantic attention reservation mechanism
CN109299657A