Systems and methods for anomaly recognition and detection using a lifetime deep neural network
Lifelong deep neural network (L-DNN) technology addresses the challenge of combining supervised and unsupervised learning by enabling the detection of unknown defects and anomalies in real-time, enhancing the efficiency and flexibility of industrial quality control processes.
Patent Information
- Application Number
- JP2022542288
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-01-12
- Filing Date
- 2021-01-12
- Publication Date
- 2025-06-26
- Estimated Expiration
- 2041-01-12
AI Technical Summary
Conventional artificial neural networks (ANNs) and deep neural networks (DNNs) struggle to effectively combine supervised and unsupervised learning, particularly in scenarios where the number of known classes cannot be defined a priori, such as in real-world quality inspection.
Lifelong deep neural network (L-DNN) technology enables quick recognition of acceptable objects and detection of unknown defects or anomalies without prior training on these unknowns, allowing for high-speed visual inspection and quality control in industrial environments.
L-DNN technology reduces AI latency, minimizes data transfer and network traffic, optimizes memory usage, enhances data privacy, and provides flexible operation, enabling real-time anomaly detection and classification in industrial settings.
Smart Images

Figure 0007699133000002 
Figure 0007699133000003 
Figure 0007699133000004
Abstract
Description
Technical Field
[0001] <Cross - Reference to Related Applications> This application claims the benefit of priority under 35 U.S.C. § 119(e) of U.S. Patent Application No. 62 / 960,132, filed on January 12, 2020, which is hereby incorporated by reference in its entirety for all purposes.
Background Art
[0002] Often, conventional artificial neural networks (ANNs) and deep neural networks (DNNs), which include many layers of neurons intervening between an input layer and an output layer, are typically designed to function either in a supervised manner or an unsupervised manner. Supervised learning is highly accurate but creates a rigid structure that does not tolerate "unknown" classes beyond the classes it was trained to recognize and cannot be easily updated without costly retraining. Unsupervised learning can accommodate and discover unknown input - output mappings but usually does not provide the typical high performance (accuracy / memory) of supervised systems. However, many real - world scenarios often require a combination of supervised ANN capabilities and unsupervised ANN capabilities. In many such scenarios, the number of known classes cannot be defined a priori.
[0003] For example, in real - world quality inspection, the number of potential defects that can occur in a production cycle may not be realistically or feasibly determined a priori. Thus, conventional DNNs trained on known mappings between input data and output classes cannot be used to detect all possible defects in quality inspection. In these scenarios, unsupervised methods (e.g., clustering) can be used to estimate defects, but these methods generally do not achieve performance equivalent to that of supervised DNNs.
[0004] Specifically, a supervised ANN is trained to infer a function from labeled training data (a set of training examples), where each example includes an input object and a desired output label. The use of backpropagation to train a supervised ANN requires large amounts of training data and many learning iterations and cannot be generated in real time as soon as the data becomes available. ANNs that learn in an unsupervised fashion do not require desired output values and are used to preprocess, compress, and search for unknown patterns in the data. Unsupervised learning may take longer than supervised learning, but some unsupervised learning processes are much faster and computationally simpler than the supervised learning process and can occur incrementally in real time, i.e., for each data presentation and with a computational cost that is only slightly more than the cost of inference.
[0005] A supervised DNN is trained to map an input (e.g., an image) to an output (e.g., an image label), where the set of labels corresponds to the number of output classes that are defined a priori and provided in the supervised learning process. The DNN designer needs to provide a priori all the classes that the DNN may encounter during inference (after learning), the training data should be sufficient for slow backpropagation, and should be balanced among some of the examples within the classes that the supervised DNN should identify. If no training examples are provided for the class of a given object, a conventional neural network has no reason to know that this class of object exists.
[0006] This requirement to provide training examples for all object classes means that in traditional neural networks, high accuracy is required and the training data contains a large number of consistent examples of correct products, but very few (inconsistent) examples of incorrect products, making it difficult to implement successfully for problems such as industrial quality control. Furthermore, machine manufacturers may prefer to delay training for customers so that the training incorporates the latest examples for each customer. Therefore, ANN and DNN are not widely adopted for industrial quality control applications because of the concentrated nature of the data and calculations in the training process, their non-updateability once deployed, and the difficulty of recognizing all possible initially unknown defects.
Summary of the Invention
[0007] Unlike traditional ANN and DNN, lifelong deep neural network (L-DNN) technology can be quickly trained to recognize acceptable or conforming objects and can be used to detect those unknown defects and other anomalies without being trained to recognize unknown defects or anomalies (or irregularities). Thus, L-DNN technology is suitable for high-speed visual inspection and quality control in manufacturing assembly lines and other industrial environments. Edge computers that execute L-DNN can use camera data or other data to identify defects or misaligned objects in real time, without human intervention, for removal, repair, or replacement.
[0008] L-DNN technology incorporates (a) the ability of DNNs trained with high accuracy for known classes and, when available, (b) is simultaneously sensitive to any number of unknown classes or class changes that cannot be determined a priori. Furthermore, to handle real-world scenarios, L-DNN technology (c) enables learning to occur when there is little data. L-DNN technology can also effectively learn from imbalanced data (e.g., data with very different numbers of examples for different classes). Finally, L-DNN technology enables (d) real-time learning without slow backpropagation.
[0009] The L-DNN technology extends artificial intelligence (AI), artificial neural networks (ANN), deep neural networks (DNN), and other machine vision processes, enabling training at the time of introduction with minimal unbalanced data and further updates as new data becomes available. The technical benefits of L-DNN include, but are not limited to, the following. ● Reduction of AI latency: Since the L-DNN AI can be located at the computing edge rather than in the cloud or on a remote server, real-time or near-real-time data analysis can be achieved because the data is processed on a processor directly located at the local device level rather than at the cloud or data center resources. ● Reduction of data transfer and network traffic: Since the data is processed at the device level, raw data does not need to be sent from the device to the data center or cloud via a computer network, thereby reducing network traffic. ● Reduction of memory usage: The L-DNN technology enables one-shot learning and eliminates the need to store data for future use. ● Data privacy: The processing of data (both inference and learning) on the device enhances user privacy and data security because user data is generated, processed, and discarded at the device level and is never transmitted from the device and never permanently stored in its raw form on the device. ● Flexible operation: L-DNN can be operated in supervised, unsupervised, or semi-supervised modes. ● Handling of unknown object classes: L-DNN can identify and report "Nothing I know" about objects not encountered during training, including defective or abnormal parts. For details of the L-DNN technology, see, for example, U.S. Patent Application Publication No. 2018 / 0330238A1 and U.S. Patent Application No. 16 / 952,250, which are hereby incorporated by reference in their entirety.
[0010] Unlike conventional ANNs and DNNs, L-DNN can identify abnormal images without being trained to identify their anomalies, making it particularly suitable for industrial visual inspection tasks. In visual inspection, the automated inspection system receives data corresponding to very similar and correct situations (e.g., good quality products on a conveyor belt, normal processing, machines operating within an appropriate regime) most of the time. However, rarely, the input data may be different or there may be anomalies (e.g., defective products, abnormal processing, machines outside the operating regime). Generally, the task of identifying data that is different from or abnormal to correct data can be referred to as anomaly recognition (e.g., recognizing that there is an anomaly such as a 12-pack of aluminum cans containing defective cans) or anomaly detection (e.g., indicating where an anomaly has occurred in a multi-dimensional space, such as which cans in a 12-pack are defective). In anomaly recognition and detection, time can be one of the inputs to the system.
[0011] A real-time operating system implementing L-DNN can learn new products and new anomalies almost immediately so that it can respond to new knowledge almost instantaneously. By using semi-supervised L-DNN, a real-time operating system for anomaly recognition and detection can be enabled to perform the following. a. Perform rapid training at the introduction site using only examples of normal situations. b. Recognize or detect anomalies and alert the user. c. Optionally, update its knowledge with each piece of data with an anomaly without interrupting the operation of the system.
[0012] The L-DNN can be used in a system for inspecting objects on an assembly line. This system may include a sensor and a processor operably coupled to the sensor. During operation, the sensor acquires data representing an object on the assembly line. And the processor executes (1) a pre-trained backbone that extracts features from the data, and (2) a fast learning head trained only on data representing normal objects that receives features from the pre-trained backbone and discriminates whether the object is normal or abnormal based on those features. The system may also include a programmable logic controller (PLC) or a human-machine interface (HMI) that removes the object from the assembly line in response to the fast learning head recognizing that the object is abnormal, or warns the operator of the abnormality in response to an appropriate signal from the processor.
[0013] The sensor may be a camera, in which case the data includes an image of the object. In this case, the fast learning head can be trained with a relatively small number of images (e.g., 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, or even 1 image). The sensor may also be a strain, vibration, or temperature sensor configured to acquire time-series data representing the object or the environment.
[0014] The pre-trained backbone may include convolutional layers of a neural network trained to recognize normal objects. Also, the features can be configured to be extracted by performing a Fourier transform or a wavelet transform on the data.
[0015] The high-speed learning head can identify that an object has an anomaly of a type that has not been seen before. In these cases, an operator can use an interface operably coupled to the processor to label the anomaly of a type that has not been seen before as an anomaly of the first type. When the sensor acquires data of other objects on the assembly line, a pre-trained backbone can also extract features of the data representing those objects for automatic anomaly detection by the high-speed learning head. The high-speed learning head can classify a second object as having an anomaly of the first type, if applicable. If the second object has an anomaly of a type that has not been seen before, the operator can use the interface to label the anomaly as an anomaly of the second type.
[0016] The processor can implement multiple high-speed learning heads, all of which receive features extracted by the same pre-trained backbone and operate on those features. This reduces the latency compared to operating the same number of backbone and head pairs sequentially. Different heads can identify different objects, different images of the same object, or different regions of interest in the same image or dataset from the same set of features extracted by the backbone. The heads can also be trained to do different things or specialize in different tasks. For example, one head can monitor for anomalies while another head can classify objects based on size, shape, position, or orientation from the same extracted features.
[0017] All combinations of the foregoing concepts, and additional concepts discussed in more detail below (assuming such concepts are not mutually inconsistent), are considered to be part of the subject matter of the invention disclosed herein. Specifically, all combinations of the claimed subject matter that appear at the end of this disclosure are considered to be part of the subject matter of the invention disclosed herein. Also, of course, terms explicitly used herein that may appear in any disclosure incorporated by reference should be given the meaning that most closely matches the particular concepts disclosed herein.
[0018] Upon consideration of the following diagrams and detailed description, other systems, processes, and features will become apparent to those skilled in the art. All such additional systems, processes, and features are intended to be included within this description, within the scope of the present invention, and to be protected by the appended claims.
Brief Description of the Drawings
[0019] Those skilled in the art will understand that the drawings are for illustrative purposes primarily and are not intended to limit the scope of the subject matter of the present invention described herein. The drawings are not necessarily to scale, and in some instances, various aspects of the subject matter of the present invention disclosed herein may be shown exaggerated or enlarged in the drawings to facilitate understanding of different features. In the drawings, like reference characters generally mean like features (e.g., functionally similar and / or structurally similar elements).
[0020]
Figure 1
[0021]
Figure 2
[0022]
Figure 3A
[0023]
Figure 3B
[0024]
Figure 3C
[0025]
Figure 4
[0026]
Figure 5A
[0027]
Figure 5B
[0028]
Figure 6A
[0029]
Figure 6B
[0030]
Figure 6C
DETAILED DESCRIPTION OF THE INVENTION
[0031] For years, manufacturers have used automated optical inspection (AOI), or computer vision, to inspect products at various points in the manufacturing process, such as during processing, labeling, and final inspection. Conventional AOI systems, which consist of cameras and software, enable rapid inspection while significantly improving quality and reducing waste. However, one limitation of conventional AOI systems is their inability to detect changes or anomalies (or irregularities) such as surface-level defects, deformations, or perform assembly verification. To counter this limitation, manufacturers may rely on human inspectors. However, human inspectors have their own limitations, including incomplete accuracy and / or human inconsistency. Also, human inspectors are not machines and cannot inspect every single part, especially on high-speed lines.
[0032] To overcome these limitations, engineers have begun combining conventional computer vision inspection with deep learning. Deep learning, a subset of machine learning, enables the execution of calculations of deep neural networks (DNNs). Superficially, computer vision and deep learning appear similar. That is, both offer several advantages, including automating the visual inspection process, increasing inspection speed, reducing line-to-line variation, and increasing the proportion of products that can be inspected efficiently and scalably. However, there are some differences between these approaches. These differences are characterized by the type of problem each is best suited to solve. Put simply, computer vision is suitable for objective or quantitative use cases that can be logically programmed (e.g., if a part measures 100 ± 0.1 mm, that's fine), while deep learning is better suited for subjective or qualitative inspections involving generalization or variability (e.g., that the weld is good).
[0033] Computer vision can work well for visual inspection when it can apply simple logic-based rules to evaluate the fitness of objects with desired criteria. For example, a computer vision system can verify how, where, and in what orientation components and / or elements are oriented or located within the field of view (e.g., for assistance during or as a precursor to automatic picking for secondary inspection such as the following component measurements). A computer vision system can also verify the presence or absence of critical components or the integrity of components / parts (e.g., whether a part is malfunctioning or a component is missing). A computer vision system can measure or estimate the distance between points within a part or component and compare that measurement to specifications or tolerances for pass / fail (e.g., to confirm the dimensions of critical parts after milling / molding and before further assembly). Also, a computer vision system can extract data from product labels (e.g., barcodes or Quick Response (QR) codes for identification of parts / components, text from labels, etc.).
[0034] Because the strengths of deep learning are generalizable, deep learning inspection systems are suitable for inspections in various situations. For example, a deep learning inspection system can identify changes in the type or location of products, parts, or defects such as changes in surface scratches of metals or plastics, welding quality, products to be packaged, etc. A deep learning inspection system can also inspect items in environments with changes such as lighting, reflection, color, etc. (A computer vision system generally has poor performance in environments with variable lighting and generally cannot evaluate changes by objects.)
[0035] Conventional deep learning inspection systems must be trained to recognize the objects to be inspected. The training needs to be performed using data such as labeled or tagged images that are very similar to what the deep learning model will actually process during production introduction. This means that the training data for the deep learning inspection system should be acquired with the same camera, the same viewpoint, and the same spatial resolution, and ideally, with the same lighting as that used during production introduction, although some variation is generalized by the system.
[0036] Furthermore, the training data needs to include approximately an equal number of labeled images for each type or class of object that the deep learning inspection system is to identify. For example, if the deep learning inspection system is to classify "good" objects and "bad" objects, the training set should have approximately equal numbers of examples (labeled images) of objects in each of these classes. The amount of training data required to train a deep learning inspection system depends on several variables, including the desired accuracy, variability, and number of classes of objects. Generally, higher accuracy, higher variability, and more classes translate to more training data. It is not uncommon for the training data to include over 1,000 labeled images per class.
[0037] However, in industrial and manufacturing use cases, it can be difficult or impossible to generate sufficient training data to train deep learning systems for defect recognition for at least two reasons. First, it is difficult to generate a sufficient number (e.g., 2,000) of images of defective products. This is because modern manufacturing processes have significantly improved yields, reducing errors and waste. Procuring a balanced and representative training dataset is itself an expensive proposition. Second, it can be difficult or impossible to identify all types of recognizable defects. While some types of defects are predictable, other types may be unpredictable because the conditions that cause them are unknown in advance. The lack of data (labeled images) for different defects makes it difficult, if not impossible, to train traditional deep learning inspection systems to recognize defects (of different types).
[0038] Traditional deep learning inspection systems classify input data into one of the classes they have been taught to recognize. Training traditional deep learning inspection systems on unbalanced training datasets (e.g., those that contain few or no labeled images of defective parts) can result in a lack of bias and accuracy when the system makes decisions regarding classifying the input data. As a result, traditional deep learning inspection systems trained only on good / normal data may be able to identify abnormal objects as good or normal, albeit with likely lower reliability.
[0039] Unlike conventional deep learning inspection systems, an L-DNN-based inspection system can identify different types of defects and other anomalies without being trained to specifically recognize them. Further, an L-DNN-based inspection system can be trained to recognize normal or good objects on a smaller training dataset than conventional deep learning inspection systems. Also, unlike conventional deep learning inspection systems, an L-DNN-based inspection system can learn to recognize new objects, including objects not previously seen, after being introduced and while running on a computing edge (e.g., on a processor coupled to a factory floor inspection camera).
[0040] In summary, an L-DNN-based inspection system can perform anomaly recognition. Anomaly recognition functions by training the model with only "good" image data (an unbalanced training dataset). By establishing a threshold for deviation from this "good" data, the system can identify images that deviate from the norm. Over time, as the system identifies and captures an increasing number of images with anomalies, the system can, with the help of an operator, label these images to form new anomaly classes and transition from anomaly recognition to more detailed anomaly classification.
[0041] The L-DNN, also referred to as the "Brain" or model that provides the artificial intelligence for the L-DNN-based inspection system, is created and optimized on a training system (e.g., a local or cloud server running the Neurala Brain Builder Software-as-a-Service (SaaS) product) and can be trained on images or other data representing only acceptable or normal objects. Next, the trained L-DNN can be introduced into one or more nodes, including an edge computer connected to a camera, microphone, meter, thermometer, or other sensor that acquires input data. For example, each edge computer may be connected to a camera compliant with the GigE Vision standard (e.g., Basler, Baumer, or Teledyne DELSA cameras). The training system is easily integrated into industrial machinery or manufacturing lines using standard interfaces such as PLC I / O and Modbus TCP.
[0042] As discussed in more detail below, the training system can train the L-DNN to recognize one or more categories or classes of "good" input data. For example, the L-DNN can be trained on images of desired surface quality, packaging, or kitting. Since the L-DNN is trained only on "good" images, the introduced L-DNN-based inspection system identifies an image as having an anomaly if it does not sufficiently match a good image.
[0043] The L-DNN-based inspection system offers several technical advantages over conventional automated inspection systems. First, it can recognize anomalies without being trained on data representing anomalies. This eliminates the time-consuming, and often impractical or difficult, task of collecting images that depict defective, unacceptable parts. Second, it can be trained to recognize acceptable parts with a smaller training data set than conventional deep learning inspection systems. This training data set may be imbalanced and, for example, may consist only of data representing acceptable parts and not defective or non-acceptable parts. This makes training faster, easier, and less expensive. Faster and easier training makes it easier to adapt the L-DNN-based inspection system to inspect different products. Third, as the L-DNN-based inspection system operates, it can learn to more accurately identify (and optionally classify with human supervision) different types of defects. Fourth, since the L-DNN is executed on an edge device, there is no need to send data to other devices. This reduces network bandwidth consumption and keeps the data private and secure.
[0044] <Lifetime Deep Neural Network (L-DNN) Technology> In the L-DNN technology, a subsystem based on a richly expressive DNN (Module A) is combined with a fast learning subsystem (Module B) so as to achieve fast yet stable learning of features representing entities or events of interest. These feature sets can be pre-trained by a slow learning methodology such as backpropagation. In the case of DNN-based scenarios described herein (descriptions regarding other features are possible by adopting non-DNN methodologies for Module A), the high-level feature extraction layer of the DNN functions as an input to the fast learning system of Module B to classify known entities and events and add knowledge of unknown entities and events on-the-fly. Module B can learn important information, capture descriptive and highly predictable features of correct inputs or behaviors without the drawbacks of slow learning, and identify previously unseen incorrect inputs and behaviors.
[0045] Also in L-DNN, it becomes possible to learn new knowledge without forgetting old knowledge, thereby reducing or eliminating catastrophic forgetting. In other words, the present technology enables a machine that operates in real time, such as a computer, smartphone, or other computing device, to adjust its behavior at the edge (continuously and / or intermittently) based on user input without (a) the need to transmit or store input images, (b) time-consuming training, or (c) significant computational resources. Learning after the introduction of L-DNN enables a machine that operates in real time to adapt to changes in its environment and user interactions, address deficiencies in the original dataset, and provide a customized experience to each user. In the case of anomaly recognition or detection, the system is particularly useful because it can start from a highly imbalanced dataset of extreme cases without examples of anomalies and discover and add anomalies to the knowledge of what constitutes a good product or correct behavior during operation.
[0046] Here, the L-DNN technology is used for monitoring manufacturing lines and packaging lines. One or more different versions of L-DNNs or brains can be created and introduced using L-DNN creation software, such as Neurala Brain Builder, each of which monitors data from at least one sensor. This sensor data can be one-dimensional (e.g., time-series data from a sensor monitoring the movement of a rotating cam) or two-dimensional (e.g., an image from a camera taken before sealing a package photo). In either case, the feature extractor of the L-DNN (implemented as module A, e.g., a feed-forward DNN, recurrent DNN, fast Fourier transform (FFT) module, or wavelet transform module) extracts features from the data. The module B classifier classifies the extracted features as normal or abnormal. If the extracted features are abnormal (e.g., the machine is wobbling outside the operating limit range or the package contains broken components), the L-DNN-based quality monitoring system can trigger some corrective action (e.g., slow down or stop the machine or remove the package from the packaging line before it is sealed).
[0047] The L-DNN implements a heterogeneous neural network architecture to combine a high-speed learning mode and a low-speed learning mode. In the high-speed learning mode, a real-time operating machine implementing the L-DNN learns new knowledge and new experiences (e.g., anomalies) so that it can respond to new knowledge almost immediately. In this mode, the learning rate of the high-speed learning subsystem is set high enough to work favorably for new knowledge and corresponding new experiences, while the learning rate of the low-speed learning subsystem is set to a low value or zero to preserve old knowledge and corresponding old experiences. The high-speed learning mode is the main operating mode of the anomaly detection system for industrial inspection. This is because it may be costly to take the system offline for slow training updates.
[0048] Figure 1 provides an overview of the L-DNN 106 that can be implemented in a machine 100 operating in real time. The machine 100 operating in real time receives sensory input data 101, such as video or audio data, from sensors 130, such as cameras or microphones, and supplies it to an L-DNN 106 that includes a slow learning module A 102 and a fast learning module B 104, which can be implemented as a neural network classifier. Each module A 102 can be based on a pre-trained (fixed-weight) DNN 110 or some other system and functions as a feature extractor. Module A 102 receives the sensory input 101. A stack of pre-trained neural network layers 112, also called a backbone 112, identifies relevant features, extracts them into a compressed representation of the object 115, and supplies these representations 115 to module B 104. Often, when using a DNN-based module A, features are read from some intermediate layers 114 of the DNN at the end of the feature extraction backbone 112. The upper layers of the DNN, which may include one or more fully connected, averaging, pooling layers 116 and a cost layer 108, may be ignored or completely removed from the DNN 110.
[0049] Module B 104 rapidly learns these object representations 115 by forming an association between the signals of the input layer 120 and the output layer 124 within a set of one or more associative layers 122. In the supervised mode, through interaction with the user, module B 104 receives the correct label for an unknown object at layer 124 and rapidly learns the association between each feature vector of layer 120 and the corresponding label, and as a result, can immediately recognize these new objects. In the unsupervised mode, the system uses the similarity of features between vectors to assign these vectors to the same class or different classes.
[0050] L-DNN 106 takes advantage of the fact that DNN 110 is an excellent feature extractor. Module B104 continuously processes the features extracted by module A102 when the input source (sensor) 130 provides data 101. Module B104 uses fast one-shot learning to associate these features with object classes.
[0051] In the fast learning mode, when a new set of features corresponding to normal (not abnormal) cases is presented as input 101, module B104 associates these features with class labels that are given by the user in the supervised mode or generated internally in the unsupervised mode. In either case, module B104 knows this input at this time and can recognize it in the next presentation. This classification by module B104 functions either as the output of L-DNN 106 itself in terms of the work that L-DNN 106 is performing or as a combination with the output from a specific DNN layer from module A102.
[0052] Since general L-DNN and specific module B are designed to operate in real time on continuous sensory inputs, the neural network of module B should not be confused by unknown objects. Conventional neural networks target a given dataset that usually includes labeled objects in the input. As a result, when there are no known objects, there is no need to handle the input. Therefore, to use such a network in module B of L-DNN, an additional special category of "knowing nothing" should be added to the network to reduce the attempts (false positives) of module B to misclassify unknown objects as known.
[0053] This concept of "unknown" is particularly useful when processing raw sensory streams that can (exclusively) include unlabeled objects that have not been seen before. This allows module B and L-DNN to potentially identify unknown objects as "unknown" or "not previously seen" instead of misidentifying them as known objects. Extending the conventional design with the "unknown" implementation can be as simple as adding a bias node to the network. "Unknown" can also be implemented in a version of L-DNN that automatically increases or decreases its effect based on the number of known object classes and their corresponding activations. Since the L-DNN system can raise an anomaly flag 118 in response to an "unknown" classification that warns the user every time it does not recognize its input as something it knows, "unknown" is very useful for anomaly detection applications.
[0054] For details of L-DNN, see U.S. Patent Application Publication No. 2018 / 0330238A1, entitled "Systems and Methods to Enable Continual, Memory-Bounded Learning in Artificial Intelligence and Deep Learning Continuously Operating Applications Across Compute Edges", and PCT Publication No. WO2019 / 226670, entitled "Systems and Methods for Deep Neural Networks on Device Learning (Online and Offline) with and without Supervision", which are hereby incorporated by reference in their entirety for all purposes.
[0055] <Industrial Inspection and Monitoring Using the L-DNN Platform> Figure 2 shows a Visual Inspection Automation (VIA) platform 200 based on L-DNN technology as an example of a system that enables visual inspection on the factory floor. The VIA platform 200 includes software executed on one or more processors having a graphical user interface (GUI) that enables a user to configure, train, deploy, and further refine the L-DNN system. The VIA platform 200 includes a local industrial PC (local server) 202 and L-DNN training (e.g., Brain Builder) software that is executed on a laptop 201 or on a computer connected to an industrial PC 202 for initial setup and training. The VIA system 200 also includes edge computers 204. Each edge computer 204 is coupled to a corresponding camera 208 or to other sensors such as a microphone, counter, timer, or pressure sensor that collects data for evaluation or inspection to form a corresponding node 209. Each node 209 may include different types of sensors and may implement an L-DNN trained to recognize normal and abnormal data acquired by the sensors.
[0056] The VIA system 200 is integrated with industrial machines or manufacturing lines using standard interfaces such as programmable logic controllers (PLCs) 210, I / O, and Modbus TCP via a switch 206 and may include one or more human-machine interfaces (HMIs) 212. In Figure 2, the HMI 212 is connected to the PLC 210. Customer-defined logic is implemented in the PLC 210 and / or the HMI 212 to perform any actions resulting from the output of the VIA system 200. For example, if the VIA system 200 recognizes an instance of a product as abnormal, the PLC 210 can serve as a trigger to automatically remove this instance from the production line.
[0057] The local server 202 and the L-DNN training software running on the laptop 201 build and deploy the brain to the edge computer 204. The edge computer 204 runs software (referred to as an inspector) that uses the trained brain, which can consider a set of parameters (e.g., neural network weights) representing the trained L-DNN to process data from the camera 208 and / or other sensors. Once the brain is introduced, these can be used to recognize and identify anomalies in parts or other objects being manufactured, sorted, and / or packaged within the factory. They can also communicate with each other and share information about newly detected types of anomalies via the network switch 206 (e.g., as disclosed in U.S. Patent Application Publication No. 2018 / 0330238A1, which is hereby incorporated by reference in its entirety).
[0058] During operation, the L-DNN-based VIA system 200 of FIG. 2 can inspect components, sub-assemblies, and / or final products in a synchronized manner determined by the cycle time of the production process. The cycle time is the time to manufacture a part divided by the number of parts required in that time interval. In theory, if the requirement is for 60 parts in one hour, the cycle time is one minute. The L-DNN inference time (the time it takes for the L-DNN to classify an image as normal or abnormal) must be shorter than this cycle time.
[0059] In FIG. 2, the VIA system 200 inspects normal objects 21 as well as abnormal objects 23 and 25 that move past camera 208 on the conveyor belt of inspection line or production line 220. Camera 208 acquires still or video images of the objects and transmits them via network switch 206 to respective edge computers 204 that implement respective L-DNNs. These L-DNNs are trained to recognize normal objects 21, but do not recognize abnormal objects 23 or 25. When an L-DNN is presented with an image of a normal object 21, the image is classified as an image of a normal object 21 (e.g., a "good" image). When an L-DNN is presented with an image of a deviant or abnormal object 23 or 25, the image indicates that it is a deviant or abnormal object.
[0060] The L-DNN can be trained to recognize images of any type of object or patterns of other types of data including spectral data, audio data, or other electrical signals. For example, the L-DNN can be trained to recognize bread or other processed foods at different points in the baking process. The first camera 208 may be positioned to image a mass of bread before it is baked, for example, to ensure that the bread is properly filled, has a correct surface (e.g., no air bubbles against air bubbles), is properly baked, and is properly proven. The second camera 208 can image the bread as it comes out of the oven, for example, on a conveyor belt running through the oven. The L-DNN on the edge computer 204 coupled to the second camera 208 can inspect the image of the baked bread for color (to determine doneness), distribution of materials sprinkled on top (e.g., chocolate chips, raisins, sesame seeds, etc.), and shape.
[0061] L-DNN is particularly suitable for such mainly subjective inspections that cannot be effectively implemented in conventional computer vision systems. Also, since the possible types of anomalies are almost infinite, it can be difficult or impossible to train conventional deep learning networks to recognize all of them. Manual inspection is time-consuming and usually requires several pans. Then, it is inconsistent and may be difficult to achieve the desired uniformity. In contrast, the VIA system can automatically inspect all clusters, resulting in higher yields, better consistency, and waste reduction, and can identify anomalies without being trained to recognize them.
[0062] The VIA system 200 of FIG. 2 may also be configured to inspect different manufactured products, including injection molded plastic parts. Plastic molded parts are used in a variety of manufactured products, from household appliances to automobiles and beyond. Contract manufacturers producing plastic molded parts may mold and deburr the parts before bulk shipping them to customers. Conventionally, a person can inspect one out of, for example, every ten plastic molded parts for fragments of these parts. However, this only generally guarantees that good parts are being sent. Customers manually inspect all parts before using them in assembly.
[0063] The GigE camera 208 installed on the assembly line images the parts after deburring. The camera 208 sends the images to the edge computer 204 directly or via the network switch 206. The L-DNN executed on the edge computer 204 classifies the parts imaged in the images as either normal (properly shaped and deburred) or abnormal (not properly shaped and / or not deburred). If a part is abnormal, the edge computer 204 can trigger a switch (not shown) on the conveyor belt to the PLC via a gateway (e.g., from Modbus TCP to OPC integrated architecture) to indicate defective parts and remove them from the line. A human operator can inspect the defective parts and determine which parts are to be discarded and which are to be reworked.
[0064] The L-DNN-based VIA system 200 is particularly suitable for inspecting injection-molded parts performed on an assembly line where the part types are frequently changed, for example, for prototypes or small quantities. This is because the L-DNN can be quickly trained on a small training set to recognize new parts. Furthermore, the L-DNN can flag non-standard or abnormal parts without having to recognize them beforehand.
[0065] Using the L-DNN-based VIA system 200, continuous data and / or 1D data may be monitored. For example, the VIA system can monitor continuous or sample temperature, strain, or vibration data for patterns indicating that a particular part or component is worn, or should be maintained or replaced. The feature extractor may take the fast Fourier transform or wavelet transform of the 1D vibration sensor data and extract different spectral components or wavelets from that data for analysis by a classifier, which recognizes a particular spectral or wavelet distribution as normal and other spectral or wavelet distributions as abnormal or irregular. In response to the detection of an abnormal spectral or wavelet distribution, the classifier sets a maintenance alert for human intervention.
[0066] <Training of L-DNN for VIA System Introduction> Unlike conventional deep learning networks, the L-DNN can be trained quickly on small, unbalanced datasets. Further, unlike conventional deep learning networks that learn through slower backpropagation, the L-DNN can learn to quickly recognize new objects without forgetting how to recognize previously learned objects. Conventional deep learning networks are also prone to catastrophic forgetting, or lack of recognition of objects they were trained to recognize. This makes it easier and faster to reconfigure the L-DNN inspection system than deep learning network inspection for inspecting new or different objects.
[0067] Figures 3A - 3C show an exemplary workflow of the VIA platform system 200 of FIG. 2, broken down into three stages of user interaction. This workflow is a process for building and introducing a brain (trained L - DNN) to the edge computer 204. Before investing time and effort in mass - producing the hardware setup shown in FIG. 2, the user should prototype the planned system using the brain builder component of the VIA system 200 and prove its feasibility in the "Prototype and Prove" stage 300. This stage 300 shown in FIG. 3A involves creating and introducing a prototype brain using the local computer 202 and / or the local server 204 and includes the following steps.
[0068] In the first step 302 of the "Prototype and Prove" stage 300, a person sets up a camera and / or other sensors to acquire raw data from the assembly line / production environment. When acquiring image data with a camera, the person sets up the image capture environment to ensure proper lighting, focus, and quality of the images captured by the camera 208 (FIG. 2). Once the environmental parameters are set, the local computer 202 and / or the local server 204 can collect images or other data directly from the camera 208 and / or other sensors. For example, these images may be from a camera monitoring a single production line in a factory with multiple production lines.
[0069] The brain builder training software executed on the local computer 202 and / or the local server 204 creates feedback in the form of predictions when collecting data, and notifies the user of the learning progress of the L-DNN and the success probability of the use case. It trains the head of the L-DNN (the backbone or feature extractor can be pre-trained). When the user sees the consistency of the predictions (for example, after about 10 "normal" images are collected), it is possible to evaluate and analyze how well the L-DNN, also called the "brain", is learning (step 304). At this point, the user can test the brain against images of normal and / or abnormal objects (step 306) and adjust the training parameters accordingly. The user can also optionally use the training workspace to apply regions of interest (described later) (308), retraining (310), and / or re-evaluate the brain (304) with more inspections (306) as desired.
[0070] When the brain scores acceptable inspection results, the workflow can proceed to the training, fine-tuning, and introduction phase 320. (The "prototype and proof" phase 300 may be only for proof of concept, and thus may be the first training in this phase 320.) This phase 320 involves creating, training, and introducing the brain (L-DNN) to the edge computer 204. First, the user uses the same live view and collection function 302 as in the "prototype and proof" phase 300 to collect a larger and more extensive dataset of normal / good images. These images are automatically annotated as good / normal and added to the training dataset during collection. If the prototype scenario is the same as or similar to the actual scenario, the brain builder can automatically annotate the images based on the prototype creation session (and most are likely to be correct) (human supervision can provide additional certainty).
[0071] Once appropriate data has been collected, the user can evaluate (322) and test the trained L-DNN. To do this, the user may add examples of "anomalies" available in the validation set used for testing and run the L-DNN on the validation set to obtain an accuracy score. The user can then adjust any parameters related to the effectiveness threshold and perform optimization 324 on the endpoints before the introduction of the brain 326. The effectiveness threshold is the distance between what is considered normal and what is considered anomalous. The effectiveness threshold slider allows the user to adequately adjust the specific point at which the object deviates from normal to be considered anomalous.
[0072] Training can be relatively fast. In one real-world scenario, Module A extracted features from 1D vectors for classification by Module B (the classifier), which detected abnormal changes in sensory readings for constant real-time monitoring. Module A (the feature extractor) used the fast Fourier transform to convert 1D time-series data from a single sensor from the time domain to the spectral domain. For classification, the resulting spectrum was fed to Module B, creating two cases of training on 100 examples of normal operation and predicting on 100 examples of normal operation, 100 anomalous examples where anomalies were classified as normal operation (false positives), and 8 cases where normal operation was classified as anomalous (false negatives). The other examples in this trial were correctly classified. Module B has a dominant parameter (discussed below) that allows the user to fine-tune the balance between false positives and false negatives to match the user's preferences.
[0073] The L-DNN can be trained with thousands fewer samples than conventional deep learning networks. In one case, the L-DNN was trained with 5 samples per class, as opposed to 10,000 samples per class for an equivalent conventional deep learning network. On average, the L-DNN can be trained with 20 - 50 samples per class. More generally, fewer than 100 normal samples (e.g., 100, 90, 80, 70, 60, 50, 40, 30, 20, 10, 5, or one normal image) may be sufficient to train the head of the L-DNN. Under restricted conditions such as those found in manufacturing inspection, where the object, lighting, and camera angle are considered the same between images, the L-DNN can be trained with as few as one sample per class.
[0074] In the final stage, the inference behavior, recipe, and workflow stage 340, the L-DNN system (brain) is connected to the actions of the quality management system. Once the user has achieved the accuracy goal, the action called the inference behavior is assigned to the inference by the trained L-DNN that estimates whether the object is normal or abnormal (342). For example, the inference behavior may be to remove abnormal (defective) parts from the assembly line by triggering a pin via the PLC210. Generally, the inference may not include actions for normal part / operating conditions or corrective measures (e.g., warning maintenance, stopping the machine, or pushing the part / package out of the inspection line by triggering a pin via the PLC210) in response to the detection of an abnormality by the L-DNN / VIA system 200. The inference behavior can be assigned to the calculation node 209 via the recipe (344), which can be selected and executed via the HMI212.
[0075] Figure 4 shows the relationship between Project 400, the dataset, Inference Behavior 410, Recipe 420, and the use of the Brain (trained L-DNN). Project 400 has a group of datasets 402 that can include sample data such as images of normal objects used to train a model (Brain) 404. A copy or version of the Brain is introduced to an endpoint (e.g., edge computer 204 accessible via an application programming interface (API)). Each Brain version 404 (Brain v.X) is associated with Inference Behavior 410. Inference Behavior 410, or simply Behavior, is a group of one or more actions assigned to Brain version 404, generated by assigning relationships between one or more recognized objects and one or more pins to one or more classes of abnormal objects within input / output transmitter 212 used to communicate with PLC 210. When a Brain version recognizes an object of a particular class or flags an abnormal object, the corresponding edge node 204 activates HMI 212 and causes PLC 210 to cause processing on or by production line 220 (e.g., excluding abnormal parts), such as automatic path change of defective parts, triggering of an alarm, or stoppage of the production line.
[0076] Recipe 420 is the collection or succession of Behavior 410 assigned to one or more nodes 209 (e.g., nodes 209-1 and 209-2 in Figure 4). The user can create Recipe 420 by defining the succession of Behavior 410 executed by different compute nodes 209, each of which is assigned one Behavior that correlates its Brain. Additional workflows can be set up to save inference images for review and send them to training data.
[0077] <Multi-Head L-DNN for Industrial Visual Inspection> In order to reduce the inference time when inspecting images from multiple cameras or multiple regions of interest (ROIs) in a single image, the L-DNN can be implemented as a multi-head L-DNN with a single module A having a pre-trained slow-learning backbone that feeds multiple module Bs (classifiers or heads). The multi-head L-DNN can analyze images from multiple inspection points in a single assembly, inspection, or production line and / or multiple ROIs in a single image. Similar to the single-head L-DNN, the multi-head L-DNN may be trained to operate on other types of data, including audio data and signals from meters, thermometers, counters, and other sensors.
[0078] The multi-head L-DNN takes advantage of the fact that the characteristics of the data are consistent within the sensory region. For example, an image can exhibit edges, corners, color changes, etc. Similarly, acoustic information captured by different microphones in the same environment can be decomposed into features based on the correlation frequencies and amplitudes, regardless of the sound source. Therefore, it is not necessary to train multiple module As for the same region. Because they all produce qualitatively similar feature vectors. For each region (sometimes for each accuracy or speed requirement), a single module A can extract the features for processing by one module B for each sensor. Considering that the processing time of module A constitutes more than 75% of the total processing time of the L-DNN, the multi-head L-DNN enables significant reduction in processing costs for inspection systems with multiple sensors within each region.
[0079] Even if the images come from different cameras at different inspection points, they can all be processed through the same module A for feature extraction and then proceed to different module Bs for classification and final analysis. The multi-head L-DNN inherits all the advantages of the L-DNN, such as reduction of AI waiting time, reduction of data transfer and network traffic, reduction of memory usage, enhancement of data privacy and protection, flexible operation, and the ability to handle unknown object classes.
[0080] The multi-head L-DNN adds the following benefits or advantages to the single-head L-DNN. First, the multi-head L-DNN can support multiple inspection points with a single computing host connected to multiple cameras. Second, the multi-head L-DNN can support multiple regions of interest (ROIs) in a single image with a single computing host connected to a single high-resolution camera with a wider-angle lens. Third, the multi-head L-DNN can execute multiple models in parallel on the same data (image) input (e.g., one model for predicting orientation, a model for predicting product variants, and a model for identifying defects). These multi-head L-DNNs reduce the computational and camera hardware costs required to perform complex visual inspection tasks using AI, while the inference latency is within the tact time range.
[0081] FIG. 5A shows a multi-head L-DNN 506 running on a machine 500 operating in real time, such as the edge computer 204 (FIG. 2). The multi-head L-DNN 506 includes a single module A 502 created from the original DNN 510. Module A 502 includes a stacked neural network layer 512, also called a backbone, and ends with a feature extraction layer 514. The remaining parts of the DNN 510, such as one or more pooling layers and fully connected layers 516 and a cost layer 518, are used to pre-train the backbone 512. They do not participate during L-DNN operation. The backbone 502 is trained to extract features 515 from input data 501 from one or more sensors 530, such as a GigE camera 208, or other devices described above with respect to FIG. 1.
[0082] The backbone 512 of module A502 sends the extracted features to several heads 504 (module B), and each head is trained on a different dataset to recognize different features. Similar to module B of a single-head L-DNN, each head 504 rapidly learns these object representations or feature sets 515 by forming an association between the signal of its input layer 520 and the signal of its output layer 524 within a set of one or more associative layers 522. Each head 504, as described above, can operate in a supervised or unsupervised mode and output an inference 531 indicating whether the input data is expected (normal) or unexpected (abnormal). Each combination of the backbone 512 and a head 504 is called a network, and the multi-head L-DNN effectively makes inferences simultaneously on multiple networks. When each head 504 is trained to recognize a different set of features, the multi-head L-DNN 506 can process inputs from multiple checkpoints without requiring a specific complete DNN trained for each of these checkpoints.
[0083] Figure 5B shows the estimated speed of the multi-head L-DNN executed on three computing platforms in frames per second. Intel's OpenVINO, Nvidia's TensorRT, and Neurala's custom software development kit (SDK). In all three cases, there is a sublinear decrease in the inference speed with respect to the number of heads. For two heads, Figure 5B shows a 30% speed improvement with TensorRT, a 25% speed improvement with the Neurala SDK, and an approximately 20% speed improvement with OpenVINO compared to when two single-head L-DNNs are sequentially executed on the same data. For eight heads, the speed improvements with TensorRT and the Neurala SDK increase to 55% and the speed improvement with OpenVINO increases to approximately 40% compared to when eight single-head L-DNNs are sequentially executed on the same data.
[0084] FIG. 6A shows a VIA platform 600 having a multi-head L-DNN 506 that is executed on an edge computer 604 and is configured to perform visual inspections of objects 21, 23, and 25 on a production line 620. The system 600 includes a plurality of input sources (here, cameras 608-1 through 608-3) coupled to the edge computer 604. Each input source (camera 608) creates an input array (image) 611-1 through 611-3 corresponding to each tact of the production line 620. These images 611 are batched together and supplied to the DNN backbone 510 of the multi-head L-DNN 506. During the batching process, the edge computer 604 may resize or pad the input arrays so that all of the input arrays are the same size. This enables modern DNN frameworks to process the entire batch in parallel, increasing the processing speed of the multi-head L-DNN 506. The backbone 112 outputs a batch of feature vectors 621-1 through 621-3. This batch is split into individual feature vectors, each of which is supplied to a corresponding module B head 504-1 through 504-3. Module B 504 is computationally lightweight and can be executed in parallel on any system powerful enough to execute module A 502, which includes the edge computer 604. The heads 504 generate inferences 531-1 through 531-3, or predictions, based on features 521 regarding whether objects 21, 23, and 25 are good, or good or not good (abnormal) for each input (camera or ROI) at each tact of the production pipeline.
[0085] The type of head 504 (module B) suitable for the multi-head L-DNN 506 includes, but is not limited to, the fast learning L-DNN classification head described in U.S. Patent Application Publication No. 2018 / 0330238A1, which is incorporated herein by reference, the anomaly recognition fast learning L-DNN head as described above, the backpropagation-based classification head in a conventional DNN or a transfer learning-based DNN, and the detection head or segmentation head based entirely or partially on L-DNN fast learning or conventional backpropagation learning. The heads of a single multi-head L-DNN 506 can be homogeneous (the same type) or heterogeneous (different types) as long as they are all trained using the same module A.
[0086] Figures 6B and 6C show interfaces for a multi-head L-DNN VIA system that inspects different components. These interfaces can be rendered on the display of a laptop 201, an HMI 212, or a display of another computer or edge computer 208 operably coupled thereto. An operator can use these interfaces to monitor the performance of the VIA system, respond to alarms set by the VIA system, and also label or classify anomalies flagged by the VIA system in a supervised learning mode. Thereby, the operator can teach the VIA system how to recognize and classify anomalies on-the-fly that have not been seen before.
[0087] In FIG. 6B, the interface shows an image of a printed circuit board (PCB) with five different ROIs, each containing different components. A single camera images the entire PCB at high resolution. Next, each ROI is cut out from the original image, the five ROIs are resized to the same size, batched together, and fed to the backbone for parallel processing. Five different heads evaluate the respective features corresponding to the five ROIs and return a normal or abnormal estimate according to the image. In FIG. 6C, the interface characterizes photographs of different parts of the kit (labeled ROIs (1) through (5)). Images of the components can be collected by one camera at a time, by five separate cameras operating simultaneously or at different times, or by a group of 1 to 5 cameras. The backbone extracts features from the five images batched together and sends the corresponding extracted features to the five corresponding heads. This head is lightweight enough to evaluate each ROI in parallel for a faster throughput.
[0088] The multi-head L-DNN is particularly useful for prototyping, custom work, and low-volume situations, i.e., where the L-DNN can be quickly trained on a small dataset. Visual inspection can flag surface defects, poor soldering, and bent pins and can replace or complement conventional functional tests. If sufficient defective parts are available or the heads operate in a supervised mode, the L-DNN can also be trained to classify types of faults or defects, including incomplete or missing solder traces or improperly oriented components.
[0089] A multi-head L-DNN-based VIA system can also inspect kits for assembling vehicles with customized features on an automobile assembly line. This can be done using any complex assembly process, which is becoming more common in automobile manufacturing, where each assembly may have different parts. Kitting can be customized and can be used when many small parts are used in one assembly. At the same time, it reduces the complexity on the line and improves the efficiency of material handling. Kitting also reduces the possibility that incorrect parts are picked up and used on the line side, ensuring that the operator has the materials necessary to complete the assembly.
[0090] Automation of the kitting process is advancing, but as the most common practice, people are made to create the kits. The form factor of the kits can vary depending on the type of parts, but is usually a box or a rack. Kits are often assembled so that each type of part is in the same place in each kit. Minimal inspection is carried out. Processes are often created to reduce or minimize errors (e.g., picking order, use of picking lights, and computer control systems), but after the kits are created, there is often no inspection to ensure that the correct parts are in the kits. Parts are assumed to be correctly selected by the kitting operator and placed in the appropriate bins. This can lead to incorrect parts being used or the operator on the assembly line not having the correct parts, which can affect productivity.
[0091] The kit can be inspected by a VIA system that executes a single-head L-DNN, where each kit has its own anomaly model. The operator scans the barcode of the kitting container to inform the system which parts are being withdrawn for which kit, and the barcode indicates which kit is being assembled and which parts are required. If a part is missing from the kit or if the wrong part is placed in the kit, the VIA system detects the anomaly and sends an error message to the HMI via Modbus TCP through the ISP to instruct the operator to check the work.
[0092] Constructed for single-head L-DNN implementation, the multi-head L-DNN VIA system can simplify model construction using multiple ROIs and provide more useful quality metrics. Each area of the kitting fixture can be its own region of interest, meaning that each part has its own ROI. This simplifies model construction, and each head is trained to evaluate different ROIs for each possible part variant. There is no need to build a model for each kit (reducing the amount of training required), and kits can be composed by selecting relevant parts with each model. Furthermore, using multi-ROI inspection, VIA can identify not only that a kit is inaccurate but also which parts are inaccurate (if any). This improves quality inspection, yield, and reduces rework.
[0093] Similarly, a multi-head L-DNN can be used to inspect the packaged cases. There are various types of case packagings, including putting cans in cardboard boxes and then shrink-wrapping the packaged sealed cardboard boxes. These case packagings are often used for pet food, canned vegetables, canned soups, etc. This packaging can only be used for transportation (e.g., to grocery stores where the cans are unpacked and placed on shelves) or for sale as units to customers (e.g., inside a warehouse for bulk purchase). Each is treated as a different ROI and dents, etc. can be inspected. As another method, different heads of the multi-head L-DNN can be trained to evaluate different aspects of the image of the packaged case, including whether the shrink-wrap is torn, the shape of the can (without bulges or dents), the folded shape of the packaging, etc., from all the same image or set of images.
[0094] <Dominant Parameter: False Positive / False Negative Anomaly Detection Threshold> How accurately the L-DNN identifies anomalies depends, in part, on a dominant parameter that can be adjusted by the user. The L-DNN head (Module B) uses this dominant parameter to determine how well the feature vector extracted by the L-DNN backbone (Module A) matches a specific representation of a known (good) object. This is done by determining how close the extracted feature is to this specific representation relative to all other representations. A dominant value of 10 means that for an input feature to be accepted as an example of a particular class, it should be approximately 10 times closer to the prototype of this class than to the prototypes of any other class. If this condition is not met, the input is recognized as an anomaly. If the dominant parameter is too small, the L-DNN may report more false positives, i.e., report more anomalies as normal objects (the upper part of Table 1). Also, if the dominant parameter is too large, the L-DNN can report more false negatives, i.e., report more normal objects as anomalies (the lower part of Table 1). The operator can adjust the false negative / false positive ratio by changing this dominant coefficient in response to the system performance.
[0095] The following examples show how changes in the dominance coefficient affect the performance of the L-DNN. The L-DNN dominance parameter can be adjusted to bias the system towards removing either false positives or false negatives, depending on the user's choice. In a human-assisted quality control system, false negatives (i.e., identifying non-existent anomalies) are more tolerable because the output will anyway pass human review, while false positives (missed anomalies) are less tolerable. In the case of a fully automated anomaly detection system without human teachers, the user has the option to set the balance.
[0096] This is the result of a pilot study examining packs of chewing gum where some pieces have been removed, crushed, or replaced, for two different values of the classifier dominance parameter d.
Table 1
[0097] In Table 1, the dominance parameter describes how the classifier's prediction is compared to all other possible predictions for the system to accept the prediction. Generally, a higher dominance parameter makes the system more detailed. All "normal" values are truly normal, but some normal values may be included among the abnormal values. With a lower dominance parameter, the system is less detailed: all "abnormal" values are truly abnormal, but some abnormal values may be included among the normal values. Here, the parameter d shifts the balance from false positives (at the top of Table 1, three anomalies missed) to false negatives (at the bottom of Table 1, only one anomaly missed, but three normal cases classified as abnormal).
[0098] <Conclusion> Although various embodiments of the invention have been described and illustrated herein, those skilled in the art will readily envision various other means and / or structures for performing the functions described herein and / or for obtaining one or more of the results and / or advantages thereof, and each such variation and / or modification is to be regarded as within the scope of the embodiments of the invention described herein. More generally, those skilled in the art will readily understand that all parameters, dimensions, materials, and configurations described herein are meant to be illustrative and that the actual parameters, dimensions, materials, and / or configurations will depend on the particular application for which the teachings of the invention are used. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments of the invention described herein. Accordingly, the foregoing embodiments are presented by way of example only and it is understood that embodiments of the invention may be practiced otherwise than as specifically described and claimed within the scope of the appended claims and their equivalents. Embodiments of the invention disclosed herein are directed to each and every individual feature, system, article, material, kit, and / or method described herein. Additionally, any combination of two or more such features, systems, articles, materials, kits, and / or methods is included within the scope of the invention disclosed herein if such features, systems, articles, materials, kits, and / or methods are not mutually inconsistent.
[0099] The above embodiments can be implemented by any of a number of means. For example, the embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code may be executed by any suitable processor or collection of processors, whether provided on a single computer or distributed among multiple computers.
[0100] Furthermore, it should be understood that the computer can be embodied in any of a number of forms, such as a rack-mounted computer, a desktop computer, a laptop computer, or a tablet computer. Additionally, the computer can be embedded within a device having suitable processing capabilities, which generally is not considered a computer, but rather a personal digital assistant (PDA), a smartphone, or any other suitable portable or fixed electronic device.
[0101] Also, the computer can have one or more input devices and output devices. These devices can be used, among other things, to present a user interface. Examples of output devices that can be used to provide a user interface include a printer or a display screen for visual representation of output, and a speaker or other audio generating device for audible representation of output. Examples of input devices that can be used for a user interface include a keyboard, as well as pointing devices such as a mouse, a touchpad, and a digitizer tablet. As another example, the computer may receive input information by voice recognition or in other audible formats.
[0102] Such computers may be interconnected by one or more networks in any suitable form, including a local area network, or a wide area network such as an enterprise network, and an intelligent network (IN) or the Internet. Such networks may be based on any suitable technology, may operate according to any suitable protocol, and may include wireless networks, wired networks, or fiber optic networks.
[0103] The various methods or processes outlined in this specification may be encoded as software executable on one or more processors using any one of a variety of operating systems or platforms. Additionally, such software may be described using any of a number of suitable programming languages and / or programming or scripting tools, and may be compiled as executable machine code or intermediate code that runs on a framework or virtual machine.
[0104] Also, concepts related to the various inventions may be embodied as one or more methods, and examples have been provided. The acts performed as part of a method may be ordered in any suitable manner. As a result, embodiments may be constructed in which the acts are performed in an order different from that illustrated, including performing some acts simultaneously, even if the acts are shown as consecutive acts in an exemplary embodiment.
[0105] All publications, patent applications, patents, and other references mentioned in this specification are hereby incorporated by reference in their entirety.
[0106] All definitions defined and used in this specification are to be understood as controlling over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of defined terms.
[0107] The indefinite articles "a" and "an" as used in this specification and the claims are to be understood to mean "at least one" unless clearly indicated otherwise.
[0108] As used in this specification and the claims, the phrase "and / or" means "any or both" of the associated elements, i.e., elements that may be present conjunctively in some cases and disjunctively in other cases. Multiple elements listed with "and / or" should be construed in the same manner, i.e., as "one or more" of the associated elements that are connected in parallel. Regardless of whether other elements are specifically identified or not, other elements may optionally exist in addition to the elements specifically identified by the "and / or" clause. Therefore, by way of non-limiting example, reference to "A and / or B", when used in conjunction with a syntax without limitations such as "comprising", may refer to, in one embodiment, only A (optionally including elements other than B), in another embodiment, only B (optionally including elements other than A), and in yet another embodiment, both A and B (optionally including other elements).
[0109] As used in this specification and the claims, "or" should be understood to have the same meaning as "and / or" as defined above. For example, when separating items in a list, "or" or "and / or" is inclusive, i.e., it includes at least one of a number of elements or a list of elements, and optionally additional items not in the list, including two or more. Terms that clearly indicate the contrary, such as "only one of" or "exactly one of", or terms such as "consisting of" when used in the claims, refer to including exactly one of a number of or enumerated elements. Generally, the term "or" as used in this specification is to be construed as indicating an exclusive alternative (i.e., "not both but either one or the other") only when preceded by an exclusive term such as "any", "one of", "only one of", or "exactly one of". "Consisting essentially of", when used in the claims, shall have the ordinary meaning as used in the field of patent law.
[0110] As used in this specification and the claims, the phrase "at least one" in relation to a list of one or more elements means at least one element selected from any one or more of the elements in the list of elements, but does not necessarily include at least one of every element specifically recited in the list of elements, and is not to be construed as excluding any combinations of elements in the list of elements. Also by this definition, it is allowed that elements other than those specifically identified in the list of elements referred to by the phrase "at least one" may optionally exist, whether or not they are related to the specifically identified elements. Therefore, by way of non-limiting example, "at least one of A and B" (or equivalently "at least one of A or B", or equivalently "at least one of A and / or B") can refer to, in one embodiment, at least one A (optionally including elements other than B) where B does not exist and optionally includes two or more A's, in another embodiment, at least one B (optionally including elements other than A) where A does not exist and optionally includes two or more B's, and in yet another embodiment, at least one A that optionally includes two or more A's, and at least one B that optionally includes two or more B's (optionally including other elements).
[0111] In the claims, as well as in the above specification, all transitional phrases such as "comprising", "including", "carrying", "having", "containing", "involving", "holding", "composed of", and the like are to be understood to be open-ended, i.e., to mean including but not limited to. Only the transitional phrases "consisting of" and "consisting essentially of" are to be considered closed or semi-closed transitional phrases as set forth in the United States Patent and Trademark Office's Manual of Patent Examining Procedure, Section 2111.03.
Claims
1. A method for identifying abnormal objects on an assembly line using a pre-trained backbone and a fast learning head executed on a processor connected to a sensor, comprising: training the fast learning head to recognize normal objects on the assembly line without training the fast learning head to recognize abnormal objects; acquiring, by the sensor, data representing an object on the assembly line; extracting, by the pre-trained backbone, features of the object data; automatically recognizing that the object is abnormal based on the features of the data extracted by the pre-trained backbone by the fast learning head; including; wherein the fast learning head is a first fast learning head and the sensor is a first sensor; furthermore, training the second fast learning head to recognize normal objects on the assembly line without training the second fast learning head to recognize abnormal objects; acquiring, by a second sensor, data representing another view of the object on the assembly line; extracting, by the pre-trained backbone, features of the data of the other view of the object; automatically recognizing that the object is abnormal based on the features of the data representing the other view of the object extracted by the pre-trained backbone by the second fast learning head; A method including.
2. The method according to claim 1, wherein the first sensor is a camera, and acquiring the data includes acquiring an image of the object.
3. The method according to claim 2, wherein training the first fast learning head includes training the first fast learning head with fewer than 100 images of normal objects.
4. The method according to claim 1, wherein acquiring the data representing the object includes acquiring time-series data.
5. The method according to claim 1, wherein extracting the features of the data includes propagating the data through a convolutional layer of a neural network trained to recognize normal objects.
6. The method according to claim 1, wherein extracting the features of the data includes at least one of performing a Fourier transform or a wavelet transform on the data.
7. The method according to claim 1, further comprising identifying, by the first fast learning head, that there is an anomaly of a type not previously seen in the object.
8. The method according to claim 7, further comprising labeling, in response to an input from a user, the anomaly of the type not previously seen as an anomaly of a first type.
9. The object is a first object, obtaining data representing a second object on the assembly line, extracting features of the data representing the second object by the pre-trained backbone, automatically recognizing, by the first fast learning head, that there is an anomaly in the second object based on the features extracted from the data representing the second object The method according to claim 8, further comprising.
10. The method according to claim 9, further comprising classifying, by the first fast learning head, the second object as having the anomaly of the first type.
11. The anomaly of the type not previously seen is a first anomaly of a type not previously seen, identifying that the second object has a second anomaly of a type not previously seen, labeling, in response to an input from a user, the second anomaly of the type not previously seen as an anomaly of a second type The method according to claim 9, further comprising.
12. A method of identifying an abnormal object on an assembly line using a pre-trained backbone and a fast learning head executed on a processor connected to a sensor, training the fast learning head to recognize normal objects on the assembly line without training the fast learning head to recognize abnormal objects, obtaining, by the sensor, data representing an object on the assembly line, extracting features of the data of the object by the pre-trained backbone, automatically recognizing, by the fast learning head, that there is an anomaly in the object based on the features of the data extracted by the pre-trained backbone including the fast learning head is a first fast learning head, the sensor is a first sensor, and the object is a first object, furthermore, Training the second fast learning head to recognize normal objects on the assembly line without training the second fast learning head to recognize abnormal objects, Obtaining data representing a second object on the assembly line by a second sensor, Extracting features of the data of the second object by the pre-trained backbone, Automatically recognizing that the second object is abnormal based on the features of the data extracted by the pre-trained backbone by the second fast learning head A method comprising.
13. A method for identifying abnormal objects on an assembly line using a pre-trained backbone and a fast learning head executed on a processor connected to a sensor, Training the fast learning head to recognize normal objects on the assembly line without training the fast learning head to recognize abnormal objects, Obtaining data representing an object on the assembly line by the sensor, Extracting features of the data of the object by the pre-trained backbone, Automatically recognizing that the object is abnormal based on the features of the data extracted by the pre-trained backbone by the fast learning head Including, The fast learning head is a first fast learning head, the data includes an image, and training the first fast learning head to recognize normal objects includes training the first fast learning head to recognize normal first regions of interest in training images, Furthermore, Training the second fast learning head to recognize normal second regions of interest in the training images without training the second fast learning head to recognize abnormal second regions of interest, Automatically recognizing that the second region of interest in the image is abnormal based on the features of the data extracted by the pre-trained backbone by the second fast learning head A method comprising.
14. The method according to claim 1, further comprising automatically identifying an abnormal position within the image of the object by the first fast learning head.
15. The method according to claim 1, further comprising automatically removing the object from the assembly line in response to the first high-speed learning head recognizing that there is an abnormality in the object.
16. The method according to claim 1, further comprising triggering an alarm in response to the high-speed learning head recognizing that there is an abnormality in the object.
17. A system for inspecting an object on an assembly line, comprising: a sensor for acquiring data representing the object on the assembly line; operatively connected to the sensor, causing a pre-trained backbone to extract features of the object from the data; a high-speed learning head trained only on data representing normal objects, receiving the features from the pre-trained backbone, and identifying whether the object is normal or abnormal based on the features; a processor; and wherein the sensor is a camera, the data is an image, the high-speed learning head is a first high-speed learning head trained to recognize only a first region of interest in a normal training image, the processor causes a second high-speed learning head trained to recognize only a second region of interest in the normal training image to recognize that there is an abnormality in the second region of interest of the image based on the features of the data extracted by the pre-trained backbone; a system.
18. The system according to claim 17, wherein the first high-speed learning head is trained with fewer than 100 images of normal objects.
19. The system according to claim 17, wherein the sensor acquires time-series data representing the object.
20. The system according to claim 17, wherein the pre-trained backbone includes a convolutional layer of a neural network trained to recognize normal objects.
21. The system according to claim 17, wherein the pre-trained backbone extracts the features by at least one of performing a Fourier transform or a wavelet transform on the data.
22. The system according to claim 17, wherein the first high-speed learning head identifies the object as having an abnormality of a type not previously seen.
23. The system according to claim 22, further comprising an interface operably connected to the processor and enabling a user to label an abnormality of a type not previously seen by the user as a first type of abnormality. **Claim 24** The system according to claim 22, wherein the object is a first object, the sensor acquires data representing a second object on the assembly line, the pre-trained backbone extracts features of the data representing the second object, and the first fast learning head automatically recognizes that there is an abnormality in the second object based on the features extracted from the data representing the second object. **Claim 25** The system according to claim 24, wherein the first fast learning head classifies the second object as having a first type of abnormality. **Claim 26** The abnormality of a type not previously seen is a first type of abnormality, the first fast learning head identifies the second object as having a second abnormality of a type not previously seen, The system according to claim 24, further comprising an interface operably connected to the processor and enabling a user to label the second abnormality of a type not previously seen as a second type of abnormality. **Claim 27** A system for inspecting an object on an assembly line, a sensor that acquires data representing the object on the assembly line; operatively connected to the sensor, causing the pre-trained backbone to extract features of the object from the data; a processor that receives the features from the pre-trained backbone and causes a fast learning head trained only on data representing normal objects to identify whether the object is normal or abnormal based on the features; and comprising wherein the fast learning head is a first fast learning head and the sensor is a first sensor, further comprising a second sensor operably connected to the processor and acquiring data representing another view of the object on the assembly line, wherein the processor causes the second fast learning head trained only on data representing normal objects to recognize that there is an abnormality in the object based on the features of the data representing the another view of the object extracted by the pre-trained backbone; a system. **Claim 28** A system for inspecting an object on an assembly line, A sensor that acquires data representing an object on the assembly line, operatively connected to the sensor, causes a pre-trained backbone to extract features of the object from the data, receives the features from the pre-trained backbone for a fast learning head trained only on data representing normal objects, and causes the fast learning head to identify whether the object is normal or abnormal based on the features a processor comprising wherein the object is a first object, the fast learning head is a first fast learning head, the sensor is a first sensor, and the object is a first object, further comprising a second sensor operatively connected to the processor and acquiring data of a second object on the assembly line, the processor causes the second object to be recognized as having an abnormality based on the features of the data of the second object extracted by the pre-trained backbone for a second fast learning head trained only on data representing normal objects, a system.
29. The system according to claim 17, wherein the first fast learning head identifies a position with an abnormality in the image of the object.
30. The system according to claim 17, further comprising a programmable logic controller operatively connected to the processor and removing the object from the assembly line in response to the first fast learning head recognizing that the object has an abnormality.
31. The system according to claim 17, wherein the processor triggers an alarm in response to the first fast learning head recognizing that the object has an abnormality.
Citation Information
Patent Citations
Detection of Textural Defects Using a One Class Support Vector Machine
US20110026804A1
Systems and methods to enable continual, memory-bounded learning in artificial intelligence and deep learning continuously operating applications across networked compute edges
WO2018208939A1