System for and method of recognizing and detecting abnormality with using lifetime deep neural network
Patent Information
- Application Number
- JP2025100271
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2020-01-12
- Filing Date
- 2025-06-16
- Publication Date
- 2025-12-17
AI Technical Summary
Conventional neural networks struggle with real-world scenarios requiring supervised learning for unknown classes, as they require extensive training data and are inflexible, while unsupervised learning lacks accuracy and efficiency, making them unsuitable for industrial quality control applications.
Lifelong Deep Neural Network (L-DNN) technology combines supervised and unsupervised learning to adapt to new data, allowing real-time recognition of known and unknown classes with minimal training data, reducing latency and memory usage, and enhancing data privacy by processing data locally.
L-DNN enables accurate, real-time anomaly detection and recognition of unknown defects without extensive training, reducing computational and storage needs, and ensuring data privacy by processing data at the edge, making it suitable for industrial quality control.
Smart Images

Figure 00000026_0000 
Figure 00000026_0001 
Figure 00000026_0002
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application is a continuation of U.S. patent application Ser. No. 62 / 960,132, filed Jan. 12, 2020. Claim the benefit of priority under 35 U.S.C. § 119(e) and incorporate by reference in its entirety for all purposes. and is hereby incorporated by reference in its entirety for the purpose of this application. [Background technology]
[0002] A conventional neural network often includes many layers of neurons interposed between the input and output layers. Artificial neural networks (ANN) and deep neural networks (DNN) , which are typically designed to work in either a supervised or unsupervised fashion. Supervised learning produces highly accurate, yet rigid structures that are difficult to recognize. It does not tolerate "unknown" classes beyond those trained to Unsupervised learning is a method to accommodate unknown input-output mappings and Although it can be called and discovered, it usually lacks the typical high performance (accuracy / memory) of supervised systems. However, many real-world scenarios often require supervised learning. Many such scenarios require a combination of supervised and unsupervised ANN features. In Rio, the number of known classes cannot be defined a priori.
[0003] For example, in a real-world quality inspection, the number of potential defects that may occur during a production cycle may be impractical or impossible to determine a priori. Traditional DNNs, which are trained on a known mapping between the input and output classes, require quality inspection. These cannot be used to detect all possible defects in an inspection. In this scenario, unsupervised methods (e.g., clustering) are used to estimate defects. However, these methods generally do not reach the same performance as supervised DNNs.
[0004] Specifically, a supervised ANN derives a set of labeled training data (training examples) from It is trained to infer a function, where each example contains an input object and a desired output label. The use of backpropagation to train ANNs requires large amounts of training data and many learning iterations. It requires a teacher to be able to generate data in real time as it becomes available. ANNs that learn in an unstructured fashion do not require desired output values and rely on prior knowledge of unknown patterns in the data. It is used to process, compress, and search. Unsupervised learning takes less time than supervised learning. It may take some time for unsupervised learning processes to achieve this, but it is still possible to achieve it in a few more steps than with supervised learning processes. is much faster and computationally simpler and can be performed in real time, i.e., per data presentation and The computational cost may be incremental, slightly more than the cost of inference.
[0005] Supervised DNNs map inputs (e.g., images) to outputs (e.g., image labels). The set of labels is defined a priori and is trained to perform a supervised learning process. The DNN designer must decide how many output classes are provided during inference (after training). We need to provide a priori all the classes that N may encounter, and the training data is , which is sufficient for slow backpropagation and some of the classes that a supervised DNN should identify. There should be a balance among the examples. If no training examples are provided for a given object class, If not, how does a traditional neural network know that this class of object exists? There is no.
[0006] This requirement to provide training examples for all object classes is a drawback of traditional neural networks. The work must be highly accurate, and the training data must contain many consistent examples of the correct product. However, there are very few examples of (inconsistent) incorrect products, and no industrial quality control Such problems make it difficult to implement successfully. You may prefer to postpone training to your customers so that the system incorporates the most recent examples for each customer. Therefore, ANNs and DNNs are data and computation intensive in their training process. Recognizing the inherent nature of software, its inability to be updated once deployed, and the possibility of initially unknown flaws Because of the difficulty in achieving this, it has not been widely adopted for use in industrial quality control applications. Summary of the Invention
[0007] Unlike traditional ANNs and DNNs, Lifelong Deep Neural Network (L-DNN) technology can be quickly trained to recognize acceptable or conforming objects and can be used to detect unknown defects and other anomalies without being trained to recognize those defects or anomalies. This makes L-DNN technology suitable for high-speed visual inspection and quality control in manufacturing assembly lines and other industrial environments. Edge computers running L-DNNs can use camera data or other data to identify defective or misaligned objects for removal, repair, or replacement in real time, without human intervention.
[0008] L-DNN technology combines (a) the ability of DNNs to be trained with high accuracy for known classes; (b) simultaneously, any number of unknown classes or Furthermore, to deal with real-world scenarios, LD NN technology allows (c) learning when there is little data. N techniques also work well with imbalanced data (e.g., a large number of very different examples for different classes). Finally, L-DNN technology allows for slow inverse neural network (SDNN) learning. Without propagation, (d) real-time learning becomes possible.
[0009] L-DNN technology is a technology that combines artificial intelligence (AI), artificial neural networks (ANN), and deep learning. Extend neural networks (DNNs) and other machine vision processes This allows training at the outset with minimal imbalanced data and adapts as new data becomes available. The technical benefits of L-DNN include, but are not limited to: Not limited to these. Reduced AI latency: Computing instead of on cloud or remote servers Being able to reside at the edge allows L-DNN AI to move data from the cloud or desktop to the cloud. on processors located directly at the local device level, rather than on data center resources. Because the data is processed, real-time or near real-time data analysis can be achieved. ● Reduced data transfer and network traffic: Data is processed at the device level Therefore, raw data is transported from the device to the data center via a computer network. does not need to be sent to the cloud or to a third party, thereby reducing network traffic. can be. ● Reduced memory usage: L-DNN technology enables one-shot learning, and in the future Eliminates the need to store data for use. Data privacy: The processing of data (both inference and learning) on the device is controlled by the user. Data is generated, processed, and destroyed at the device level, and leaves the device no matter what It is not stored persistently in raw form on the device, so it is not Enhance privacy and data security. Flexible operation: L-DNN can be operated in supervised, unsupervised, or semi-supervised modes. Cut. Handling unknown object classes: L-DNN can handle classes of objects, including defective or abnormal parts, during training. "Nothing I know" about objects I haven't encountered in can be identified and reported. For more information on L-DNN technology, see, for example, No. 2018 / 0330238A1 and U.S. Patent Application No. 16 / See No. 952,250.
[0010] Unlike traditional ANNs and DNNs, L-DNNs are trained to identify anomalies. It can identify images with abnormalities even if they have not been inspected, making it ideal for industrial visual inspection work. In visual inspection, automated inspection systems are very similar most of the time. , operating under the right conditions (e.g., good quality product on the conveyor belt, normal processing, appropriate regime) However, in rare cases, the input data may differ. or may be out of order (e.g., defective products, abnormal processes, machinery out of operation) In general, the task of distinguishing between correct data and incorrect or abnormal data is called anomaly recognition. (For example, recognizing that a 12-pack of aluminum cans contains a defective can ), or anomaly detection (e.g., which cans in a 12-pack are defective, in a multidimensional space) In anomaly recognition and detection, time is of the essence. The time can be one of the inputs to the system.
[0011] Real-time operating systems that implement L-DNN can respond almost immediately to new knowledge. It learns new products and new anomalies on the fly so that it can be applied to a wide range of applications. A real-time action system for anomaly recognition and detection that can be used to: This system becomes possible. a. Conduct rapid training at the deployment site using only examples of normal situations. b. Recognize or detect anomalies and alert the user. c. Optionally, without interrupting the operation of the system, for each abnormal data Update your knowledge.
[0012] L-DNNs can be used in systems for inspecting objects on assembly lines. The system may include a sensor and a processor operably coupled to the sensor. In the story, sensors acquire data representing objects on an assembly line, and a processor calculates the 1) a pre-trained backbone that extracts object features from data, and (2) a pre-trained It receives features from the trained backbone and identifies objects as normal or abnormal based on those features. We run a fast learning head trained only on data representing normal objects. The system also responds to the high-speed learning head recognizing an object as abnormal by moving the object to the assembly line. or respond to an appropriate signal from the processor to alert the operator to the anomaly. Programmable logic controller (PLC) or human machine interface (HMI) may also be included.
[0013] The sensor may be a camera, in which case the data includes an image of the object. The fast learning head is trained on a relatively small number of images (e.g., 100, 90, 80, 70, 60, 50, The sensor can be trained on a number of images (40, 30, 20, 10, 5, or even 1 image). and strain, vibration, or It may also be a temperature sensor.
[0014] The pre-trained backbone is a neural network trained to recognize normal objects. It may also include a convolutional layer of the network. It can be configured to extract features by transforming them.
[0015] A fast learning head can identify objects as having types of anomalies not previously seen In these cases, the operator may configure an interface operably coupled to the processor. Use it to label a previously unseen type of anomaly as the first type of anomaly If the sensor acquires data from other objects on the assembly line, pre-training can be performed. The resulting backbone maps these objects for automatic anomaly detection by a fast learning head. The fast learning head can also extract features of the data representing the second object. If the second object is not previously seen, it can be classified as having an anomaly of the first type. If the operator has an anomaly of a type that was not detected, the operator may classify the anomaly as a second type anomaly. An interface can be used to label it as such.
[0016] A processor can implement many fast learning heads, all of which use the same prior It receives the features extracted by the trained backbone and operates on them. This reduces latency compared to running the same number of backbone and head pairs sequentially. Different heads can be used to display different objects, different images of the same object, or different backgrounds. Different regions of interest in the same image or dataset from the same set of features extracted Heads can also be trained to do different things or to perform different tasks. For example, one head can monitor for abnormalities while another head can , objects may be classified based on size, shape, position, or orientation from the same extracted features. do.
[0017] All combinations of the above concepts, and additional concepts discussed in more detail below, are (provided that such concepts are not mutually exclusive) In particular, all of the claimed subject matter appearing at the end of this disclosure is considered to be part of the All combinations are considered to be part of the inventive subject matter disclosed herein. It is understood that the present invention is not limited to the above and may be modified in any way without departing from the spirit or scope of the present invention. Terms expressly used in this specification are designated by the meanings most consistent with the particular concepts disclosed herein. should be given.
[0018] Upon review of the following diagrams and detailed descriptions, other systems, processes, and features may be applicable. Such additional systems, processes, and features will become apparent to those skilled in the art. All such modifications and variations are intended to be included within this description, be within the scope of the present invention, and be protected by the accompanying claims. It is intended that [Brief explanation of the drawings]
[0019] Those skilled in the art will appreciate that the drawings are primarily for illustrative purposes and that the It will be understood that no limitation of the scope of the present subject matter is intended. In some instances, various of the inventive subject matter disclosed herein may be used in various proportions. Some aspects may be shown exaggerated or enlarged in the drawings to facilitate understanding of the different features. In the drawings, like reference characters generally indicate like features (e.g., functionally similar). "similar and / or structurally similar elements" means
[0020] [Figure 1] Figure 1 shows a Lifelong Deep Neural Network (L-DNN) implemented on a machine running in real time.
[0021] [Figure 2] Figure 2 shows the Visual Inspection Automation (VIA) platform for creating and deploying manufacturing inspection brains (L-DNNs) and an on-premise solution tailored for visual inspection tasks.
[0022] [Figure 3A] Figure 3A shows the process for prototyping and delivering an L-DNN for deployment on the VIA platform.
[0023] [Figure 3B] Figure 3B shows the process for training, fine-tuning (tuning), and deployment of L-DNN on the VIA platform.
[0024] [Figure 3C]FIG. 3C shows a process for generating inference behavior that executes on the VIA platform in response to inference by the L-DNN.
[0025] [Figure 4] Figure 4 shows the relationships between projects, datasets, inferences, and recipes for creating and deploying versions of a brain in the environment of Figure 2 and the workflow of Figure 3.
[0026] [Figure 5A] Figure 5A shows a multi-head L-DNN implemented on a machine running in real time.
[0027] [Figure 5B] Figure 5B shows example performance results of a multi-head L-DNN in inference frames per second versus the number of heads in the system.
[0028] [Figure 6A] FIG. 6A shows the operation of a multi-head L-DNN in a production line with three image sources and corresponding three separate tasks.
[0029] [Figure 6B] FIG. 6B shows a dashboard of a multi-head L-DNN configured to inspect and identify anomalies in different regions of interest (ROIs) of an integrated circuit.
[0030] [Figure 6C] FIG. 6C shows a dashboard of a multi-head L-DNN configured to inspect and identify anomalies in different parts packed together in a kit. DETAILED DESCRIPTION OF THE INVENTION
[0031] For many years, manufacturers have used automated optical inspection (AOI), or computer vision, to to inspect products at various points in the manufacturing process, for example, during processing, labeling, and final inspection. Traditional AOI systems, consisting of cameras and software, have dramatically improved quality. This allows for rapid inspection while improving accuracy and reducing waste. One limitation of the system is the ability to detect surface level defects, deformations, and other changes or anomalies (or irregularities). The inability to detect defects or perform assembly verification. To do this, manufacturers may rely on human inspectors. However, human inspectors are subject to imperfect accuracy. It has its own limitations, including gender and / or human inconsistencies. It is not possible to inspect every single part, especially on a high speed line.
[0032] To overcome these limitations, engineers are looking to extend traditional computer vision inspection to Deep learning, a subset of machine learning, is a method of learning neural networks. It makes computations of deep learning networks (DNNs) feasible. Ostensibly, it is used in computer vision. and deep learning are similar: they both automate the visual inspection process and improve the speed of inspection. Increased productivity, reduced line-to-line deviations, improved scalability and efficient inspection However, these approaches offer several advantages, including a higher percentage of There are several differences between the two. These differences affect the types of problems they are best suited to solving. In simple terms, computer vision is a logical, programmable Objective or quantitative use cases (e.g., parts measure 100 ± 0.1 mm) that can be Deep learning is well suited to subjective or deterministic learning with generalization or variability. It is more suitable for inspection (e.g., that the weld is good).
[0033] Computer vision applies simple logic-based rules to find the desired criteria Visual inspection works well when the object's compatibility can be assessed. The computer vision system determines how, where, and how the parts and / or components are viewed. Oriented or located within the field (e.g., following part measurement) Computer vision systems can be used to verify the effectiveness of automated picking (or as a precursor to it). The system also determines the presence or absence of critical components or the integrity of components / parts (e.g. For example, whether a part is faulty or a component is missing. A computer vision system measures or estimates the distance between points within a part or component and then Measurements can be compared to specifications or tolerances for pass / fail (e.g., milling / forming). (Check the dimensions of critical parts after assembly and before further assembly). The Vision System uses product labels (e.g., barcodes for part / component identification) Or extract data from Quick Response (QR) codes, text from labels, etc. do.
[0034] The strength of deep learning is generalization, so deep learning inspection systems can be used in a variety of situations. For example, deep learning inspection systems can detect surface scratches on metal or plastic. Product, part, or defect types, such as clutch variations, welding quality, or packaged products Deep learning inspection systems can identify changes in type or location of objects. (Computer vision systems are generally In addition, it performs poorly in environments with variable lighting and is generally unable to assess changes due to objects. I can't.)
[0035] Traditional deep learning inspection systems must be trained to recognize the objects they inspect. Training is done on labels that closely resemble what the deep learning model will actually process when deployed in production. This must be done using data such as labeled or tagged images. The training data for the system is ideally raw data from the same camera, viewpoint, and spatial resolution. This means that the images should be taken under the same lighting as that used during production deployment, but Some changes are generalized by the system.
[0036] Additionally, the training data includes each type of object that the deep learning inspection system is supposed to identify. It should contain roughly equal numbers of labeled images per type or class. For example, ,When a deep learning inspection system classifies "good" and "bad" objects, the training set has roughly equal numbers of examples (labeled images) of objects in each of these classes. The amount of training data required to train a deep learning inspection system should be It depends on several variables, including accuracy, variability, and number of object classes. Higher accuracy, higher variability, and more classes are associated with more training data. Containing 1,000 or more labeled images per class is a significant advantage over training data. That's not uncommon.
[0037] However, in industrial and manufacturing use cases, defect recognition is a key challenge for at least two reasons. It is difficult to generate enough training data to train a deep learning system to First, it may be impossible to generate enough images of the defective product (e.g., 2,000 images). This is difficult because modern manufacturing processes have dramatically improved yields and reduced errors and This is because waste is reduced. Procurement of a balanced and representative training dataset is Second, it is difficult to identify all the recognizable defect types. Some types of defects are predictable, but the process that causes them can be difficult or impossible. Other types of defects can be unpredictable because the conditions under which they occur are unknown in advance. Due to the lack of defect data (labeled images), traditional deep learning inspection systems are trained to It is difficult, if not impossible, to recognize defects (of different types).
[0038] Traditional deep learning inspection systems process input data as one of the classes they are taught to recognize. Imbalanced training datasets (e.g., those that do not contain labeled images of defective parts) (those with few or no labeled images) Training a system is the process of training the system to make decisions about classifying input data. This leads to bias and lack of accuracy. As a result, the model is trained only on good / normal data. Traditional deep learning inspection systems can easily identify abnormal objects, although perhaps with lower reliability. It can be distinguished as good or normal.
[0039] Unlike traditional deep learning inspection systems, L-DNN-based inspection systems They are unable to recognize different types of defects and other anomalies without being trained to specifically recognize them. Furthermore, L-DNN-based inspection systems are more accurate than traditional deep learning inspection systems. It can be trained to recognize normal or good objects on a smaller training dataset. Moreover, unlike traditional deep learning inspection systems, L-DNN-based inspection systems and then coupled to the computing edge (e.g., factory floor inspection cameras) While running on a computer (on a processor with a different processor), new objects, including previously unseen objects, are discovered. You can learn to recognize your body.
[0040] In summary, the L-DNN-based inspection system is capable of performing anomaly recognition. Anomaly recognition is achieved by training the model only on "good" image data (an imbalanced training dataset). It works by establishing a threshold for deviation from this "good" data. The system can identify images that deviate from the norm. As the system identifies and captures an increasing number of images with anomalies, the system With the help of the data, we can label these images to form new anomaly classes and It is possible to move from recognition to more detailed anomaly classification.
[0041] L-DNN is a "brain" that provides artificial intelligence for L-DNN-based inspection systems. The training system (e.g., Neurala Brain) in Builder Software-as-a-Service (SaaS) products It is created and optimized on a local or cloud server running can be trained on images or other data that represent only normal objects. L-DNN can be applied to cameras, microphones, scales, thermometers, or other sensors that capture input data. It can be deployed on one or more nodes, including edge computers connected to the server. For example, each edge computer is equipped with a GigE Vision standard (e.g., Basler Connects to compatible cameras (such as NI, Baumer, or Teledyne DELSA cameras) The training system may be configured with standard I / O such as PLC I / O and Modbus TCP. Using the interface, it is easily integrated into industrial machines or production lines.
[0042] As discussed in more detail below, the training system trains the L-DNN to produce a "good" One or more categories or classes of input data can be recognized. For example, Training an L-DNN on images of surface quality, packaging, or kitting Since L-DNN is trained only on "good" images, L-DNN-based inspection systems flag images as anomalies if they do not match well with good images. Identify it as such.
[0043] L-DNN-based inspection systems offer several technical advantages over traditional automated inspection systems. First, it can recognize anomalies without being trained on data that represents them. This allows for the collection of images depicting defective, unacceptable parts. This eliminates the time-consuming and often impractical or difficult tasks of Second, it requires a smaller training dataset than traditional deep learning inspection systems. This training dataset is imbalanced and can be trained to recognize parts that are not For example, they may represent only acceptable parts and not defective or acceptable parts. can be constructed from data, making training faster, easier, and cheaper. Faster and easier training allows L-DNN-based inspection systems to adapt to different Third, the L-DNN-based inspection system works well. This allows for more accurate identification (and, optionally, classification with human supervision) of different types of defects. Fourth, L-DNN runs on edge devices. This saves network bandwidth because the data does not need to be sent to another device. Your data is kept private and secure, while your spending is reduced.
[0044] <Lifelong Deep Neural Network (L-DNN) Technology> L-DNN technology allows for fast yet stable learning of features that represent entities or events of interest. To achieve this, we developed a subsystem (Module A) based on a representation-rich DNN for rapid learning. Combine with the subsystem (Module B). These feature sets are used in low-level programming, such as backpropagation. It can be pre-trained using fast learning methodologies. This is possible by adopting non-DNN methodologies in Module A) In the case of the example, the high-level feature extraction layer of the DNN classifies known entities and events and extracts unknown ones. Module B's fast learning system allows new entities and events to be added on the fly. Module B serves as the input to the system. Module B learns important information and avoids the drawbacks of slow learning. Instead, it captures descriptive and highly predictive characteristics of correct input or behavior, and It can identify incorrect inputs and behaviors that were not detected.
[0045] L-DNN also makes it possible to learn new knowledge without forgetting old knowledge. In other words, the technique: (a) (b) the need to transmit or store input images, (b) time-consuming training, or (c) significant computation Resource-free behavior adjustment at the edge based on user input (continuous and / or intermittent) To adjust the Real-time machines (e.g., computers running on production line clocks) After the introduction of L-DNN, learning can be performed in real time. The machine adapts to changes in its environment and user interactions, and addresses imperfections in the original dataset. This allows for a customized experience for each user. In the case of detection, the system is run on a highly imbalanced dataset of extreme cases with no anomaly examples. It starts from the beginning and varies from knowledge of what is a good product or correct behavior during operation. This is particularly useful as it allows you to discover and add new features.
[0046] Here, L-DNN technology is used to monitor manufacturing and packaging lines. L-DNN creation software, such as Neurala Brain Bu Use ilder to create one or more different versions of an L-DNN or brain. and may be deployed, each of which monitors data from at least one sensor. This sensor data is one-dimensional (e.g., from a sensor monitoring the movement of a rotating cam). time-series data) or two-dimensional (e.g., a camera that takes a photo of a package before sealing it) In either case, the L-DNN feature extractor (module A ,e.g., feedforward DNN, iterative DNN, Fast Fourier Transform (FFT) module The wavelet transform module (implemented as a linear or wavelet transform module) extracts features from the data. Module B classifier classifies the extracted features as normal or abnormal. If the extracted features indicate an abnormality (e.g., a machine is wobbling outside its operating limits), (including components with broken or corrupted packages), L-DNN-based quality monitoring system may take any corrective action (e.g., slowing or stopping the machine or This may trigger a malfunction (removing the package from the packing line before the package is sealed).
[0047] L-DNN is a heterogeneous neural network that combines fast and slow learning modes. Implements neural network architecture. In fast learning mode, L-DNN is implemented. Just as real-time machines can respond to new knowledge almost instantly, In this mode, the system learns new knowledge and new experiences (e.g., anomalies). The learning rate of the stem is high enough to favor new knowledge and corresponding new experiences, while The learning rate of the slow learning subsystem is set to preserve old knowledge and corresponding old experiences. It is set to a low value or zero. Fast Learn mode is primarily used in anomaly detection systems for industrial inspection. This is a required operational mode, meaning that the system cannot be taken offline for slow training updates. This is because it can be costly.
[0048] FIG. 1 provides an overview of an L-DNN 106 that can be implemented in a machine 100 operating in real time. The machine 100 that operates in real time uses a sensor 130 such as a camera or a microphone. Receives sensory input data 101, such as video or audio data, from a slow learning module and feeding it to the L-DNN 106, which includes a module A 102 and a fast learning module B 104. This can be implemented as a neural network classifier. Each module A102 It can be based on the trained (constant weight) DNN110 or some other system. , which acts as a feature extractor. Module A 102 receives the sensory input 101. A stack of pre-trained neural network layers 112, also known as cloud 112. The algorithm identifies and extracts relevant features into a condensed representation of the object 115 and then converts these representations 115 into a model. Module B104 is fed to the DNN-based module A. The features are read from several hidden layers 114 of the DNN at the end of the feature extraction backbone 112. One or more fully connected, averaging, pooling layers 116 and cost The upper layers of the DNN, which may include layer 108, are ignored or removed entirely from the DNN 110. may be removed.
[0049] Module B 104 combines the signals of the input layer 120 with the signals in a set of one or more associative layers 122. These object representations 115 are quickly generated by forming associations between the object representations 115 and the output layer 124. In supervised mode, through user interaction, module B 104 , layer 124 receives the correct label for the unknown object, and each feature vector in layer 120 and It quickly learns the association between the object and its corresponding label, and as a result, can immediately recognize these new objects. In unsupervised mode, the system uses feature similarities between vectors. and assign these vectors to the same or different classes.
[0050] The L-DNN 106 takes advantage of the fact that the DNN 110 is a good feature extractor. Module B 104 receives data 101 from input sources (sensors) 130. Module B 104 continuously processes the features extracted by module A 102. Rapid one-shot learning is used to associate these features with object classes.
[0051] In fast training mode, a new set of features corresponding to normal (non-anomalous) cases is generated. When presented as input 101, module B 104 identifies these features as classes given in supervised mode or generated internally in unsupervised mode. In either case, module B 104 now knows about this input. This classification of module B104 is shown in the following presentation: Depending on the task it is performing, the classification itself can be used as the output of the L-DNN106 or as a module In combination with the output from a specific DNN layer from the model A102, do.
[0052] The general L-DNN and specific module B operate in real time in response to continuous sensory input. The neural network in Module B is designed to recognize unknown objects. Traditional neural networks use labeled inputs. As a result, the target dataset is a set of known objects that typically contain If there is no object to be detected, there is no need to deal with the input. To use the work in module B of the L-DNN, we need to add an additional special feature called "knowing nothing." Adding categories to the network reduces the likelihood of incorrectly classifying unknown objects as known. This should mitigate against module B attempts (false positives).
[0053] This concept of "knowing nothing" is a previously unseen and unlabeled It is particularly useful when processing raw sensory streams, which may (exclusively) involve objects. This potentially causes Module B and L-DNN to mistakenly recognize unknown objects. Instead of identifying the unfamiliar object as a familiar one, they would label the unfamiliar object as "unfamiliar" or "I've seen it before." It becomes possible to distinguish between "what has not been seen before" and "what has not been seen before." Extending the conventional design with a "no idea" implementation is to add bias nodes to the network. "Nothing Known" can also be a known object class. The L-DNN automatically scales its influence depending on the number of nodes and their corresponding activations. This version can be implemented as follows: the L-DNN system recognizes its input as something it knows. Whenever the user does not know anything, an abnormal flag 118 is set in response to the classification of "I don't know anything." Therefore, anomaly detection applications have a very "no idea" approach. It is useful.
[0054] For more information on L-DNN, see the attached document, which is incorporated herein by reference for all purposes. , “Systems and Methods to Enable Continua l,Memory-Bounded Learning in Artificial Intelligence and Deep Learning Continuou sly Operating Applications Across Compute U.S. Patent Application Publication No. 2018 / 0330238A1, entitled "e Edges", and also refer to PCT Publication No. WO2019 / 226670, entitled "Systems and Methods for Deep Neural Ne tworks on Device Learning (Online and Of fline) with and without Supervision". Please refer to the above.
[0055] <Industrial Inspection and Monitoring Using the L-DNN Platform> FIG. 2 shows a Visual Inspection Automation (VIA) platform 200 based on L-DNN technology as an example of a system that enables visual inspection on a factory floor. The VIA platform 200 includes software executed on one or more processors and has a graphical user interface (GUI) that enables a user to configure, train, deploy, and further refine the L-DNN system. The VIA platform 20 0 is executed on a local industrial PC (local server) 202 connected to a laptop 201 or an industrial PC 202 for initial setup and training, and also includes L-DNN training (e.g., brain builder) software. The VIA system 20 0 also includes edge computers 204. Each edge computer 204 is coupled to a corresponding camera 208 or other sensors such as a microphone, counter, timer, or pressure sensor that collects data for evaluation or inspection to form a corresponding node 209. Each node 209 may include different types of sensors, and the sensors by which form a corresponding node 209. Each node 209 may include different types of sensors, and the sensors collect data for evaluation or inspection. Even if we implement an L-DNN trained to recognize normal and abnormal data obtained, good.
[0056] The VIA system 200 is connected to a programmable logic controller 206 via a switch 206. Uses standard interfaces such as (PLC)210, I / O and Modbus TCP It is integrated into industrial machines or production lines and uses one or more human-machine interfaces. In FIG. 2, the HMI 212 is connected to the PLC 210. Customer-defined logic is implemented in the PLC210 and / or HMI212 and is controlled by the VIA system. For example, if the VIA system 200 If the PLC210 recognizes a product instance as abnormal, it This can trigger automatic removal of the device from the production line.
[0057] A local server 202 and L-DNN training software running on a laptop 201 The software builds and deploys the brain on the edge computer 204. Data 204 is software (called Inspector) that uses trained brains. to process data from the camera 208 and / or other sensors. Then, we provide a set of parameters (e.g., neural network weights) that represent the trained L-DNN. Once the brain is introduced, it can be used to control the manufacturing process in the factory. Recognize and identify anomalies in parts or other objects being packed, sorted, and / or packaged They can also communicate with each other and be connected to a network switch 206 (e.g. See, for example, U.S. Pre-Grant Publication No. 2018 / 004024, which is incorporated herein by reference in its entirety. Information on newly detected types of anomalies via the You can also share.
[0058] In operation, the L-DNN-based VIA system 200 of FIG. 2 reduces the takt time of a production process. components, subassemblies, and / or final components in a synchronized manner determined by The product can be inspected. Takt time is the time required to manufacture a part at that time interval. Divide by the number of parts. Inferentially, if the demand is for 60 parts in 1 hour, In this case, the takt time is 1 minute. L-DNN inference time (the time it takes for L-DNN to determine whether an image is normal or abnormal) The time it takes to sort must always be less than this takt time.
[0059] In FIG. 2, a VIA system 200 is shown mounted on a conveyor belt of an inspection or production line 220. Normal objects 21 and abnormal objects 23 and 24 move past the camera 208 on the route. The camera 208 captures still or video images of the object and Each edge controller implementing each L-DNN is connected to the network switch 206. These L-DNNs are trained to recognize normal objects 21. The L-DNN is trained to recognize normal objects 21 and 22, but the abnormal objects 23 and 25 are not recognized. When presented together with an image of a normal object 21 (e.g., a "good" image), the image ) are classified as deviant or abnormal objects. When presented, the image shows a deviant or abnormal object.
[0060] L-DNN can be adapted to recognize images of any kind of object, or spectral data, Recognizing patterns in audio data or other types of data, including other electrical signals For example, L-DNNs can be trained to process bread or other processed foods in a way that is consistent with the baking process. The first camera 208 can be trained to recognize different points in the image. Check the pan to ensure it is properly filled and has the correct surface (e.g., bubbles vs. no bubbles). To ensure that the bread is properly baked and certified, the loaf is imaged before it is baked. The second camera 208 may be positioned to capture images of the food as it emerges from the oven, e.g. For example, bread may be imaged on a conveyor belt running through an oven. The L-DNN on the edge computer 204 connected to the 8 to determine doneness), sprinkled ingredients on top (e.g., chocolate chips, The distribution and shape of the seeds (poppy seeds, sesame seeds, etc.) can be examined.
[0061] L-DNN is a machine learning algorithm that cannot be effectively implemented in traditional computer vision systems. It is particularly suited to such largely subjective tests, and the types of abnormalities that can occur are almost limitless. Therefore, it is difficult to train a traditional deep learning network to recognize all of them. Manual inspection is time-consuming and usually requires several parts This can be inconsistent and difficult to achieve the desired uniformity. In contrast, the VIA system automatically inspects every batch, resulting in a higher yield. , resulting in greater consistency and less waste, and trained to recognize them. Abnormalities can be identified without any need for a
[0062] The VIA system 200 of FIG. 2 also supports different manufactured products, including injection molded plastic parts. Plastic molded parts are used in a wide range of products from home appliances to automobiles. and other various manufactured products. Contract manufacturers may mold and deburr parts before sending them in bulk to customers. Traditionally, people have been collecting fragments of these parts, for example, every 10 plastic molded parts. We can inspect one part at a time, but this generally means we are sending you a good part. This is only a guarantee. The customer will manually inspect all parts before using them in an assembly.
[0063] A GigE camera 208 installed on the assembly line captures images of the parts after deburring. The image sensor 208 transmits the image to the edge computer either directly or through the network switch 206. The L-DNN executed on the edge computer 204 sends the (not properly formed and / or deburred) or abnormal (not properly formed and / or Classify the parts that appear in the image as either deburred or not deburred. If abnormal, the edge computer 204 may Switch on conveyor belt to PLC via TCP to OPC Integrated Architecture A switch (not shown) can be triggered to indicate defective parts and remove them from the line. The operator can inspect the defective parts and determine which parts to scrap and which to rework.
[0064] The L-DNN-based VIA system 200 is suitable for prototypes or small numbers of Inspection of injection molded parts on an assembly line where part types change frequently due to volume It is particularly suitable for inspection. This is because the L-DNN can be quickly trained on a small training set. Furthermore, the L-DNN can flag non-standard or abnormal parts without having recognized them beforehand. .
[0065] Continuous data and / or 1D data may be monitored using the L-DNN-based VIA system 200. For example, the VIA system can monitor continuous or sample temperature, strain, or vibration data for parameters indicating or suggesting that a particular part or component is worn, or that maintenance or replacement is due. For example, the VIA system can monitor continuous or sample temperature, strain, or vibration data for parameters indicating or suggesting that a particular part or component is worn, or that maintenance or replacement is due. For example, the VIA system can monitor continuous or sample temperature, strain, or vibration data for parameters indicating or suggesting that a particular part or component is worn, or that maintenance or replacement is due. For example, the VIA system can monitor continuous or sample temperature, strain, or vibration data for parameters indicating or suggesting that a particular part or component is worn, or that maintenance or replacement is due. The feature extractor may perform a fast Fourier transform or wavelet transform on the 1D vibration sensor data and extract different spectral components or wavelets from that data for analysis by a classifier, which recognizes a particular spectrum or wavelet distribution as normal and other spectra or wavelet distributions as abnormal or irregular. In response to the detection of an abnormal spectrum or wavelet distribution, the classifier sets a maintenance alert for human intervention. The feature extractor may perform a fast Fourier transform or wavelet transform on the 1D vibration sensor data and extract different spectral components or wavelets from that data for analysis by a classifier, which recognizes a particular spectrum or wavelet distribution as normal and other spectra or wavelet distributions as abnormal or irregular. In response to the detection of an abnormal spectrum or wavelet distribution, the classifier sets a maintenance alert for human intervention. The feature extractor may perform a fast Fourier transform or wavelet transform on the 1D vibration sensor data and extract different spectral components or wavelets from that data for analysis by a classifier, which recognizes a particular spectrum or wavelet distribution as normal and other spectra or wavelet distributions as abnormal or irregular. In response to the detection of an abnormal spectrum or wavelet distribution, the classifier sets a maintenance alert for human intervention. The feature extractor may perform a fast Fourier transform or wavelet transform on the 1D vibration sensor data and extract different spectral components or wavelets from that data for analysis by a classifier, which recognizes a particular spectrum or wavelet distribution as normal and other spectra or wavelet distributions as abnormal or irregular. In response to the detection of an abnormal spectrum or wavelet distribution, the classifier sets a maintenance alert for human intervention. The feature extractor may perform a fast Fourier transform or wavelet transform on the 1D vibration sensor data and extract different spectral components or wavelets from that data for analysis by a classifier, which recognizes a particular spectrum or wavelet distribution as normal and other spectra or wavelet distributions as abnormal or irregular. In response to the detection of an abnormal spectrum or wavelet distribution, the classifier sets a maintenance alert for human intervention. In response to the detection of an abnormal spectrum or wavelet distribution, the classifier sets a maintenance alert for human intervention.
[0066] <Training of L-DNN for VIA System Introduction> Unlike conventional deep learning networks, the L-DNN can be quickly trained on a small, unbalanced dataset. Furthermore, unlike conventional deep learning networks that learn through slower backpropagation and forget how to recognize previously learned objects, the L-DNN can learn to quickly recognize new objects without forgetting how to recognize previously learned objects. Unlike conventional deep learning networks, the L-DNN can be quickly trained on a small, unbalanced dataset. Furthermore, unlike conventional deep learning networks that learn through slower backpropagation and forget how to recognize previously learned objects, the L-DNN can learn to quickly recognize new objects without forgetting how to recognize previously learned objects. Unlike conventional deep learning networks that learn through slower backpropagation and forget how to recognize previously learned objects, the L-DNN can learn to quickly recognize new objects without forgetting how to recognize previously learned objects. Unlike conventional deep learning networks that learn through slower backpropagation and forget how to recognize previously learned objects, the L-DNN can learn to quickly recognize new objects without forgetting how to recognize previously learned objects. The network also shows catastrophic forgetting, or the lack of recognition of objects that they were trained to recognize. This allows for the development of deep learning networks to inspect new or different objects. This makes it easier and faster to reconfigure the L-DNN testing system than testing.
[0067] Figures 3A-3C show the VIA platform of Figure 2, broken down into three stages of user interaction. 2 shows an example workflow of the form system 200. This workflow is A process for building and deploying (trained L-DNN) on edge computers 204 Before investing time and effort into mass production of the hardware configuration shown in Figure 2, The VIA System 200's Brain Builder component was used to Prototype the system and verify its feasibility in the "Prototype and Prove" phase30 This step 300 shown in FIG. 3A is performed by the local computer 202 and and / or a local server 204 to create and deploy prototype brains. The method includes inserting the information into the computer, and includes the steps of:
[0068] In the first step 302 of the "Prototype and Proof" phase 300, one and / or other sensors to capture raw data from the assembly line / production environment. When acquiring image data with a camera, a person sets up an image capture environment and uses a camera 208 ( Ensure proper lighting, focus, and quality of the image captured by the microscope (Figure 2). Once the environmental parameters are set, the local computer 202 and / or the local server The server 204 collects images or other data directly from the camera 208 and / or other sensors. For example, these images can be collected from a single production line in a factory with many production lines. This may be from a camera monitoring the area.
[0069] The processor 202 executes the program on the local computer 202 and / or the local server 204. Inbuilda training software provides feedback in the form of predictions as it collects data. It notifies the user of the L-DNN training progress and the success probability of the use case. trains the head of the L-DNN (the backbone or feature extractor can be pre-trained) The user can check the consistency of the predictions (e.g., about 10 "normal" images were collected). (Later) shows how well the L-DNN, also known as the "brain," is learning. At this point, the user can evaluate and analyze whether the object is normal and / or not. Alternatively, the brain can be tested against images of unusual objects (step 306) and trained accordingly. The user can also optionally select a region of interest (described below) (308), Use the training workspace to apply retraining (310), and / or The patient may reassess the brain (304) with more tests (306) if desired. do.
[0070] Once the brain scores acceptable results, the workflow continues with training, fine-tuning, and guidance. (The "Prototype and Proof" stage 300 is a conceptual demonstration.) It may be for demonstration purposes only, and therefore the first training in this stage 320. This step 320 involves creating and training the brain (L-DNN). and introducing it to the edge computer 204. Using the same live view and collection function 302 as in the "Type and Certify" stage 300, Collect a larger and more extensive dataset of normal / good images. These images are automatically are annotated as good / normal and added to the training dataset upon collection. If the prototype scenario is the same or similar to the real scenario, Brain Builder will Images can be automatically annotated based on typing sessions (and more) Most likely correct) (human oversight can provide additional certainty).
[0071] Once the appropriate data has been collected, the user evaluates the trained L-DNN (322), To do this, the user must have a validating center that will be used to test. Add the available "anomalous" examples to the test set and run L-DNN on the validation set to score the accuracy score. The user may then adjust any parameters related to the validity threshold. However, optimization324 should be performed on the endpoint before the brain is deployed326. The validity threshold is the distance between what is considered normal and what is considered abnormal. The validity threshold slider allows the user to identify the specific point at which an object deviates from normal. It can be adjusted enough to be considered abnormal.
[0072] Training can be relatively quick. In some real-world scenarios, module A may Extract features from 1D vectors for classification by Rule B (classifier), which is real-time Module A (Feature Extractor) detects abnormal changes in sensory readings for constant monitoring. ) is obtained from a single sensor from the time domain to the spectral domain using a fast Fourier transform. The resulting spectra were then converted into a 1D time series. The training data is fed to Rule B, which is a 100-example normal operation and a 100-example normal operation. Predictions about 100 cases that produced two cases where an anomaly was classified as normal (false positive). 8 abnormal examples were generated, as well as 8 cases where normal behavior was classified as abnormal (false negatives). The other examples in the trial were correctly classified. Module B contains the user Allows you to fine-tune the balance between false positives and false negatives to match your preferences , there are dominant parameters (discussed below).
[0073] L-DNN can be trained with thousands of fewer examples than traditional deep learning networks. In one example, L-DNN can perform a class-specific training on five samples per class. Trained against an equivalent traditional deep learning network with 10,000 examples per node On average, L-DNN is trained on 20-50 examples per class. More commonly, 100 or fewer normal samples (e.g., 100, 90 , 80, 70, 60, 50, 40, 30, 20, 10, 5, or one normal image) It may be sufficient to train the L-DNN head. Under limited conditions, such as those found in manufacturing tests, L- A DNN can be trained with as little as one example per class.
[0074] The final stage, the inference behavior, recipe and workflow stage 340, involves the L-DNN Connecting the system (brain) to the behavior of a quality control system. Users achieve their accuracy goals. Once the object is detected, the behavior, called inferential behavior, is used to infer whether the object is normal or abnormal (342). Assign inference to the trained L-DNN. For example, the inference behavior is By triggering the pin with a Generally, inference can be performed by the L-DNN / VIA system 200. In response to the detection of abnormalities, normal component / operating conditions or corrective action (e.g., via the PLC210) can be taken. to trigger the pin to warn maintenance, stop the machine, or request parts / packs. Inference behavior does not necessarily have to include actions for the recipe (pushing the cage off the inspection line). 344, which can be assigned to the compute node 209. can be selected and executed via
[0075] FIG. 4 illustrates a project 400, a dataset, an inference behavior 410, a recipe 420, and Project 400 demonstrates the relationship between the use of the Brain (trained L-DNN). Examples, such as images of normal objects, used to train Del(Brain)404 A copy or backup of a brain may be included. The application is configured to communicate with endpoints (e.g., application programming interfaces). Each break is deployed on an edge computer 204, which is accessible via an API. The inversion 404 (brain vX) is associated with an inference behavior 410. Behavior 410, or simply behavior, is the input / output traffic used to communicate with PLC 210. one or more recognized objects and one or more classes of anomalous objects in the transmitter 212. A brain version is generated by assigning relationships between the A brain version is a group of one or more actions that are assigned to a specific When recognizing an object of class ,or flagging an abnormal object, the corresponding edge The node 204 activates the HMI 212 and notifies the PLC 210 of the automatic rerouting of the defective part. changes, triggering an alarm, or stopping the production line, on or along the production line 220. 220 (e.g., removing abnormal parts) occurs.
[0076] The recipe 420 may be implemented by one or more nodes 209 (e.g., nodes 209-1 and 209-2 in FIG. 4). 09-2). The user can select different calculations. A recipe 420 is created by defining a sequence of behaviors 410 to be executed by a node 209. You can create a Brain, each of which is assigned a behavior that correlates to that Brain. Additional workflows can be configured to save inference images for review and to Data can be sent to
[0077] <L-DNN multi-head for industrial visual inspection> Inspect images from multiple cameras or multiple regions of interest (ROIs) in a single image To reduce the inference time when performing the L-DNN, multiple modules B (classifiers or A single module with a pre-trained slow learning backbone that feeds the neural network (the head) It can be implemented as a multi-head L-DNN with A. N is the number of images and / or Or analyze multiple ROIs in a single image. Similar to N, the multi-head L-DNN can handle voice data and scales, thermometers, counters, as well as being trained to work with other types of data, including signals from other sensors. Good too.
[0078] Multi-head L-DNN takes advantage of the fact that data features are consistent within the sensory domain. For example, the image may show edges, corners, color changes, etc. Similarly, different masks in the same environment may Acoustic information captured by the microphone is divided into features based on correlation frequency and amplitude, regardless of the sound source. Therefore, there is no need to train multiple modules A on the same domain. Because they all produce qualitatively similar feature vectors. A single module A is used for each sensor (sometimes per accuracy or speed requirement). Module A can extract features for processing by Module B. Considering that the processing time of the multi-head L-DNN significantly reduces processing costs for inspection systems with multiple sensors within each region. Enables reduction.
[0079] Even if the images come from different cameras at different inspection points, they all undergo feature extraction. are processed through the same module A for extraction and then sorted into different modules for classification and final analysis. The multi-head L-DNN reduces AI latency. , reduced data transfer and network traffic, reduced memory usage, data replication LD features include enhanced privacy and protection, flexible operation, and the ability to handle unknown object classes. It inherits all the advantages of NN.
[0080] Multi-head L-DNN adds the following benefits or advantages over single-head L-DNN: First, multi-head L-DNN uses a single compute host connected to multiple cameras. Second, the multi-head L-DNN can support multiple inspection points with a wider angle. A single image sensor with a single computing host connected to a single high-resolution camera with multiple lenses. It can support multiple ROIs in an image. N allows multiple models to be run in parallel on the same data (image) input (e.g. ,One model to predict orientation, one to predict product variants, and one to identify defects. These multi-head L-DNNs use AI to perform complex visual inspection tasks. This reduces the computational and camera hardware costs required to perform the inference, but increases inference latency. It is within the takt time range.
[0081] FIG. 5A illustrates a machine 5 operating in real time, such as an edge computer 204 (FIG. 2). Multi-head L-DNN506 running on 00. Multi-head L-DNN50 6 includes a single module A 502 created from the original DNN 510. Module A 502 includes a layered neural network layer 512, also called a backbone, and It ends with a readout layer 514. One or more pooling layers and a fully connected layer 516 and a cost The rest of the DNN 510, such as layer 518, is used to pre-train the backbone 512. They do not participate during L-DNN operation. Backbone 502 is a GigE camera. One or more sensors 530, such as sensor 208 or other devices described above with respect to FIG. The system is trained to extract features 515 from input data 501 from a database.
[0082] The backbone 512 of module A 502 distributes the extracted features to several heads 50 4 (Module B), and each head uses a different data set to recognize different features. Similar to module B of the single-head L-DNN, each head 504 is trained with The signals of its input layer 520 and its output layer 524 in a set of one or more associative layers 522 These object representations or feature sets 515 are then used to form associations between the object representations or feature sets 515 Each head 504 can be trained in either a supervised or unsupervised mode, as described above. It operates in a mode that indicates whether the input data is expected (normal) or unexpected (abnormal). The backbone 512 and head 504 can output an inference 531. The combination is called a network, and a multi-head L-DNN is a network Each head 504 is designed to recognize a different set of features. When trained together, the multi-head L-DNN506 performs well on each of these test points. Processes inputs from multiple test points without requiring a specific full DNN trained on it can.
[0083] Figure 5B shows the estimation results of the multi-head L-DNN running on three computing platforms. Speeds are shown in frames per second. Intel OpenVINO, Nvid ia's TensorRT, and Neurala's Custom Software Development Kit (S DK). In all three cases, there is a sublinear decrease in inference speed with the number of heads. For two heads, Figure 5B shows two single-head L-DNNs on the same data. Compared to sequential execution, TensorRT provides a 30% speedup, while Neurala SDK shows a 25% speed improvement, and OpenVINO shows about a 20% speed improvement. In the head, eight single-head L-DNNs are run sequentially on the same data. TensorRT and Neurala SDK speedup of 55%, OpenV The speedup for INO increases to about 40%.
[0084] FIG. 6A is a diagram illustrating a process executed on an edge computer 604 to process an object 21, The multi-head L-DNN506 is configured to perform visual inspections of 23 and 25. The system 600 includes an edge computer 6 608-1 to 608-3) are connected to the Each input source (camera 608) has a corresponding input adapter on each tact of the production line 620. Rays (images) 611-1 to 611-3 are created. These images 611 are batched together. and fed to the DNN backbone 510 of the multi-head L-DNN 506. During the process, the edge computer 604 ensures that the input arrays are all the same size. In addition, the input array can be resized or padded. This allows for modern DNN frameworks The multi-head L-DNN506 allows the entire batch to be processed in parallel. The backbone 112 is a set of feature vectors 621-1 to 621-3. This batch is split into individual feature vectors, each of which corresponds to The module B heads 504-1 to 504-3 are supplied with the It is computationally lightweight and can be used to run module A 502, including edge computer 604. It can be run in parallel on any sufficiently powerful system. and 25 are good for each input (camera or ROI) for each takt of the production pipeline. Inference 53 based on characteristics 521 regarding whether the 1-1 to 531-3, or generate a prediction.
[0085] The head 504 (module B) suitable for the multi-head L-DNN506 has the following features: U.S. Pre-Grant Publication No. 2018 / 0330238A, which is incorporated herein by reference. The fast-learning L-DNN classification head described in Issue 1, as mentioned above, is an anomaly recognition fast-learning L -DNN head, backpropagation-based analysis in traditional DNN or transfer learning-based DNN Based entirely or partially on class heads and L-DNN fast learning or traditional backpropagation learning These include, but are not limited to, a detection head or a segmentation head. All heads of the multi-head L-DNN506 train using the same module A. The components may be homogeneous (same type) or heterogeneous (different types), so long as they are homogeneous.
[0086] Figures 6B and 6C show a multi-head L-DNN VIA system testing different components. These interfaces include the laptop 201, Operable with the HMI 212, or another computer or edge computer 208 The operator can then render the image on the display of the combined display. uses these interfaces to monitor the performance of the VIA system and respond to alerts set by the system and in supervised learning mode to the VIA system This allows operators to label or classify flagged anomalies. The data enables the VIA system to recognize and classify previously unseen anomalies on the fly. It can teach you how to do this.
[0087] In FIG. 6B, the interface is a printed circuit board (PCB) with five different ROIs. B) shows images of the PCB, each containing a different component. A single camera can capture the entire PCB. Then, each ROI is cut out from the original image, and five ROIs are cut out in the same size. resize them into chunks, batch them together, and feed them to the backbone for parallel processing Five different heads evaluated each feature corresponding to the five ROIs, and the results were analyzed according to the image. In Figure 6C, the interface allows you to select different parts of the kit and return a normal or abnormal estimate. Characterize the photos (labeled ROIs (1) to (5)). Images of parts are taken one at a time. by five separate cameras operating simultaneously or at different times, It can be collected by a group of 1 to 5 cameras. Extract features from the five images and map the corresponding extracted features to the five corresponding heads. This head evaluates the ROI of each of the parallel faster throughputs. It is lightweight enough to
[0088] Multi-head L-DNN is suitable for prototypes, custom work, and small volumes, i.e., This is particularly useful in situations where L-DNNs can be trained quickly on small datasets. can flag surface defects, bad welds, and bent pins, eliminating the need for traditional can replace or complement performance testing; there are sufficient defective parts available In this case, or when the head operates in supervised mode, the L-DNN is Faults or defects, including broken solder traces or improperly oriented components It can also be trained to classify types of
[0089] The multi-head L-DNN-based VIA system is also being customized for use on automotive assembly lines. Kits for assembling vehicles with customized features can be inspected because each assembly is different. Any complex component may have components that are becoming more common in automotive manufacturing. This can be done using the assembly process. Kitting is customizable and many It can be used when small parts are used in one assembly, while reducing the complexity on the line. Kitting reduces the number of incorrect parts shipped lineside and improves material handling efficiency. It also reduces the chance of being picked up and used, reducing the time required for the operator to complete the assembly. Make sure you have the materials.
[0090] Although the kitting process is becoming increasingly automated, the most common practice is to have people administer the kits. The kit form factor can vary depending on the type of part, but is typically The kits are often boxes or racks with each type of part in the same location in each kit. Minimal inspection is performed. The process reduces errors. or to minimize (e.g., picking order, use of picking lights, and Computer Control Systems) are often created, but after the kit is created, the correct parts There is often no testing to ensure that the parts are in the kit. It is assumed that the samples are correctly selected and placed in the appropriate bins by the feeding operator. This can result in incorrect parts being used or affecting productivity. This can lead to operators not having the correct parts on the assembly line.
[0091] The kits run a single-head L-DNN, with each kit having its own anomaly model. The operator can inspect the kitting container by the VIA system. Scan the code to tell the system which kit you're pulling the part for and The barcode indicates which kit is being built and which parts are required. If the VIA system is not in the kit or if the wrong parts are placed in the kit, M through ISP to detect abnormalities and prompt operators to check their work. Sending error messages to HMI via odbus TCP.
[0092] Building on a single-head L-DNN implementation, multi-head L-DNN VIA systems can use multiple ROIs to simplify model building and provide more informative quality metrics. Each area of the kitting fixture can be its own area of interest, with each part being its own This simplifies model construction and each head has its own ROI. , which is trained to evaluate different ROIs for each possible part variant. There is no need to build a model for each kit (reducing the amount of training required), and each model can be used to You can configure the kit by selecting the items. Using multi-ROI inspection, VIA not only shows you which parts are inaccurate, but also This allows for improved quality inspection and improved yield. Improved performance and reduced rework.
[0093] Similarly, a multi-head L-DNN can be used to inspect packed cases. This involves placing the cans in a cardboard box and then shrink-wrapping the enclosed, sealed cardboard box. There are various types of case packaging, including pet food, canned food, and It is often used for vegetables, canned soups, etc. This packaging is used for transportation (e.g., when opening the can). to a grocery store and put on the shelf) or for sale as a unit to customers (e.g. in a warehouse) , for bulk purchases). Each is treated as a different ROI, Another method is to use a multi-head L-DNN with different heads. The items may all be from the same image or set of images, with broken shrink wrap or different can shape ( Images of the packaged case, including the shape of the package (without bulges or dents), the folded shape of the package, etc. It can be trained to evaluate different aspects.
[0094] <Dominance parameter: false positive / false negative anomaly detection threshold> How accurately the L-DNN identifies anomalies is partly tunable by the user. The L-DNN head (module B) uses this dominant parameter. Features extracted by the L-DNN backbone (module A) using the parameters Determine how well the vector matches a particular representation of a known (good) object. This is how close the extracted features are to this particular representation relative to all other representations. A dominance value of 10 means that a particular class is accepted as an example. To be considered a candidate, the input features must be more similar to this class than to the prototypes of any other class. This means that the prototype should be 10 times closer to the original. If the dominance parameter is too small, the input is recognized as abnormal. N may report more false positives, i.e., more anomalies than normal objects (Table 1). Also, if the dominance parameter is too large, L-DNN will perform better. It can report more false negatives, i.e., it can identify more normal objects as abnormal. The operator can change this dominance factor in response to system performance. By changing the ratio, the false negative / false positive ratio can be adjusted.
[0095] The following example shows how changing the dominance coefficient affects the performance of L-DNN. The DNN dominance parameter is slanted towards eliminating false positives or false negatives, depending on the user's choice. The human-assisted quality control system can be adjusted to bias the A negative (i.e., identifying anomalies that do not exist) means that the output would pass human review anyway. It is more tolerant of false positives (missed anomalies) but less tolerant of unsupervised human learning. In the case of a fully automated anomaly detection system, the user has the option to set the balance.
[0096] This means that for two different values of the classifier dominance parameter d, some pieces are removed. The results of a pilot study examining chewing gum packs that had been broken, crushed, or replaced do. [Table 1]
[0097] In Table 1, the dominance parameter is the parameter that determines whether the prediction of the classifier is acceptable to other systems. Describe how to compare all possible predictions of . In general, the higher the dominance parameter, the , making the system very detailed. All "normal" values are truly normal, but abnormal values Some normal values may be included. Lower dominant parameters may reduce the system to less detail. No: All "abnormal" values are truly abnormal, but normal values contain some abnormal values. Here, the parameter d balances the false positives (the upper part of Table 1, the three anomalies that were overlooked). ) to false negatives (the bottom of Table 1, where only one abnormality was missed but three normal cases were not) shift to the normal (classified as normal).
[0098] <Conclusion> While various embodiments of the invention have been described and illustrated herein, those skilled in the art will recognize that , for performing the functions described herein, and / or results and / or advantages Various other means and / or structures for obtaining one or more of the following are readily envisioned, and Each of these variations and / or modifications is incorporated herein by reference to the embodiments of the invention described herein. More generally, those skilled in the art will be able to understand all of the methods described herein. Please note that the parameters, dimensions, materials, and configurations are meant to be examples and may not be exact. The data, dimensions, materials, and / or configuration may vary depending on the particular application in which the teachings of the present invention are used. Those skilled in the art will readily understand that the specific inventions described herein are Many equivalents to the embodiments will be recognized or ascertainable using no more than routine experimentation. Accordingly, the foregoing embodiments are presented by way of example only and should not be construed as limiting the scope of the appended claims. Within the scope of the present invention and its equivalents, embodiments of the present invention are as specifically described and patented. It is understood that the invention may be practiced otherwise than as claimed. , each individual feature, system, article, material, kit, and / or method described herein In addition, any combination of two or more such features, systems, articles, materials, kits, and Any combination of such features, systems, articles, materials, kits, and / or methods may be used. and / or methods are included within the inventive scope of the present disclosure, if they are not mutually inconsistent.
[0099] The above-described embodiments can be implemented in any of numerous ways. For example, embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code may be provided to a single computer. Any suitable processor, whether centralized or distributed among multiple computers Or it may be executed on a collection of processors.
[0100] Furthermore, the computer may be a rack-mounted computer, a desktop computer, It can be used in a variety of forms, including laptops and tablets. It should be understood that a computer may be embodied in any of the following: A device that is not considered a computer, but a personal digital assistant (PDA) , smartphone, or any other suitable portable or fixed electronic device , can be embedded in a device with appropriate processing capabilities.
[0101] A computer may also have one or more input and output devices. The device can be used, among other things, to present a user interface. Examples of output devices that can be used to provide an interface include a printer or A display screen for visual representation of the force, and a speaker or speakers for audible representation of the output Other voice generating devices include: Examples of devices include keyboards, mice, touchpads, and digitizer tablets. As another example, a computer may use voice recognition Input information may be received in audible or other audible format.
[0102] Such computers can be part of a local area network or an enterprise network. Wide area networks such as the Internet, and Intelligent Networks (IN) or are communicated over one or more networks of any suitable form, including the Internet. Such networks may be based on any suitable technology, It may operate according to any suitable protocol and may be used in wireless networks, wired networks, , or a fiber optic network.
[0103] The various methods or processes outlined herein may be implemented using a variety of operating systems. on one or more processors using one of the following systems or platforms: In addition, such software may be coded as executable software. The software can be written in a number of suitable programming languages and / or programming or scripting languages. It may be written using any of the available tools, frameworks, or Compiled into executable machine code or intermediate code that runs on a virtual machine This may be done.
[0104] Also, various inventive concepts may be embodied as one or more methods, The actions performed as part of the method may be ordered in any suitable way. As a result, embodiments may be constructed in which the actions are performed in a different order than illustrated. The process may be constructed in a sequential manner, even if shown as a series of acts in the exemplary embodiment. Even if the two actions are performed simultaneously, the two actions may involve some actions being performed simultaneously.
[0105] All publications, patent applications, patents, and other references mentioned herein are incorporated by reference. It incorporates the whole of them.
[0106] All definitions defined and used herein are dictionary definitions, which are incorporated by reference. The document's definitions and / or understandings govern the ordinary meaning of defined terms. It should be.
[0107] As used in this specification and claims, the indefinite articles "a" and "an" are used to clearly indicate that Unless otherwise indicated, "one" should be understood to mean "at least one."
[0108] As used in this specification and claims, the term "and / or" is conjunctive. means "either or both" of the elements, i.e., conjunctively present in some cases. , and in other cases should be understood to mean disjunctively present elements. The elements listed with " / or" are connected in the same manner, i.e., among the coordinated elements. "one or more" shall be interpreted as relating to or relating to the specifically identified elements. other than those elements specifically identified by the "and / or" clause, whether or not Other elements may optionally be present. Thus, as a non-limiting example, "A and / or B" may be used. ", when used in conjunction with open-ended language such as "including," means that in one embodiment In some embodiments, only A (optionally including elements other than B) and in other embodiments, only B (optionally including elements other than B) and in yet another embodiment both A and B (optionally including elements other than A). It can refer to a group of entities (including other elements).
[0109] As used in this specification and claims, "or" means "or" as defined above. and / or" should be understood to have the same meaning as "and / or." For example, items in a list When separating elements, "or" or "and / or" is inclusive, i.e., applies to multiple elements. A list of elements or components, and optionally additional items not in the list, at least "inclusive" means one, but shall be construed to include two or more unless expressly indicated to the contrary. Terms such as "only one of" or "exactly one of," or Only terms such as "consisting of," when used in the claims, shall be used to refer to a number or enumeration of Generally, as used herein, "contains" refers to the inclusion of exactly one of the elements. The term "or" does not mean "either," "one of," "only one of," or When preceded by an exclusive term, such as "exactly one of," "from basic to When used in the claims, "consists of" shall have its ordinary meaning as used in the field of patent law. It shall have a taste.
[0110] As used in this specification and claims, "list" refers to a list of one or more elements. The phrase "at least one" refers to a selection from any one or more of the elements in a list of elements. means at least one element, but not all, specifically enumerated within the list of elements. It does not necessarily contain at least one of every element, but It should be understood that this definition does not exclude any combination of elements. The phrase "at least one" refers to all elements other than those specifically identified in the list of elements. Any element of the Therefore, as a non-limiting example, "A and B" may be at least one of" (or equivalently "at least one of A or B", or (equivalently, "at least one of A and / or B"), in one embodiment B is absent, optionally containing two or more As, at least one A (optionally other than B) In another embodiment, A is absent and optionally contains two or more Bs. , at least one B (optionally including elements other than A), and in another embodiment At least one A, optionally including two or more A's, and optionally two or more It can contain B, refer to at least one B (optionally including other elements), etc. .
[0111] In the claims, as well as in the above specification, all transitional phrases, such as "comprises" "comprising," "including," "carrying" ), "having", "containing", "accompanying" involving, holding, and consisting of "sed of" and the like are understood to be without limitation, i.e., including but not limited to: "consisting of" means that the term "consisting of" includes, but is not limited to, the following: " and "consisting essentially of " is the only transitional phrase set forth in the United States Patent Office Manual of Patent Examining Procedure, Section 2111.03. The phrases in the preceding paragraphs are closed or semi-closed transitional phrases, respectively.
Claims
1. 1. A method for real-time learning of an anomaly detection system in an assembly line during operation without interrupting the operation, comprising: training a fast learning head of a lifelong deep neural network to recognize normal objects on the assembly line using only examples of normal objects; deploying the lifetime deep neural network and running it continuously on the assembly line; acquiring data representative of objects on the assembly line using sensors during operation; using the fast learning head to automatically recognize that a first object is anomalous based on features extracted by a pre-trained backbone; updating the knowledge of the high speed learning head with data representative of the first object without stopping operation of the assembly line; A method comprising:
2. The method of claim 1 , wherein updating the knowledge of the fast learning head is performed locally at an edge device without transmitting data representing the first object to a remote computer.
3. The method of claim 1 , wherein the knowledge update comprises one-shot learning that incorporates data representing the first object into the fast learning head in real time.
4. 2. The method of claim 1, wherein the fast learning head has a classification category of "unknown" to identify that the first object has a type of anomaly that has not been previously seen.
5. acquiring data representative of a second object on the assembly line; automatically recognizing, using the high-speed learning head, that the second object has the same type of anomaly as the first object based on the updated knowledge; The method of claim 1 further comprising:
6. 2. The method of claim 1, wherein the knowledge update occurs within a takt time of the assembly line operation, the takt time comprising a time to produce a part divided by the number of parts required in a given time.
7. The method of claim 1 , wherein the fast learning head uses associative learning to form associations between extracted features and object classifications, without using backpropagation.
8. The method of claim 1 , further comprising incorporating the updated knowledge into previously learned representations without catastrophic forgetting of previously learned anomaly types.
9. The method of claim 1 , wherein the lifetime deep neural network operates in a semi-supervised mode.
10. The method of claim 1 , wherein the knowledge update comprises adjusting a dominance parameter to balance false positive and false negative detection rates.
11. The method of claim 1 , wherein the lifetime deep neural network has multiple fast learning heads.
12. The method of claim 1 , further comprising, after obtaining data representing a plurality of said objects, processing a plurality of said objects in parallel by batch processing.
13. The method of claim 1 , further comprising accepting user input for labeling the first object through a human-machine interface without stopping an assembly line.
14. 1. A system for performing real-time learning during operation of an anomaly detection system in an assembly line without interrupting the operation, comprising: a sensor that acquires data representative of objects on the assembly line; a processor operatively connected to the sensor; Equipped with The processor: a pre-trained backbone that extracts features from the data; a high-speed learning head that is trained only with data representing normal objects and automatically recognizes whether an object is abnormal based on features extracted by the pre-trained backbone; Run The processor further updates the knowledge of the high speed learning head with data representing abnormal objects during operation without stopping operation of the assembly line. system.
15. 15. The system of claim 14, wherein the processor is located on an edge device and updates the knowledge of the fast learning head locally without transmitting data representing anomalous objects to a remote computer.
16. 15. The system of claim 14, wherein the fast learning head has a classification category of "unknown" for identifying objects with previously unseen types of anomalies.
17. The system of claim 14 , wherein the processor performs one-shot learning to update the fast learning head in real time during operation.
18. The system of claim 14 , wherein the processor incorporates updated knowledge into previously learned representations without catastrophic forgetting.
19. The system of claim 14 , wherein the processor runs multiple high speed learning heads that learn different types of anomalies during operation without interference.
20. a human-machine interface operatively connected to the processor for receiving real-time labeling of abnormal objects without stopping operation of the assembly line; the processor automatically incorporates the real-time labeling into the high-speed learning head during operation; The system of claim 14.