Adaptive Learning for Image Classification
By using multiple classification models and weight empowerment methods in the image classification tool, the problem of customers deploying autonomous, self-learning image classification tools on site and maintaining prediction accuracy is solved, and image classification with high accuracy and stability is achieved.
Patent Information
- Application Number
- JP2022557888
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-03-23
- Filing Date
- 2021-03-04
- Publication Date
- 2025-05-12
- Estimated Expiration
- 2041-03-04
AI Technical Summary
It is difficult for the prior art to deploy an autonomous, self-learning image classification tool on the customer site and maintain prediction accuracy in the face of data changes.
By obtaining a set of classification models, each model is used to predict image labels, these models are applied to the calibration dataset to calculate the matching measurements, and applied on the production dataset to detect data drift. At the same time, the classification model is trained by empowering multiple image data sets and selecting a subset of training data.
It realizes the deployment of autonomous and self-learning image classification tools on the customer site, which can maintain prediction accuracy when data changes, reduce dependence on developer information, and improve the accuracy and stability of image classification models.
Smart Images

Figure 0007675093000001 
Figure 0007675093000002 
Figure 0007675093000003
Abstract
Description
[Technical field]
[0001] FIELD OF THE DISCLOSURE This disclosure relates generally to image classification, and more particularly to adaptive learning for image classification. [Background technology]
[0002] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to provisional patent application filed on March 23, 2020, and assigned to U.S. Application No. 62 / 993112, the disclosure of which is incorporated herein by reference.
[0003] Artificial intelligence (AI) is a fundamental tool for many tasks in a variety of computational systems. AI simulates the human intelligence process through machines, computer systems, learning algorithms, etc. The intelligence process may involve acquiring information, learning rules for using that information, using the rules to reason and self-correct to reach near-accurate or definite conclusions. Specific applications of AI include expert systems, speech recognition, machine vision, autonomous driving, intelligent routing in content delivery networks, military simulations, etc.
[0004] The use of AI has become very popular in inspection systems, especially those aimed at identifying and classifying defects in items, or items and products, etc. AI techniques have become an integral part of the technology industry, helping to solve many difficult problems in the manufacturing process. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] International Publication No. 2017 / 201107 [Patent Document 2] US Patent Application Publication No. 2018 / 068224 [Patent Document 3] US Patent Application Publication No. 2005 / 278322 Summary of the Invention [Problem to be solved by the invention]
[0006] One technical problem addressed by the disclosed subject matter is to provide an autonomous, self-learning image classification tool that can be deployed at a customer site and detached from its developer. [Means for solving the problem]
[0007] One exemplary embodiment of the disclosed subject matter is a method that includes obtaining a set of classification models, each classification model configured to predict a label for an image, where the predicted label indicates a class of the image, and each classification model in the set of classification models configured to predict a label from the same set of labels; applying the set of classification models to a calibration dataset of images, thereby imparting an array of predicted labels to each image in the calibration dataset of images, thereby imparting the set of predicted label arrays to the calibration dataset; computing a discrepancy measure for the set of classification models in the calibration dataset, where the discrepancy measure is computed based on the set of predicted label arrays for the calibration dataset, the discrepancy measure being affected by a difference between predictions of the classification models in the set of classification models; applying the set of classification models to a production dataset of images, thereby imparting the set of predicted label arrays to the production dataset; computing a production discrepancy measure for the set of classification models in the production dataset; determining a similarity measure between the production discrepancy measure and the discrepancy measure; and indicating data drift in the production dataset in response to the similarity measure being below a predetermined threshold.
[0008] Another example embodiment of the disclosed subject matter is a method including: acquiring an image dataset, where the image dataset includes a plurality of ordered image sets, where each image set includes images acquired over a time interval, the image sets being ordered according to their respective time intervals; determining a training dataset including images from the image dataset, where the training dataset includes determining weights for the plurality of ordered image sets, where each image set is associated with a weight, where at least two image sets are associated with different weights; and selecting a subset of images from the plurality of ordered image sets for inclusion in the training dataset, where the selection is based on the weights; and training a classification model using the training dataset.
[0009] Yet another exemplary embodiment of the disclosed subject matter is a computing device having a processor adapted to perform the following steps: obtaining a set of classification models, each classification model configured to predict a label for an image, where the predicted label indicates a class of the image, and each classification model in the set of classification models configured to predict a label from the same set of labels; applying the set of classification models to a calibration dataset of images, thereby imparting an array of predicted labels to each image in the calibration dataset of images, thereby imparting the set of predicted label arrays to the calibration dataset; calculating a mismatch measure of the set of classification models in the calibration dataset, where the mismatch measure is calculated based on the set of predicted label arrays for the calibration dataset, the mismatch measure being affected by differences between predictions of classification models in the set of classification models; applying the set of classification models to a production dataset of images, thereby imparting the set of predicted label arrays to the production dataset; calculating a production mismatch measure of the set of classification models in the production dataset; determining a similarity measure between the production mismatch measure and the mismatch measure; and indicating data drift in the production dataset in response to the similarity measure being below a predetermined threshold.
[0010] Yet another exemplary embodiment of the disclosed subject matter is a computer program product comprising a non-transitory computer readable storage medium bearing program instructions that, when read by a processor, cause the processor to execute a method, the method including obtaining a set of classification models, each classification model configured to predict a label for an image, where the predicted label indicates a class of the image, each classification model of the set of classification models configured to predict a label from a same set of labels; applying the set of classification models to a calibration dataset of images, thereby providing an array of predicted labels for each image of the calibration dataset of images, thereby providing the set of arrays of predicted labels to the calibration dataset. applying the set of classification models to a production dataset of images, thereby applying the set of predicted label sequences to the production dataset, calculating a production mismatch measure for the set of classification models in the production dataset, determining a similarity measure between the production mismatch measure and the mismatch measure, and indicating data drift in the production dataset in response to the similarity measure falling below a predetermined threshold.
[0011] Yet another exemplary embodiment of the disclosed subject matter is a computing device having a processor adapted to perform the steps of acquiring an image dataset, where the image dataset includes a plurality of ordered image sets, each image set including images acquired over a time interval, the image sets being ordered according to their respective time intervals; determining a training dataset including images from the image dataset, the training dataset including images, the training dataset including: determining weights for the plurality of ordered image sets, each image set being associated with a weight, at least two image sets being associated with different weights; and selecting a subset of images from the plurality of ordered image sets for inclusion in the training dataset, the selection based on the weights; and training a classification model using the training dataset.
[0012] Yet another exemplary embodiment of the disclosed subject matter is a computer program product comprising a non-transitory computer-readable storage medium bearing program instructions that, when read by a processor, cause the processor to execute a method including: acquiring an image dataset, wherein the image dataset includes a plurality of ordered image sets, each image set including images acquired over a time interval, the image sets being ordered according to their respective time intervals; determining a training dataset including images from the image dataset, wherein determining weights for the plurality of ordered image sets, each image set being associated with a weight, at least two image sets being associated with different weights; and selecting a subset of images from the plurality of ordered image sets for inclusion in the training dataset, wherein the selection is based on the weights; and training a classification model using the training dataset. [Brief description of the drawings]
[0013] The subject matter of the present disclosure will be more fully understood and appreciated in conjunction with the following detailed description, in which corresponding or like numerals or characters indicate corresponding or like components, and in which, unless otherwise indicated, the drawings provide illustrative embodiments or aspects of the present disclosure and are not intended to limit the scope of the present disclosure. [Figure 1] 1 shows a flowchart diagram of a method according to some example embodiments of the disclosed subject matter. [Figure 2A] 1 shows a schematic diagram of an example disparity measurement, in accordance with some example embodiments of the disclosed subject matter; [Figure 2B] 1 shows a schematic diagram of an example disparity measurement, in accordance with some example embodiments of the disclosed subject matter; [Figure 2C] 1 shows a schematic diagram of an example disparity measurement, in accordance with some example embodiments of the disclosed subject matter; [Figure 2D] 1 shows a schematic diagram of an example disparity measurement, in accordance with some example embodiments of the disclosed subject matter; [Diagram 3] 1 shows a flowchart diagram of a method according to some example embodiments of the disclosed subject matter. [Figure 4A] 1 shows a schematic diagram of an example architecture in accordance with some example embodiments of the disclosed subject matter. [Figure 4B] 1 shows a schematic diagram of an example architecture in accordance with some example embodiments of the disclosed subject matter. [Diagram 5] 1 shows a schematic diagram of an exemplary data selection for a dataset, according to some exemplary embodiments of the disclosed subject matter. [Figure 6A] 1 shows a schematic diagram of an exemplary training data selection pattern, in accordance with some exemplary embodiments of the disclosed subject matter. [Figure 6B] 1 shows a schematic diagram of an exemplary training data selection pattern, in accordance with some exemplary embodiments of the disclosed subject matter. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0014] One technical problem addressed by the disclosed subject matter is to provide an autonomous, self-learning image classification tool that can be deployed at a customer site and detached from its developer. In some exemplary embodiments, the self-learning image classification tool can maintain predictive accuracy over time despite potential changes in the data.
[0015] In some exemplary embodiments, production processes in factories, manufacturing plants, and the like may be monitored at customer sites using AI-based tools that utilize image classification. Such AI-based tools may be utilized in inspection systems for identification and classification of defects in items, or items, products, and the like, in the manufacturing process. In some exemplary embodiments, the AI-based tools may be configured to determine, based on visual input, whether a machine is functioning properly, whether the items produced are in accordance with a production plan, and the like. The visual input may include images of products produced by the machine, at different production stages, from different angles, and the like. The AI-based tools may be configured to perform image classification on such images, classify items in the images, and identify defects in the items or products. The AI-based tools may be configured to apply deep learning or other machine learning methods to solve image classification learning tasks. As an example, the AI-based software may be used in automated optical inspection (AOI) in flat panel display (FPD) manufacturing, printed circuit board (PCB) manufacturing, and the like. The AOI process may be based on image classification and other AI techniques.
[0016] In some exemplary embodiments, monitoring the production process, such as by AOI, may be performed at a detached customer site, e.g., at the customer's factory, without providing complete information to the developer of the AI-based tool or its image classification model. The information provided to the developer may be limited to limit information that may be disclosed to other customers, development centers, or other competitive parties associated with the developer. The AI-based tool may be required to operate and adapt without the ability to transmit or share information with other AI software, its developer, research and development centers, and the like. Additionally or alternatively, the AI-based tool may operate in a connected environment, but may be subject to limitations regarding the amount of data it transfers and shares with third parties, including its developer. As a result, training a classification model in such an AI-based tool may be a relatively difficult task due to the relatively little updated training data available at the development site and limited computational resources at the customer site that cannot handle large amounts of training data.
[0017] Another technical problem addressed by the disclosed subject matter is that the predictive accuracy and performance of the image classification model are maintained despite changes that may occur at detached customer sites. In some exemplary embodiments, an AI-based model or image classification model may be developed and trained based on an initial training database, e.g., obtained from a customer or from other sources, and provided to the customer to solve the image classification learning task. Such AI-based models and image classification models may be continuously monitored to maintain predictive quality. The predictive quality of such models may be affected by several factors, e.g., changes in imaging technology, degradation of cameras or other light sensors, changes in lighting conditions, changes in customer processes, use of different materials, changes in data generation processes, changes in the products being produced, etc. As an example, in FPD or PCB manufacturing, image pixels may change due to changes in the color of the manufactured FPD or PCB, by using different materials in production, etc. Such changes may affect the reflectance, transmittance, etc. of the image being classified. As another example, in FPD or PCB manufacturing, rapid changes, even if small, may be continuously made to the products being produced, e.g., to adapt the products to customer needs, to accommodate new design features, etc. As a result, new defects may appear in the product that need to be identified and corrected.
[0018] In some exemplary embodiments, changes in the data may not show up in statistical measurements of the data, but may still need to be distinguished. The new image may have similar statistical characteristics, e.g., distribution of pixels, colors, etc., as the old image. However, the new image may capture new manufacturing defects that may not show up in previously acquired images, recently acquired images, etc.
[0019] In some exemplary embodiments, methods to maintain the predictive accuracy of an AI-based model, such as training a new model every few days, may not provide a sufficient solution. As an example, such retraining may not take into account underlying changes that affect predictive accuracy. As a result, the newly trained model may overfit to a given dataset and may have increasingly poor generalization performance. Thus, the predictive accuracy of the AI-based tool may decrease.
[0020] Yet another technical problem addressed by the disclosed subject matter is to provide continuous AI learning to deal with continuous changes in data. On the one hand, AI-based models may be required to be sensitive to changes in the data to be classified, but on the other hand, they may be required to be stable, e.g., not perturbed by new data. Referring to the above example, in the FPD or PCB industry, deep learning classification models for image defect detection may be utilized. In such industries, rapid changes between manufactured devices may be common, e.g., new devices or parts thereof may be substituted into the scanned designs utilized for training every two or three days. As a result, accurate classification may need to be performed on recent designs and less on older scanned designs. However, older scanned designs may still be relevant for classification tasks with respect to basic components of devices, frequently occurring defects, etc.
[0021] In some exemplary embodiments, AI-based classification models may be consciously retrained on new data to maintain the accuracy rate of their predictions. In some cases, new types of data may be acquired every few hours, every few days, etc. AI learning techniques may be required to handle the new types of data. However, the amount of data increases over time, and the techniques may not be able to train or retrain the classification model using such a large amount of data, which may include overlapping or nearly overlapping images. In some exemplary embodiments, the AI learning techniques may be designed to completely "forget" previously learned data when learning new data. As an example, in response to acquiring a new dataset of data, retraining of the classification model may be performed on the new dataset, while completely forgetting the last training dataset, its oldest samples, being trained using only a portion of the last training dataset, etc. Although such a solution may reduce the amount of training data for retraining, it may result in losing important data for training, such as old data that is more related to the current samples than the new data. Other naive AI retraining algorithms may waste expensive computational resources and time while retraining using the entire amount of data. Additionally or alternatively, an AI retraining algorithm may uniformly retrain the model using a predetermined percentage, e.g., 0%-100%, of all available datasets, but such an algorithm may miss important data relevant to training, use data irrelevant to the current classification task, etc.
[0022] One technical solution is to apply multiple classification models, which can determine whether data drift has occurred based on the consensus of the multiple classification models. In some cases, additional or alternative training may be required to account for data drift. In some exemplary embodiments, the solution may be implemented at a customer site. Additionally or alternatively, the solution may be implemented autonomously.
[0023] In some exemplary embodiments, each classification model may be configured to predict, for an image, a label indicative of its class from the same set of labels. The multiple classification models may include different classification models from different types, several classification models from the same type, etc. As an example, the multiple classification models may include one or more supervised learning models, one or more unsupervised learning models, one or more semi-supervised learning models, one or more automated machine learning (AutoML) generative models, one or more lifelong learning models, one or more ensembles of classification models, hypothetical models trained with one or more noisy labels, etc. In some exemplary embodiments, training of the classification models may be performed using the same training dataset, different training datasets, etc. It should be noted that the set of labels of the classification models may or may not be identical to the set of labels of the predictors utilized in the image classification task. Each classification model may be configured to predict, for an image, a label from the same set of labels of the predictors.
[0024] Additionally or alternatively, the multiple classification models may include one or more in-distribution classification models. The in-distribution classification models may be trained in an unsupervised manner to determine whether the data is from a training distribution. Such in-distribution classification models may be used to determine whether a set of images exhibits the same or similar statistical properties as the training distribution (e.g., "in-distribution"), or exhibits sufficiently different statistical properties (e.g., "out-of-distribution"). In some exemplary embodiments, the in-distribution classification model may be a binary classification model. The binary classification model may be a detector, g(x):X→{0, 1}, that assigns label 1 if the data is from in-distribution, and label 0 otherwise.
[0025] In some exemplary embodiments, the multiple classification models may be applied to a calibration dataset of labeled images. The calibration dataset may include multiple images from different training stages, different production rounds, etc. Each image from the calibration dataset may be given an array of predicted labels, each provided by a different classification model. In some exemplary embodiments, a mismatch measure of the multiple classification models may be calculated in the calibration dataset. The mismatch measure may be calculated based on the multiple arrays of predicted labels in the calibration dataset. The mismatch measure may be affected by differences between the predictions of the classification models of the multiple classification models. The multiple classification models may then be applied to a production dataset of unlabeled images, e.g., a newly acquired set of images. A production mismatch measure of the multiple classification models in the production dataset may be calculated and compared to the mismatch measures of the multiple classification models in the calibration dataset. Data drift in the production dataset may be determined based on a similarity measure between the production mismatch measure and the training stage mismatch measure. In some exemplary embodiments, a similarity measure below a threshold may be indicative of data drift.
[0026] In some exemplary embodiments, a predictor may be utilized to predict a label for an image. The predictor may be applied to a production dataset to predict the label. For example, the predictor may classify an image into a category indicative of a proper production output, an output caused by a malfunction (and the type of malfunction), etc. In some exemplary embodiments, multiple classification models may include a predictor. The predictor may be applied to the calibration dataset and the production dataset along with other classification models, and its predicted labels may be compared as in an array of predicted labels. Additionally or alternatively, the predictor may be excluded from the multiple classification models. Assuming the predictor has been validated against the calibration dataset, the predictor may be assumed to remain valid when applied to the production dataset as long as the production dataset shows similar discrepancy measures of the multiple classification models across the calibration and production datasets.
[0027] In some exemplary embodiments, in response to exhibiting drift in the production dataset, a predictor utilized for the classification task may be retrained. The predictor may be retrained based on at least a portion of the production dataset. The baseline discrepancy measure may be recalculated based on at least a portion of the production dataset utilized for retraining. The retrained predictor may replace a predictor if the retrained predictor has an improved accuracy rate over the predictor on the validation dataset.
[0028] Note that in some cases, lifelong learning of an image classification model may use a fixed set of output classes, meaning that the learning problem does not change over time in terms of the number of classes, reducing training complexity. In some exemplary embodiments, irrelevant classes, e.g., classes for which no samples have been classified, classes for which no new samples have been classified, etc., may be removed from the output classes. Additionally or alternatively, new classes may be created and added over time.
[0029] Another technical solution is to retrain the classifier using an adaptive training dataset selected in a non-uniform manner from data acquired at different time intervals. In some exemplary embodiments, the training data may be acquired over time. Each subset of data may be acquired at a different time interval, e.g., every few hours, every few days, etc. When a new dataset is acquired, e.g., new samples to be classified, a new training dataset may be determined. The new training dataset may include a different proportion of samples from different time intervals.
[0030] In some exemplary embodiments, a plurality of ordered sets of images may be obtained that are utilized to determine the training dataset. In some exemplary embodiments, the plurality of ordered sets may be arranged in ascending order of date. In some exemplary embodiments, the extent of the significance or informativeness of the dataset may be determined or estimated, for example, based on the knowledge of a domain expert or based on the accuracy rate of training a classification model based thereon. In some exemplary embodiments, weights may be determined that indicate the number or percentage of images to select from each set. The weights may be non-uniform, for example, different weights may be assigned to at least two sets. The selection of images to be included in the training dataset may be performed according to the weights, for example, by randomly or pseudo-randomly obtaining samples of each set according to the associated weight. As an example, the weight of the first set may be 10%, indicating the selection of 10% of samples from the population of sets included in the training dataset. The second set may be associated with a weight of 5%, indicating that the training dataset includes 5% of samples selected from the population of the second set.
[0031] As an example, the weight of each set may be determined based on the accuracy rate of a classifier trained on it, the accuracy rate of a classifier trained on a predefined number of consecutive sets ending with the set, etc. The accuracy rate of the classifier may be calculated based on application of the classifier to the most recent set, the set following the related set, etc. As another example, the weight of each set may be associated with the difference between the accuracy rate metric of the set and the accuracy rate metric of the set from a previous time interval. As yet another example, the weights may be determined based on a monotonic pattern associated with the ordering of the sets over time, e.g., a "past means more" pattern where a larger weight may be given to an older set than a recent set, a "past means less" pattern where a smaller weight may be given to an older set than a recent set. On such a principle, a customer may be offered the option of having a trained model that performs better on examples similar to recent samples and less accurately on examples similar to past samples, or vice versa. It should be noted that the sample content curve may have a linear growth or decay, a non-linear growth or decay, an exponential growth or decay, or the like.
[0032] In some exemplary embodiments, a validation set for training may be determined. The validation set may be utilized to calculate the accuracy rate of a classifier, such as a predictor. In some exemplary embodiments, the validation set may be determined inversely to the selected training set. As an example, if the proportion of samples selected from the set for the training set increases from the past to the present, the proportion of samples selected from the set for the validation set may decrease from the past to the present. As an example, if the training set selection pattern is nonlinearly growing past-means-less, the validation set pattern selection may be nonlinearly growing past-means-more. In some exemplary embodiments, weights for the training data set may be determined over the order of the set (e.g., x indicates time, etc.) using a function f(x), while weights for the validation data set may be determined based on the function -f(x). So, if the training data set is determined using a "past-means-more" pattern, the validation data set is determined using a "past-means-less" pattern, and vice versa.
[0033] One technical effect of utilizing the disclosed subject matter is to provide an autonomous solution for adaptive autonomous learning, which is decoupled from external computing and databases and adapts to conditions at customer sites. The disclosed subject matter allows for increasing the accuracy rate of an image classification AI model, without exposing the factory or manufacturing plant data utilizing the AI classifier to the AI model developer or other external parties. The disclosed subject matter provides an automated capability that uses adaptive training, judged by a large and diverse AI model, to stabilize the accuracy rate and user-perceived performance of the image classification AI model under a series of large changes, which may be applied to customer sites via a set of similar models, ensemble models, contradicting models, and data distribution change detection models packaged into a large binary execution AI model. Furthermore, the disclosed subject matter supports a complete learning regularization and moderation process, which is a set of control models that act as a set of external judgements on the predicted results of the main classification predictors.
[0034] Another technical effect of utilizing the disclosed subject matter is shortening the time to market (TTM) required for a product to be conceived and available for sale. TTM can be important in industries where products become obsolete quickly, especially in the world of microelectronics such as FPD, PCB, etc. The disclosed subject matter provides adaptive learning with more accurate training, which improves accuracy rate, robustness, and stability of AI-based software, and enables automatic model performance tracking and enhancement with lower time complexity. The disclosed subject matter shortens TTM by reducing the amount of training data required after each change in the production line.
[0035] Yet another technical effect of utilizing the disclosed subject matter is to increase the proficiency and accuracy rate of an AI-based, lifelong learning classification model that requires continuous retraining. The disclosed subject matter provides a mechanism to "remember" and "forget" samples that go into a retraining dataset, thereby reducing the amount of training data and increasing the accuracy rate of the training. Such a mechanism enables continuous, adaptive, lifelong learning of the AI-based classification model without increasing the computational and time complexity required over time. Furthermore, such a mechanism avoids the retrained classification model from overfitting a given dataset, such as the latest dataset. The disclosed subject matter allows for introducing more generalization performance into the retrained classification model by retraining according to a wide variety of data from different production stages over time. Furthermore, the disclosed subject matter provides an effective validation mechanism, where the selection of validation data is performed in an opposite pattern to the selection of training data. Such a mechanism may reduce the convergence to the classification model and may allow for validation of whether the retraining of the classification model was successful and generalization of the classifier.
[0036] The disclosed subject matter may provide one or more technical improvements over any existing technology, and over any technology previously routine or customary in the art. Additional technical problems, solutions, and advantages will be apparent to those skilled in the art in view of the present disclosure.
[0037] Reference is now made to FIG. 1, which illustrates a flowchart diagram of a method according to some example embodiments of the disclosed subject matter.
[0038] At step 100, a predictor is trained. The predictor may be utilized for image classification tasks. The predictor may be configured to predict a label from a set of labels for an image. In some example embodiments, the predictor may be trained using a training dataset that includes images and their labels indicating the class of each image. The training dataset may include images that have been manually labeled by a human expert, images that have been labeled using a predictor utilized in a previous iteration, images that have been labeled by a classification model and validated by a human expert, etc.
[0039] At step 110, a set of classification models is obtained. Each classification model may be configured to predict a label for an image. The predicted label may indicate a class of the image. Each classification model in the set of classification models may be configured to predict a label from the same set of labels as the predictor.
[0040] In some exemplary embodiments, the set of classification models may be selected from multiple classification models from different types having different parameters, different conditions, different types of learning, etc. The classification models may be trained using a training dataset that may include both labeled and unlabeled images. The training dataset may include a training dataset utilized to train the predictor or a portion thereof. Additionally or alternatively, the training dataset may include both labeled and unlabeled images, such as images believed to be mislabeled, new images from a new production stage, etc. Some learning types may be performed only on labeled images, others on unlabeled image subsets, and still others using both labeled and unlabeled images. As one example, supervised learning using class conditional probabilities may be performed on the labeled image subset. As another example, unsupervised learning of an intradistribution classification model may be performed on the unlabeled image subset. As yet another example, lifelong learning using network plasticity may be performed using the entire training dataset including both labeled and unlabeled images.
[0041] The composition of the set of classification models may be determined randomly, arbitrarily, based on the decisions of a domain expert, based on previous iteration accuracy measurements, based on the composition of training data, etc. As an example, the set of classification models may include three lifetime classification models, each utilizing a neural network with different parameters, two classification models with noise labels, each assuming a different number of noise labels, one within-distribution classification model, etc.
[0042] It should be noted that in some exemplary embodiments, the set of classification models may include the predictors utilized in the classification task. Additionally or alternatively, the predictors may not be members of the set of classification models. The predictors may or may not be applied individually to the calibration data set.
[0043] At step 120, a set of classification models may be applied to the calibration dataset of images.
[0044] In some exemplary embodiments, the calibration dataset may include multiple images from different training stages, different production rounds, etc. The calibration dataset may be a subset of the training dataset utilized to train the classification model, a validation set determined based on the training dataset, a recently acquired subset of images, etc. In some exemplary embodiments, the calibration dataset may be the same dataset used to train a predictor used for the intra-site classification task. Additionally or alternatively, the calibration dataset may be different from the training dataset. It should also be noted that in some cases, the calibration dataset may be unlabeled and the labels of the images therein may not be utilized. Instead, only the predicted labels may be utilized.
[0045] In some example embodiments, each classification model may assign a label to each image in the production dataset, and each image in the image calibration dataset may be assigned an array of predicted labels, resulting in a set of arrays of predicted labels for the calibration dataset.
[0046] As an example, a set of classification models may be named cls1, ..., cls n Each image in the calibration dataset is assigned a value (l1, ..., l n ), where l iare labels from the same set of labels (L), and cls i where (l1, ..., l n ) can potentially be non-uniform.
[0047] In step 130, a mismatch measure for the set of classification models in the calibration dataset may be calculated. The mismatch measure may be calculated based on a set of predicted label sequences in the calibration dataset. The mismatch measure may be calculated based on a percentage of mismatch between the classification models in the set of classification models. Thus, the mismatch measure may be affected by differences between predictions of the classification models in the set of classification models for images in the calibration dataset.
[0048] In some example embodiments, the disparity measure may be calculated based on portions of images to which tuples of classification models of the set of classification models assign different labels. The disparity measure may be determined as the difference between predictions of a pair of classification models of the set of classification models for images in the calibration dataset, as the difference between predictions of a triplet of classification models of the set of classification models for images in the calibration dataset, or as the difference between any other tuples of classification models of the set of classification models for images in the calibration dataset.
[0049] Additionally or alternatively, the disparity measure may be calculated based on a portion of images to which tuples of classification models in the set of classification models assign the same label. The disparity measure may indicate disparity between predictions of tuples of classification models at a particular label, a predetermined number of labels, etc.
[0050] Additionally or alternatively, the discrepancy measure may indicate anomalies in the prediction of at least one classification model compared to the predictions of other classification models in the set of classification models.
[0051] At step 140, the set of classification models may be applied to a production dataset of images. In some example embodiments, each classification model may assign a label to each image in the production dataset. As a result, a set of predicted label sequences may be generated for the production dataset.
[0052] If the set of classification models does not include a predictor utilized for a classification task, the predictor can be applied separately to a production dataset to perform a desired function of the system (e.g., perform an AOI process), for example, in accordance with the disclosed subject matter.
[0053] A production discrepancy measure for the set of classification models in the production dataset may be calculated at step 150. In some exemplary embodiments, the production discrepancy measure for the set of classification models in the production dataset may be calculated similarly to the calculation of the discrepancy measure for the set of classification models in the calibration dataset.
[0054] At step 160, a similarity measure between the production mismatch measure and the mismatch measure may be determined. In some exemplary embodiments, the similarity measure may indicate how similar the production mismatch measure and the mismatch measure are. Without loss of generality and for clarity of explanation, a low number may indicate low similarity, while a high number may indicate high similarity. As an example, the measure may be a number between 0 and 1, with 1 meaning that the two mismatch measures are identical and 0 meaning that they are as different from each other as possible. It is noted that the disclosed subject matter may be implemented using a distance metric that indicates a distance measure between the mismatch measures. In such an embodiment, a low value may indicate similarity, while a high value may indicate dissimilarity.
[0055] At step 170, in response to the similarity measure falling below a predetermined threshold (eg, a minimum similarity threshold not being met), data drift may be indicated in the production data set.
[0056] Note that a decrease in the mismatch measure may indicate data drift, as may an increase in the mismatch measure, for example, when a set of classification models are multiply matched. Note that if the classification models exhibit different mismatch patterns, such differences may indicate that the production dataset is significantly different from the calibration dataset, regardless of whether they are more or less matched with each other. Thus, a predictor accuracy measure determined on the calibration dataset may no longer apply to the production dataset.
[0057] In some exemplary embodiments, a determination may be made whether the similarity measure is below a predetermined threshold, which may be a value of the similarity measure that indicates a minimum similarity between the measurements that may indicate data drift, such as about 20%, 30%, etc.
[0058] In some exemplary embodiments, a similarity measure below a predefined threshold may indicate that the set of classification models provides different mismatch patterns. In some cases, the classification models may match less than before in terms of data drift, while in other cases the classification models match more. Additionally or alternatively, the characteristics of the examples where the different classification models match or mismatch may be changed to account for data drift. As a result, data drift may be identified autonomously, taking into account the change in mismatch measures in a production environment compared to the time of calibration.
[0059] At step 180, in response to determining drift in the production dataset, the predictor utilized in the classification task may be retrained. In some exemplary embodiments, data drift may indicate that training using an existing training dataset is not suitable for the production dataset. A new training dataset for retraining the predictor may be determined according to the production dataset. The new training dataset may be determined based on the production dataset, based on a calibration dataset, based on a combination thereof, etc. Additionally or alternatively, the new training dataset may be determined individually based on a similarity measure, a production discrepancy measure, outlier detection, etc. As an example, the training dataset may be determined using the method of FIG. 3.
[0060] In some exemplary embodiments, once a predictor has been retrained and reaches a sufficient accuracy score, the training set used to train the predictor, the validation set used to validate the predictor, samples thereof, a combination thereof, etc. may be utilized as a new calibration data set to simultaneously re-run steps 120-160.
[0061] Reference is now made to Figures 2A-2D, which show schematic diagrams of example mismatch measurements, in accordance with some example embodiments of the disclosed subject matter.
[0062] In some exemplary embodiments, the set of predicted label sequences 220 of the calibration dataset is represented by the classification models cls1, ..., cls m Let the set of sample1, …, sample n Each image in the calibration dataset contains a value (l1, ..., l k ), where l i are labels from the same set of labels, and cls i where (l1, ..., l n) may potentially have non-uniform values. As an example, the array 223 may include classification models cls1, ..., cls m , the predicted label for the third image.
[0063] In some exemplary embodiments, the calibration datasets {sample1, ..., sample n}, where classification models cls1, ..., cls m A disparity measure for the set of sequences 220 may be calculated. The disparity measure may be calculated based on the set of sequences 220. Thus, the disparity measure may be affected by differences between predictions of the classification models of the set of classification models for the images in the calibration dataset. The disparity measure may be calculated based on statistical measures for the values of the set of sequences 220, their aggregation, the disparity between two or more respective classification models, etc.
[0064] In some exemplary embodiments, the disparity measure may be calculated based on a portion of an image to which tuples of a classification model in a set of classification models assign different labels. As an example, the disparity measure may be calculated based on a portion of an image, e.g., sample, by a subset of classification models, e.g., cls1, cls4, and cls6. 15 , sample 18 , sample 22 , and sample 36 The difference between the labels given to the
[0065] Additionally or alternatively, the mismatch measure may be determined as the difference between predictions of a pair of classification models of the set of classification models for images in the calibration data set, as the difference between predictions of a triplet of classification models of the set of classification models for images in the calibration data set, or as the difference between any other tuples of classification models of the set of classification models for images in the calibration data set. As an example, the graph 240 may represent predictions of cls1 and cls2 on the calibration set. Each point (p1, p2) of the graph 240 may represent diagrammatically a tuple of labels assigned to an image by cls1 and cls2, where p1 is a value representing the label assigned to the image by cls1 and p2 is a value representing the label assigned to the image by cls2. As an example, the point 241 located at (l2, l3) indicates that cls1 assigned the label l2 to an image and cls2 assigned the label l3 to the same image (see also FIG. 2C).
[0066] It should be noted that graph 240 is merely a schematic diagram, showing different examples using different points such that the labeling is continuous and not one of a set of enumerated possible values. Note that in some embodiments, the labels may be continuous values, such as numbers between 0 and 1. Furthermore, such diagrams are useful for clarity of explanation in showing different examples that are similarly labeled by two classifiers.
[0067] The points on the line 242 may represent images where cls1 and cls2 agree on their classification. Each cluster of points, such as clusters 244, 246, and 248, may represent a pattern of agreement or disagreement between cls1 and cls2. As an example, clusters 246 and 248, which are substantially on the line 242, may represent an agreement between a pair of classification models cls1 and cls2 with respect to l2 (246) and l4 (248), while cluster 244 may represent a disagreement between them (e.g., with respect to an example classified as l2 by cls1 and as l4 by cls2). Statistical measures may be performed on the different clusters to determine a disagreement measure. Additionally or alternatively, the disagreement measure may indicate an anomaly in the prediction of at least one classification model compared to the predictions of other classification models in the set of classification models. Cluster 244 may represent such an anomaly between cls1 and cls2.
[0068] Additionally or alternatively, the mismatch measure may be calculated based on a portion of the image where tuples of the classification models of the set of classification models assign the same label. The mismatch measure may indicate the mismatch between the predictions of the tuples of the classification models with respect to a particular label, a predetermined number of labels, etc. As an example, the mismatch measure may be calculated based on the index 280 as shown in FIG. 2D. Each cell of the index 280 includes a match value representing the number or percentage of times that a pair of classifiers assigned the pair of labels. As an example, the value of cell 282 indicates that for 0.2 of the image, the first classification model assigned the label l4 and the second classification model assigned the label l3. The mismatch measure may be calculated based on cells that indicate different labels, e.g., cells that are not on the diagonal of the index 280. It is noted that the index 280 may be exploited in multiple dimensions to represent the match values of multiple classifiers for each pair of labels. Additionally or alternatively, multiple different indices may be utilized to compare, for example, different tuples of classifiers.
[0069] Reference is now made to FIG. 3, which illustrates a flowchart diagram of a method according to some example embodiments of the disclosed subject matter.
[0070] In step 310, an image dataset may be acquired that includes a plurality of ordered image sets. Each set of the plurality of image sets may include images acquired over a time interval. The image sets may be ordered according to their respective time intervals. As an example, a first image set in the ordered plurality may include images acquired in a first time interval, such as Monday. A second image set in the ordered plurality may include images acquired in a second time interval that immediately follows the first time interval, such as Tuesday. It should be noted that a time interval is considered to immediately follow if there are no additional sets with the time interval in between. For example, assuming images are captured only during weekdays, the Monday time interval may be considered to immediately follow the Friday time interval. In some exemplary embodiments, the successor relationship may be based on the ordered plurality of sets and may be determined based on the relative order between the two sets and whether there are additional sets between the two sets.
[0071] In some example embodiments, the images in the dataset may be labeled with a label indicating its class. Additionally or alternatively, some images in the dataset may be unlabeled.
[0072] In some exemplary embodiments, the last image set of the ordered plurality of image sets may be a recent production dataset. The recent production dataset may be obtained during application of the classification model to unlabeled images in a production environment, for example during an AOI process such as FPD manufacturing, PCB manufacturing, etc. In some exemplary embodiments, the plurality of production datasets may be included in the ordered plurality of image sets, each of which may be ordered according to a time interval during which the production dataset was obtained. The plurality of production datasets may be ordered last in the ordered plurality of sets, the recent production dataset may be ordered last among the plurality of production datasets, and the initial production dataset may be ordered first among the plurality of production datasets.
[0073] At step 320, a training dataset may be determined. The training dataset may include images from the image dataset acquired at step 310. The training dataset may be configured to be utilized to train a classification model utilized for the task of classifying images having similar notions as the images in the image dataset.
[0074] At step 322, a weight may be determined for each set of the ordered plurality of image sets. In some exemplary embodiments, the weights may be non-uniform, e.g., different sets having different characteristics may be assigned different weights. However, it should be noted that some of the sets may be weighted similarly or the same.
[0075] In some exemplary embodiments, the weights may define the number of images to take from the set. The weights may define a relative number of images (e.g., a percentage of images), an absolute number (e.g., an exact number of images), etc. As an example, each set of images may be associated with a percentage indicating the proportion of images from the image set that are selected from the image set. A minimum weight value, e.g., zero, a negative value, etc., may indicate to avoid selecting images from the associated image set that is assigned the minimum weight value, e.g., to discard the associated image set. Additionally or alternatively, a maximum weight value, such as 100%, may indicate to keep all images of the associated image set, e.g., to keep the associated set in its entirety.
[0076] In some exemplary embodiments, the weights may be determined to correlate with the order of the image set as defined based on time intervals. In such a case, smaller weights (e.g., weights indicative of a lower percentage) may be determined for older training data sets, and a "past-means-less" pattern may be implemented. Additionally or alternatively, the determined weights may be correlated inversely to the order of the time intervals of the image set. In such a case, higher weights (e.g., weights indicative of a higher percentage) may be determined for older training data sets, and a "past-means-more" pattern may be implemented.
[0077] In some exemplary embodiments, weights for the ordered sets of images may be determined according to a non-uniform monotonic function, such as a strictly monotonically increasing function, a strictly monotonically decreasing function, a linear function, a non-linear function, or the like.
[0078] Additionally or alternatively, weights for the ordered sets of images may be determined based on an expected accuracy rate of a classification model trained thereon. The accuracy rate may be determined by training a classifier on each set of images and determining its accuracy rate. Additionally or alternatively, weights for the ordered sets of images may be determined based on a difference between an expected accuracy rate of a classification model trained on training data including the set of images and training data excluding the set of images. A classifier may be trained for each set of images based on a training dataset, the training dataset including a series of sets of images preceding and including the set of images. A weight for each set of images may be determined based on a difference between an accuracy rate measure of the associated classifier and an accurate measure of a classifier trained using the same training set but without the associated set of images. If the accuracy rate measure improves when the set of images is in the training data, the weight of the associated set of images may be associated with the improved measure. If the accuracy rate measure decreases when the set of images is in the training data, the smallest weight value may be assigned to the second set of associated images.
[0079] Additionally or alternatively, a classifier may be trained using a subseries of a set and tested on successive sets to determine the impact of training with such sets on future sets. Each subseries may include a predetermined number of successive image sets from the ordered set of images, e.g., 5 successive sets, 10 successive sets, 15 successive sets, etc. A classifier trained on a subseries may be applied to one or more image sets from the ordered set of images, where the one or more image sets are ordered immediately after the last set of the subseries according to the order of the image dataset. The classifier may be applied to a different number of different plurality of image sets, e.g., the next image set ordered immediately after the subseries, the next two image sets, the next four image sets, etc. A classifier accuracy measure for each different plurality may be calculated. A selection pattern may be determined based on the different accuracy measures of a classifier trained on each subseries. As an example, the weight of an image set in a subseries may be determined based on the accuracy measure of applying the associated classifier to the next set of images, based on an average of the different plurality of different accuracy measures, etc. As another example, the determination of the number of image sets ordered after a subseries to omit during training may be determined based on an accuracy measure. A minimum weight value may be assigned to the sets ordered after the subseries where the accuracy measure of applying the classifier to them is above a threshold, e.g., above 80%, above 90%, etc. As yet another example, a minimum weight value may be assigned to a sequence of several sets of image sets if the image sets achieve the highest accuracy measure that exceeds a predefined threshold.
[0080] In step 324, a subset of images may be selected from the ordered sets of images. In some exemplary embodiments, the selection of images from each set of images may be performed based on a weight attached to the set. In some exemplary embodiments, a sample of images from each set may be selected. The size of the sample may be determined based on the weight. As an example, the selection may be performed according to a percentage of images selected from the set of images, the percentage being determined based on the weight. A portion of the set of images may be selected according to a ratio defined by a percentage associated with each set. In some exemplary embodiments, if a set is not discarded entirely, a minimum number of examples may be selected therefrom, e.g., at least 10 images, at least 30 images, at least 100 images, etc.
[0081] In some exemplary embodiments, the selection of samples from each image set may be performed in a random manner, a pseudorandom manner, etc. The random sample selected is a narrow subset of the set from which it was selected and has a reduced size compared thereto. Random selection is likely to result in unbiased, statistically valid samples. Additionally or alternatively, in some cases, the selection may be biased, such as to purposefully implement a desired bias.
[0082] In some exemplary embodiments, the training dataset may include exceptional data from a particular time interval, which may not be selected according to a training data selection pattern, but may be relevant to the classification task. As an example, the dataset of time intervals may introduce one or more new classes, remove one or more classes, etc. Such training sets may be important to the classification task despite not being selected according to a selection pattern, because they may adapt the set of classes that the classifier classifies into, thereby improving the accuracy rate of the classification. As another example, datasets that exceed a certain number of consecutive time intervals, e.g., about 10 time intervals, 15 time intervals, etc., may contain samples classified into less prevalent classes and be removed from the training dataset. In some cases, the number of samples classified into a particular category may be reduced in consecutive time intervals (e.g., in the last time interval, fewer samples are labeled with this category). This may indicate that such a category may not exist in the near future and may be forgotten from the entire dataset. Additionally or alternatively, duplicate or similar samples may be removed from the training dataset. To determine such similarity, Deep Clustering, Entropy Maximization, Normalized Cross Correlation, and other algorithms may be applied to data sets from different intervals.
[0083] At step 330, a validation data set and a test data set may be determined. In some exemplary embodiments, the validation data set may include images from the image data set. The validation data set may be configured to be utilized to validate the classification model during training, which is trained at step 330. In some exemplary embodiments, the validation data set may be determined in a manner similar to the determination of the training data set (step 320).
[0084] In some exemplary embodiments, the validation images may be selected from a set of images whose order is used to select the training dataset in step 320. Additionally or alternatively, the validation dataset may be selected from another dataset, such as a previously utilized validation dataset.
[0085] In some exemplary embodiments, the validation data set may be selected by selecting a random sample from each of the ordered sets of images. In some exemplary embodiments, the size of the random sample selected from the set may be determined based on a validation weight in a manner similar to the selection made in step 320. In some exemplary embodiments, the validation weight may be uniform for all sets, such as 10%, 20%, etc. Additionally or alternatively, a possibly different validation weight may be determined for each of the ordered sets of images. In some exemplary embodiments, the validation weight may be determined as opposed to the weight utilized to select the training data set. If the weighting of the ordered sets of images in step 322 is performed based on a function f(x), the validation weight may be determined based on a function g(x)=-f(x). For example, if the weights are monotonically increasing in the training data set, implementing a "past-means-less" pattern, the use of g(x) may cause the weights to correspondingly decrease monotonically in the validation data set, implementing a "past-means-more" pattern.
[0086] In step 340 , a classification model may be trained using the training data set determined in step 320 and the validation data set determined in step 330 .
[0087] In step 350, the test data set may be utilized to test and validate the classification model trained using the training data set. In some example embodiments, the test data set may be used to determine an accuracy measure of the classification model trained in step 340.
[0088] 4A and 4B, schematic diagrams of an example environment and architecture in which the disclosed subject matter may be utilized, in accordance with some example embodiments of the disclosed subject matter, are shown.
[0089] In some demonstrative embodiments, a system including subsystem 400a shown in FIG. 4A and subsystem 400b shown in FIG. 4B may be configured to perform adaptive learning for image classification in accordance with the disclosed subject matter.
[0090] In some example embodiments, the classification model may be initially trained at a development site based on initial training data provided by a customer. The classification model may be utilized in a classification task at the customer site to detect defects based on visual input, such as images of products at different production stages. The classification model may be configured to predict, for each obtained image, a label indicative of its class from a set of predefined classes. Each class may be associated with or indicative of a product defect. The system may be configured to continuously monitor the classification model to improve the accuracy rate of the classification model, retrain the classification model based on updated training data, etc.
[0091] In some demonstrative embodiments, the classification model may be trained based on an initial training dataset 402. The dataset 402 may include sample images, where each image is labeled with a label indicating its class. Additionally or alternatively, the dataset 402 may include unlabeled images. The dataset 402 may include images of production defects, images from previous production rounds, images of different product components, etc.
[0092] In some exemplary embodiments, the data cleansing module 410 may be configured to clean the collected samples and evaluate their labels. The data cleansing module 410 may be configured to select samples from the dataset 402 and apply an “out-of-distribution” classifier to the selected samples. The out-of-distribution classifier may be configured to apply unsupervised learning to the selected samples and determine the distance of the image from the distribution of the data. Images with a distribution similar to the distribution of the data are considered to be in-distribution data X in On the other hand, images with a clear distribution can be classified as out-of-distribution data X out Data in the distribution X in X may be classified and labeled using a classification model, may be labeled by a human expert, etc. The labeled data may be reviewed to determine whether it has been tagged with the correct label. out may be reviewed by human experts to determine whether such outliers are labeled, provided as unlabeled data, removed, etc. In some example embodiments, corrupted images may be indicated as outliers, erroneously taken images (e.g., images taken at the wrong time when the product being inspected is not visible), etc.
[0093] In some demonstrative embodiments, the selected data, which may include both labeled data (x, y) and unlabeled data (X), may be provided to a learning module 420. In the learning module 420, a set of classification models may be trained based on the selected data. Each classification model may be configured to predict a label for an image exhibiting its class with respect to a set of labels in the dataset 402.
[0094] In some exemplary embodiments, the set of classification models may include multiple classification models, such as models 421-426. Each classification model may be trained using a different training scheme, with different parameters, different conditions, etc. As an example, model 423 may be trained by supervised learning using conditional probabilities of classes. Each classification model may be trained using labeled data (x, y), using unlabeled data (X), or a combination thereof, etc. As an example, model 425 performing intra-distribution classification may be trained using unlabeled data (X). The composition of the set of classification models may be determined randomly, arbitrarily, based on previous iteration accuracy measurements, based on the composition of the training data, etc. The composition of the classification models may be different for each iteration.
[0095] In some exemplary embodiments, a mismatch measure for a set of classification models may be calculated. The mismatch measure may be calculated based on the gap (e.g., difference) between the predictions of each pair or tuple of classification models. As an example, gap 430 may be calculated between the prediction results of model 422 and model 423. In some exemplary embodiments, the calculated gap may help understand how robust a set of classification models is. The larger the gap, the larger the mismatch measure and the more useful the classification models may be as discriminants. The mismatch measure may later be compared to a mismatch measure of real-time predicted data to determine how different the images of the production data are from the data used in training. Additionally or alternatively, the mismatch measure may later be compared to a mismatch measure of real-time predicted data to determine the predictive resilience of the classification models to these images.
[0096] Additionally or alternatively, the discrepancy measure may be determined based on other distances between the predicted outcomes, or based on the predicted outcomes of one or more classification models, etc. As an example, the predicted outcome model 425 may indicate a distribution of the data and may be compared to previous predictions of similar classification models, against a predefined threshold, etc.
[0097] Additionally or alternatively, a set of classification models, e.g., models 421, 422, 423, 424, 425, 426, may be aggregated together into a packaged model 435. The model 435 may be aggregated along with the training datasets used to train the set of classification models and other information associated therewith, e.g., its similarity measures, data distribution measures, etc. In some exemplary embodiments, the model 435 may be compared to a model generated in a previous iteration by the learning module 420. If its accuracy rate has improved, the model 435 may be deployed. If not, the previous model may be deployed.
[0098] In some exemplary embodiments, the set of classification models may include a main classification model utilized by the classification module 440 for the classification task. The main classification model may be referred to as a predictor, and may be an ensemble classifier, such as model 422, or other classifiers that predict labels for unlabeled images, such as model 423, model 424, etc. It should be appreciated that it may be preferable for the main classification model to run fast, consume a relatively small amount of resources, etc. In some exemplary embodiments, the main classification model may be configured to determine the probability that each class label is selected in the dataset, for example using a Sofmax layer. For a new production set of images X new In response to obtaining the classification model, the classification module 440 may apply the master classification model to the new production dataset. The images labeled as indicating the presence of defects, e.g., defects in a product, may be provided to the model tracking module 450. In some demonstrative embodiments, all or substantially all of the images of defects may be provided to the model tracking module 450 because the majority of the images may be labeled as not containing defects, such as would be expected in an operational manufacturing plant where most products meet a desired quality level.
[0099] In some exemplary embodiments, a new production set of images X newA predetermined percentage K% of images from the calibration dataset may be provided to the model tracking module 450. The predetermined percentage K% may be 5%, 10%, etc. Such images are utilized as a production dataset that is compared against the calibration dataset to generate a new production set of images X new It may be determined whether the data includes data drift, for example from different image generation process distributions, which may reduce the level of assurance in the accuracy rate of the main classification model, which was trained and validated using the original calibration dataset from which the production dataset drifted.
[0100] In some exemplary embodiments, the model tracking module 450 may new Anomaly in the image of the new production set X new and the main classification model, alerting on a drop in model accuracy, tagging special outlier images, etc. The model tracking module 450 may be configured to apply each of the models 421-426 to a production dataset as a "discriminative" classifier. A production discrepancy measure for the classification model on the production dataset may be calculated and compared to the discrepancy measure for the classification model calculated by the learning module 420 using the calibration dataset. The production discrepancy measure may be calculated in a similar manner to the discrepancy measure, e.g., based on the gap between the predictions of each pair or tuple of classification models, based on a comparison between the results of the in-distribution classifiers, etc.
[0101] In some demonstrative embodiments, if data drift is determined, the production dataset along with the subset of defects may be provided to the training dataset selection module 460. The training dataset selection module 460 may be configured to determine which data to select for retraining the main classification model.
[0102] In some demonstrative embodiments, images that are determined to be special, such as images suspected of being involved in defects, outliers, drifting images, images that do not fit known classes, etc., may be selected for retraining. Such images may be provided to the data cleansing module 410 to be relabeled by human experts.
[0103] Additionally or alternatively, the training dataset selection module 460 may be configured to select images from available data over time. The available data may be stored in several image sets, each image set including images acquired over a certain time interval. As an example, the dataset DS i-k may include images taken in the time interval between time point i and k points back from i (e.g., 8 days back from today). The image sets are ordered according to their respective time intervals, e.g., {DS i-k , D.S. i-k+1 , … , D.S. i-1}, etc.
[0104] In some exemplary embodiments, the training dataset selection module 460 may be configured to apply a dataset forgetting algorithm 462 to the image sets. The dataset forgetting algorithm 462 may be configured to determine which sets should be selected for retraining and which sets should be forgotten, e.g., excluded from the retraining dataset. The dataset forgetting algorithm 462 may be configured to apply different hypotheses to select which datasets to forget, e.g., forgetting old image sets, forgetting image sets that are out of distribution, forgetting image sets that are believed to reduce the overall accuracy rate of the classification model, forgetting image sets that are noisy compared to other sets, etc. Additionally or alternatively, the dataset forgetting algorithm 462 may be configured to apply best practice hypotheses based on industry type, imaging technology, statistical measurements, etc.
[0105] Additionally or alternatively, the training dataset selection module 460 may be configured to apply a dataset mixing algorithm 464 to the set of images selected by the dataset forgetting algorithm 462. The dataset mixing algorithm 464 may be configured to determine which portion of the selected set should be selected for retraining. The dataset mixing algorithm 464 may be configured to apply different selection hypotheses, e.g., a "past-means-more" hypothesis where more images are selected from the old dataset, a "past-means-less" hypothesis where more images are selected from the new dataset, selection based on the possible accuracy rate of training, etc.
[0106] In some exemplary embodiments, the selected images may be provided to the learning module 420 to retrain a set of classification models. Such selected images are a new calibration data set and may be utilized to determine new discrepancy measures.
[0107] Reference is now made to FIG. 5, which illustrates a schematic diagram of an exemplary data selection for a dataset, in accordance with some exemplary embodiments of the disclosed subject matter.
[0108] In some exemplary embodiments, the image dataset 500 may include multiple ordered image sets, databases 510, 520, ..., 590. Each image set may include images acquired over a time interval. Each time interval may be a predefined period related to the pace of a production environment, e.g., a period indicative of a production round, a period indicative of a classification round, etc. Additionally or alternatively, the period may be determined based on the pace of change of the sampled images, e.g., an average period over past samples during which images with different statistical characteristics were obtained, etc. As an example, in the PCB industry, it is common to have rapid changes between manufactured devices, e.g., swapping out scanned designs every 2 days, or every 3 days. Thus, each time interval may consist of a few hours, such as 8 hours, 12 hours, 24 hours, or a few days, such as 2 days, 5 days, or work days.
[0109] In some exemplary embodiments, the multiple image sets may be ordered according to the order of the time intervals in which the images in each set are acquired. As an example, database 510 may include images acquired in the oldest time interval, database 520 may include images acquired in the next time interval immediately following the oldest time interval, etc. Database 590 may include images acquired in a recent (e.g., most recent) time interval. Database 590 may be a production dataset, e.g., a set of images acquired in the last time interval while applying a classification model to unlabeled images in a production environment. Such images may be unlabeled, partially labeled, labeled by a classification model that has not been updated, etc.
[0110] In some exemplary embodiments, the training dataset 505 may be determined based on the image dataset 500. The training dataset 505 may be utilized to train a classification model configured to predict, for each given image, a label indicative of its class. In some exemplary embodiments, the classification model may be utilized in ongoing AI classification tasks with potential changes in the data being classified in each round, such as defect detection in the FPD or PCB industry. The classification model may be periodically retrained to maintain accuracy rates, but may also need to be unconfused by new data. Thus, the training data may need to include both recent and older samples. The amount of samples from each time interval may be modified based on the requirements of the classification task, the nature of the data changes, etc.
[0111] In some exemplary embodiments, a weight may be determined for each image set in the image dataset 500. The weights may be non-uniform and may vary from set to set. The weights may be determined according to different selection patterns, such as, but not limited to, the patterns described in FIG. 6A. In some exemplary embodiments, the selection pattern may be determined based on customer desires, such as providing more accurate classification results for recent designs and less accurate classification results for older scanned designs. The weights may indicate the amount of images selected from each image set (e.g., each time interval). As an example, each weight may be equal to a percentage indicating a portion of the associated image set according to a ratio defined by the percentage. As another example, the weights may be the proportion of images from the image set selected, the maximum number of images selected from the image set, the minimum number of images selected from the image set, etc.
[0112] In some exemplary embodiments, a subset of each image set from the image dataset 500 may be selected to include in the training dataset 505. Each subset may be selected based on a weight determined for the associated set. As an example, a proportion (or ratio) α may be selected from the database 510, a proportion β may be selected from the database 520, etc. Thus, the training dataset 505 may include multiple image subsets, such as subset 515, subset 525, ..., subset 595, etc. Size(S) may indicate the number of images in the image set S. In some exemplary embodiments, Size(Subset 515)=α·Size(Database 510), Size(Subset 525)=β·Size(Database 525), ..., Size(Subset 595)=γ·Size(Database 590).
[0113] Note that at least one pair of weights (or percentages) α, β...γ are different (e.g., α ≠ β). Note further that some percentages may be 0, e.g., no image may be selected from the associated image set.
[0114] Reference is now made to FIG. 6A, which illustrates a schematic diagram of an exemplary training data selection pattern, in accordance with some exemplary embodiments of the disclosed subject matter.
[0115] In some exemplary embodiments, weights determined for an image set (such as the image set of image dataset 500 of FIG. 5) may be associated with the order of the time interval associated therewith. In some exemplary embodiments, smaller weights, indicating a lower percentage of images are selected, may be determined for older training datasets, e.g., patterns 610, 615 ("past means more" patterns). In such a selection pattern, more images are selected from the recent dataset and fewer images are selected from the older dataset. Thus, the classification model may be stronger (e.g., more accurate) for recent samples and less accurate for past samples. Additionally or alternatively, higher weights, indicating a higher percentage of images are selected, may be determined for older training datasets, e.g., patterns 620, 625 ("past means less" patterns). In such a selection pattern, more images are selected from the older dataset and fewer images are selected from the recent dataset. Such patterns may be relevant when data from older datasets is more informative for recent samples than new data, e.g., when new defects are introduced, when defects are repeated from older time intervals, etc. In some exemplary embodiments, the weights may be determined according to a monotonic function of time, which may be a non-uniform function, such as a strictly monotonically increasing function (e.g., patterns 620, 625), a strictly monotonically decreasing function (e.g., patterns 610, 615), a linear function (e.g., patterns 615, 625), a non-linear function (e.g., patterns 610, 620), etc.
[0116] Additionally or alternatively, the weights of an image set may be determined based on an importance measure of the image set, such as the accuracy rate of training based on the image set, the variability between samples, the introduction of new classes, etc. (as in pattern 630).
[0117] In some exemplary embodiments, the weight of each image set may be determined based on the potential accuracy rate of a classification model trained on the image set when a future data set is applied. The potential accuracy rate may be determined by training a classifier on a training data set including the image set and applying the trained classifier to a later obtained image set. The obtained accuracy rate may be compared to the accuracy rate of other classifiers trained without the associated image set. As an example, a first classifier may be trained on a first data set including a consecutive image set. A second classifier may be trained on a second data set including, in addition to the image set included in the first data set, an additional image set ordered immediately after the last image set of the first data set. The weight of the additional image set may be determined based on the difference in the accuracy rate measure of the second classifier compared to the accuracy rate measure of the first classifier. If the difference in the accuracy rate measure of the second classifier compared to the accuracy rate measure of the first classifier indicates an improvement, the weight of the additional image set may be associated with the improvement measure. If the difference in the accuracy rate measure of the second classifier compared to the accuracy rate measure of the first classifier indicates a decrease in accuracy rate, the weights of the additional image set may be assigned a minimum weight value, indicating that selection of any images may be avoided.
[0118] Additionally or alternatively, a classifier may be trained based on each image set and all preceding image sets. The accuracy measure of each classifier may be compared to the accuracy measure of classifiers trained based on previous image sets and preceding sets. As an example, for each new image set acquired in time interval t, a new classifier Cls t may be trained based on all image sets up to the new image set (e.g., the accumulated training data set including the previous data sets (up to time interval t-1)) and the new image set (the set acquired in time interval t). t The accuracy measure of the classifier C1s t-1(e.g., a classifier trained on all image sets up to the image set acquired in time interval t-1) may be compared to a measure of accuracy. The weights determined for the new image set acquired in time interval t may be associated with the difference between the measures of accuracy.
[0119] Additionally or alternatively, weights may be attached to a tuple of image sets (e.g., a series of image sets) as a unit, e.g., a tuple of 5 image sets, a tuple of 10 image sets, etc. Each image set in the tuple may be assigned the same weight. Each tuple of the set may be utilized to train a classifier that is applied to images from future (e.g., successive) image sets. The classifier may be applied to a subseries of image sets ordered immediately after the last set in the tuple according to the ordering of the multiple image sets to determine its accuracy measure. The weight of each tuple may be determined based on an accuracy measure that applies the classifier trained on each tuple to the subseries of successive image sets. Tuples whose accuracy measure exceeds a minimum accuracy threshold may be assigned a minimum weight value, indicating to avoid selecting images from the image sets in the tuple (e.g., discard such image sets from the training data set). Additionally or alternatively, a number of time intervals (e.g., number of image sets) that may be omitted in the retraining process may be determined. A high accuracy measure (e.g., above a predetermined threshold) may indicate the predictive power of the classification model and the lack of need for further training on such data. As an example, a minimum weight value may be assigned to the subseries of image sets that have the same number of sets as the subseries that provide an accuracy measure that exceeds a predetermined similarity threshold.
[0120] In some exemplary embodiments, one or more exceptions may be applied to the selection pattern. As an example, an image set that introduces a new classification category (e.g., a new class) may be selected to be in the training dataset as a whole and may be weighted the most, introducing the new category with retraining. As another example, an image set that precludes classification into one or more classification categories (e.g., no images classified into one or more classes) may be discarded from the training dataset. Additionally or alternatively, images from all image sets classified into excluded classes may be removed from the image set and thus not selected to be in the training dataset. As yet another example, classes with an imbalanced number of classified images over time, classes with a decreasing number of classified images over time, etc. may also be considered deleted and images classified therein may be removed.
[0121] Additionally or alternatively, duplicate or similar images may be removed from the training dataset. Statistical techniques such as deep clustering, entropy maximization, normalized cross-correlation between dataset images may be applied to selected images to determine the similarity between them. Images having a similarity measure above a predefined threshold, e.g., above 90%, above 95%, etc., may be removed from the training dataset.
[0122] Referring now to FIG. 6B, a schematic diagram of an exemplary training and validation dataset selection pattern is shown, in accordance with some exemplary embodiments of the disclosed subject matter.
[0123] In some exemplary embodiments, a validation data set of images may be selected based on the ordered sets of images. A validation weight may be determined for each set of images. A validation set of images may be selected from the ordered sets of images based on its validation weight.
[0124] In some exemplary embodiments, the validation weights may be determined in a selection pattern opposite to that of the training set selection.
[0125] As an example, graph 650 may represent a selection pattern for training and validation data sets. The training selection pattern may follow a function f(t) shown in curve 660, which may represent weights of image sets over time, percentage of images selected from each image set over time, etc. The validation data set selection pattern may follow a function g(t) shown in curve 670, where g(x)=-f(x). The training selection pattern shown in curve 660 is a non-linear "past-means-more" pattern. The validation data set selection pattern shown in curve 670 is a non-linear "past-means-less" pattern.
[0126] Additionally or alternatively, a test dataset for testing the retrained classification model may be selected based on multiple ordered image sets. As an example, the test dataset may include a predetermined percentage of images from each image set, e.g., 10%, 15%, etc. from the image sets.
[0127] The present invention may be a system, a method, or a computer program product. The computer program product may include a computer readable storage medium having instructions for causing a processor to perform aspects of the present invention.
[0128] A computer readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. The computer readable storage medium may be, for example, but not limited to, electronic, magneto-optical storage devices, including but not limited to hard disks, random access memory (RAM), read only memory (ROM), erasable programmable read only memory (EPROM or flash memory), etc. In some cases, instructions may be downloadable to the storage medium from a server, a remote computer, a remote storage device, etc.
[0129] The computer readable program instructions for carrying out the operations of the present invention may be either assembler instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, or source or object code written in any combination of one or more programming languages, including object oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as the "C" programming language or similar programming languages. The program instructions may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be to an external computer (e.g., via the Internet using an Internet Service Provider).
[0130] Aspects of the present invention are described herein with reference to flowcharts and block diagrams of methods, apparatus, systems, and computer program products. It will be understood that each block of the diagrams, and combinations of blocks in the diagrams, can be implemented by computer readable program instructions.
[0131] The computer readable program instructions can be loaded into a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be executed on the computer, other programmable apparatus, or other device to produce a computer-implemented process, such that the instructions executing on the computer, other programmable apparatus, or other device implement the functions specified in the diagram blocks.
[0132] The flowcharts and block diagrams in the figures illustrate possible implementations of various embodiments of the present invention. In this regard, each block in the flowcharts or block diagrams may represent a module, segment, or portion of instructions, including one or more executable instructions for implementing a specified logical function(s). In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may in fact be executed substantially simultaneously, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved. It should also be noted that each block and combination of blocks in the figures may be implemented by a special-purpose hardware-based system.
[0133] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting of the disclosure. As used herein, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used herein, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0134] Corresponding structures, materials, acts, and equivalents of all means or step-plus-function elements in the following claims are intended to include any structure, material, or act for performing a function in combination with other claimed elements as specifically claimed. The description of the present invention has been presented for purposes of illustration and description, and is not intended to be exhaustive or to limit the invention in the disclosed form. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the invention. The embodiments have been selected and described in order to best explain the principles and practical application of the invention, and to enable those skilled in the art to understand the invention in terms of various embodiments with various modifications suitable for the particular use envisaged.
Claims
1. obtaining a set of classification models, each configured to predict a label for an image, where the predicted label indicates a class of the image, and where each classification model in the set of classification models is configured to predict a label from a same set of labels; applying the set of classification models to a calibration data set of images, thereby assigning an array of predicted labels to each image in the calibration data set of images, and assigning the set of predicted label arrays to the calibration data set; calculating a disparity measure for the set of classification models in the calibration data set, where the disparity measure is calculated based on a set of sequences of predicted labels in the calibration data set, and the disparity measure is influenced by differences between predictions of classification models in the set of classification models; applying the set of classification models to a production dataset of images, thereby assigning a set of predicted label sequences to the production dataset; calculating a production discrepancy measure for the set of classification models on the production dataset; and determining a similarity measure between the production mismatch measure and the mismatch measure, a lower value indicating lower similarity and a higher value indicating higher similarity; indicating data drift in the production data set in response to the similarity measure falling below a predetermined threshold, the value being indicative of a minimum similarity between the measures; A method comprising:
2. 10. The method of claim 1, further comprising: applying a predictor to the production dataset, where the predictor is configured to predict a label from the same set of labels for an image.
3. The method of claim 2 , wherein the set of classification models excludes the predictor.
4. The method of claim 2 , wherein the set of classification models includes the predictor.
5. 3. The method of claim 2, further comprising, in response to the indication of the data drift in the production dataset, retraining the predictor based on at least a portion of the production dataset.
6. The method of claim 5 , further comprising recalculating the disparity measure based on the at least a portion of the production data set.
7. The method of claim 1 , wherein the disparity measure is calculated based on portions of an image to which tuples of classification models in the set of classification models assign different labels.
8. The method of claim 1 , wherein the disparity measure is calculated based on portions of images to which tuples of classification models in the set of classification models assign the same label.
9. The set of classification models is 1 , …, cls n where the discrepancy measure is determined by the calibration data set being a value (l 1 , …, l n ), where l i is a label from the same set of labels, and cls i where (l 1 , …, l n 2. The method of claim 1, wherein the values of may be non-uniform.
10. 13. A computer program product comprising a non-transitory computer readable storage medium carrying program instructions, the computer program product being configured to perform the method of claim 1.
11. 1. A method comprising: acquiring an image dataset, wherein the image dataset includes a plurality of ordered image sets, each image set including images acquired over a time interval, the image sets being ordered according to their respective time intervals; determining a training dataset, wherein the training dataset includes images from the image dataset, said determining the training dataset comprising: determining weights for a plurality of image sets having the order, wherein each image set is associated with a weight, and wherein at least two image sets are associated with different weights; and selecting a subset of images from the ordered set of images to be included in the training data set, said selection being based on said weights; training a classification model using the training dataset; Including, a final set of images in the ordered plurality of image sets are a plurality of production datasets, wherein the production datasets are obtained during application of the classification model to unlabeled images in a production environment, and each of the plurality of production datasets is ordered according to a time interval during which the production dataset was obtained.
12. 12. The method of claim 11, wherein the weights are percentages whereby each image set is associated with a percentage, and wherein the selecting of the image subsets is performed by selecting from each image set a portion of the image set according to a ratio defined by the percentage associated with each image set.
13. The determining the proportion of images of each selected image set that are included in the training data set comprises: selecting a ratio associated with an order of said time intervals, whereby a lower ratio is determined for older training data sets; The method of claim 11 , comprising:
14. The method of claim 11 , wherein the determining the weights for the ordered sets of images comprises determining the weights according to a monotonic function.
15. The monotonic function is Strictly monotonically increasing function, Strictly monotonically decreasing function, A linear function, or 15. The method of claim 14, wherein the non-uniform function is selected from the group consisting of:
16. Determining the weights for the ordered sets of images includes: training a first classifier on a first set of images, wherein the first classifier is trained based on a first dataset, the first dataset including the first set of images; training a second classifier on a second set of images, wherein the second classifier is trained based on a second data set, the second data set consisting of the first data set and the second set of images, the second set of images being ordered immediately after the first set of images according to the ordered plurality of image sets; determining the weights for the second set of images based on a difference in the accuracy measure of the second classifier compared to the accuracy measure of the first classifier; The method of claim 11 , comprising:
17. 17. The method of claim 16, wherein the difference in the accuracy measure of the second classifier compared to the accuracy measure of the first classifier indicates an improvement, and the weights for the second set of images are associated with the improvement measure.
18. 17. The method of claim 16, wherein the difference in the accuracy measure of the second classifier compared to the accuracy measure of the first classifier indicates a decrease in accuracy, and the weights of the second set of images are minimum weight values, and wherein the selecting the subset of images comprises avoiding selecting any images from the set of images assigned the minimum weight value, thereby discarding the second set of images.
19. The determining of the weights for the ordered set of images is performed based on a function f(x), and the method further comprises: determining verification weights for the ordered set of images, the verification weights being determined based on a function g(x), where g(x)=-f(x); selecting a validation data set of images from the ordered set of images, wherein the selecting of the validation data set of images is based on the validation weights; validating the classification model utilizing a validation dataset of images; The method of claim 11 , comprising:
20. determining a first sub-series of image sets from the ordered plurality of image sets, wherein the first sub-series includes a predetermined number of consecutive image sets; training a classifier configured to predict a label for an image using the first subseries, where the label indicates a class of the image; applying the classifier to a second sub-series of sets of images from the ordered set of images, where the second sub-series has a number of sets, and a first set of the second sub-series is ordered immediately after a last set of the first sub-series according to the ordered set of images; determining an accuracy measure of the classifier in the second subseries; and wherein said determining weights for said ordered sets of images is based on said accuracy measure; The method of claim 11 , comprising:
21. 21. The method of claim 20, wherein the determining weights for the ordered set of images includes assigning a minimum weight value to the set of the second sub-series based on the accuracy measure being above a threshold, and the selecting the subset of images includes avoiding selecting any image from the set of images assigned the minimum weight value, thereby discarding the second sub-series.
22. 21. The method of claim 20, wherein the determining weights for the ordered plurality of image sets comprises assigning a minimum weight value to a plurality of sequences of image sets, each sequence of the plurality of sequences having a number of sets.
23. A computer program product comprising a non-transitory computer readable storage medium carrying program instructions, the computer program product being configured to perform the method of claim 11.
24. 1. A computing device having a processor, the processor comprising: obtaining a set of classification models, each classification model configured to predict a label for an image, where the predicted label indicates a class of the image, and each classification model in the set of classification models configured to predict a label from the same set of labels; applying said set of classification models to a calibration data set of images, thereby assigning an array of predicted labels to each image in said calibration data set of images, and assigning a set of predicted label arrays to said calibration data set; calculating a discrepancy measure for the set of classification models for the calibration data set, where the discrepancy measure is calculated based on the set of predicted label sequences for the calibration data set, and the discrepancy measure is influenced by differences between predictions of classification models in the set of classification models; applying the set of classification models to a production dataset of images, thereby assigning a set of predicted label sequences to the production dataset; calculating a production discrepancy measure for the set of classification models on the production dataset; determining a similarity measure between the production mismatch measure and the mismatch measure, a lower value indicating less similarity and a higher value indicating more similarity; indicating data drift in the production data set in response to the similarity measure falling below a predetermined threshold, the value being indicative of a minimum similarity between the measures; A computing device adapted to execute the
25. 1. A computing device having a processor, the processor comprising: acquiring an image dataset, wherein the image dataset comprises a plurality of ordered image sets, each image set comprising images acquired over a time interval, the image sets being ordered according to their respective time intervals; determining a training dataset, wherein the training dataset includes images from the image dataset, said determining the training dataset comprising: determining weights for a plurality of image sets having the order, wherein each image set is associated with a weight, and wherein at least two image sets are associated with different weights; and selecting a subset of images from the ordered set of images to be included in the training data set, said selection being based on said weights; training a classification model using the training dataset; [0033] a final set of images in the ordered plurality of image sets are a plurality of production datasets, wherein the production datasets are obtained during application of the classification model to unlabeled images in a production environment, and each of the plurality of production datasets is ordered according to a time interval during which the production dataset was obtained.
Citation Information
Patent Citations
Imaging apparatus and control method of the same
JP2019106694A
System and method for mining time-changing data streams
US20050278322A1
Model based data processing
US20180068224A1
Predictive drift detection and correction
WO2017201107A1