Computer-implemented method and system for optimizing clustering of a plurality of input data
Patent Information
- Application Number
- DE102024202027
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-05
- Publication Date
- 2025-09-11
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] The invention relates to a computer-implemented method and / or a system for optimizing clustering of a plurality of input data, which are preferably generated during the automatic optical inspection of at least one manufacturing component. State of the art
[0002] Image classification is well known in the art and involves extracting information from an image to enable the image to be assigned to a specific image category and / or image class. The resulting cluster from such image classification can be used, for example, to create thematic categories.
[0003] While the classification of image data can in principle be performed by a human analyst, it is increasingly being automated and performed using machine learning approaches and / or artificial intelligence methods. With existing methods for classifying image data, misclassifications repeatedly occur due to image-specific, environmental, and / or process-specific conditions. Images that should actually be assigned to a specific category and / or cluster are assigned to a different category by the classification algorithm.
[0004] Automated optical inspection (AOI) is a process that uses image processing systems and machine learning to verify product quality. This process is particularly useful in component manufacturing because it is fast, precise, and cost-effective. Using AOI can reduce defect rates and improve product quality. AOI can be used in a wide variety of industries.
[0005] Semiconductor manufacturing is an exemplary, rapidly growing industry necessary to further advance developments in the field of vehicle electrification. Demand for high-performance and complex semiconductor chips is increasing worldwide. The individual process steps are precisely controlled, and defects must be detected as quickly and reliably as possible to minimize scrap. Due to the large number of chips produced in a manufacturing facility, automated defect detection during AOI is desirable, which can be supported by machine learning. The positions and / or types of defects found during various production steps of individual chips on a wafer can be represented as an image known as a wafer map.
[0006] This enables the detection and / or classification of patterns where multiple defects (possibly of the same type) have been found, forming a signature on the wafer map. These patterns provide insights that can be used to analyze the root cause of defects. For example, if a line is visible at certain locations on the wafer, a device within the production process may have scratched the wafer during a production step. The semiconductor manufacturing process is highly complex and variable. Therefore, new defects can constantly arise for which no or insufficiently labeled (image) data is available. Furthermore, the labeling and / or classification of defects, even for known defect patterns, is an extremely time-consuming task that can often only be performed with the aid of expert knowledge.
[0007] Several approaches for classifying error patterns are known from the state of the art. These include classical supervised learning approaches, transfer learning approaches using neural networks with similarity search, and more novel approaches using autoencoders. This typically requires extensive and well-tuned training of the network architecture. Furthermore, it requires the use of a specific set of known error patterns, which must be precisely characterized or labeled.
[0008] In their scientific publication "CNN features are also great at unsupervised classification," arXiv:1707.01700, Tech. Rep., 2018. [Online]. Available: https: / / arxiv.orgfabs / 1707.01700," Guerin et al. analyzed a two-step approach to image clustering. A convolutional neural network (CNN) is used as the feature extractor. A clustering algorithm is then applied, investigating different combinations of feature extractor and clustering algorithm. In the scientific publication "P. Napoletano, F. Piccoli, and R. Schettini, "Anomaly detection in nanofibrous materials by CNN-based self-similarity," Sensors, vol. 18, no. 1, 2018," a three-step method is investigated with its application to images of nanofibrous materials. A ResNet-18 CNN is used as a feature extractor along with principal component analysis (PCA) and k-means clustering.
[0009] A variety of methods for performing iterative clustering have also been investigated. For example, an iterative Gaussian Mixture Model (GMM) clustering method for scene images was presented in the scientific paper "KN Doan, TT Do, and TH Le, "Scene image clustering based on boosting and gmm," in Proceedings of the Second Symposium on Information and Communication Technology, ser. SolCT '11. New York, NY, USA: Association for Computing Machinery, 2011, p. 226232. [Online]. Available: https: / / doi.org / 10.1145 / 2069216.2069258" is presented, in which the weights of the original dataset are calculated iteratively. Real scenes were evaluated along semantic axes (e.g., degree of naturalness of a scene, degree of verticality, and degree of openness), using predefined filters to extract features from the images. Thus, this method is not adaptively applicable to other use cases.
[0010] In view of the state of the art, which is shown here only as an example and in part, there is still potential for optimization in order to optimize the classification and / or clustering of images of production components, which are captured in particular during an automatic optical inspection, for defect detection.
[0011] The invention is therefore based on the object of further developing a method and / or a system for clustering images of manufactured components to detect component defects in such a way that, in particular, novel and / or unforeseeable defects can be assigned to a cluster, and the method and / or the system, in particular, can also be transferred and / or applied as adaptively as possible to other (image) data as input variables. In other words, one object of the invention is to provide a method and / or a system for optimizing the clustering of a large number of input data.
[0012] The object is achieved by a computer-implemented method for optimizing clustering of a plurality of input data according to the features of patent claim 1. The object is alternatively or additionally achieved by a system for optimizing clustering of a plurality of input data according to the features of patent claim 10. Disclosure of the invention
[0013] According to the invention, a computer-implemented method for optimizing a clustering of a plurality of input data, which are preferably generated during the automatic optical inspection of at least one manufacturing component, is proposed.The method comprises the steps of: providing S1 a plurality of input data items of a manufacturing component to be inspected; extracting S2 at least one feature from the plurality of input data items by applying an extraction algorithm; clustering S3 the plurality of input data items based on the at least one feature by applying a clustering algorithm; evaluating S4 clusters of the clustered plurality of input data items by applying a cluster evaluation algorithm; sorting out S5 input data items from the plurality of input data items that are assigned to at least one cluster with a high rating and / or that exceed a predetermined limit number of input data items within the at least one cluster with the high rating; and in particular iteratively repeating steps S3 to S5 until the plurality of input data items is completely clustered and / or evaluated, and / or until a predetermined termination criterion is reached.
[0014] Furthermore, the invention proposes a system for optimizing the clustering of a plurality of input data items, which are preferably generated by an automatic optical inspection of at least one manufacturing component. The system comprises a provision device configured to provide a plurality of input data items of the manufacturing component to be inspected.The system further comprises an evaluation and computing device configured to extract at least one feature from the plurality of input data items by applying an extraction algorithm; to cluster the plurality of input data items based on the at least one feature by applying a clustering algorithm; to evaluate the clusters of the clustered plurality of input data items by applying a cluster evaluation algorithm; to sort out input data from the plurality of input data items that are assigned to at least one cluster with a high rating and / or that exceed a predetermined limit number of input data items within the at least one cluster with the high rating; and to perform at least the clustering, the evaluation, and the sorting until the plurality of input data items is completely clustered and / or evaluated, and / or until a predetermined termination criterion is reached.If the predetermined termination criterion is reached, it is preferable if the remaining data points are further processed manually or with another (clustering) algorithm.
[0015] The method and / or system according to the invention can increase the accuracy of clustering input data. Furthermore, the method according to the invention offers an unsupervised learning approach, so that labeling of input data or training data is no longer necessary. The method according to the invention can be implemented efficiently, thus achieving increased computing speed and reduced computational effort. The invention also enables the automatic identification of data points that are difficult to cluster, in particular, since these remain after the termination criterion is reached or form only very small clusters after the input data has been fully clustered.By iteratively removing input data that has been assigned to a cluster with a high rating, even input data that is inherently difficult to cluster can be efficiently clustered, since each iteration step "zooms in" on the multitude of input data still to be clustered, with the already clustered data points being excluded. Furthermore, the invention enables faster identification of, in particular, new types of defects. Likewise, the invention can achieve a reduction in production component scrap, since less time is required for defect detection. Particularly preferably, the input data from the cluster is assigned to exactly one cluster. This cluster with the "highest rating" is preferably eliminated.The threshold number preferably only becomes relevant at the end of a process, especially when the iterative process is complete and / or the individual clusters have been sorted, for example, into a "large enough" or "too small" category. The data points from the "too small" category clusters are preferably assigned to the "large enough" category clusters using a neural network.
[0016] According to the invention, a computer-implemented approach is proposed that preferably does not require expert knowledge regarding the input data to be analyzed and is particularly flexible for new and / or unknown defect patterns. According to the invention, no complex training setup is required. The inventive method for optimized clustering of input data is particularly preferably used for the optimized clustering of defect patterns on wafer maps in the field of semiconductor device manufacturing. The inventive method is preferably completely unsupervised.
[0017] A wafer map is preferably a visual tool used to represent data on a semiconductor wafer (e.g., a silicon wafer). A wafer is preferably a thin slice of semiconductor material on which integrated circuits are fabricated. A wafer map preferably shows a location and / or quality of individual chips on a wafer. Each chip is preferably represented by a specific color and / or a symbol indicating its performance data (e.g., voltage, current) and / or its test status (e.g., pass, fail). The wafer map preferably makes it possible to quickly identify problems with individual chips and understand their distribution on the wafer. Wafer maps are preferably used in semiconductor manufacturing to monitor and / or improve the performance and quality of the chips on a wafer.They are a preferred component of process control and / or process characterization and preferably contribute to increasing the efficiency and / or quality of semiconductor manufacturing.
[0018] The method according to the invention preferably uses a cluster evaluation algorithm, such as a silhouette metric, to evaluate individual clusters. The method according to the invention is preferably used to cluster a real industrial wafer map dataset. The results of the method according to the invention can, for example, be compared with a dataset containing expert labels for known, recurring defect patterns. Compared to a non-iterative method from the prior art, the method according to the invention allows for the formation of more homogeneous clusters. This enables pre-sorting of defects in (semiconductor) manufacturing components without requiring prior knowledge and / or the effort required to label a dataset and train a supervised classifier.Thus, the method according to the invention can be used to obtain information about the various defect patterns in a wafer map dataset and / or to use the clustering result as an initial suggestion for manual labeling. It is also possible according to the invention to identify subclasses within wafer map datasets that have been identified and / or labeled by experts as one of the known classes.
[0019] The method according to the invention essentially comprises three steps: On the one hand, feature extraction, for example with the help of a pre-trained convolutional neural network (CNN) together with a dimensionality reduction, for example with the help of PCA, and clustering, for example with the help of active clustering (AC).
[0020] For the unsupervised clustering of input image data, the use of features from a pre-trained CNN is preferred, as this provides an abstract representation of images. The use of CNNs as a feature extraction method has also proven highly effective in many pattern recognition applications, particularly in the field of semiconductor device manufacturing. According to the invention, visually recognizable components of a wafer map are preferably recognized as input data and can preferably be represented as a feature vector. The use of a CNN as a feature extractor preferably provides a large number of abstract features useful for image classification.Subsequently, in a preferred embodiment of the invention, the initially high number of (feature) dimensions is reduced, preferably by performing a principal component analysis (PCA) to identify the features responsible for the greatest variance in the data set. In this case, several, in particular freely selectable, parameters can preferably be set for the dimensionality reduction. These parameters can preferably be set dataset-dependent and / or problem-specific. The dimensionality reduction method itself is preferably merely one component or a preferred sub-step within the clustering method according to the invention.
[0021] For the purposes of this disclosure, the term "plurality" refers to a plurality of input data items. In other words, the phrase "a plurality of input data items" refers to at least two input data items.
[0022] It is understood that the steps according to the invention, as well as other optional steps, do not necessarily have to be performed in the order shown, but can also be performed in a different order. Furthermore, additional intermediate steps can be provided. The individual steps can also comprise one or more substeps without thereby departing from the scope of the method according to the invention.
[0023] In a preferred embodiment, the extraction algorithm comprises a deep neural network, preferably an autoencoder, and / or a particularly pre-trained convolutional neural network, particularly in combination with PCA, and / or an autoencoder in combination with PCA, and / or a supervised, preferably pre-trained, convolutional neural network. For example, an autoencoder, preferably without PCA, can be used for feature extraction. The autoencoder can be retrained based on a reduced data set after each iteration step. The autoencoder can function as a feature extractor. Furthermore, a deep neural network with a dimensionality reduction function can be used. The use of a pre-trained, convolutional neural network (CNN) is particularly preferred.This has the advantage that CNN features are always identical per data point and per iteration step and therefore do not always need to be recalculated. This makes the application of a CNN computationally efficient. Furthermore, a CNN can be supplemented with dimensionality reduction metrics, with PCA being a suitable approach for this, as it is fast and efficient. A combination of an autoencoder and PCA may also be preferred. The use of a supervised learning approach, for example, a supervised pre-trained CNN with supporting dimensionality reduction, may also be preferred. The CNN is preferably trained based on labeled input data and / or by solving an auxiliary classifier problem (transfer learning). Alternatively, dimensionality reduction alone can be used for feature extraction.
[0024] Autoencoders are a type of artificial neural network used to compress and reproduce data. They consist of two main parts: an encoder, which transforms the input data into a compressed coding space, and a decoder, which reproduces the compressed data back into its original form. There are different types of autoencoders, such as the simple autoencoder, the convolutional autoencoder, the variational autoencoder, and the generative adversarial autoencoder.
[0025] Principal component analysis (PCA) is a statistical technique used to capture the structure of complex data. It is designed to reduce the dimensions of the data without losing important information. PCA is primarily used to transform data into a smaller set of new variables, known as principal components, which may exhibit greater variance than the original variables. Preferably, a relatively large portion of the variance is represented by relatively few variables. PCA therefore uses a smaller number of features to analyze than the original number to obtain residuals that can provide information about an anomaly in the semiconductor device under investigation. Inverse PCA may be used to reconstruct data.Then, it is determined whether and where reconstruction errors are particularly high compared to the original data, thereby identifying an anomaly. Preferably, a back-calculation is performed using the feature extractor to depict anomalies in the original domain.
[0026] In a preferred embodiment, the clustering algorithm comprises a GMM algorithm and / or an AC algorithm and / or a DBSCAN algorithm and / or a k-means algorithm. In principle, however, any type of clustering algorithm is conceivable, so the aforementioned algorithms are merely examples.
[0027] GMM stands for Gaussian Mixture Modeling and is a data cluster analysis method. It is a supervised learning method and is used to divide data points into homogeneous groups or clusters. Unlike other clustering methods, such as k-means clustering, GMM assumes that the data points come from multiple Gaussian distributions. Each of these distributions or a combination of individual distributions represents a cluster. Particularly preferably, a value of the probability density function can be determined directly from the GMM model. Alternatively, GMM preferably uses an estimation method to calculate the probability of each data point belonging to a particular cluster. By optimizing the parameters of the Gaussian distributions and the probabilities, GMM can then detect the cluster structure in the data.GMM is particularly well-suited for datasets where clusters are not clearly defined or where clusters are non-spherical. It is also well-suited for datasets with an overlapping structure, where a data point can be part of multiple clusters.
[0028] AC clustering is an abbreviation for "agglomerative clustering." It is a hierarchical clustering analysis technique used for unsupervised learning. Unlike other clustering techniques, such as k-means clustering, agglomerative clustering starts with each data point as a separate cluster and then gradually joins neighboring clusters until only a few large clusters remain. This process preferably continues until only a single large cluster remains, encompassing the entire dataset. AC clustering is easy to implement and well-suited for datasets with a complex cluster structure.
[0029] DBSCAN stands for Density-Based Spatial Clustering of Applications with Noise and is a data cluster analysis method used for unsupervised learning. DBSCAN is a method in which data points are grouped into a cluster if they are close together and have a certain minimum number of neighbors. It can automatically detect the number of clusters and is therefore particularly well suited for datasets with complex cluster structures. The method works primarily by searching the dataset for points with high density and then forming clusters around these points. Points that are not part of a cluster are preferably referred to as noise. DBSCAN is a versatile algorithm and is well suited for datasets with clusters that do not have a spherical shape, as well as for datasets with an overlapping structure where a data point can be part of multiple clusters.
[0030] In a preferred embodiment, the cluster evaluation algorithm comprises a silhouette coefficient metric or silhouette metric and / or another distance metric. The silhouette metric is primarily used to determine the number of clusters and generally to compare different clustering results. Using the silhouette metric for automation in an iterative clustering process is unknown in the prior art.
[0031] The silhouette coefficient is a metric preferably used to evaluate the quality of clustering. The idea behind it is to measure, for each data point, the similarity to its neighbors within the same cluster and preferably to compare it with the similarity to neighbors in neighboring clusters. The silhouette coefficient for a given data point is preferably calculated as follows: Calculating an average similarity (e.g., using a distance metric) between a data point and all other points within the same cluster. Calculating an average similarity between the data point and all points in the nearest neighboring cluster. Calculating a silhouette coefficient for the data point by dividing a difference between the average similarity within the cluster and the average similarity to neighboring clusters by the larger of these two.The silhouette coefficient can preferably take a value between -1 and 1. A value of 1 preferably means that the data point fits very well to its cluster and very poorly to neighboring clusters. A value of -1 preferably means that the data point fits better to a neighboring cluster than to its own cluster. A value of 0 preferably means that the data point fits equally well to its own cluster and to neighboring clusters. The average silhouette coefficient for all data points in a cluster can be used to assess the quality of the overall clustering. A higher average silhouette coefficient preferably means better clustering.
[0032] Distance metrics can include: Euclidean distance: This is the most commonly used distance metric and is calculated by taking the square root of the sum of the squared differences between two points. Manhattan distance: This distance is calculated by calculating the sum of the absolute differences between two points. Chebyshev distance: This distance is calculated by calculating the largest absolute difference between two points. Minkowski distance: This is a more general form of the distance metric that includes both Euclidean and Manhattan distances. Hamming distance: This distance is used to measure the similarity of two binary vectors and calculates the number of positions where the two vectors differ.Jaccard distance: This distance is used to measure the similarity of two nominal or categorical features and calculates the ratio of the number of common features to the number of different features. Mahalanobis distance: This distance takes feature covariance into account and can be used to measure the similarity of two multivariate data points. It should be noted that the choice of distance metric depends on the specific requirements of the problem, and the use of a particular distance metric can influence the outcome of a data clustering analysis.
[0033] In a preferred embodiment, the predetermined termination criterion comprises a minimum number of input data and / or features within a cluster. Other termination criteria are also conceivable.
[0034] In a preferred embodiment, the plurality of input data comprises RGB image data and / or grayscale image data and / or image depth-related image data and / or multispectral image data and / or time series data and / or process curves in which, in particular, physical quantities are plotted against one another (e.g., force versus displacement, which, in contrast to time series, are preferably not dependent on time). In other words, any input data set can be clustered according to the invention, with a data set generated during an automatic and / or automated optical inspection being preferred. The method according to the invention can preferably also be adapted to other high-dimensional data sources.
[0035] In a preferred embodiment, the plurality of input data is acquired by at least one imaging sensor, in particular a camera and / or a multi-camera and / or an ultrasonic sensor and / or a lidar sensor and / or an infrared sensor.
[0036] In a preferred embodiment, after the extraction step S2, a reduction S2' of the dimensionality is performed by applying a dimensionality reduction algorithm, in particular a PCA and / or an autoencoder. Particularly preferably, the step of reducing the dimensionality of the extracted features is also repeated in each iterative loop during the repetition step S6.
[0037] In a preferred embodiment, a computer-implemented method is proposed for verifying a clustering result for a plurality of input data to be verified, which was generated by a machine learning algorithm. The plurality of input data to be verified is clustered on the basis of the inventive method according to one of its embodiments in order to provide a verification result that is compared with the clustering result. The inventive method can thus also be used to verify classification results of a supervised and / or alternative clustering approach, since the invention makes it possible to gradually sort even input data that is difficult to cluster into a suitable cluster based on iteration and / or refinement of the input data by masking out already well-clustered data.In addition, the method according to the invention provides an indication of an underlying class and / or clustering logic due to the iterative approach.
[0038] The invention also claims a computer program with program code for executing at least parts of the method according to the invention in one of its embodiments when the computer program is executed on a computer. In other words, the invention provides a computer program (product) comprising instructions that, when executed by a computer, cause the computer to execute the method / steps of the method according to the invention in one of its embodiments.
[0039] According to the invention, a computer-readable data carrier with program code of a computer program is also proposed for executing at least parts of the method according to the invention in one of its embodiments when the computer program is executed on a computer. In other words, the invention relates to a computer-readable (storage) medium comprising instructions which, when executed by a computer, cause the computer to execute the method / steps of the method according to the invention in one of its embodiments.
[0040] The described designs and further training courses can be combined as desired.
[0041] Further possible embodiments, developments and implementations of the invention also include combinations of features of the invention described previously or below with regard to the embodiments that are not explicitly mentioned. Short description of the drawings
[0042] The accompanying drawings are intended to provide a further understanding of embodiments of the invention. They illustrate embodiments and, in conjunction with the description, serve to explain principles and concepts of the invention.
[0043] Other embodiments and many of the aforementioned advantages will become apparent upon review of the drawings. The elements illustrated in the drawings are not necessarily drawn to scale.
[0044] They show: Fig. 1 is a schematic flow diagram of an embodiment of the method according to the invention; and Fig. 2 a schematic flow diagram of an embodiment of the method according to the invention.
[0045] In the figures of the drawings, the same reference symbols designate the same or functionally equivalent elements, parts or components, unless otherwise stated.
[0046] Fig. 1 shows a schematic flow diagram of a computer-implemented method for optimizing clustering of a plurality of input data, which are preferably generated during the automatic optical inspection of at least one manufacturing component.
[0047] In any embodiment, the method can be carried out at least partially by a system 1, which for this purpose can comprise several components not shown in detail, for example, one or more provision devices and / or at least one evaluation and computing device. It is understood that the provision device can be designed jointly with the evaluation and computing device or can be different from it. Furthermore, the system can comprise a storage device and / or an output device and / or a display device and / or an input device.
[0048] According to the invention, the computer-implemented method comprises at least the following steps: In a step S1, a plurality of input data of a production component to be inspected is provided.
[0049] In a step S2, at least one feature is extracted from the plurality of input data by applying an extraction algorithm.
[0050] In a step S3, the plurality of input data is clustered on the basis of the at least one feature by applying a clustering algorithm.
[0051] In a step S4, clusters of the clustered plurality of input data are evaluated by applying a cluster evaluation algorithm.
[0052] In a step S5, input data that is assigned to at least one cluster with a high rating and / or that exceeds a predetermined limit number of input data within the at least one cluster with the high rating is sorted out from the plurality of input data.
[0053] In a step S6, the steps S3 to S5 are repeated, in particular iteratively, until the plurality of input data n is completely clustered and / or evaluated (|c i | = n), and / or until a predetermined termination criterion A is reached. Here, c describes i the power of cluster i, where the calculation is preferably performed for all i clusters.
[0054] In Fig. 2 shows a further schematic flow diagram of an embodiment of the method according to the invention.
[0055] In contrast to Fig. 1 is carried out in accordance with Fig. 2, after step S2, a further step S2' is performed. In step S2', the dimensionality of the plurality of input data and / or the features extracted therefrom is reduced by applying a dimensionality reduction algorithm, in particular a PCA and / or an autoencoder.
[0056] In step S5, according to Fig. 2 explicitly an advantageous check which cluster is the cluster with the highest rating s max Furthermore, it is advantageous to check whether the cluster with the highest rating s max also a predetermined minimum number n min of data points or input data. The query results in the predetermined minimum number n min that this is not reached, the corresponding input data remain in the data set of the plurality of input data to be clustered and / or are transferred to another data set U. If the query for the predetermined minimum number nmin that this is achieved, the data for this cluster are removed from the data set to be clustered, so that a reduced feature vector F' D is generated.
[0057] Preferably, on the basis of the reduced feature vector F' D a check whether a number of remaining features |F' D | in the reduced feature vector F' D greater than a limit number n PCA of the dimensionality reduction algorithm. If the number of remaining features |F' D | greater than the limit number n PCA , preferably steps S2' to S5 are repeated iteratively. If the number of remaining features |F' D | not greater than the limit number n PCA, the underlying input data is preferably assigned to the dataset U. The dataset U is preferably assigned to a cluster of nearest neighbors (preferably determined using the k-nearest neighbor method).
[0058] The iterative process flow as it is described in Fig. 2, preferably proceeds as follows: First, the described three-stage clustering procedure is carried out. In a first step, the feature vectors F Dfor example, by an Xception CNN for all input data, e.g., all wafer maps, in a dataset D. Subsequently, dimensionality reduction (e.g., PCA) and clustering (e.g., AC) are preferably performed. After the wafer maps have been sorted into clusters, the clusters are preferably evaluated using a silhouette metric. The metric preferably compares a data point with all other data points in the resulting cluster and calculates the distance to the nearest cluster. Thus, with this metric, dense clusters that are far away from other clusters are highly rated. Possible values for the silhouette metric range from -1 to 1. For the cluster with the highest silhouette value, it can be assumed that the wafer maps in this cluster have the highest separability compared to the rest of the dataset and, at the same time, a high similarity within the cluster.If the number of wafer maps in this cluster is greater than the minimum n. min , the preferred cluster is excluded and stacked or placed next to the result data set of the assigned wafer maps A. If the cluster is small (|c i | < n min ), it is preferably not counted as a "good" cluster, but is stacked next to D to form a data set of unassigned input data or wafer maps U and preferably not considered further until the end of the iterations according to the invention. The number of required wafer maps n min in a cluster can preferably be adjusted depending on an underlying use case and / or the own definition of the minimum size of a cluster. If the stopping criterion |F' D | < n PCAof the iterations is met, the remaining input data or wafer maps are preferably assigned to data set U. At this point in the method, data set D is preferably divided into two data sets. Data set A preferably comprises all input data or wafer maps that are sorted into specific clusters determined by the iterative process. Data set U preferably comprises specific input data or wafer maps whose number was classified as too small by the iterative process.
[0059] At the end of the iterative method according to the invention, preferably all input data or wafer maps in the data set U are assigned to an identified cluster of the data set A. This can be done, for example, by identifying a nearest neighbor in the CNN feature vector space of F D∈ A is searched and assigned. Note that the nearest neighbor search is performed in the CNN feature vector space and not in the PCA space, where the components and thus the distances can change from iteration to iteration.
[0060] This final assignment, which occurs after iteratively filtering out good clusters of sufficient size, allows each input file or wafer map to be assigned to a cluster while simultaneously preventing the number of resulting clusters from becoming too large and resulting in some clusters containing only very few input data or wafer maps. This final assignment of input data or wafer maps of the data set U that have ended up in clusters that are too small is not necessarily required. Depending on the application of the proposed iterative clustering method, a separate decision can be made as to whether or not the wafer maps of U are still to be assigned to specific clusters. For example, if it is more important to find relevant pure clusters, this final assignment may not be desired.
[0061] The main idea of this iterative clustering method is to enable PCA, in particular, for different variances within the CNN feature vectors. The visual appearance of defect classes in wafer maps can vary. Therefore, there are defect classes that are expected to be clearly separated using PCA, and others for which PCA may not be able to easily achieve separation, as long as there are defect classes in the dataset that allow a clear separation of the principal components. By iteratively removing input data or wafer maps that are already clustered into good clusters and repeatedly reinitializing PCA on the shrunken dataset, it is possible to analyze variances in the CNN feature space that were not captured in the initial iterations.In this way, corresponding smaller or different variances in the data set can be mapped to the principal components, which somewhat loosens up very closely spaced class distributions and thus simplifies clustering.
[0062] The presented iterative clustering method preferably has certain hyperparameters that can be adapted depending on the dataset and application. These include, for example, the choice of the CNN network architecture, the dimensionality reduction and / or clustering method, and the choice of the distance metric. The freely selectable parameters for the inventive method include, for example, the number of permissible principal components (n PCA ), a number of allowed clusters per iteration (n c ) and / or a minimum number from which a formed cluster is considered large enough (n min ). QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited non-patent literature
[0000] CNN features are also great at unsupervised classification," arXiv:1707.01700, Tech. Rep., 2018. [Online]. Available: https: / / arxiv.orgfabs / 1707.01700
[0008] P. Napoletano, F. Piccoli, and R. Schettini, “Anomaly detection in nanofibrous materials by CNN-based self-similarity,” Sensors, vol. 18, no. 1, 2018
[0008] K. N. Doan, T. T. Do, and T. H. Le, „Scene image clustering based on boosting and gmm," in Proceedings of the Second Symposium on Information and Communication Technology, ser. SolCT '11. New York, NY, USA: Association for Computing Machinery, 2011, p. 226232. [Online]. Available: https: / / doi.org / 10.1145 / 2069216.2069258
[0009]
Claims
[1] Computer-implemented method for optimizing a clustering of a plurality of input data, which are preferably generated during the automatic optical inspection of at least one manufacturing component, the method comprising the steps: - Providing (S1) a plurality of input data of a production component to be inspected; - extracting (S2) at least one feature from the plurality of input data by applying an extraction algorithm; - clustering (S3) the plurality of input data based on the at least one feature by applying a clustering algorithm; - Evaluating (S4) clusters of the clustered plurality of input data by applying a cluster evaluation algorithm; - sorting out (S5) input data from the plurality of input data that belong to at least one cluster c i with a high rating s maxare assigned and / or which have a predetermined limit number n min of input data within the at least one cluster with the high rating; and - in particular iteratively repeating (S6) the steps (S3) to (S5) until the plurality of input data is completely clustered and / or evaluated, and / or until a predetermined termination criterion is reached. [2] Computer-implemented method according to claim 1, wherein the extraction algorithm comprises a deep neural network, preferably an autoencoder, and / or a particularly pre-trained convolutional neural network, in particular in combination with a PCA, and / or an autoencoder in combination with a PCA, and / or a supervised, preferably pre-trained, convolutional neural network. [3] A computer-implemented method according to claim 1 or 2, wherein the clustering algorithm comprises a GMM algorithm and / or an AC algorithm and / or a DBSCAN algorithm. [4] A computer-implemented method according to any preceding claim, wherein the cluster evaluation algorithm comprises a silhouette coefficient metric and / or another distance metric. [5] A computer-implemented method according to any preceding claim, wherein the predetermined termination criterion comprises a minimum number of input data within a cluster. [6] Computer-implemented method according to one of the preceding claims, wherein the plurality of input data comprises RGB image data and / or grayscale image data and / or image depth-related image data and / or multispectral image data and / or time series data and / or process curves. [7] Computer-implemented method according to one of the preceding claims, wherein the plurality of input data is acquired by at least one imaging sensor, in particular a camera and / or a multi-camera and / or an ultrasonic sensor and / or a lidar sensor and / or an infrared sensor. [8] Computer-implemented method according to one of the preceding claims, wherein after the step of extraction (S2) a reduction (S2') of the dimensionality takes place by applying a dimensionality reduction algorithm, in particular a PCA and / or an autoencoder. [9] A computer-implemented method for verifying a clustering result for a plurality of input data to be verified, which was generated by a machine learning algorithm, wherein the plurality of input data to be verified is clustered on the basis of the method according to any one of claims 1 to 8 so as to provide a verification result which is compared with the clustering result. [10] System (1) for optimising a clustering of a plurality of input data, which are preferably generated by an automatic optical inspection of at least one manufacturing component, the system (1) comprising: - a provision device which is designed to provide a plurality of input data of the production component to be inspected; - an evaluation and computing device which is designed ◯ extract at least one feature from the plurality of input data by applying an extraction algorithm; ◯ cluster the plurality of input data based on the at least one feature by applying a clustering algorithm; ◯ to evaluate the clusters of the clustered plurality of input data by applying a cluster evaluation algorithm; ◯ to sort out input data from the plurality of input data that are assigned to at least one cluster with a high rating and / or that exceed a predetermined threshold number of input data within the at least one cluster with the high rating; and ◯ at least carry out the clustering, the scoring and the sorting until the plurality of input data is completely clustered and / or scored, and / or until a predetermined termination criterion is reached. [11] Computer program with program code to carry out at least parts of a method according to one of claims 1 to 9 when the computer program is executed on a computer. [12] Computer-readable data carrier with program code of a computer program for carrying out at least parts of a method according to one of claims 1 to 9 when the computer program is executed on a computer.