Computer-Implemented Method for Defect Analysis, Device for Defect Analysis, Computer Storage Medium, and Defect Analysis System

Through the computer-implemented defect analysis method, composite images are generated and clustered analysis is performed, which solves the problem of low defect tracking efficiency in display panel manufacturing, and realizes rapid identification of defect causes and device correlation analysis.

CN114930385BActive Publication Date: 2025-07-18BOE TECHNOLOGY GROUP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080003638.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-12-03
Publication Date
2025-07-18
Estimated Expiration
2040-12-03

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently analyze the root causes of defects in display panel manufacturing process, resulting in inefficient defect tracking and cause analysis.

Method used

Through computer-implemented methods, multiple sets of defect point coordinates are obtained, composite images are generated, and clustering analysis is performed to classify defect points into multiple point clusters. The fitting algorithm is used to generate occlusion areas and feature vectors to identify potential devices that cause defects.

Benefits of technology

Improves the efficiency and accuracy of defect analysis, enables rapid identification of the correlation between defects and manufacturing equipment, and supports process improvement and equipment maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114930385B_ABST
    Figure CN114930385B_ABST
Patent Text Reader

Abstract

A computer-implemented method for defect analysis is provided. The computer-implemented method includes: obtaining multiple sets of defect point coordinates, each set of the multiple sets of defect point coordinates including the coordinates of defect points in each of multiple substrates, and the coordinates of the defect points in each substrate being coordinates in an image coordinate system; combining the multiple sets of defect point coordinates according to the image coordinate system into a composite set of coordinates to generate a composite image; and performing clustering analysis to classify the defect points in the composite set in the composite image into multiple point clusters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to display technology, and more particularly, to a computer-implemented method for defect analysis, a device for defect analysis, a computer storage medium, and a defect analysis system. Background Art

[0002] Due to improved performance and load capacity, high availability and failover, and faster access to data, distributed computing and distributed algorithms have become prevalent in various contexts. With the development of technologies such as big data, cloud computing, and artificial intelligence, technologies related to big data analysis have been widely applied in various fields of manufacturing. Summary of the Invention

[0003] In one aspect, the present disclosure provides a computer-implemented method for defect analysis, including: obtaining multiple sets of defect point coordinates, each set of the multiple sets of defect point coordinates including the coordinates of defect points in each of multiple substrates, and the coordinates of defect points in each substrate being coordinates in an image coordinate system; combining the multiple sets of defect point coordinates according to the image coordinate system into a composite set of coordinates to generate a composite image; and performing clustering analysis to classify the defect points in the composite set in the composite image into multiple point clusters.

[0004] Optionally, the computer-implemented method further includes obtaining multiple selected point clusters from the multiple point clusters; wherein the number of defect points in each of the multiple selected point clusters is greater than a threshold number.

[0005] Optionally, the computer-implemented method further includes: respectively determining multiple contours of at least multiple selected point clusters among the multiple point clusters, each contour of the multiple contours including multiple edge defect points in each of the multiple selected point clusters; applying a fitting algorithm to the edge defect points of the multiple selected point clusters to generate multiple occlusion regions respectively corresponding to the multiple selected point clusters; and respectively generating multiple feature vectors of the multiple occlusion regions.

[0006] Optionally, generating the multiple feature vectors includes: generating the Hu geometric moment m i,j and the center-to-center distance M i,j of each of the multiple occlusion regions, where m i,j = ∑ (x,y)∈A x i y j ; calculating the defect point density ρ, area a, centroid O(O x , O y ) and direction θ of each of the multiple occlusion regions; and generating each feature vector of the multiple feature vectors of each of the multiple occlusion regions.

[0007] Optionally, each of the multiple feature vectors is represented as:

[0008] F = [ρ, a, Ox, Oy, θ, L, W, r] T ,

[0009] where N represents the number of defective points in each of the multiple occlusion regions, and a represents the area of each of the multiple occlusion regions;

[0010]

[0011]

[0012]

[0013] M ij = ∑ (x,y)∈A (x - O x ) i (y - O y ) j ;

[0014]

[0015] L represents the length of the minimum bounding rectangle of each of the multiple occlusion regions; and W represents the width of the minimum bounding rectangle of each of the multiple occlusion regions.

[0016] Optionally, a method based on α Shapes is used to determine the multiple contours.

[0017] Optionally, the computer-implemented method further includes: allocating one or more selected occlusion regions among the multiple occlusion regions as multiple defect aggregation regions; wherein, the feature vectors of the one or more selected occlusion regions satisfy a threshold condition.

[0018] Optionally, the computer-implemented method further includes: comparing the parameters of the first defective points within the one or more selected occlusion regions with the parameters of the second defective points outside the one or more selected occlusion regions; and identifying a potential device causing the first defective points based on the comparison.

[0019] Optionally, the computer-implemented method further includes: obtaining multiple sets of substrate defective point coordinates, each set of the multiple sets of substrate defective point coordinates including the coordinates of the substrate defective points in each substrate, and the coordinates of the substrate defective points in each substrate being coordinates in a substrate coordinate system; and converting the coordinates of the substrate defective points in the substrate coordinate system into the coordinates of the defective points in the image coordinate system.

[0020] Optionally, obtaining the multiple sets of defect point coordinates includes: obtaining multiple sets of original defect point coordinates; and selecting, from the multiple sets of original defect point coordinates, multiple sets of defect point coordinates each including more than a threshold number of defect point coordinates as the multiple sets of defect point coordinates.

[0021] Optionally, a hierarchical clustering method is used to perform the clustering analysis.

[0022] Optionally, the hierarchical clustering method is a single-linkage clustering method.

[0023] Optionally, for each point cluster among the multiple point clusters, the Euclidean distance between adjacent defect points is less than or equal to a threshold, and the Euclidean distance between any two defect points from two different point clusters among the multiple point clusters is greater than the threshold.

[0024] In another aspect, the present disclosure provides a device for defect analysis, including: a memory; one or more processors; wherein, the memory and the one or more processors are connected to each other; and the memory stores computer-executable instructions for controlling the one or more processors to: obtain multiple sets of defect point coordinates, where each set of the multiple sets of defect point coordinates includes the coordinates of defect points in each of multiple substrates, and the coordinates of defect points in each substrate are coordinates in an image coordinate system; combine the multiple sets of defect point coordinates according to the image coordinate system into a composite set of coordinates to generate a composite image; and perform clustering analysis to classify the defect points in the composite set in the composite image into multiple point clusters.

[0025] Optionally, the memory further stores computer-executable instructions for controlling the one or more processors to: respectively determine multiple contours of at least multiple selected point clusters among the multiple point clusters, where each of the multiple contours includes multiple edge defect points in each of the multiple selected point clusters; apply a fitting algorithm to the edge defect points of the multiple selected point clusters to generate multiple occlusion regions respectively corresponding to the multiple selected point clusters; and respectively generate multiple feature vectors of the multiple occlusion regions.

[0026] Optionally, the memory further stores computer-executable instructions for controlling the one or more processors to: generate the Hu geometric moment m i,j and the center-to-center distance M i,j for each of the multiple occlusion regions, where, m i,j = ∑ (x,v)∈A x i y j ; calculate the defect point density ρ, area a, and centroid O(O x , Oy ) and the direction θ; and each of the plurality of feature vectors that generate each of the plurality of occluded regions.

[0027] In another aspect, the present disclosure provides a computer program product, including a non-transitory tangible computer-readable medium having computer-readable instructions thereon, the computer-readable instructions being executable by a processor to cause the processor to perform: obtaining multiple sets of defect point coordinates, each set of the multiple sets of defect point coordinates including the coordinates of defect points in each of a plurality of substrates, and the coordinates of the defect points in each substrate being coordinates in an image coordinate system; combining the multiple sets of defect point coordinates according to the image coordinate system into a composite set of coordinates to generate a composite image; and performing clustering analysis to classify the defect points in the composite set in the composite image into a plurality of point clusters.

[0028] Optionally, the computer-readable instructions may further be executable by the processor to cause the processor to perform: respectively determining multiple contours of at least a plurality of selected point clusters among the plurality of point clusters, each of the multiple contours including multiple edge defect points in each of the plurality of selected point clusters; applying a fitting algorithm to the edge defect points of the plurality of selected point clusters to generate a plurality of occluded regions respectively corresponding to the plurality of selected point clusters; and respectively generating a plurality of feature vectors of the plurality of occluded regions.

[0029] Optionally, the computer-readable instructions may further be executable by the processor to cause the processor to perform: generating the Hu geometric moment m i,j and the center-to-center distance M i,j , where, m i,j = ∑ (x,v)∈ A x i y j ; calculating the defect point density ρ, area a, centroid O(O x , O y ) and the direction θ of each of the plurality of occluded regions; and generating each of the plurality of feature vectors of each of the plurality of occluded regions.

[0030] In another aspect, the present disclosure provides a defect analysis system, comprising: a distributed computing system including one or more networked computers configured to execute in parallel to perform at least one common task; one or more computer-readable storage media storing instructions that, when executed by the distributed computing system, cause the distributed computing system to execute software modules; wherein the software modules include: a data manager configured to store data and extract, transform, or load the data; a query engine connected to the data manager and configured to query the data directly from the data manager; an analyzer connected to the query engine and configured to perform defect analysis upon receiving a task request, the analyzer including a plurality of business servers and a plurality of algorithm servers, the plurality of algorithm servers being configured to query the data directly from the data manager; and a data visualization and interaction interface configured to generate the task request; wherein one or more of the plurality of algorithm servers are configured to execute the computer-implemented method described herein. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] According to various disclosed embodiments, the following drawings are merely examples for illustrative purposes and are not intended to limit the scope of the present invention.

[0032] Figure 1 are multiple defect point images according to some embodiments of the present disclosure.

[0033] Figure 2A illustrates a computer-implemented method for defect analysis according to some embodiments of the present disclosure.

[0034] Figure 2B illustrates a computer-implemented method for defect analysis according to some embodiments of the present disclosure.

[0035] Figure 3A and Figure 3B illustrates the conversion process from a separate panel coordinate system to a unified substrate coordinate system.

[0036] Figure 3C illustrates the concatenation process of multiple defect point images according to the same image coordinate system.

[0037] Figure 4 illustrates a computer-implemented method for defect analysis according to some embodiments of the present disclosure.

[0038] Figure 5A is a schematic diagram of the structure of a device according to some embodiments of the present disclosure.

[0039] Figure 5B is a schematic diagram showing the structure of a device according to some embodiments of the present disclosure.

[0040] Figure 6 Shows a distributed computing environment in some embodiments according to the present disclosure.

[0041] Figure 7 Shows software modules in a defect analysis system in some embodiments according to the present disclosure.

[0042] Figure 8 Shows software modules in a defect analysis system in some embodiments according to the present disclosure.

[0043] Figure 9 Shows a defect analysis method using a defect analysis system in some embodiments according to the present disclosure.

[0044] Figure 10 Shows a defect analysis method using a defect analysis system in some embodiments according to the present disclosure.

[0045] Figure 11 Shows a defect analysis method using a defect analysis system in some embodiments according to the present disclosure.

[0046] Figure 12 Shows a defect analysis method using a defect analysis system in some embodiments according to the present disclosure.

[0047] Figure 13 Shows a data management platform in some embodiments according to the present disclosure.

[0048] Figure 14 Depicts a plurality of sub-tables segmented from a data table stored in a general data layer in some embodiments according to the present disclosure.

[0049] Figure 15 Shows a defect analysis method in some embodiments according to the present disclosure.

[0050] Figure 16 Shows a defect analysis method in some embodiments according to the present disclosure.

[0051] Figure 17 Is an exemplary defect point map image.

[0052] Figure 18A Shows an exemplary discrete point cluster.

[0053] Figure 18B Shows an example of a cluster region map fitted by a convex hull.

[0054] Figure 18C Shows an example of a cluster region map fitted by α Shapes.

[0055] Figure 19 An algorithm for determining a clustering region map is shown.

[0056] Figure 20A An exemplary defect point map is shown.

[0057] Figure 20B Exemplary results of classification and selection are shown.

[0058] Figure 21A An exemplary set of image regions of point clusters is shown.

[0059] Figure 21B A defect point aggregation region is shown. Detailed implementation manners

[0060] The present disclosure will now be described more specifically with reference to the following embodiments. It should be noted that the following description of some embodiments presented herein is for illustrative and descriptive purposes only. It is not exhaustive or limited to the exact forms disclosed.

[0061] The manufacture of a display panel (especially an organic light emitting diode display panel) involves highly complex and integrated processes, which involve a large number of processes, technologies, and devices. Defects occurring in such integrated processes are difficult to track. For example, engineers may have to rely on manual data classification to analyze the root cause of defects based on experience.

[0062] Accordingly, the present disclosure particularly provides a computer-implemented method for defect analysis, a device for defect analysis, a computer program product, and a defect analysis system, which substantially eliminate one or more problems caused by the limitations and disadvantages of the related art. In one aspect, the present disclosure provides a computer-implemented method for defect analysis. In some embodiments, the computer-implemented method for defect analysis includes obtaining multiple sets of defect point coordinates, each set of the multiple sets of defect point coordinates including the coordinates of defect points in each of a plurality of substrates, and the coordinates of the defect points in each substrate being coordinates in an image coordinate system; combining the multiple sets of defect point coordinates according to the image coordinate system into a composite set of coordinates to generate a composite image; and performing clustering analysis to classify the defect points in the composite set in the composite image into a plurality of point clusters.

[0063] Figure 1 are multiple defect point images in some embodiments according to the present disclosure. Referring to Figure 1 , the multiple defect point images include a large number of defect points. The distribution and quantity of the defect points in each of the multiple defect point images are different, making it difficult to find out the root cause or correlation of the defects in the display panel. Various defect point images in various appropriate formats can be used in combination with the present disclosure. In one example, the defect point image is a preprocessed image of defect points with specified coordinates, such asFigure 1 As shown. In another example, the multiple defect point images include defect point images from different batches. In another example, the multiple defect point images include defect point images from different substrates (but for the same type of product). In another example, the multiple defect point images include defect point images from different substrates of different types of products. In another example, the multiple defect point images include defect point images from the same batch over a period of time. In another example, the multiple defect point images include defect point images generated during the same time period.

[0064] Figure 2A illustrates a computer-implemented method for defect analysis according to some embodiments of the present disclosure. Referring to Figure 2A , in some embodiments, the computer-implemented method includes obtaining multiple defect point images, each defect point image in the multiple defect point images including a defect point in a substrate, the defect points being respectively assigned coordinates in an image coordinate system; cascading the multiple defect point images according to the image coordinate system, so as to align points from different defect point images having the same coordinates in the image coordinate system relative to each other in a cascaded manner; after aligning the points from different defect point images having the same coordinates in the image coordinate system in a cascaded manner, combining the multiple defect point images into a composite image including multiple defect points from the multiple defect point images; and performing a clustering analysis on the multiple defect points in the composite image to classify the multiple defect points into multiple point clusters. In this method, by cascading multiple defect point images, the distribution of defect points can be analyzed according to the specific steps outlined in the present disclosure. In one example, the multiple defect point images are defect point images of all substrates of the same type of display product (for example, each mother substrate is a substrate to be cut into multiple display panels). In another example, the multiple defect point images are defect point images of all substrates of the same type of organic light-emitting diode display panel. In one example, the defect points include defective sub-pixels. In another example, the defect points include open circuits. In another example, the defect points include short circuits. In another example, the defect points include dark lines. In another example, the defect points include bright lines.

[0065] Based on the coordinates in the image coordinate system, multiple defect point images are cascaded such that, for example, defect points having the same coordinates in the image coordinate system or coordinate points having the same coordinates in the image coordinate system but respectively from different defect point images are all aligned relative to each other in a cascaded manner. When the multiple defect point images are not pre-aligned according to the image coordinate system, the multiple defect point images must be cascaded according to the image coordinate system before combining the multiple defect point images into a composite image. In the composite image, defect points having the same coordinates in the image coordinate system or coordinate points having the same coordinates in the image coordinate system but respectively from different defect point images are located at the same coordinates in the image coordinate system. Since one of the purposes of this defect analysis method is to determine the correlation between the manufacturing equipment and the defect points, it is advantageous to cascade the multiple defect point images according to the image coordinate system.

[0066] Optionally, when cascading the multiple defect point images, the multiple defect point images are arranged in descending order, where the multiple defect point images are arranged in descending order according to the number of defect points therein. For example, the multiple defect point images are arranged such that the defect point image with more defect points is at the top and the defect point image with fewer defect points is at the bottom.

[0067] By combining the multiple defect point images into a composite image, the defect points from the multiple defect point images can be revealed and analyzed in a single composite image. The composite image adopts the image coordinate system such that the defect points from different defect point images can be located in a consistent coordinate system.

[0068] In some embodiments, defect points are detected on respective display panels. To analyze the correlation between the defects and the equipment, in some embodiments, this defect analysis method is performed at the substrate level. Therefore, the coordinates of the defect points in each display panel (cut from the substrate) are converted to the coordinates in the substrate. In some embodiments, the computer-implemented method further includes obtaining multiple panel defect point images, where each panel defect point image in the multiple panel defect point images includes the panel defect points in each panel, and the panel defect points are respectively assigned coordinates in the panel coordinate system; and converting the coordinates of the panel defect points in the panel coordinate system to the coordinates of the defect points in the substrate coordinate system. Optionally, each panel is a panel cut from the substrate.

[0069] The computer-implemented method can have various suitable embodiments. Figure 2A An embodiment using multiple defect point images is described. However, various other embodiments can be practiced. For example, the method can be practiced directly using the defect point coordinates. Figure 2B A computer-implemented method for defect analysis according to some embodiments of the present disclosure is shown. Refer to Figure 2B, in some embodiments, the computer-implemented method includes obtaining multiple sets of coordinates of defect points, where each set of the multiple sets of coordinates of defect points includes the coordinates of defect points in each of multiple substrates, and the coordinates of defect points in each substrate are coordinates in an image coordinate system; combining the multiple sets of coordinates of defect points according to the image coordinate system into a composite set of coordinates to generate a composite image; and performing clustering analysis to classify the defect points in the composite set in the composite image into multiple point clusters.

[0070] Figure 3A and Figure 3B illustrates the conversion process from individual panel coordinate systems to a unified substrate coordinate system. Refer to Figure 3A , the panel defect points in multiple panel defect point images (A to N) have panel coordinates assigned according to respective panel coordinate systems (pcs1 to pcs14). For example, panel coordinates are assigned to the panel defect points in panel defect image A according to the first panel coordinate system pcs1, panel coordinates are assigned to the panel defect points in panel defect image B according to the second panel coordinate system pcs2, panel coordinates are assigned to the panel defect points in panel defect image C according to the third panel coordinate system pcs3, and so on. Since panels are cut from substrates during the manufacturing process, the coordinates in each panel coordinate system can be converted to a unified substrate coordinate system. In one example, the defect analysis is for the same type of product (e.g., an organic light-emitting diode display panel), so the multiple substrates used to manufacture the panels have the same substrate coordinate system. Refer to Figure 3B , the coordinates of the panel defect points in each panel coordinate system are converted to the coordinates of defect points in the substrate coordinate system scs. Figure 3C illustrates the cascading process of multiple defect point images according to the same substrate coordinate system. As Figure 3C shown, the panel defect point images are converted into substrate defect point images using the unified substrate coordinate system. Before combining the substrate defect point images into a composite image, the substrate defect point images are cascaded in the unified substrate coordinate system. In some embodiments of the present disclosure, the substrate coordinate system and the image coordinate system may be the same or different. That is, the conversion between the substrate coordinate system and the image coordinate system can be performed, or the conversion between the substrate coordinate system and the image coordinate system may not be performed.

[0071] Therefore, in some embodiments, the method further includes obtaining multiple sets of coordinates of substrate defect points, where each set of the multiple sets of coordinates of substrate defect points includes the coordinates of substrate defect points in each substrate, and the coordinates of substrate defect points in each substrate are coordinates in the substrate coordinate system; converting the coordinates of the substrate defect points in the substrate coordinate system to the coordinates of defect points in the image coordinate system.

[0072] In one example, the computer-implemented method further includes establishing a mapping relationship between coordinates in a substrate coordinate system and coordinates in an image coordinate system. In another example, the mapping relationship can be expressed as:

[0073]

[0074] where which is a rotation matrix; and which is a translation matrix.

[0075] Optionally, the coordinate transformation can be expressed as:

[0076] where X i and Y i represent coordinates in the image coordinate system; x i and y i represent coordinates in the substrate coordinate system.

[0077] In one example, the origin of the substrate coordinate system is located at the center of the substrate, while the origin of the image coordinate system is located at the lower left corner of the image. In another example, the substrate coordinate system has a larger range than the image coordinate system (e.g., the x-axis or the y-axis or both). During the process of converting coordinates in the substrate coordinate system to the image coordinate system, the substrate coordinate system can be scaled according to the image coordinate system, thereby reducing the amount of calculation.

[0078] In some embodiments, outliers are excluded from multiple defect point images. In one example, a "defect" image with a number of defect points less than or equal to a threshold can be considered a normal product without defects, and thus such a "defect" image can be excluded. In some embodiments, the method includes obtaining multiple sets of original defect point coordinates; and selecting, from the multiple sets of original defect point coordinates, multiple sets of defect point coordinates including more than the threshold number of defect point coordinates as the multiple sets of defect point coordinates. In one example, the method includes obtaining multiple original defect point images; and selecting, from the multiple original defect point images, defect point images including more than the threshold number of defect points as the multiple defect point images. In one example, the threshold number is a positive integer, such as 2 (or 3, or 4, or 5, or 10). Optionally, the multiple original defect point images are substrate images, and the multiple defect point images are also substrate images.

[0079] Various suitable conditions can be used to perform clustering analysis. Examples of suitable conditions include Euclidean distance, chi-square, and correlation.

[0080] In some embodiments, the Euclidean distance between defect points can be used as a condition for performing clustering analysis. In some embodiments, the Euclidean distance between adjacent defect points in each of the plurality of point clusters is less than or equal to a threshold, and the Euclidean distance between any two defect points respectively from two of the plurality of point clusters is greater than the threshold. The method generates a plurality of point clusters C = {C1, C2, C3, ..., C n}, 1 ≤ i ≤ n, n ≥ 1, R represents the set of all defect points in the composite image, and Ci represents a set of defect points in each of the plurality of point clusters, where the number of defect points is greater than or equal to 1. The Euclidean distance can be expressed as: x i and x i-1 represent the x - coordinates of two defect points; y i and y i-1 represent the y - coordinates of two defect points.

[0081] In one example, the clustering process starts from a candidate defect point. The method determines whether the defect points adjacent to the candidate defect point (e.g., adjacent defect sub - pixels) have a Euclidean distance less than or equal to the threshold. When it is determined that the Euclidean distance is less than or equal to the threshold, the method adds the adjacent defect points to the candidate point cluster containing the candidate defect point. For example, using any defect point already included in the candidate point cluster as the starting defect point, determine whether the Euclidean distance from its adjacent defect points to the starting defect point is less than or equal to the threshold, and when it is determined that the Euclidean distance is less than or equal to the threshold, add its adjacent defect points to the candidate point cluster, and repeat the process. When the Euclidean distance between any defect point in the candidate point cluster and any defect point outside the candidate point cluster is greater than the threshold, stop the repeating process.

[0082] Various suitable methods can be used for clustering analysis. Examples of suitable clustering methods include hierarchical clustering, identification of connected components, connectivity - based clustering, distribution - based clustering, density - based clustering, single - link clustering, Markov clustering (MCL), and centroid clustering. Optionally, the clustering method is hierarchical clustering. Examples of hierarchical clustering methods include complete - link clustering; average - link clustering and single - link clustering. Optionally, the hierarchical clustering method is single - link clustering.

[0083] In some embodiments, the computer - implemented method further includes obtaining a plurality of selected point clusters from the plurality of point clusters. Optionally, the number of defect points in each of the plurality of selected point clusters is greater than a threshold number. This step generates a plurality of selected point clusters C’ = {C’1, C’2, C’3, ..., C’ n’}, If C' is an empty set, it indicates that the defect points are not clustered. If C' is not an empty set, the method proceeds to the subsequent steps. Optionally, the threshold is greater than or equal to 1, such as 2, 5, 10, 20, 30, 40, 50, 60, 70, 80, 90, or 100.

[0084] Figure 4 FIG. shows a computer-implemented method for defect analysis according to some embodiments of the present disclosure. Referring to Figure 4 , in some embodiments, the computer-implemented method further includes determining, respectively, a plurality of contours of at least a plurality of selected point clusters among the plurality of point clusters, each of the plurality of contours including a plurality of edge defect points in each of the plurality of point clusters of the plurality of selected point clusters; applying a fitting algorithm to the edge defect points of the plurality of selected point clusters to generate a plurality of mask areas corresponding to the plurality of selected point clusters, respectively; and generating a plurality of feature vectors of the plurality of mask areas, respectively.

[0085] Various suitable methods can be used to generate, respectively, multiple contours of at least a plurality of selected point clusters out of a plurality of point clusters. Examples of contour generation methods include alpha Shape-based methods (see, e.g., “Alpha shapes: definition and software” by N. Akkiraju, H. Edelsbrunner, M. Facello, P. Fu, E. P. Mucke, and C. Varela, Proc. Internat. Comput. Geom. Software Workshop 1995; “Three-dimensional alpha shapes” by H. Edelsbrunner and E. P. Mucke, ACM Trans. Graphics 13 (1994), 43 - 72, and “Computing Dirichlet Tesselations” by A. Bowyer, The Computer Journal, 24(2), pp 162 - 166, February 1981; the contents of which are incorporated herein by reference in their entirety). Other suitable examples of contour generation methods include weighted minimum path search (see, e.g., “A note on two problems in connexion with graphs” by Edsger W. Dijkstra, Numerical Mathematics, 1, 1959, p. 269 - 271; the contents of which are incorporated herein by reference in their entirety); and proximity search using KD - Tree FLANN (see, e.g., “Scalable Nearest Neighbor Algorithms for High Dimensional Data” by Marius Muja and David G. Lowe, Pattern Analysis and Machine Intelligence (PAMI), volume 36, 2014; the contents of which are incorporated herein by reference in their entirety).

[0086] Various suitable methods can be used to apply a fitting algorithm to the edge defect points of multiple selected clusters of points to generate multiple occluded regions corresponding to the multiple selected clusters of points, respectively. Examples of fitting algorithms include region fitting algorithms (see, for example, "Pattern Classification" by Richard O. Duda, Peter E. Hart, and David G. Stork, John Wiley and Sons, Inc., New York, 2001, and "Understanding Synthetic Aperture Radar Images" by C. Oliver, S. Quegan; The Duda Reference, pages 548 and 549; the contents of which are incorporated herein by reference in their entirety).

[0087] In one example, for each of the multiple selected clusters of points C’ i ={p1, p2, p3, ..., p n}, the αShapes algorithm is used to extract its intuitive outer shape from a discrete and unordered set of points and obtain a set of points C c i ={p c 1, p c 2, p c 3, ..., p c n}, The set of outer shape contour points corresponds to the set of multiple selected clusters of points, respectively. In another example, a fitting algorithm (e.g., a region fitting algorithm) is applied to the edge defect points of multiple selected clusters of points to generate multiple occluded regions corresponding to the multiple selected clusters of points, respectively. In another example, the multiple occluded regions are a set of minimally enclosed occluded regions A = {A1, A2, A3, ..., A n}.

[0088] In some embodiments, the computer-implemented method further includes generating multiple feature vectors for the minimally enclosed occluded regions A = {A1, A2, A3, ..., A n}, respectively. For Ai, the method includes generating the Hu geometric moments m i,j and the center-to-center distance M i,j of each occluded region in the multiple occluded regions, where m i,j = ∑ (x,v)∈A x i y j ; calculating the defect point density ρ, area a, and centroid O(O x , O y) and direction θ; and generating each feature vector among the multiple feature vectors of each of the multiple occlusion regions.

[0089] In some embodiments, each of the multiple feature vectors (e.g., for a corresponding one A in the smallest enclosed occlusion region A = {A1, A2, A3,..., A n}) is represented as F = [ρ, a, Ox, Oy, θ, L, W, r] i , where T N represents the number of defective points in each of the multiple occlusion regions, and a represents the area of each of the multiple occlusion regions; M M ij = ∑ (x,y)∈A (x - O x ) i (y - O y ) j ; L represents the length of the minimum bounding rectangle of each of the multiple occlusion regions; W represents the width of the minimum bounding rectangle of each of the multiple occlusion regions.

[0090] In some embodiments, one or more feature vectors among the multiple feature vectors that do not meet the threshold condition are removed from the subsequent steps. Thus, in some embodiments, the computer-implemented method further includes assigning one or more selected occlusion regions among the multiple occlusion regions as multiple defect aggregation regions, wherein the feature vectors of the one or more selected occlusion regions meet the threshold condition. In one example, the threshold condition can be represented as [α i , β i ,...]. In another example, α i can be the condition that ρ > 0.5. In another example, β i can be the condition that a > 200.

[0091] In one example, the multiple defect aggregation regions can be represented as a set A' = {A'1, A'2, A'3,..., A' n}. When A′ is an empty set, the defect aggregation region may not exist in the multiple defective point images.

[0092] In some embodiments, when A′ is not an empty set, the computer-implemented method further includes comparing parameters of first defect points within one or more selected occluded regions with parameters of second defect points outside one or more selected occluded regions; and identifying a potential device that causes the first defect points based on the comparison. By comparing negative samples (parameters of the first defect points within one or more selected occluded regions) with positive samples (parameters of the second defect points outside one or more selected occluded regions), the root cause of the defect can be traced back to one or more potentially problematic devices responsible for one or more manufacturing processes of the display panel.

[0093] In another aspect, the present disclosure provides a device for defect analysis. In some embodiments, the device for defect analysis includes a memory; and one or more processors. The memory and the one or more processors are connected to each other. In some embodiments, the memory stores computer-executable instructions for controlling the one or more processors to obtain multiple sets of defect point coordinates, each set of the multiple sets of defect point coordinates including coordinates of defect points in respective substrates among multiple substrates, and the coordinates of the defect points in each substrate being coordinates in an image coordinate system; combining the multiple sets of defect point coordinates according to the image coordinate system into a composite set of coordinates to generate a composite image; and performing clustering analysis to classify the defect points in the composite set in the composite image into multiple point clusters.

[0094] In some embodiments, the memory stores computer-executable instructions for controlling the one or more processors to obtain multiple defect point images, each defect point image of the multiple defect point images including a defect point in a substrate, and the defect points being respectively assigned coordinates in an image coordinate system; cascading the multiple defect point images according to the image coordinate system, so as to align points from different defect point images having the same coordinates in the image coordinate system relative to each other in a cascaded manner; after aligning the points from different defect point images having the same coordinates in the image coordinate system in a cascaded manner, combining the multiple defect point images into a composite image including multiple defect points from the multiple defect point images; and performing clustering analysis on the multiple defect points in the composite image to classify the multiple defect points into multiple point clusters.

[0095] In some embodiments, the memory stores computer-executable instructions for controlling the one or more processors to obtain multiple selected point clusters from the multiple point clusters. Optionally, the number of defect points in each of the multiple selected point clusters is greater than a threshold number.

[0096] In some embodiments, the memory stores computer-executable instructions for controlling one or more processors to respectively determine a plurality of contours of at least a plurality of selected point clusters among a plurality of point clusters, each contour among the plurality of contours including a plurality of edge defect points in each of the plurality of selected point clusters; apply a fitting algorithm to the edge defect points of the plurality of selected point clusters to generate a plurality of occlusion regions respectively corresponding to the plurality of selected point clusters; and respectively generate a plurality of feature vectors of the plurality of occlusion regions.

[0097] In some embodiments, the memory stores computer-executable instructions for controlling one or more processors to generate the Hu geometric moment m i,j and the center-to-center distance M i,j of each of the plurality of occlusion regions, where m i,j = ∑ (x,v)∈A x i y j ; calculate the defect point density ρ, area a, centroid O(O x , O y ) and orientation θ of each of the plurality of occlusion regions; and generate each of the plurality of feature vectors of each of the plurality of occlusion regions. Optionally, a corresponding one of the plurality of feature vectors is represented as F = [ρ, a, Ox, Oy, θ, L, W, r] T , where N represents the number of defect points in each of the plurality of occlusion regions, and a represents the area of each of the plurality of occlusion regions; M ij = ∑ (x,y)∈A (x - O x ) i (y - O y ) j ; L represents the length of the minimum bounding rectangle of each of the plurality of occlusion regions; W represents the width of the minimum bounding rectangle of each of the plurality of occlusion regions.

[0098] Optionally, a method based on α Shapes is used to determine the plurality of contours.

[0099] In some embodiments, the memory stores computer-executable instructions for controlling one or more processors to assign one or more selected occlusion regions among the plurality of occlusion regions as a plurality of defect aggregation regions. Optionally, the feature vectors of the one or more selected occlusion regions satisfy a threshold condition.

[0100] In some embodiments, the memory stores computer-executable instructions for controlling one or more processors to compare parameters of first defect points within one or more selected occluded regions with parameters of second defect points outside the one or more selected occluded regions; and to identify a potential device causing the first defect points based on the comparison.

[0101] In some embodiments, the memory stores computer-executable instructions for controlling one or more processors to obtain a plurality of substrate defect point images, each of the plurality of substrate defect point images including a substrate defect point in each substrate, and the substrate defect points being respectively assigned coordinates in a substrate coordinate system; and to convert the coordinates of the substrate defect points in the substrate coordinate system into coordinates of defect points in an image coordinate system.

[0102] In some embodiments, to obtain a plurality of defect point images, the memory stores computer-executable instructions for controlling one or more processors to obtain multiple sets of raw defect point coordinates; and to select, as the multiple sets of defect point coordinates, multiple sets of defect point coordinates that include more than a threshold number of defect point coordinates from the multiple sets of raw defect point coordinates. In one example, to obtain a plurality of defect point images, the memory stores computer-executable instructions for controlling one or more processors to obtain a plurality of raw defect point images; and to select, as the plurality of defect point images, defect point images that include more than a threshold number of defect points from the plurality of raw defect point images.

[0103] Optionally, clustering analysis is performed using a hierarchical clustering method.

[0104] Optionally, the hierarchical clustering method is a single-linkage clustering method.

[0105] Optionally, the Euclidean distance between adjacent defect points in each point cluster among the multiple point clusters is less than or equal to a threshold, and the Euclidean distance between any two defect points from two of the multiple point clusters is greater than the threshold.

[0106] Figure 5A is a schematic diagram of the structure of a device according to some embodiments of the present disclosure. Referring to Figure 5A , in some embodiments, the device includes a central processing unit (CPU) configured to perform actions according to computer-executable instructions stored in a ROM or a RAM. Optionally, data and programs required by the computer system are stored in the RAM. Optionally, the CPU, ROM, and RAM are electrically connected to each other through a bus. Optionally, an input / output interface is electrically connected to the bus.

[0107] Figure 5B is a schematic diagram showing the structure of a device according to some embodiments of the present disclosure. Referring to Figure 5B, in some embodiments, the device includes a display panel DP; an integrated circuit IC connected to the display panel DP; a memory M; and one or more processors P. The memory M and the one or more processors P are connected to each other. In some embodiments, the memory M stores computer-executable instructions for controlling the one or more processors P to execute the method steps described herein.

[0108] In another aspect, the present disclosure provides a computer program product including a non-transitory tangible computer-readable medium having computer-readable instructions thereon. In some embodiments, the computer-readable instructions are executable by a processor to cause the processor to obtain multiple sets of coordinates of defect points, each set of the multiple sets of coordinates of defect points including the coordinates of defect points in each of a plurality of substrates, and the coordinates of the defect points in each substrate being coordinates in an image coordinate system; combine the multiple sets of coordinates of defect points according to the image coordinate system into a composite set of coordinates to generate a composite image; and perform clustering analysis to classify the defect points in the composite set in the composite image into a plurality of point clusters.

[0109] In some embodiments, the computer-readable instructions are executable by a processor to cause the processor to obtain multiple defect point images, each of the multiple defect point images including a defect point in a substrate, and the defect points being respectively assigned coordinates in an image coordinate system; cascade the multiple defect point images according to the image coordinate system, so as to align the points respectively from different defect point images having the same coordinates in the image coordinate system relative to each other in a cascaded manner; after aligning the points respectively from different defect point images having the same coordinates in the image coordinate system in a cascaded manner, combine the multiple defect point images into a composite image including multiple defect points from the multiple defect point images; and perform clustering analysis on the multiple defect points in the composite image to classify the multiple defect points into a plurality of point clusters.

[0110] In some embodiments, the computer-readable instructions are further executable by a processor to cause the processor to obtain multiple selected point clusters from the plurality of point clusters. Optionally, the number of defect points in each of the multiple selected point clusters is greater than a threshold number.

[0111] In some embodiments, the computer-readable instructions are further executable by a processor to cause the processor to respectively determine multiple contours of at least multiple selected point clusters among the plurality of point clusters, each of the multiple contours including multiple edge defect points in each of the multiple selected point clusters; apply a fitting algorithm to the edge defect points of the multiple selected point clusters to generate multiple occlusion regions respectively corresponding to the multiple selected point clusters; and respectively generate multiple feature vectors of the multiple occlusion regions.

[0112] In some embodiments, the computer-readable instructions may also be executed by a processor to cause the processor to execute generating the Hu geometric moment m of each of the plurality of occlusion regions i,j and the center-to-center distance M i,j , wherein, m i,j =∑ (x,y)∈A x i y j ; calculating the defect point density ρ, area a, centroid O(O x , O y ) and orientation θ of each of the plurality of occlusion regions; and generating each of the plurality of feature vectors of each of the plurality of occlusion regions. Optionally, a corresponding one of the plurality of feature vectors is represented as F = [ρ, a, Ox, Oy, θ, L, W, r] T , wherein N represents the number of defect points in each of the plurality of occlusion regions, and a represents the area of each of the plurality of occlusion regions; M ij =∑ (x,y)∈A (x - O x ) i (y - O y ) j ; L represents the length of the minimum bounding rectangle of each of the plurality of occlusion regions; W represents the width of the minimum bounding rectangle of each of the plurality of occlusion regions.

[0113] Optionally, a method based on α Shapes is used to determine the plurality of contours.

[0114] In some embodiments, the computer-readable instructions may also be executed by a processor to cause the processor to execute assigning one or more selected occlusion regions of the plurality of occlusion regions as a plurality of defect aggregation regions. Optionally, the feature vectors of the one or more selected occlusion regions satisfy a threshold condition.

[0115] In some embodiments, the computer-readable instructions may also be executed by a processor to cause the processor to execute comparing the parameters of the first defect points within one or more selected occlusion regions with the parameters of the second defect points outside the one or more selected occlusion regions; and identifying a potential device causing the first defect points based on the comparison.

[0116] In some embodiments, the computer-readable instructions may also be executed by a processor to cause the processor to execute obtaining a plurality of substrate defect point images, each of the plurality of substrate defect point images including substrate defect points in each substrate, and the substrate defect points are respectively assigned coordinates in a substrate coordinate system; and converting the coordinates of the substrate defect points in the substrate coordinate system into the coordinates of the defect points in an image coordinate system.

[0117] In some embodiments, to obtain the plurality of defect point maps, the computer-readable instructions may further be executed by a processor to cause the processor to obtain multiple sets of original defect point coordinates; and select, from the multiple sets of original defect point coordinates, multiple sets of defect point coordinates that include more than a threshold number of defect point coordinates as the multiple sets of defect point coordinates. In one example, to obtain multiple defect point images, the computer-readable instructions may further be executed by a processor to cause the processor to obtain multiple original defect point images; and select, from the multiple original defect point images, defect point images that include more than a threshold number of defect points as the multiple defect point images.

[0118] Optionally, clustering analysis is performed using a hierarchical clustering method.

[0119] Optionally, the hierarchical clustering method is a single-linkage clustering method.

[0120] Optionally, the Euclidean distance between adjacent defect points in each point cluster among the multiple point clusters is less than or equal to a threshold, and the Euclidean distance between any two defect points from two of the multiple point clusters is greater than the threshold.

[0121] Various defects may occur in the manufacture of semiconductor electronic devices. Examples of defects include particles, residues, line defects, holes, splashes, wrinkles, discoloration, and bubbles. Defects that occur in the manufacture of semiconductor electronic devices are difficult to track. For example, engineers may have to rely on manual data classification to analyze the root cause of defects based on experience.

[0122] When manufacturing a liquid crystal display panel, the manufacture of the display panel at least includes an array stage, a color filter (CF) stage, a cell stage, and a module stage. In the array stage, a thin film transistor array substrate is manufactured. In one example, in the array stage, a material layer is deposited, and the material layer is subjected to photolithography. For example, a photoresist is deposited on the material layer, and the photoresist is subjected to exposure and then developed. Subsequently, the material layer is etched and the remaining photoresist is removed ("stripping"). In the CF stage, a color filter substrate is manufactured, which involves several steps, including coating, exposure, and development. In the cell stage, the array substrate and the color filter substrate are assembled to form a cell. The cell stage includes several steps, including coating and rubbing an alignment layer, injecting a liquid crystal material, coating a cell sealant, boxing under vacuum, cutting, grinding, and cell inspection. In the module stage, peripheral components and circuits are assembled onto the panel. In one example, the module level includes several steps, including the assembly of a backlight, the assembly of a printed circuit board, the attachment of a polarizer, the assembly of a chip on film, the assembly of an integrated circuit, aging, and final inspection.

[0123] When manufacturing an organic light-emitting diode (OLED) display panel, the manufacture of the display panel includes at least four device processes, including an array stage, an OLED stage, an EAC2 stage, and a module stage. In the array stage, a backplane of the display panel is manufactured, for example, including manufacturing a plurality of thin-film transistors. In the OLED stage, a plurality of light-emitting elements (e.g., organic light-emitting diodes) are manufactured, an encapsulation layer is formed to encapsulate the plurality of light-emitting elements, and optionally, a protective film is formed on the encapsulation layer. In the EAC2 stage, a large glass is first cut into half-glasses (hglass), and then further cut into panels (panel). In addition, in the EAC2 stage, inspection equipment is used to inspect the panel to detect defects therein, such as dark spots and bright lines. In the module stage, for example, a flexible printed circuit is bonded to the panel using chip-on-film technology. A cover glass is formed on the surface of the panel. Optionally, further inspection is performed to detect defects in the panel. Data from the manufacture of the display panel includes biographical information, parameter information, and defect information, and this information is stored in a plurality of data sources. The biographical information is record information uploaded to a database from the array stage to the module stage through each processing device, including glass ID, device model, station information, etc. The parameter information includes data generated by the device when processing the glass. Defects may occur in each stage. Inspection information can be generated in each of the stages discussed above. Only after the inspection is completed can the inspection information be uploaded to the database in real time. The inspection information can include defect type and defect location.

[0124] In summary, various sensors and inspection devices are used to obtain biographical information, parameter information, and defect information. A defect analysis method or system is used to analyze the biographical information, parameter information, and defect information, and the defect analysis method or system can quickly determine the device, station, and / or stage where the defect occurs, thereby providing key information for subsequent process improvement and equipment repair or maintenance, thus greatly improving the yield.

[0125] Accordingly, the present disclosure particularly provides a data management platform, a defect analysis system, a defect analysis method, a computer storage medium, and a method for defect analysis thereof, which substantially eliminate one or more problems caused by the limitations and disadvantages of the prior art. The present disclosure provides an improved data management platform with superior functions. Based on this data management platform (or other suitable databases or data management platforms), the inventors of the present disclosure have further developed a novel and unique defect analysis system, defect analysis method, computer storage medium, and method for defect analysis.

[0126] In one aspect, the present disclosure provides a defect analysis system. In some embodiments, the defect analysis system includes a distributed computing system that includes one or more networked computers configured to execute in parallel to perform at least one common task; and one or more computer-readable storage media that store instructions that, when executed by the distributed computing system, cause the distributed computing system to execute software modules. In some embodiments, the software modules include: a data management platform configured to store data and extract, transform, or load data, where the data includes at least one of resume data information, parameter information, or defect information; an analyzer configured to perform defect analysis when receiving a task request, the analyzer including a plurality of business servers and a plurality of algorithm servers, the plurality of algorithm servers being configured to directly obtain data from the data management platform and perform algorithm analysis on the data to obtain result data regarding potential causes of defects; and a data visualization and interaction interface configured to generate the task request. Optionally, the defect analysis system is used for defect analysis in panel manufacturing. As used herein, the term "distributed computing system" generally refers to an interconnected computer network having multiple network nodes that connect multiple servers or hosts to each other or to an external network (e.g., the Internet). The term "network node" generally refers to a physical network device. Example network nodes include routers, switches, hubs, bridges, load balancers, security gateways, or firewalls. A "host" generally refers to a physical computing device configured to implement, for example, one or more virtual machines or other suitable virtualized components. For example, a host can include a server having a hypervisor configured to support one or more virtual machines or other suitable types of virtual components.

[0127] Figure 6 illustrates a distributed computing environment in some embodiments according to the present disclosure. Referring to Figure 6 , in a distributed computing environment, a plurality of autonomous computers / workstations, called nodes, communicate with each other in a network such as a LAN (local area network) to solve tasks, such as executing applications. Each computer node typically includes its own processor(s), memory, and communication links to other nodes. The computers can be located within a specific location (e.g., a cluster network of points) or can be connected via a wide area network (LAN) such as the Internet. In such a distributed computing environment, different applications can share information and resources.

[0128] The network in a distributed computing environment can include a local area network (LAN) and a wide area network (WAN). The network can include wired technologies (e.g., Ethernet) and wireless technologies (e.g., Code Division Multiple Access (CDMA), Global System for Mobile Communications (GSM), Universal Mobile Telecommunications Service (UMTS), Bluetooth, etc.).

[0129] Multiple computing nodes are configured to join a resource group to provide a distributed service. Computing nodes in a distributed network can include any computing device, such as a computing device or a user device. Computing nodes can also include a data center. As used herein, a computing node can refer to any computing device or multiple computing devices (i.e., a data center). Software modules can be executed on a single computing node (e.g., a server) or distributed across multiple nodes in any suitable manner.

[0130] The distributed computing environment can also include one or more storage nodes for storing information related to the execution of software modules and / or outputs generated by the execution of software modules and / or other functions. The one or more storage nodes communicate with each other in the network and with one or more computing nodes in the network.

[0131] Figure 7 A software module in a defect analysis system according to some embodiments of the present disclosure is shown. Referring to Figure 7 , the defect analysis system includes a distributed computing system that includes one or more networked computers configured to execute in parallel to perform at least one common task; one or more computer-readable storage media storing instructions that, when executed by the distributed computing system, cause the distributed computing system to execute the software module. In some embodiments, the software module includes a data management platform DM configured to store data and extract, transform, or load data; a query engine QE connected to the data management platform DM and configured to obtain data directly from the data management platform DM; an analyzer AZ connected to the query engine QE and configured to perform defect analysis when receiving a task request, the analyzer AZ including multiple business servers BS (similar to backend servers) and multiple algorithm servers AS, the multiple algorithm servers AS being configured to obtain data directly from the data management platform DM; and a data visualization and interaction interface DI configured to generate a task request. Optionally, the query engine QE is a query engine based on Impala TM technology. As used herein, in the context of the present disclosure, the term "connected to" refers to having a relationship of direct information or data flow from a first component of a system to a second component and / or from a second component of the system to the first component.

[0132] Figure 8 A software module in a defect analysis system according to some embodiments of the present disclosure is shown. Referring to Figure 8, in some embodiments, the data management platform DM includes an ETL module ETLP configured to extract, transform, or load data from multiple data sources DS onto a data mart DMT and a general data layer GDL. When receiving an assigned task, each of the multiple algorithm servers AS is configured to directly obtain first data from the data mart DMT. When performing defect analysis, each of the multiple algorithm servers AS is configured to directly send second data to the general data layer GDL. It can be understood that the first data is the input data of the algorithm server, and the second data is the result data obtained by the algorithm server through calculation. The multiple algorithm servers AS are deployed with various general algorithms for defect analysis, such as algorithms based on big data analysis, which can be algorithms based on specific machine learning models, such as one or more of decision tree, random forest, GBDT, XGBoost, naive Bayes, support vector machine, Adaboost, neural network model, etc., or other statistical algorithm models, such as WOE&IV, Apriori, etc. The multiple algorithm servers AS are configured to analyze data to identify the causes of defects. As used herein, the term "ETL module" refers to computer program logic configured to provide functions such as extracting, transforming, or loading data. In some embodiments, the ETL module is stored on a storage node, loaded into a memory, and executed by a processor. In some embodiments, the ETL module is stored on one or more storage nodes in a distributed network, loaded into one or more memories in the distributed network, and executed by one or more processors in the distributed network.

[0133] The data management platform DM stores data for the defect analysis system. For example, the data management platform DM stores the data required for algorithm analysis by the multiple algorithm servers AS. In another example, the data management platform DM stores the results of algorithm analysis. In some embodiments, the data management platform DM includes multiple data sources DS (e.g., data stored in an Oracle database), an ETL module ETLP, a data mart DMT (e.g., a data mart based on Apache Hbase TM technology) and a general data layer GDL (e.g., based on Apache Hive TMTechnical data storage). For algorithm analysis and interactive display for users, data from multiple data sources DS is cleaned and merged into verification data by the ETL module ETLP. Examples of useful data for defect analysis include trace history data, data variable (dv) parameter data, mapped defect location data, etc. The amount of data in a typical manufacturing process (e.g., the manufacturing process of a display panel) is huge. For example, there may be more than 30 million dv parameter data per day at a typical site. To meet the user's demand for defect analysis, it is necessary to improve the speed at which the algorithm server reads production data. In one example, the data required for algorithm analysis is stored in a data mart based on Apache Hbase TM to improve efficiency and save storage space. In another example, the results of algorithm analysis and other auxiliary data are stored in a general data layer based on Apache Hive TM technology.

[0134] Apache Hive TM is an open-source data warehouse system built on top of Hadoop, which is used to query and analyze big data stored in structured and semi-structured forms in Hadoop files. Apache Hive TM is mainly used for batch processing and is therefore called OLAP. Apache Hive TM is not a database but has a schema model.

[0135] Apache Hbase TM is a non-relational column-oriented distributed database running on top of the Hadoop Distributed File System (HDFS). In addition, it is a NoSQL open-source database that stores data by column. Apache Hbase TM is mainly used for transaction processing and is called OLTP. However, in the case of Apache Hbase TM real-time processing is possible. ApacheHbase TM is a NoSQL database and does not have a schema model.

[0136] In one example, various components of the data management platform (e.g., general data layer, data warehouse, data source) can be in the form of a distributed data storage based on, for example, Apache Hadoop TM and / or Apache Hive TM .

[0137] Figure 13 shows a data management platform in some embodiments according to the present disclosure. Refer to Figure 13, in some embodiments, the data management platform includes a distributed storage system (DFS), such as the Hadoop Distributed File System (HDFS). The data management platform is configured to collect data generated during the factory production process from multiple data sources DS. For example, the data generated during the factory production process is stored in a relational database (such as Oracle) using RDBMS (Relational Database Management System) grid computing technology. In RDBMS grid computing, the problem that requires a very large amount of computer power is divided into many small parts, and these small parts are assigned to many computers for processing. The results of distributed computing are combined to obtain the final result. For example, in Oracle RAC (Real Application Clusters), all servers can directly access all data in the database. However, applications based on RDBMS grid computing have limited hardware scalability. When the data volume reaches a certain order of magnitude, the input / output bottleneck of the hard disk makes it very inefficient to process a large amount of data. The parallel processing of the distributed file system can meet the challenges posed by the increasing demands for data storage and computing. During the defect analysis process, first, the data in multiple data sources DS is extracted into the data management platform, greatly accelerating the processing process.

[0138] In some embodiments, the data management platform includes multiple sets of data with different contents and / or storage structures. In some embodiments, the ETL module ETLP is configured to extract raw data from multiple data sources DS into the data management platform, forming a first data layer (e.g., data lake DL). The data lake DL is a centralized HDFS or KUDU database configured to store any structured or unstructured data. Optionally, the data lake DL is configured to store the first set of data extracted by the ETL module ETLP from multiple data sources DS. Optionally, the first set of data and the raw data have the same content. The dimensions and attributes of the raw data are preserved in the first set of data. In some embodiments, the first set of data stored in the data lake includes dynamically updated data. Optionally, the dynamically updated data includes database real-time update data based on Kudu, or periodic update data in the Hadoop Distributed File System. In one example, the periodic update data stored in the Hadoop Distributed File System is the periodic update data stored in the memory based on Apache Hive TM of the memory. In one example, the dynamically updated data includes real-time update data and periodic update data. In one example, real-time update means updates below the minute level, excluding updates at the minute level; periodic update means updates at the minute level and above, including updates at the minute level.

[0139] In some embodiments, the data management platform includes a second data layer, such as a data warehouse DW. The data warehouse DW includes an internal storage system configured to provide data in an abstract manner, such as in a table format or a view format, without exposing the file system. The data warehouse DW may be based on Apache Hive TM . The ETL module ETLP is configured to extract, clean, transform, or load a first set of data to form a second set of data. Optionally, the second set of data is formed by cleaning and normalizing the first set of data.

[0140] In some embodiments, the data management platform includes a third data layer (e.g., a general data layer GDL). The general data layer GDL may be based on Apache Hive TM . The ETL module ETLP is configured to perform data fusion on the second set of data to form a third set of data. In one example, the third set of data is the data obtained by performing data fusion on the second set of data. Examples of data fusion include concatenation based on the same fields in multiple tables. Examples of data fusion also include generating statistical data of the same fields or records (e.g., summation and ratio calculation). In one example, the generation of statistical data includes counting the number of defective panels in a glass and the ratio of defective panels in multiple panels of the same glass. Optionally, the general data layer GDL is based on Apache Hive TM . Optionally, the general data layer GDL is used for data querying.

[0141] In some embodiments, the data management platform includes a fourth data layer (e.g., at least one data mart). In some embodiments, the at least one data mart includes a data mart DMT. Optionally, the data mart DMT is a NoSQL type database storing data available for computing processes. Optionally, the data mart DMT is based on Apache Hbase TM . Optionally, the data mart DMT is used for computing processes. The ETL module ETLP is configured to transform the third data layer to form a fourth set of data. Optionally, the fourth set of data classifies the data based on different types and / or rules to form a multi-level index structure. Those skilled in the art can understand that the fourth set of data of the multi-level index structure may be multiple sub-data tables with index relationships. In one example of the present disclosure, classifying the data based on different types and / or rules refers to classifying the data according to different environmental factors, keys / values, column families, etc. The first index in the multi-level index structure corresponds to the filtering criteria of the front-end interface, e.g., corresponding to the user-defined analysis criteria in the interactive task sub-interface communicating with the data management platform, thereby facilitating a faster data querying and computing process.

[0142] Those skilled in the art can understand that the first group of data, the second group of data, the third group of data, and the fourth group of data can be stored and queried in the form of one or more data tables.

[0143] In the process of converting the third group of data into the fourth group of data, in some embodiments, the data in the general data layer GDL can be imported into the data mart DMT. In one example, a first table is generated in the data mart DMT, and a second table (e.g., an external table) is generated in the general data layer GDL. The first table and the second table are configured to be synchronized so that when data is written to the second table, the first table will be updated simultaneously to include the corresponding data.

[0144] In another example, a distributed computing processing module can be used to read the data written to the general data layer GDL. The MapReduce module in Hadoop can be used as the distributed computing processing module to read the data written to the general data layer GDL. Then, the data written to the general data layer GDL can be written to the data mart DMT. In one example, the data can be written to the data mart DMT using the HBase API. In another example, once the MapReduce module reads the data written to the data mart DMT, it can generate an HFile file and bulk load it into the data mart DMT.

[0145] In some embodiments, the data flow, data conversion, and data structure among various components of the data management platform are described herein. In some embodiments, the raw data collected by multiple data sources DS includes at least one of resume data information, parameter information, or defect information. Optionally, the raw data may include dimension information (time, factory, equipment, operator, Map, chamber, card slot, etc.) and attribute information (factory location, equipment age, number of bad points, abnormal parameters, energy consumption parameters, processing duration, etc.).

[0146] The resume data information includes information on specific processes that a product (e.g., a panel or glass) has undergone during manufacturing. Examples of specific processes that a product has undergone during manufacturing include factory, process, site, equipment, chamber, card slot, and operator.

[0147] The parameter information includes information on specific environmental parameters and their changes that a product (e.g., a panel or glass) has experienced during manufacturing. Examples of specific environmental parameters and their changes that a product has experienced during manufacturing include environmental particle conditions, equipment temperature, and equipment pressure, etc.

[0148] The defect information includes information on the quality of the product based on inspection. Examples of product quality information include defect type, defect location, and defect size, etc.

[0149] In some embodiments, the parameter information includes device parameter information. Optionally, the device parameter information includes at least three types of data, which can be output from a General Model (GEM) interface for communication and control in the manufacture of the device. The first type of data that can be output from the GEM interface is a Data Variable (DV), which can be collected when an event occurs. Thus, the data variable is only valid in the case of an event. In one example, the GEM interface can provide an event called PPChanged, which is triggered when the recipe changes; and a data variable called "Changed Recipe", which is only valid in the case of the PPChanged event. Polling the value at other times may result in invalid or unexpected data. The second type of data that can be output from the GEM interface is a Status Variable (SV), which contains device-specific information that is valid at any time. In one example, the device can be a temperature sensor, and the GEM interface provides temperature status variables for one or more modules. The host can request the value of the status variable at any time, and can expect the value to be true. The third type of data that can be output from the GEM interface is an Equipment Constant (EC), which contains data items set by the device. The equipment constant determines the behavior of the device. In one example, the GEM interface provides an equipment constant called "MaxSimultousTrace", which specifies the maximum number of traces that can be requested from the host simultaneously. The value of the equipment constant is always guaranteed to be valid and up-to-date.

[0150] In some embodiments, the data lake DL is configured to store a first set of data formed by extracting raw data from multiple data sources by the ETL module ETLP. The first set of data has the same content as the raw data. The ETL module ETLP is configured to extract raw data from multiple data sources DS while maintaining dimensional information (e.g., dimensional columns) and attribute information (e.g., attribute columns). The data lake DL is configured to store the extracted data sorted by extraction time. The data can be stored in the data lake DL, which has a new name indicating the (one or more) attributes of the "data lake" and / or each data source, while maintaining the dimensions and attributes of the raw data. The first set of data and the raw data are stored in different forms. The first set of data is stored in a distributed file system, while the raw data is stored in a relational database such as an Oracle database. In one example, the business data collected by multiple data sources DS includes data from various business systems, such as a yield management system (YMS), a fault detection and classification (FDC) system, and a manufacturing execution system (MES). The data in these business systems has their respective signatures, such as product models, production parameters, and equipment model data. The ETL module ETLP uses tools (e.g., sqoop commands, data stack tools, Pentaho tools) to extract the raw production data from each business system into a hadoop in the raw data format, thus achieving the fusion of data from multiple business systems. The extracted data is stored in the data lake DL. In another example, the data lake DL is based on technologies such as Hive TM and Kudu TM . The data lake DL contains dimensional columns (time, factory, equipment, operator, Map, chamber, card slot, etc.) and attribute columns (factory location, equipment age, number of defective points, abnormal parameters, energy consumption parameters, processing duration, etc.) involved in the factory automation process.

[0151] In one example, this data management platform integrates various business data (e.g., data related to semiconductor electronic device manufacturing) into multiple data sources DS (e.g., Oracle databases). The ETL module ETLP extracts the data from multiple data sources DS into the data lake DL using, for example, data stack tools, SQOOP tools, kettle tools, Pentaho tools, or DataX tools. Then, the data is cleaned, transformed, and loaded into the data warehouse DW and the general data layer GDL. The data warehouse DW, the general data layer GDL, and the data mart DMT use tools such as Kudu TM , Hive TM and Hbase TM to store large amounts of data and analysis results.

[0152] The information generated at various stages of the manufacturing process is obtained by various sensors and inspection devices and is subsequently stored in multiple data sources DS. The calculation and analysis results generated by this defect analysis system are also stored in multiple data sources DS. Data synchronization (data flow) between the various components of the data management platform is achieved through the ETL module ETLP. For example, the ETL module ETLP is configured to obtain a parameter configuration template for the synchronization process, including network permissions and database port configurations, incoming database names and table names, outgoing database names and table names, field correspondence relationships, task types, scheduling cycles, etc. The ETL module ETLP configures the parameters of the synchronization process based on the parameter configuration template. The ETL module ETLP synchronizes data and cleans the synchronized data based on the process configuration template. The ETL module ETLP cleans the data through SQL statements to remove null values, remove outliers, and establish correlations between related tables. Data synchronization tasks include data synchronization between multiple data sources DS and the data management platform, as well as data synchronization between the various layers of the data management platform (e.g., data lake DL, data warehouse DW, general data layer GDL, or data mart DMT).

[0153] In another example, data extraction to the data lake DL can be done either in real time or offline. In the offline mode, the data extraction task is scheduled periodically. Optionally, in the offline mode, the extracted data can be stored in a storage device based on the Hadoop Distributed File System (e.g., a database based on Hive TM ). In the real-time mode, the data extraction task can be executed by OGG (Oracle GoldenGate) in combination with Apache Kafka. Optionally, in the real-time mode, the extracted data can be stored in a database based on Kudu TM . OGG reads the log files in multiple data sources (e.g., Oracle databases) to obtain added / removed data. In another example, the topic information is read by Flink, and Json is selected as the synchronization field type. The data is parsed using a JAR package, and the parsed information is sent to the Kudu API to achieve the addition / removal of Kudu table data. In one example, the front-end interface can perform display, query, and / or analysis based on the data stored in a database based on Kudu TM . In another example, the front-end interface can be based on the data stored in a database based on Kudu TM , the Hadoop Distributed File System (e.g., a database based on Apache Hive TM ) and / or based on Apache Hbase TMto perform display, query, and / or analysis on data in any one or any combination of databases. In another example, short-term data (e.g., generated over several months) is stored in a database based on Kudu TM while long-term data (e.g., all data generated over all cycles) is stored in a Hadoop Distributed File System (e.g., a database based on Apache Hive TM In another example, the ETL module ETLP is configured to extract data stored in a database based on Kudu TM to a Hadoop Distributed File System (e.g., a database based on Apache Hive TM .

[0154] By combining data from various business systems (MDW, YMS, MES, FDC, etc.), a data warehouse DW is built based on the data lake DL. The data extracted from the data lake DL is divided according to the task execution time, and the task execution time does not exactly match the timestamp in the original data. Additionally, there is a possibility of data duplication. Therefore, it is necessary to build the data warehouse DW based on the data lake DL by cleaning and standardizing the data in the data lake DL to meet the needs of upper-layer applications for data accuracy and division. The data tables stored in the data warehouse DW are obtained by cleaning and standardizing the data in the data lake DL. Based on user requirements, the field format is standardized to ensure that the data tables in the data warehouse DW are exactly the same as the data tables in multiple data sources DS. At the same time, the data is divided by date or month according to time and other fields, greatly improving the query efficiency and reducing the running memory requirements. The data warehouse DW can be one or any combination of a database based on Kudu TM and a database based on Apache Hive TM .

[0155] In some embodiments, the ETL module ETLP is configured to clean the extracted data stored in the data lake into cleaned data, and the data warehouse is configured to store the cleaned data. Examples of cleaning performed by the ETL module ETLP include removal of redundant data, removal of null-value data, removal of virtual fields, etc.

[0156] In some embodiments, the ETL module ETLP is also configured to perform standardization (e.g., field standardization and format standardization) on the extracted data stored in the data lake, and the cleaned data is data that has undergone field format standardization (e.g., format standardization of date and time information).

[0157] In some embodiments, at least a portion of the business data in the multiple data sources DS is in binary large object (blob) format. After data extraction, at least a portion of the extracted data stored in the data lake DL is in compressed hexadecimal format. Optionally, at least a portion of the cleaned data stored in the data warehouse DW is obtained by decompressing and processing the extracted data. The data in the binary large object format is converted into data in hexadecimal format when extracted and stored in the data lake; the data in the hexadecimal format is decompressed when extracted and stored in the data warehouse to form a second set of data. In one example, a business system (e.g., the above-mentioned FDC system) is configured to store a large amount of parameter data. Therefore, the data must be compressed into the blob format in the business system. During data extraction (e.g., from an Oracle database to a Hive database), the blob field will be converted into a hexadecimal (HEX) string. To retrieve the parameter data stored in the file, the HEX file is decompressed, and the content of the file can be directly obtained thereafter. The required corresponding data is encoded to form a long string, and according to the output requirements, different contents are separated by specific symbols. To obtain the data in the required format, operations such as cutting according to special characters and row-column conversion are performed on the long string. The processed data is written into the target table (e.g., the data in the table format stored in the data warehouse DW discussed above) together with the original data.

[0158] In one example, the cleaned data stored in the data warehouse DW maintains the dimension information (e.g., dimension columns) and attribute information (e.g., attribute columns) of the original data in the multiple data sources DS. In another example, the cleaned data stored in the data warehouse DW keeps the same data table name as the data table in the multiple data sources DS.

[0159] In some embodiments, the ETL module ETLP is further configured to generate a dynamically updated table with periodic updates. Optionally, as described above, the general data layer GDL is configured to store a dynamically updated table including information about high-incidence defects (one type of defect information of interest). In one example of the present disclosure, the high-incidence defect information is the defect information of the top five or top ten with a relatively high defect occurrence rate, or other defect information specified by the user. Optionally, the data mart DMT is configured to store a dynamically updated table that includes information about high-incidence defects, as described above.

[0160] The general data layer GDL is constructed based on the data warehouse DW. In some embodiments, the GDL is configured to store a third set of data formed by data fusion of the second set of data by the ETL module ETLP. Optionally, data fusion is performed based on different topics. The data in the general data layer GDL has a high degree of subject matter and a high degree of aggregation, which greatly improves the query speed. In an example, tables in the data warehouse DW can be used to construct tables with correlations constructed according to different user needs or different topics, and names are assigned to the tables according to their respective uses.

[0161] Various topics may correspond to different data analysis requirements. For example, topics may correspond to different defect analysis requirements. In one example, a topic may correspond to an analysis of defects attributed to one or more manufacturing node groups (e.g., one or more devices), and data fusion based on the topic may include data fusion of historical information about the manufacturing process and defect information. In another example, a topic may correspond to an analysis of defects attributed to one or more parameter types, and data fusion based on the topic may include data fusion of parameter feature information and defect information. In another example, a topic may correspond to an analysis of defects attributed to one or more equipment operations (e.g., equipment defined by corresponding process sites where corresponding equipment performs corresponding operations), and data fusion based on the topic may include data fusion of at least two types of information in parameter feature information, historical information about the manufacturing process, and defect information. In another example, a topic may correspond to feature extraction of at least one type of parameter information to generate parameter feature information, wherein one or more of the maximum value, minimum value, average value, and median value are extracted for at least one type of parameter information. In one example of the present disclosure, at least one type of parameter information includes data of at least one equipment parameter, such as temperature, humidity, pressure, etc., and also includes data such as environmental granularity.

[0162] In some embodiments, defect analysis includes performing feature extraction on at least one type of parameter information to generate parameter feature information; performing data fusion on at least two types of information among the parameter feature information, the resume information of the manufacturing process, and the defect information. Optionally, performing data fusion includes performing data fusion on the parameter feature information and the defect information. Optionally, performing data fusion includes performing data fusion on at least two types of information among the parameter feature information, the resume information of the manufacturing process, and the defect information. In another example, performing data fusion includes performing data fusion on the parameter feature information and the resume information of the manufacturing process to obtain first fusion data information; performing data fusion on the first fusion data information and its associated defect information to obtain second fusion data information. In one example, the second fusion data information includes glass serial number, site information, equipment information, parameter feature information, and defect information. For example, data fusion is performed in the General Data Layer (GDL) by constructing a table with relevance built according to user needs or topics. Optionally, the step of performing data fusion includes performing data fusion on the resume information and the defect information. Optionally, the step of performing data fusion includes performing data fusion on all three of the parameter feature information, the resume information of the manufacturing process, and the defect information.

[0163] In one example, the CELL_PANEL_MAIN table in the Data Warehouse (DW) stores the basic resume data of the panels in the cassette factory, and the CELL_PANEL_CT table stores the detailed data of the CT (cell test) process in the factory (such as defect information or parameter information). The General Data Layer (GDL) is configured to perform related operations based on the CELL_PANEL_MAIN table and the CELL_PANEL_CT table to create a wide table YMS_PANEL, thereby achieving data fusion. The basic resume data of the panels and the detailed data of the CT process can be queried in the YMS_PANEL table. The YMS prefix in the table name "YMS_PANEL" represents the topic for defect analysis, and the PANEL prefix represents the specific PANEL information stored in the table. By performing related operations on the tables in the Data Warehouse (DW) by the General Data Layer (GDL), the data in different tables can be fused and associated.

[0164] According to different business analysis requirements and based on glass, hglass (half glass), and panel, the tables in the General Data Layer (GDL) can be divided into the following data tags: production records, defect rate, defect MAP, DV, SV, inspection data, and test data.

[0165] Build a data mart DMT based on a data warehouse DW and / or a general data layer GDL. The data mart DMT can be used to provide various reporting data and data required for analysis, especially highly customized data. In one example, the customized data provided by the data mart DMT includes consolidated data on defect rates, frequencies of specific defects, etc. In another example, the data in the data lake DL and the general data layer GDL is stored in Hive, and the data in the data mart DMT is stored in a database based on Hbase. Optionally, the table names in the data mart DMT can be kept the same as those in the general data layer GDL. Optionally, the general data layer GDL is based on Apache Hive TM technology, and the data mart DMT is based on Apache Hbas TM technology. The general data layer GDL is used for data query through a user interface. Data in Hive can be quickly queried in Hive through Impala. The data mart DMT is used to provide computing data for an algorithm server. Based on the advantages of columnar data storage in Hbase, multiple algorithm servers AS can quickly access the data in Hbase.

[0166] In some embodiments, the data mart DMT is configured to store multiple sub-tables into which a data table in the data table stored in the general data layer GDL is divided. In some embodiments, the data stored in the data mart DMT and the data stored in the general data layer GDL have the same content. The difference between the data stored in the data mart DMT and the data stored in the general data layer GDL lies in that they are stored in different data models. Depending on the different types of NoSQL databases used for the data mart DMT, the data in the data mart DMT can be stored in different data models. Examples of data models corresponding to different NoSQL databases include key / value pair data models, column family data models, versioned document data models, and graph structure data models. In some embodiments, a query on the data mart DMT can be performed based on a specified key to quickly locate the data (e.g., value) to be queried. Thus, and as discussed more specifically below, the table stored in the general data layer GDL can be divided into at least three sub-tables in the data mart DMT. The first sub-table corresponds to the user-defined analysis criteria in the interactive task sub-interface. The second sub-table corresponds to the key (e.g., product serial number). The third sub-table corresponds to the value (e.g., the value in the table stored in the general data layer GDL, including consolidated data). It can be understood that through the first sub-table, the product range that the user needs to analyze can be determined, so as to query the corresponding data (value) in the third sub-table based on the serial number (key) of the corresponding product in the second sub-table. In one example, the data mart DMT utilizes a base on Apache Hbase TMA NoSQL database of technology; the specified key in the second subtable can be a row key; and the fused data in the third subtable (column family data corresponding to the row key) can be stored in the column family data model. Optionally, the fused data in the third subtable can be fused data from at least two of parameter feature information, historical information of the manufacturing process, and defect information. In addition, the data mart DMT may include a fourth subtable. Certain characters in the third subtable may be stored in codes, for example, due to their length or other reasons. The fourth subtable includes characters corresponding to these codes stored in the third subtable (e.g., equipment names, sites). Indexes or queries between the first subtable, the second subtable, and the third subtable may be based on the code. The fourth subtable can be used to replace codes with characters before the results are presented to the user interface.

[0167] In some embodiments, the plurality of subtables have an index relationship between at least two of the plurality of subtables. Optionally, the data in the plurality of subtables are classified based on type and / or rule. In some embodiments, the plurality of subtables include a first subtable (e.g., an attribute subtable) including a plurality of environmental factors corresponding to a user-defined analysis criterion in an interactive task subinterface communicating with a data management platform; a second subtable including a product serial number (e.g., a glass identification number or a batch identification number); and a third subtable (e.g., a master subtable) including a value corresponding to the product serial number in a third set of data. Optionally, based on different topics, the second subtable may include different designated keys, such as a glass identification number or a batch identification number (e.g., a plurality of second subtables). Optionally, the value in the third set of data corresponds to the glass identification number through an index relationship between the third subtable and the second subtable. Optionally, the plurality of subtables further include a fourth subtable (e.g., a metadata subtable) including a value corresponding to the batch identification number in the third set of data. Optionally, the second subtable further includes a batch identification number; the value corresponding to the batch identification number in the third set of data can be obtained through the index relationship between the second subtable and the fourth subtable. Optionally, the plurality of subtables further include a fifth subtable (e.g., a code generator subtable) including site information and device information. Optionally, the third subtable includes codes or abbreviations of sites and devices, and the site information and device information can be obtained from the fifth subtable through the index relationship between the third subtable and the fifth subtable.

[0168] Figure 14 Depicts multiple sub-tables derived from a data table stored in a common data layer in some embodiments according to the present disclosure. Figure 14, in some embodiments, the multiple sub-tables include one or more of the following: an attribute sub-table, which includes multiple environmental factors corresponding to user-defined analysis criteria in an interactive task sub-interface communicating with a data management platform; a context sub-table, which includes at least a first quantity of environmental factors among the multiple environmental factors and multiple manufacturing stage factors, and multiple columns corresponding to a second quantity of environmental factors among the multiple environmental factors; a metadata sub-table, which includes at least a first manufacturing stage factor among the multiple manufacturing stage factors and equipment factors associated with the first manufacturing stage, and multiple columns corresponding to parameters generated in the first manufacturing stage; a master sub-table, which includes at least a second manufacturing stage factor among the multiple manufacturing stage factors, and multiple columns corresponding to parameters generated in the second manufacturing stage; and a code generator sub-table, which includes at least a third quantity of environmental factors among the multiple environmental factors and equipment factors.

[0169] In one example, the multiple sub-tables are sub-tables stored in a column family database, including one or more of the following: an attribute sub-table, which includes a primary key composed of data label, factory information, site information, product model information, product type information, and product serial number, and based on the selection in the interactive task sub-interface, filtering conditions can be determined; a context sub-table, which includes a primary key composed of the value of the first three digits after encrypting the site with MED5, factory information, site information, data label, manufacturing end time, batch serial number, and glass serial number, and the corresponding columns are product model information, product serial number, and product type information respectively. The filtering conditions determined from the attribute sub-table are located at the corresponding batch serial number and glass serial number based on the context sub-table; a metadata sub-table, which includes a primary key composed of the value of the first three digits after encrypting the batch serial number with MED5, batch serial number, data label, site information, and equipment information, and the corresponding columns include time and manufacturing parameters, and the corresponding values include specific data values. The batch serial number determined from the context sub-table can obtain the corresponding data values based on the metadata sub-table; a master sub-table, which includes a primary key composed of the value of the first three digits after encrypting the batch serial number with MED5, serial number, and glass serial number, and the corresponding columns include time and manufacturing parameters, and the corresponding values include specific data values. The glass serial number determined from the context sub-table can obtain the corresponding data values based on the metadata sub-table; and a code generator sub-table, which includes a primary key composed of data label, site information, and equipment information, and the corresponding column is the serial number, and the corresponding site and equipment data in the code generator sub-table can be found according to the serial number in the master sub-table. Optionally, the multiple environmental factors in the attribute sub-table include data label, factory information, site information, product model information, product type information, and product serial number. Optionally, the multiple manufacturing stage factors include batch serial number and glass serial number. Optionally, the equipment factors include equipment information.

[0170] Refer to Figure 7 and Figure 8, in some embodiments, the software module further includes a load balancer LB connected to the analyzer AZ. Optionally, the load balancer LB (e.g., the first load balancer LB1) is configured to receive task requests and is configured to distribute the task requests to one or more of the multiple business servers BS to achieve load balancing among the multiple business servers BS. Optionally, the load balancer LB (e.g., the second load balancer LB2) is configured to distribute tasks from the multiple business servers BS to one or more of the multiple algorithm servers AS to achieve load balancing among the multiple algorithm servers AS. Optionally, the load balancer LB is an Nginx-based TM load balancer.

[0171] In some embodiments, the defect analysis system is configured to meet the needs of many users simultaneously. By having the load balancer LB (e.g., the first load balancer LB1), the system sends user requests to the multiple business servers AS in a balanced manner, thus keeping the overall performance of the multiple business servers AS optimal and preventing slow responses of the service due to excessive pressure on a single server.

[0172] Similarly, by having the load balancer LB (e.g., the second load balancer LB2), the system sends tasks to the multiple algorithm servers AS in a balanced manner to keep the overall performance of the multiple algorithm servers AS optimal. In some embodiments, when designing the load balancing strategy, not only the number of tasks sent to each of the multiple algorithm servers AS should be considered, but also the amount of computing load required for each task should be considered. In one example, there are three types of tasks, including defect analysis of type "glass", defect analysis of type "hglass", and defect analysis of type "panel". In another example, the average number of defect data items related to type "glass" is 1 million per week, and the average number of defect data items related to type "panel" is 30 million per week. Therefore, the amount of computing load required for defect analysis of type "panel" is much greater than that required for defect analysis of type "glass". In another example, the formula f(x, y, z) = mx + ny + oz is used to perform load balancing, where x represents the number of tasks for defect analysis of type "glass"; y represents the number of tasks for defect analysis of type "hglass"; z represents the number of tasks for defect analysis of type "panel"; m represents the weight assigned to defect analysis of type "glass"; n represents the weight assigned to defect analysis of type "hglass"; o represents the weight assigned to defect analysis of type "panel". The weights are assigned based on the amount of computing load required for each type of defect analysis. Optionally, m + n + o = 1.

[0173] In some embodiments, the ETL module ETLP is configured to generate a dynamically updated table, which is updated periodically (e.g., daily, hourly, etc.). Optionally, the General Data Layer GDL is configured to store the dynamically updated table. In one example, a dynamically updated table is generated based on the logic of the defect incidence rate in the computing factory. In another example, data from multiple tables in the Data Management Platform DM are merged and subjected to various calculations to generate a dynamically updated table. In another example, the dynamically updated table includes information such as job name, defect type, frequency of occurrence of the defect type, level of the defect type (glass / hglass / panel), factory, product model, date, and other information. The dynamically updated table is updated regularly, and when the production data in the Data Management Platform DM changes, the information in the dynamically updated table is updated accordingly to ensure that the dynamically updated table can have defect type information for all factories.

[0174] Figure 9 A defect analysis method using a defect analysis system in some embodiments according to the present disclosure is shown. Refer to Figure 9 , in some embodiments, the Data Visualization and Interaction Interface DI is configured to generate a task request; the Load Balancer LB is configured to receive the task request and is configured to assign the task request to one or more of a plurality of business servers to achieve load balancing among the plurality of business servers; one or more of the plurality of business servers are configured to send a query task request to the Query Engine QE; the Query Engine QE is configured to query the dynamically updated table when receiving the query task request from one or more of the plurality of business servers to obtain information about high-incidence defects and send the information about high-incidence defects to one or more of the plurality of business servers; one or more of the plurality of business servers are configured to send a defect analysis task to the Load Balancer LB to assign the defect analysis task to one or more of a plurality of algorithm servers, thereby achieving load balancing among the plurality of algorithm servers; when receiving the defect analysis task, one or more of the plurality of algorithm servers are configured to directly obtain data from the Data Mart DMT to perform defect analysis; and when the defect analysis is completed, one or more of the plurality of algorithm servers are configured to send the result of the defect analysis to the General Data Layer GDL.

[0175] The query engine QE can access the data management platform DM quickly. For example, it can quickly read data from or write data to the data management platform DM. Having the query engine QE is beneficial compared to direct queries through the general data layer GDL because it does not need to execute a MapReduce (MR) program to query the general data layer GDL (such as a Hive data store). Optionally, the query engine QE can be a distributed query engine that can query the general data layer GDL (HDFS or Hive) in real time, greatly reducing the waiting time and improving the responsiveness of the entire system. The query engine QE can be implemented using various suitable technologies. Examples of technologies for implementing the query engine QE include Impala TM technology, Kylin TM technology, Presto TM technology, and Greenpall TM technology.

[0176] In some embodiments, the task request is a recurring task request that defines the recurrence period during which the defect analysis will be performed. Figure 10 FIG. shows a defect analysis method using a defect analysis system in some embodiments according to the present disclosure. Refer to Figure 10 , in some embodiments, the data visualization and interaction interface DI is configured to generate a recurring task request; the load balancer LB is configured to receive the recurring task request and is configured to distribute the recurring task request to one or more of the multiple business servers to achieve load balancing among the multiple business servers; one or more of the multiple business servers are configured to send a query task request to the query engine QE; the query engine QE is configured to query a dynamically updated table when receiving the query task request from one or more of the multiple business servers to obtain information about high-incidence defects for the recurrence period and send the information about high-incidence defects to one or more of the multiple business servers; when receiving the information about the high-incidence defects during the recurrence period, one or more of the multiple business servers are configured to generate a defect analysis task based on the information about the high-incidence defects during the recurrence period; one or more of the multiple business servers are configured to send the defect analysis task to the load balancer LB to distribute the defect analysis task to one or more of the multiple algorithm servers to achieve load balancing among the multiple algorithm servers; when receiving the defect analysis task, one or more of the multiple algorithm servers are configured to directly obtain data from the data mart DMT to perform the defect analysis; and when the defect analysis is completed, one or more of the multiple algorithm servers are configured to send the result of the defect analysis to the general data layer GDL.

[0177] Refer to Figure 8, in some embodiments, the data visualization and interaction interface DI includes an automatic task sub-interface SUB1, which allows input of the repetition period for performing defect analysis. The automatic task sub-interface SUB1 is capable of automatically performing defect analysis on high-incidence defects periodically. In the automatic task mode, information about high-incidence defects is sent to multiple algorithm servers AS to analyze the potential causes of the defects. In one example, the user sets the repetition period for performing defect analysis in the automatic task sub-interface SUB1. The query engine QE captures defect information from a dynamically updated table at regular intervals based on system settings and sends this information to multiple algorithm servers AS for analysis. In this way, the system can automatically monitor high-incidence defects, and the corresponding analysis results can be stored in a cache for access and for display in the data visualization and interaction interface DI.

[0178] In some embodiments, the task request is an interactive task request. Figure 11 A defect analysis method using a defect analysis system in some embodiments according to the present disclosure is shown. Refer to Figure 11, in some embodiments, the data visualization and interaction interface DI is configured to receive user-defined analysis criteria and is configured to generate an interactive task request based on the user-defined analysis criteria; the data visualization and interaction interface DI is configured to generate an interactive task request; the load balancer LB is configured to receive the interactive task request and is configured to distribute the interactive task request to one or more of a plurality of business servers to achieve load balancing among the plurality of business servers; one or more of the plurality of business servers are configured to send a query task request to the query engine; the query engine QE is configured to, when receiving the query task request from one or more of the plurality of business servers, query a dynamically updated table to obtain information about highly-occurring defects and send the information about highly-occurring defects to one or more of the plurality of business servers; when receiving the information about highly-occurring defects, one or more of the plurality of business servers are configured to send the information to the data visualization and interaction interface; the data visualization and interaction interface DI is configured to display information about highly-occurring defects and a plurality of environmental factors associated with the highly-occurring defects and is configured to receive a user-defined selection of one or more of the plurality of environmental factors and send the user-defined selection to one or more of the plurality of business servers; one or more of the plurality of business servers are configured to generate a defect analysis task based on the information and the user-defined selection; one or more of the plurality of business servers are configured to send the defect analysis task to the load balancer LB to distribute the defect analysis task to one or more of a plurality of algorithm servers to achieve load balancing among the plurality of algorithm servers; when receiving the defect analysis task, one or more of the plurality of algorithm servers are configured to directly obtain data from the data mart DMT to perform the defect analysis; and when the defect analysis is completed, one or more of the plurality of algorithm servers are configured to send the result of the defect analysis to the general data layer GDL.

[0179] Reference Figure 8, in some embodiments, the data visualization and interaction interface DI includes an interactive task sub-interface SUB2, which allows the input of user-defined analysis criteria, including user-defined selections of one or more environmental factors. In one example, the user can stepwise filter various environmental factors in the interactive task sub-interface SUB2, including data sources, factories, sites, models, product models, batches, etc. One or more of the multiple business servers BS are configured to generate defect analysis tasks based on information about high-incidence defects and user-defined selections of one or more environmental factors. The analyzer AZ continuously interacts with the general data layer GDL and causes the selected one or more environmental factors to be displayed on the interactive task sub-interface SUB2. The interactive task sub-interface SUB2 allows the user to limit the environmental factors to a few based on the user's experience, for example, certain selected devices or certain selected parameters.

[0180] In some embodiments, the general data layer GDL is configured to generate tables based on different topics. In one example, the table includes a resume table, and the resume information includes information about the sites and devices that the glass or panel has passed through during the entire manufacturing process. In another example, the table includes a dv table, which contains parameter information uploaded by the device. In another example, if the user only wants to analyze device correlations, the user can select the resume table for analysis. In another example, if the user only wants to analyze device parameters, the user can select the dv table for analysis.

[0181] Reference Figure 8 , in some embodiments, the analyzer AZ further includes a cache server CS. The cache server CS is configured to store a part of the defect analysis task results in the cache C. In some embodiments, the data visualization and interaction interface DI further includes a defect visualization sub-interface SUB-3. In one embodiment, the main function of the defect visualization sub-interface SUB-3 is to allow the user to customize queries and display the corresponding defect analysis task results when the user clicks on the defect type. In one example, the user clicks on the defect type, and the system sends the request to one or more of the multiple business servers BS via the load balancer LB. One or more of the multiple business servers BS first query the result data cached in the cache C, and if the cached result data exists, the system directly displays the cached result data. If the result data corresponding to the selected defect type is not currently cached in the cache C, the query engine QE is configured to query the general data layer GDL for the result data corresponding to the selected defect type. Once queried, the system caches the result data corresponding to the selected defect type in the cache C, and this result data can be used for the next query of the same defect type.

[0182] Figure 12shows a defect analysis method using a defect analysis system in some embodiments according to the present disclosure. Refer to Figure 12 , in some embodiments, the defect visualization sub-interface DI is configured to receive a user-defined selection of a defect to be analyzed and generate a call request; the load balancer LB is configured to receive the call request and is configured to distribute the call request to one or more of a plurality of business servers to achieve load balancing among the plurality of business servers; one or more of the plurality of business servers are configured to send the call request to the cache server; and the cache server is configured to determine whether information about the defect to be analyzed is stored in the cache. Optionally, when it is determined that the information about the defect to be analyzed is stored in the cache, one or more of the plurality of business servers are configured to send the information about the defect to be analyzed to the defect visualization sub-interface for display. Optionally, when it is determined that the information about the defect to be analyzed is not stored in the cache, one or more of the plurality of business servers are configured to send a query task request to the query engine; the query engine is configured to, when receiving the query task request from one or more of the plurality of business servers, query a dynamically updated table to obtain information about the defect to be analyzed, and send the information about the defect to be analyzed to the cache; the cache is configured to store the information about the defect to be analyzed; one or more of the plurality of business servers are configured to send the information about the defect to be analyzed to the defect visualization sub-interface for display.

[0183] Optionally, a part of the defect analysis task result includes a defect analysis task result based on a periodic task request. Optionally, a part of the defect analysis task result includes a defect analysis task result based on a periodic task request; and a defect analysis task result obtained based on a query task request.

[0184] By having a cache server CS, a high requirement for the system response speed (e.g., displaying results associated with defect types) can be met. In one example, through periodic task requests, up to 40 tasks can be generated every half hour, where each task is associated with up to five different defect types, and each defect type is associated with up to 100 parameter information. If all analysis results are cached, a total of 40 * 5 * 100 = 20000 queries must be stored in the cache C, which will impose a great pressure on the dot cluster memory. In one example, a part of the defect analysis task result is limited to the results associated with the top three ranked defect types, and only this part is cached.

[0185] Various suitable methods for defect analysis can be implemented by one or more of the plurality of algorithm servers of the defect analysis system described herein. Figure 15 shows a defect analysis method in some embodiments according to the present disclosure. Refer toFigure 15 , in some embodiments, the method includes obtaining manufacturing data information including defect information; classifying the manufacturing data information into multiple groups of data according to a manufacturing node group, where each group of data in the multiple groups of data is associated with each manufacturing node group in the manufacturing node group; calculating the weight of evidence of the manufacturing node group to obtain multiple weights of evidence, where the weight of evidence represents the difference between the proportion of defect data in each manufacturing node group and the proportion of defect data in all manufacturing node groups; sorting the multiple groups of data based on the multiple weights of evidence; obtaining a list of the multiple groups of data sorted based on the multiple weights of evidence; and performing defect analysis on one or more selected groups of the multiple groups of data. Optionally, each manufacturing node group includes one or more selected from the group consisting of a manufacturing process, equipment, a site, and a process section. Optionally, the manufacturing data information can be obtained from a data mart DMT. Optionally, the manufacturing data information can be obtained from a general data layer GDL.

[0186] Optionally, the method includes processing manufacturing data information including history data information and defect information to obtain processed data; classifying the processed data into multiple groups of data according to an equipment group, where each group of data in the multiple groups of data is associated with each equipment group in the equipment group; calculating the weight of evidence of the equipment group to obtain multiple weights of evidence; sorting the multiple groups of data based on the multiple weights of evidence; and performing defect analysis on one or more groups with the highest ranking in the multiple groups of data. Optionally, the defect analysis is performed at a parameter level.

[0187] In some embodiments, each weight of evidence of each equipment group is calculated according to Equation (1):

[0188]

[0189] where woei represents each weight of evidence of each equipment group; P(yi) represents the ratio of the number of positive samples in each equipment group to the number of positive samples in all manufacturing node groups (e.g., equipment groups); P(ni) represents the ratio of the number of negative samples in each equipment group to the number of negative samples in all manufacturing node groups (e.g., equipment groups); a positive sample represents data including defect information associated with each equipment group; a negative sample represents data in which there is no defect information associated with each equipment group; #yi represents the number of positive samples in each equipment group; #yr represents the number of positive samples in all manufacturing node groups (e.g., equipment groups); #ni represents the number of negative samples in each equipment group; #nr represents the number of negative samples in all manufacturing node groups (e.g., equipment groups).

[0190] In some embodiments, the method further comprises processing the manufacturing data information to obtain processing data. Optionally, processing the manufacturing data information comprises performing data fusion on the history data information and the defect information to obtain fused data information.

[0191] In one example, processing manufacturing data information to obtain processing data includes obtaining raw data information of various manufacturing processes of the display panel, including historical data information, parameter information, and defect information; preprocessing the raw data to remove null data, redundant data, and virtual fields, and filtering the data based on preset conditions to obtain verification data; fusing the historical data information and defect information in the verification data to obtain third fused data information; determining whether any defect information in the fused data information contains the same machine inspection defect information and manual review defect information, identifying the manual review defect information (rather than the machine inspection defect information) as the defect information to be analyzed, thereby generating audited data; fusing the audit data and the historical data information to obtain fourth fused data information; removing non-representative data from the fourth fused data information to obtain processed data. For example, data generated in the process of glass passing through a very small number of devices can be eliminated. When the number of devices through which the glass passes only accounts for a small proportion (e.g., 10%) of the total number of devices, non-representative data will deviate the analysis, thereby affecting the accuracy of the analysis.

[0192] In one example, the historical data information (used to be merged with the audit data to obtain the fourth fused data information) includes glass data and hglass data (half glass data, i.e., historical data after the entire glass is cut in half). However, the audited data is panel data. In one example, the glass_id / hglass_id in the fab (manufacturing) stage is merged with the panel_id in the EAC2 stage, where redundant data is removed. The purpose of this step is to ensure that the historical data information in the fab stage is consistent with the defect information in the EAC2 stage. For example, the number of bits in glass_id / hglass_id is different from the number of bits in panel_id. In one example, the number of bits in panel_id is processed to be consistent with the number of bits in glass_id / hglass_id. After data fusion, complete data is obtained, including glass_id / hglass_id, site information, equipment information, and defect information. Optionally, the fused data is subjected to additional operations to remove redundant data items.

[0193] In some embodiments, performing defect analysis includes performing feature extraction on at least one type of parameter information to generate parameter feature information, wherein one or more of a maximum value, a minimum value, an average value, and a median value are extracted for at least one type of parameter information. Optionally, performing feature extraction includes performing time domain analysis to extract statistical information, the statistical information including one or more of count, average value, maximum value, minimum value, range, variance, deviation, kurtosis, and percentile. Optionally, performing feature extraction includes performing frequency domain analysis to convert time domain information obtained in the time domain analysis into frequency domain information including one or more of a power spectrum, information entropy, and a signal-to-noise ratio.

[0194] In one example, feature extraction is performed on a list of multiple groups of data ranked based on multiple weights of evidence. In another example, feature extraction is performed on one or more groups of the multiple groups of data with the highest ranking. In another example, feature extraction is performed on the group of data with the highest ranking.

[0195] In some embodiments, performing defect analysis also includes performing data fusion on at least two of parameter feature information, historical information of the manufacturing process, and defect information. Optionally, performing data fusion includes performing data fusion on parameter feature information and defect information. Optionally, performing data fusion includes performing data fusion on parameter feature information, historical information of the manufacturing process, and defect information. In another example, data fusion is performed on the parameter feature information and the historical information of the manufacturing process to obtain first fused data information; data fusion is performed on the first fused data information and the defect information associated with it to obtain second fused data information, and the second fused data information includes glass serial number, site information, equipment information, parameter feature information, and defect information. In some embodiments, for example, data fusion is performed in a general data layer GDL by constructing a table having correlations constructed according to user needs or themes as described above.

[0196] In some embodiments, the method further comprises performing a correlation analysis. Figure 16 The defect analysis method in some embodiments of the present disclosure is shown. Figure 16 In some embodiments, the method includes extracting parameter feature information and defect information from the second fused data information; performing correlation analysis on the parameter feature information and the defect information for each type of parameter; generating multiple correlation coefficients for multiple types of parameters respectively; and sorting the absolute values of the multiple correlation coefficients. In one example, the absolute values of the multiple correlation coefficients are arranged in order from largest to smallest, so that the relevant parameters that cause the defect to appear can be visually observed. The absolute value is used here because the correlation coefficient can be positive or negative, that is, there can be a positive or negative correlation between the parameter and the defect. The larger the absolute value, the stronger the correlation.

[0197] In some embodiments, the plurality of correlation coefficients are a plurality of Pearson correlation coefficients. Optionally, each Pearson correlation coefficient is calculated according to Equation (2):

[0198]

[0199] where x represents the value of a parameter feature; y represents the value of the presence or absence of a defect, where when a defect is present, y is assigned the value 1, and when no defect is present, y is assigned the value 0; μ x represents the mean of x; μ y represents the mean of y; σ x σ y represents the product of the respective standard deviations of x and y; cov(x, y) represents the covariance of x and y; and ρ(x, y) represents each Pearson correlation coefficient.

[0200] In another aspect, the present disclosure provides a defect analysis method executed by a distributed computing system, the distributed computing system including one or more networked computers configured to execute in parallel to perform at least one common task. In some embodiments, the method includes: executing a data management platform configured to store data and extract, transform, or load data; executing a query engine connected to the data management platform and configured to directly obtain data from the data management platform; executing an analyzer connected to the query engine and configured to perform defect analysis upon receiving a task request, the analyzer including a plurality of backend servers and a plurality of algorithm servers configured to obtain data from the data management platform; and executing a data visualization and interaction interface configured to generate the task request.

[0201] In some embodiments, the data management platform includes an ETL module configured to extract, transform, or load data from a plurality of data sources onto a data mart and a common data layer. In some embodiments, the method further includes: when a respective one of the plurality of algorithm servers receives an assigned task, directly querying first data from the data mart by the respective one of the plurality of algorithm servers; and when performing defect analysis, directly sending second data to the common data layer by the respective one of the plurality of algorithm servers.

[0202] In some embodiments, the method further includes generating a periodically updated and dynamically updated table by the ETL module; storing the dynamically updated table in the common data layer.

[0203] In some embodiments, the software module further includes a load balancer connected to the analyzer. In some embodiments, the method further includes receiving, by the load balancer, a task request, and distributing, by the load balancer, the task request to one or more of a plurality of backend servers to achieve load balancing among the plurality of backend servers, and distributing, by the load balancer, tasks from the plurality of backend servers to one or more of a plurality of algorithm servers to achieve load balancing among the plurality of algorithm servers.

[0204] In some embodiments, the method further includes generating, by the data visualization and interaction interface, a task request; receiving, by the load balancer, the task request, and distributing, by the load balancer, the task request to one or more of a plurality of backend servers to achieve load balancing among the plurality of backend servers; sending, by one or more of the plurality of backend servers, a query task request to the query engine; when the query engine receives the query task request from one or more of the plurality of backend servers, querying, by the query engine, a dynamically updated table to obtain information about highly-occurring defects; sending, by the query engine, the information about highly-occurring defects to one or more of the plurality of backend servers; sending, by one or more of the plurality of backend servers, a defect analysis task to the load balancer to distribute the defect analysis task to one or more of a plurality of algorithm servers to achieve load balancing among the plurality of algorithm servers; when one or more of the plurality of algorithm servers receive the defect analysis task, querying, by one or more of the plurality of algorithm servers, data directly from the data mart to perform defect analysis; and when the defect analysis is completed, sending, by one or more of the plurality of algorithm servers, the result of the defect analysis to the common data layer.

[0205] In some embodiments, the method further includes generating a periodic task request. The periodic task request defines a repeating period for performing defect analysis. Optionally, the method further includes querying, by the query engine, a dynamically updated table to obtain information about highly-occurring defects in the repeating period; and when receiving information about highly-occurring defects during the repeating period, generating, by one or more of the plurality of backend servers, a defect analysis task based on the information about highly-occurring defects during the repeating period. Optionally, the method further includes receiving, for example, an input of the repeating period for performing defect analysis through an automatic task sub-interface of the data visualization and interaction interface.

[0206] In some embodiments, the method further includes generating an interactive task request. Optionally, the method further includes receiving, via a data visualization and interaction interface, user-defined analysis criteria; generating, by the data visualization and interaction interface, an interactive task request based on the user-defined analysis criteria; sending, by one or more of the plurality of backend servers, information to the data visualization and interaction interface when information regarding a high-incidence defect is received; displaying, via the data visualization and interaction interface, information regarding the high-incidence defect and a plurality of environmental factors associated with the high-incidence defect; receiving, by the data visualization and interaction interface, a user-defined selection of one or more of the plurality of environmental factors; sending, by the data visualization and interaction interface, the user-defined selection to one or more of the plurality of backend servers; and generating, by one or more of the plurality of backend servers, a defect analysis task based on the information and the user-defined selection. Optionally, the method further includes receiving an input of user-defined analysis criteria, such as via an interactive task sub-interface of the data visualization and interaction interface, the user-defined analysis criteria including a user-defined selection of one or more environmental factors.

[0207] In some embodiments, the analyzer further includes a cache server and a cache. The cache is connected to the plurality of backend servers, the cache server, and the query engine. Optionally, the method further includes storing, by the cache, a portion of the defect analysis task results.

[0208] In some embodiments, the data visualization and interaction interface includes a defect visualization sub-interface. Optionally, the method further includes receiving, via the defect visualization sub-interface, a user-defined selection of a defect to be analyzed and generating a call request; receiving the call request by a load balancer; distributing, by the load balancer, the call request to one or more of a plurality of backend servers to achieve load balancing among the plurality of backend servers; sending, by one or more of the plurality of backend servers, the call request to a cache server; and determining, by the cache server, whether information about the defect to be analyzed is stored in the cache. Optionally, the method further includes, when determining that the information about the defect to be analyzed is stored in the cache, configuring one or more of the plurality of backend servers to send the information about the defect to be analyzed to the defect visualization sub-interface for display. Optionally, the method further includes, when determining that the information about the defect to be analyzed is not stored in the cache, sending, by one or more of the plurality of backend servers, a query task request to a query engine; querying, by the query engine upon receiving the query task request from one or more of the plurality of backend servers, a dynamically updated table to obtain information about the defect to be analyzed; sending, by the query engine, the information about the defect to be analyzed to the cache; storing the information about the defect to be analyzed in the cache; and sending, by one or more of the plurality of backend servers, the information about the defect to be analyzed to the defect visualization sub-interface for display. Optionally, the portion of the defect analysis task result includes a defect analysis task result based on a periodic task request; and a defect analysis task result obtained based on a query task request.

[0209] In another aspect, the present disclosure provides a computer program product for defect analysis. The computer program product for defect analysis includes a non-transitory tangible computer-readable medium having computer-readable instructions thereon. In some embodiments, the computer-readable instructions are executable by a processor in a distributed computing system to cause the processor to perform: executing a data management platform configured to store data and extract, transform, or load data; executing a query engine connected to the data management platform and configured to obtain the data directly from the data management platform; executing an analyzer connected to the query engine and configured to perform defect analysis upon receiving a task request, the analyzer including a plurality of backend servers and a plurality of algorithm servers configured to obtain the data directly from the data management platform; and executing a data visualization and interaction interface configured to generate a task request, wherein the distributed computing system includes one or more networked computers configured to execute in parallel to perform at least one common task.

[0210] In some embodiments, the data management platform includes an ETL module configured to extract, transform, or load data from multiple data sources onto a data mart and a common data layer. In some embodiments, computer-readable instructions may further be executed by a processor in a distributed computing system to cause the processor to perform: when a task assigned is received by a corresponding one of the multiple algorithm servers, querying, by the corresponding one of the multiple algorithm servers, a first data directly from the data mart; and when performing defect analysis, sending, by the corresponding one of the multiple algorithm servers, a second data directly to the common data layer.

[0211] In some embodiments, computer-readable instructions may also be executed by a processor in a distributed computing system to cause the processor to perform generating, by the ETL module, a dynamically updated table with periodic updates; and storing the dynamically updated table in the common data layer.

[0212] In some embodiments, the software module further includes a load balancer connected to an analyzer. In some embodiments, computer-readable instructions may further be executed by a processor in a distributed computing system to cause the processor to perform: receiving, by the load balancer, a task request, and distributing, by the load balancer, the task request to one or more of the multiple backend servers to achieve load balancing among the multiple backend servers, and distributing, by the load balancer, tasks from the multiple backend servers to one or more of the multiple algorithm servers to achieve load balancing among the multiple algorithm servers.

[0213] In some embodiments, computer-readable instructions may further be executed by a processor in a distributed computing system to cause the processor to perform: generating, by a data visualization and interaction interface, a task request; receiving, by the load balancer, the task request, and distributing, by the load balancer, the task request to one or more of the multiple backend servers to achieve load balancing among the multiple backend servers; sending, by one or more of the multiple backend servers, a query task request to a query engine; when the query engine receives the query task request from one or more of the multiple backend servers, querying, by the query engine, the dynamically updated table to obtain information about highly-occurring defects; sending, by the query engine, the information about highly-occurring defects to one or more of the multiple backend servers; sending, by one or more of the multiple backend servers, a defect analysis task to the load balancer to distribute the defect analysis task to one or more of the multiple algorithm servers to achieve load balancing among the multiple algorithm servers; when a defect analysis task is received by one or more of the multiple algorithm servers, querying, by one or more of the multiple algorithm servers, data directly from the data mart to perform defect analysis; and when the defect analysis is completed, sending, by one or more of the multiple algorithm servers, the result of the defect analysis to the common data layer.

[0214] In some embodiments, the computer-readable instructions may further be executed by a processor in a distributed computing system to cause the processor to perform: generating a periodic task request. The periodic task request defines a repeating period for performing defect analysis. Optionally, the computer-readable instructions may further be executed by a processor in a distributed computing system to cause the processor to perform: querying a dynamically updated table by a query engine to obtain information on defects with a high incidence rate during the repeating period; and when receiving information on defects with a high incidence rate during the repeating period, generating a defect analysis task by one or more of the plurality of backend servers based on the information on defects with a high incidence rate during the repeating period. Optionally, the computer-readable instructions may further be executed by a processor in a distributed computing system to cause the processor to perform: receiving, for example, via an automatic task sub-interface of a data visualization and interaction interface, an input of the repeating period for performing defect analysis.

[0215] In some embodiments, the computer-readable instructions may further be executed by a processor in a distributed computing system to cause the processor to perform: generating an interactive task request. Optionally, the computer-readable instructions may further be executed by a processor in a distributed computing system to cause the processor to perform: receiving user-defined analysis criteria by a data visualization and interaction interface; generating the interactive task request by the data visualization and interaction interface based on the user-defined analysis criteria; sending information to the data visualization and interaction interface by one or more of the plurality of backend servers when receiving information on defects with a high incidence rate; displaying, via the data visualization and interaction interface, information on defects with a high incidence rate and a plurality of environmental factors associated with the defects with a high incidence rate; receiving a user-defined selection of one or more of the plurality of environmental factors by the data visualization and interaction interface; sending the user-defined selection to one or more of the plurality of backend servers by the data visualization and interaction interface; and generating a defect analysis task by one or more of the plurality of backend servers based on the information and the user-defined selection. Optionally, the computer-readable instructions may further be executed by a processor in a distributed computing system to cause the processor to perform: receiving, for example, via an interactive task sub-interface of a data visualization and interaction interface, an input of user-defined analysis criteria that includes a user-defined selection of one or more environmental factors.

[0216] In some embodiments, the analyzer further includes a cache server and a cache. The cache is connected to the plurality of backend servers, the cache server, and the query engine. Optionally, the computer-readable instructions may further be executed by a processor in a distributed computing system to cause the processor to perform: storing, by the cache, a portion of the defect analysis task results.

[0217] In some embodiments, the data visualization and interaction interface includes a defect visualization sub-interface. Optionally, the computer-readable instructions may also be executed by a processor in a distributed computing system to cause the processor to perform: receiving a user-defined selection of a defect to be analyzed through the defect visualization sub-interface and generating a call request; receiving the call request by a load balancer; allocating the call request by the load balancer to one or more of a plurality of backend servers to achieve load balancing among the plurality of backend servers; sending the call request by one or more of the plurality of backend servers to a cache server; and determining by the cache server whether information about the defect to be analyzed is stored in the cache. Optionally, the computer-readable instructions may also be executed by a processor in a distributed computing system to cause the processor to perform: when it is determined that the information about the defect to be analyzed is stored in the cache, one or more of the plurality of backend servers are configured to send the information about the defect to be analyzed to the defect visualization sub-interface for display. Optionally, the computer-readable instructions may also be executed by a processor in a distributed computing system to cause the processor to perform: when it is determined that the information about the defect to be analyzed is not stored in the cache, sending a query task request by one or more of the plurality of backend servers to a query engine; querying, by the query engine, a dynamically updated table to obtain information about the defect to be analyzed when receiving the query task request from one or more of the plurality of backend servers; sending the information about the defect to be analyzed by the query engine to the cache; storing the information about the defect to be analyzed in the cache; and sending the information about the defect to be analyzed by one or more of the plurality of backend servers to the defect visualization sub-interface for display. Optionally, the portion of the defect analysis task result includes a defect analysis task result based on a periodic task request; and a defect analysis task result obtained based on a query task request.

[0218] Example: Research on Automatic MAP Defect Location of AMOLED Based on Machine Learning and Image Processing

[0219] Due to the complex production process, AMOLED is vulnerable to the environment, the cleanliness of chemical gases and liquids, and a large number of uneven point defects will be generated on the glass during the manufacturing process. The concentrated areas with a large number of point defects will lead to a loss of the final yield. Therefore, it is very important to detect and locate the concentrated areas in a timely manner in the field of AMOLED manufacturing. This example proposes a new algorithm to automatically locate the concentrated areas based on machine learning and image processing technologies. First, using the hierarchical clustering algorithm, the defects are divided into multiple classes through the Euclidean distance threshold of the defect points and the α Shapes algorithm to extract the outer contour points. Secondly, the interpolation algorithm is used to fit the minimum surrounding area of each class, and then the Hu moment algorithm is used to calculate the regional features, such as regional density, centroid, orientation, area, aspect ratio, etc. Finally, according to the eigenvalue of the target area, the candidate areas are filtered. The experimental results show that the proposed algorithm can analyze the defect MAP image, automatically locate the defect concentrated area, which can replace manual inspection, ensure quality, and reduce costs.

[0220] Flexible display materials can be used to make AMOLED deformable and bendable. AMOLED has the advantages of bendability, low power consumption, good display quality, long life, etc. However, due to the complexity of the AMOLED process, defect points with uneven sizes are formed on the panel due to the influence of air cleanliness, chemical gases, liquids, and equipment process parameters. In most cases, except for the clusters of the aggregated defect points, these unwanted defect points do not cause final product defects. In this example, we perform clustering analysis on the MAP map generated by the defect points detected by AOI and identify the clusters of the spots that cause yield loss for subsequent defect analysis.

[0221] In the past, the analysis of the aggregation of defect points was carried out by manual visual inspection, but the quality and consistency of the inspection could not be guaranteed. Based on the above problems, this example proposes an analysis method for the defect point map based on machine learning and image processing technologies, which can quickly locate the aggregation area of the defect points in the defect point map and analyze the aggregation area according to the characteristic parameters of the area to screen out the target defect point cluster area that causes yield loss, thus realizing the full-automatic online analysis of the defect point map. At the same time, through a series of quantitative evaluation indicators, the method disclosed in this paper reduces the risks of misjudgment and omission caused by the subjective judgment of workers, saves labor costs, and improves the inspection efficiency.

[0222] The defect point map is a mapping map, which refers to mapping a set of defect points into a digital image according to coordinates for visualization and subsequent defect analysis. During the AMOLED manufacturing process, the same panel will pass through several AOI detection points, and the AOI equipment will detect all the defect points P i and coordinates (x i , y i)Report to the storage system. The process of synthesizing the defective point map is a process of drawing the defective coordinates on the image.

[0223] First, create the coordinate system O of the glass panel glass -XY to the image coordinate system O map -XY two-dimensional mapping M 3×3 , which is expressed as follows:

[0224]

[0225] Among them, represents the rotation matrix, θ represents the rotation angle, represents the translation matrix.

[0226] Then, all the defective points p in the coordinate system O of the glass panel glass -XY i (x i , y i ) are transformed as described in formula (2) to obtain the coordinates p of this point in the image coordinate system O map -XY i (x i , y i ).

[0227]

[0228] Finally, the defective points obtained by the mapping transformation are drawn on an image with a resolution of M×N. M represents the length of the substrate, and N represents the width of the substrate. Figure 17 is an exemplary defective point map image.

[0229] In the algorithm for automatically locating the region with poor aggregation of the defective point map in the present invention, regions that satisfy the aggregation characteristics are extracted for subsequent image analysis. In this example, an unsupervised learning algorithm for hierarchical clustering is used. Hierarchical clustering is an unsupervised learning algorithm in data mining, and its purpose is to divide the data into family classes with the maximum intra-class similarity and the minimum inter-class similarity.

[0230] The hierarchical clustering algorithm is divided into a merging (bottom-up) method and a splitting (top-down) method. In the merging method, each object is treated as a separate point cluster, and then similar classes are continuously merged until all objects are merged into a single point cluster or certain termination conditions are met. Examples of the hierarchical clustering algorithm include AGNES, BIRCH, CURE, ROCK, CHAMELEON.

[0231] In this example, the AGNES hierarchical clustering algorithm is used. The Euclidean distance D between point clusters is used as a measure for point cluster analysis, and the minimum sampling distance between point clusters is used as the connection criterion for point cluster analysis. All points in point cluster R with a distance D ≤ d are considered as one type of point cluster.

[0232]

[0233] Among them, p i (x i , y i ) and p j (x j , y j ) are two closest neighboring points in any two types of point clusters.

[0234] The clustering process is as follows. 1) Consider the M defect points in the defect point map as point clusters, and calculate the distance between each pair of point clusters. 2) Randomly find two point clusters with the minimum distance and merge them to obtain M - 1 point clusters. 3) Calculate the distance between two adjacent point clusters among the M - 1 point clusters. 4) Repeat steps 2) and 3). (5) Obtain the number N of point clusters that satisfy the condition D ≤ d.

[0235] The α Shapes algorithm is an algorithm for reconstructing an image of a two - dimensional region of an unordered point cluster of points. α Shapes is based on the principle of a circle with a radius of α rolling outside an unordered point cluster S of points. If α is large enough, the circle will not fall into the interior of the point cluster of points. The trajectory of the rolling circle can be considered as the boundary line of the point cluster. In this algorithm, the radius α is the only parameter, and its size determines the fineness of the boundary region. When α is large enough, the extracted boundary line is the convex hull of point cluster S. When α is small enough, any point in point cluster S can be a boundary point. In addition, the adaptive α Shapes algorithm enables the rolling circle to adaptively adjust the value of the radius α when rolling along the boundary to ensure the fineness and integrity of the boundary, and its core algorithm process is as follows.

[0236] 1) Use the point cloud data to generate a mapping graph to obtain an image with a resolution of M x N size. The value of the pixel is equal to the number of points within that pixel.

[0237] (2) Boundary point determination: Traverse each pixel. If the value of the pixel is greater than 0 and the 8 neighboring points are greater than 0, then this point is a non - boundary point, and this pixel can be discarded.

[0238] 3) Determine the α Shapes of the remaining points.

[0239] (a) Traverse all the remaining pixels p i (x i , y i), and use the core idea of the K-Nearest Neighbor algorithm to calculate their K nearest neighbors and the average of their Euclidean distances. Designate the average of their Euclidean distances as the radius α of the rolling circle. Search for all pixels whose distance from this pixel is less than 2α to form a new point set Q.

[0240] (b) For any point p in the point set Q j (x j , y j ), two rolling circles and their centers O1 and O2 are determined based on p i 、p j and the radius α. If the distances from all other points in the set Q to the centers O1 and O2 are greater than α, then this point p i is a boundary point.

[0241] (c) If there is no point in the set Q that satisfies the condition, then this point p i is not a boundary point.

[0242] 4) Repeat steps 2) and 3) until all boundary points are identified.

[0243] In this example, we obtain clusters of points of classes that satisfy the clustering conditions, each time having a series of discrete points. Figure 18A An exemplary discrete point cluster is shown. The algorithm involved in this example needs to fit the discrete points to a polygon region and use image processing techniques to calculate its region features for aggregated region screening.

[0244] The most commonly used technique for region fitting of discrete points is convex hull fitting. The region enclosed by the discrete point cluster is concave. If the convex hull fitting algorithm is used, the obtained domain is obviously not the true shape of the point cluster. Figure 18B An example of a clustering region map by convex hull fitting is shown. In this example, the adaptive α Shapes algorithm is used to find the boundary points of the point cluster and connect the boundary points to form the aggregated region of the point cluster. Figure 18C An example of a clustering region map by α Shapes fitting is shown.

[0245] The automatic extraction of the target region of an image can include first performing threshold segmentation on the image, connecting region labeling, calculating region feature values, and then automatically locating the target region based on its features. Hu proposed the concept of geometric moments based on Cartesian coordinates in 1962 and derived a series of variables with scale invariance, translation invariance, and rotation invariance.

[0246] The geometric moments and central distances of Hu are defined as follows:

[0247]

[0248]

[0249] Among them, m pq represents the geometric moment of the (p + q)-th order, M pq represents the geometric central moment of the (p + q)-th order, A represents the target area, and (x0, y0) represents the geometric center of area A.

[0250] In this example, the eigenvalue of the clustering area includes area a, point density ρ, centroid O (O x , O y ), direction θ, length L, width W, aspect ratio r, etc., and is used to automatically filter and locate the defect point clustering area in the defect point map. For any point cluster area A i , the specific formulas of the above eigenvalues are as follows:

[0251] a = m 00 (8);

[0252]

[0253]

[0254]

[0255]

[0256] Among them, N represents the number of defect points in area A i , and L and W respectively represent the length and width of the minimum circumscribed rectangle of area A i .

[0257] In this example, the defect points reported by AOI are mapped to the defect point map, and the clustering area of the defect points is automatically located through hierarchical clustering, adaptive αShapes, regional feature extraction, and filtering algorithms. Figure 19 Shows the algorithm for determining the clustering area map.

[0258] (1) Read the coordinates (x i , y i ) of single / multiple glass substrate defect points reported by AOI, establish the mapping between the glass substrate coordinate system and the image coordinate system (Equation 2), and convert the defect point coordinates in the substrate coordinate system into the image coordinate system coordinates (x i , y i ). In this example, a batch of substrates (28 substrates) are selected for defect analysis. All the defect point coordinates from this batch of substrates are superimposed to form a defect point map. Figure 20A Shows an exemplary defect point map. In Figure 20A the shown defect point map, the defect points are clustered on the top side, the right side, and the middle and lower sides.

[0259] (2) Using the Euclidean distance D between defect points as a prerequisite for point cluster analysis, and regarding all points in the point cluster R with D ≤ d as the same type of point cluster. The specific implementation uses the hierarchical clustering algorithm in the field of machine learning, and uses single-linkage as the connection criterion for point cluster analysis to obtain the final classification result C = {C1, C2, C3,..., C n}, and n ≥ 1.

[0260] (3) For the classification result in step (2), filter the defect points in each type of point cluster and count them, excluding the point clusters with N ≤ n, where n is the minimum number of non-expected points in the point cluster. Obtain the filtered result C' = {C'1, C'2, C'3,..., C' n}, Due to the hierarchical clustering algorithm used in this example, single points or a small number of defect points can be considered as independent classes, which are discarded before subsequent analysis. Figure 20B Shows an exemplary result of classification and selection.

[0261] 4) If the filtered result in step (3) is then there is no problem of defect point aggregation in the defect point map. Otherwise, perform further analysis.

[0262] 5) For any point cluster C i ' = (p1, p2,..., p n ), use the adaptive α Shapes algorithm to extract the intuitive external shape from the discrete and disordered point cluster to obtain the set C c i = (p c 1, p c 2,..., p c n ), The external contours C c of all point clusters C' can be obtained from this process.

[0263] (6) In the field of image processing, using the interpolation fitting technique, according to the contour points C i of any point cluster C c i fit the minimum closed graphic area A of this point cluster to obtain the corresponding image area set A = {A1, A2,..., A7} of this point cluster C'. Figure 21A Shows an exemplary set of image areas of point clusters.

[0264] For any image area A, first calculate the Hu geometric moment m i,j of the graphic area and the central distance M ij(Equations 4 and 5), the dot density ρ, area a, centroid O (O x , O y ), direction θ, length L, width W, aspect ratio r and other characteristic parameters of the image region (Equations 6 to 10) are derived, and the feature vector F of the generated region is F = [ρ, a, O x , O y , θ, L, W, r] T . Figure 21A The eigenvalues of each image region in are shown in Table 1. Among them, the resolution of the defect point map is 1500 pixels * 1850 pixels, and the unit of the calculated values in Table 1 is pixels.

[0265] Table 1: Figure 21A The eigenvalues of each image region in

[0266]

[0267] 8) Based on the region feature vector calculated in step 7), remove the regions in the set that do not satisfy the condition F i ∈ [α i , β i . For example: Based on the region center coordinates (O x , O y ), sets A1 and A2 can be excluded because the peripheral region of the glass substrate has no impact on the product quality. Based on the area a, regions A5 and A7 can be excluded. The aggregation of small areas has little impact on the final quality of the product. A set of defect point aggregation regions A' = {A3, A4, A6} that affect the yield loss can be obtained. Figure 21B shows the defect point aggregation region.

[0268] The method described in this example can be used in AMOLED manufacturing, but is also applicable to other flat panel display and semiconductor industries. The method described in this embodiment innovatively combines algorithms such as hierarchical clustering, α Shapes in machine learning, and blob analysis in image processing, utilizes the defect location idea in image processing, and completes the analysis of the defect point map. Experimental results show that this method can quickly locate the aggregation region of the defect point map. Through a series of quantitative judgment indicators, this method reduces the subjective judgment risk caused by human misjudgment and omission, saves labor costs, and improves the detection efficiency.

[0269] The various illustrative operations described in connection with the configurations disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. These operations may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an ASIC or ASSP, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the configurations disclosed herein. For example, such a configuration may be implemented at least in part as hardwired circuitry, a circuit configuration fabricated into a special purpose integrated circuit, or a firmware program loaded into non-volatile storage, or a software program loaded or loadable from a data storage medium as machine-readable code that is instructions executable by an array of logic elements such as a general purpose processor or other digital signal processing unit. The general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Software modules may reside in a non-transitory storage medium such as RAM (random access memory), ROM (read only memory), non-volatile RAM (NVRAM), such as flash RAM, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, a hard disk, a removable disk, or a CD-ROM; or in any other form of storage medium known in the art. The illustrative storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral with the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.

[0270] The foregoing description of the embodiments of the present invention has been presented for purposes of illustration and description. It is not exhaustive and is not intended to limit the present invention to the precise forms or exemplary embodiments disclosed. Thus, the foregoing description should be considered illustrative rather than restrictive. Obviously, many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to explain the principles of the present invention and its best mode of practical application, so that those skilled in the art can understand the various embodiments of the present invention and the various modifications suitable for the particular use or implementation being considered. The scope of the present invention is intended to be defined by the appended claims and their equivalents, where all terms are meant in their broadest reasonable sense unless otherwise stated. Thus, terms such as "the invention", "the present invention", etc. do not necessarily limit the scope of the claims to a particular embodiment, and the reference to exemplary embodiments of the present invention does not imply a limitation of the present invention and no such limitation should be inferred. The present invention is limited only by the spirit and scope of the appended claims. In addition, these claims may refer to the use of "first", "second", etc. followed by a noun or element. These terms should be understood as nomenclature and should not be construed as limiting the number of elements modified by these nomenclatures, unless a specific number has been given. Any advantages and benefits described may not apply to all embodiments of the present invention. It should be understood that those skilled in the art can make changes to the described embodiments without departing from the scope of the present invention as defined by the appended claims. Furthermore, no element or component in this disclosure is intended to be dedicated to the public, whether or not the element or component is expressly recited in the appended claims.

Claims

1. A computer-implemented method for defect analysis, comprising: Obtaining multiple sets of defect point coordinates, where each set of the multiple sets of defect point coordinates includes the coordinates of defect points in each of multiple substrates, and the coordinates of the defect points in each substrate are coordinates in an image coordinate system; Combining the multiple sets of defect point coordinates according to the image coordinate system into a composite set of coordinates to generate a composite image; And Performing clustering analysis to classify the defect points in the composite set in the composite image into multiple point clusters; Respectively determining multiple contours of at least multiple selected point clusters among the multiple point clusters, where each contour among the multiple contours includes multiple edge defect points in each of the multiple selected point clusters; Applying a fitting algorithm to the edge defect points of the multiple selected point clusters to generate multiple occlusion regions corresponding to the multiple selected point clusters respectively; and Respectively generating multiple feature vectors of the multiple occlusion regions.

2. The computer-implemented method according to claim 1, further comprising obtaining multiple selected point clusters from the multiple point clusters; Among them, The number of defect points in each of the multiple selected point clusters is greater than a threshold number.

3. The computer-implemented method according to claim 1, wherein, Generating the multiple feature vectors includes: Generate the Hu geometric moment m of each of the multiple occlusion regions i,j and the center-to-center distance M i,j , where m i,j = ∑ (x,v)∈A x i y j ; Calculate the defect point density ρ, area a, centroid O (O x , O y ), and direction θ of each of the multiple occlusion regions; and Generating each feature vector among the multiple feature vectors of each of the multiple occlusion regions.

4. The computer-implemented method according to claim 3, wherein, Each of the multiple feature vectors is represented as: F = [ρ, a, Ox, Oy, θ, L, W, r] T , Among them N represents the number of defective points in each of the plurality of occluded regions, and a represents the area of each of the plurality of occluded regions; M ij = ∑ (x,y)∈A (x - O x ) i (y - O y ) j ; L represents the length of the minimum bounding rectangle of each of the multiple occlusion regions; And W represents the width of the minimum bounding rectangle of each of the multiple occlusion regions.

5. The computer-implemented method according to any one of claims 1 to 4, wherein, Using a method based on αShapes to determine the multiple contours.

6. The computer-implemented method according to any one of claims 1 to 4, further comprising: Assigning one or more selected occlusion regions among the multiple occlusion regions as multiple defect aggregation regions; Wherein, the feature vectors of the one or more selected occlusion regions satisfy a threshold condition.

7. The computer-implemented method according to claim 6, further comprising: Comparing the parameters of first defect points within the one or more selected occlusion regions with the parameters of second defect points outside the one or more selected occlusion regions; And Identifying a potential device causing the first defect points based on the comparison.

8. The computer-implemented method according to any one of claims 1 to 4, further comprising: Obtaining multiple sets of substrate defect point coordinates, where each set of the multiple sets of substrate defect point coordinates includes the coordinates of substrate defect points in each substrate, and the substrate defect points in each substrate are assigned coordinates in a substrate coordinate system; And Converting the coordinates of the substrate defect points in the substrate coordinate system into the coordinates of the defect points in the image coordinate system.

9. The computer-implemented method according to any one of claims 1 to 4, wherein, Obtaining the multiple sets of defect point coordinates includes: Obtaining multiple sets of original defect point coordinates; and Selecting, from the multiple sets of original defect point coordinates, multiple sets of defect point coordinates including more than a threshold number of defect point coordinates as the multiple sets of defect point coordinates.

10. The computer-implemented method according to any one of claims 1 to 4, wherein, Using a hierarchical clustering method to perform the clustering analysis.

11. The computer-implemented method according to claim 10, wherein, The hierarchical clustering method is a single-linkage clustering method.

12. The computer-implemented method according to any one of claims 1 to 4, wherein, The Euclidean distance between adjacent defect points in each of the multiple point clusters is less than or equal to a threshold, and the Euclidean distance between any two defect points from two of the multiple point clusters is greater than the threshold.

13. A device for defect analysis, comprising: A memory; One or more processors; Wherein, the memory and the one or more processors are connected to each other; and The memory stores computer-executable instructions for controlling the one or more processors to: Obtain multiple sets of defect point coordinates, each set of the multiple sets of defect point coordinates including the coordinates of defect points in each of multiple substrates, and the coordinates of defect points in each substrate being coordinates in an image coordinate system; Combine the multiple sets of defect point coordinates according to the image coordinate system into a composite set of coordinates to generate a composite image; and Perform clustering analysis to classify the defect points in the composite set in the composite image into multiple point clusters; The memory further stores computer-executable instructions for controlling the one or more processors to: Respectively determine multiple contours of at least multiple selected point clusters among the multiple point clusters, each of the multiple contours including multiple edge defect points in each of the multiple selected point clusters; Apply a fitting algorithm to the edge defect points of the multiple selected point clusters to generate multiple occlusion regions corresponding to the multiple selected point clusters respectively; and Respectively generate multiple feature vectors of the multiple occlusion regions.

14. The apparatus according to claim 13, wherein The memory further stores computer-executable instructions for controlling the one or more processors to: Generate the Hu geometric moment m of each of the multiple occlusion regions i,j and the center-to-center distance M i,j , where m i,j = ∑ (x,v)∈A x i y j ; Calculate the defect point density ρ, area a, centroid O (O x , O y ) and direction θ of each occlusion area among the multiple occlusion areas; and Generate each of the multiple feature vectors of each of the multiple occlusion regions.

15. A computer-readable medium having computer-readable instructions thereon, the computer-readable instructions being executable by a processor to cause the processor to perform: Obtain multiple sets of defect point coordinates, each set of the multiple sets of defect point coordinates including the coordinates of defect points in each of multiple substrates, and the coordinates of defect points in each substrate being coordinates in an image coordinate system; Combining the multiple sets of defect point coordinates in the image coordinate system into a composite set of coordinates to generate a composite image; And Perform clustering analysis to classify the defect points in the composite set in the composite image into multiple point clusters; The computer-readable instructions can further be executable by the processor to cause the processor to perform: Respectively determine multiple contours of at least multiple selected point clusters among the multiple point clusters, each of the multiple contours including multiple edge defect points in each of the multiple selected point clusters; Apply a fitting algorithm to the edge defect points of the multiple selected point clusters to generate multiple occlusion regions corresponding to the multiple selected point clusters respectively; and Respectively generate multiple feature vectors of the multiple occlusion regions.

16. The computer-readable medium according to claim 15, wherein, The computer-readable instructions can further be executable by the processor to cause the processor to perform: Generate the Hu geometric moment m of each of the multiple occlusion regions i,j and the center-to-center distance M i,j , where m i,j = ∑ (x,v)∈A x i y j ; Calculate the defect point density ρ, area a, centroid O (O x , O y ) and direction θ of each of the multiple occlusion regions; and Generate each of the multiple feature vectors of each of the multiple occlusion regions.

17. A defect analysis system, comprising: A distributed computing system, which includes one or more networked computers configured to execute in parallel to perform at least one common task; One or more computer-readable storage media storing instructions that, when executed by the distributed computing system, cause the distributed computing system to execute software modules; Wherein, the software modules include: A data manager configured to store data and extract, transform, or load the data; A query engine connected to the data manager and configured to query the data directly from the data manager; An analyzer connected to the query engine and configured to perform defect analysis upon receiving a task request, the analyzer including a plurality of business servers and a plurality of algorithm servers, the plurality of algorithm servers being configured to query the data directly from the data manager; and A data visualization and interaction interface configured to generate the task request; Wherein, one or more of the plurality of algorithm servers are configured to execute the computer-implemented method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Liquid crystal screen defect layered positioning method and device

    CN105842885A

  • Cluster forming device, defect classifying device, cluster forming method and program

    JP2008268232A