Method, device, storage medium and defect analysis system for defect analysis
Through computer-implemented methods, using evidence weight scores and defect point coordinate analysis, we can identify process equipment and contact equipment related to defects in the manufacturing process of display panels, solving the problem of low efficiency in defect analysis in the prior art, and achieving automated and accurate identification of defect causes.
Patent Information
- Application Number
- CN202080003659.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-03
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2040-12-03
AI Technical Summary
The root causes of defects are difficult to efficiently track and analyze in the manufacturing process of display panels, relying on manual data classification and empirical analysis.
Using a computer-implemented method, by calculating the evidence weight (WOE) score, process equipment and contact equipment that are highly related to defects are identified, and a list of defective equipment is generated based on defect point coordinates and contact area analysis.
Automatic defect analysis is realized, the efficiency and accuracy of defect cause identification is improved, the dependence on experience is reduced, and the controllability of the manufacturing process is improved.
Smart Images

Figure CN114916237B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to display technology, and more particularly, to a computer-implemented method for defect analysis, a computer-implemented method for evaluating the likelihood of a defect occurring, an apparatus for defect analysis, a computer storage medium, and a defect analysis system. Background Art
[0002] Distributed computing and distributed algorithms have become common in various contexts due to improved performance and load capacity, high availability and failover, and faster access to data. With the development of technologies such as big data, cloud computing, and artificial intelligence, technologies related to big data analysis have been widely used in various fields of manufacturing. Summary of the Invention
[0003] In one aspect, the present disclosure provides a computer-implemented method for defect analysis, comprising: calculating a plurality of weight of evidence (WOE) scores for a plurality of process tools, respectively, regarding defects occurring during a manufacturing cycle, wherein a higher WOE score indicates a higher correlation between the defect and the process tool; and sorting the plurality of WOE scores to obtain a list of selected process tools that are highly correlated with the defects occurring during the manufacturing cycle, the process tools in the selected list of process tools having a WOE score greater than a first threshold score, wherein each of the plurality of process tools is a respective tool defined by a respective process site at which the respective tool performs a respective operation.
[0004] Optionally, the computer-implemented method further includes: obtaining candidate contact areas of a plurality of contact devices, wherein each of the plurality of contact devices is a respective device that contacts the intermediate substrate in each chamber during the manufacturing cycle, and each candidate contact area includes a respective theoretical contact area and a respective edge area surrounding each theoretical contact area; respectively obtaining the coordinates of defect points during the manufacturing cycle; selecting a plurality of defective contact areas from the candidate contact areas, wherein each of the plurality of defective contact areas surrounds at least one of the defective point coordinates; and obtaining a list of selected contact devices, wherein each selected contact device is a respective contact device that contacts the intermediate product in each defective contact area.
[0005] Optionally, the computer-implemented method further includes obtaining one or more candidate defective devices based on the list of selected process equipment and the list of selected contact equipment that are highly correlated with the defect; wherein each of the one or more candidate defective devices is a device in the list of selected contact equipment and is also a device in the list of selected process equipment.
[0006] On the other hand, the present disclosure provides a device for defect analysis, comprising: a memory; one or more processors; wherein the memory and the one or more processors are connected to each other; and the memory stores computer-executable instructions for controlling the one or more processors to: calculate multiple weight of evidence (WOE) scores for multiple process tools respectively regarding defects occurring during a manufacturing cycle, a higher WOE score indicating a higher correlation between the defect and the process tool; and sort the multiple WOE scores to obtain a selected list of process tools that are highly correlated with the defects occurring during the manufacturing cycle, the process tools in the selected list of process tools having a WOE score greater than a first threshold score, wherein each of the multiple process tools is a corresponding tool defined by a corresponding process site at which the corresponding tool performs a corresponding operation.
[0007] Optionally, the memory further stores computer-executable instructions for controlling the one or more processors to: obtain candidate contact areas of a plurality of contact devices, wherein each of the plurality of contact devices is a respective device that contacts the intermediate product in each chamber during the manufacturing cycle, and each candidate contact area includes a respective theoretical contact area and a respective edge area surrounding each theoretical contact area; respectively obtain the coordinates of defect points during the manufacturing cycle; select a plurality of defective contact areas from the candidate contact areas, wherein each of the plurality of defective contact areas surrounds at least one of the defective point coordinates; and obtain a list of selected contact devices, wherein each selected contact device is a respective contact device that contacts the intermediate product in each defective contact area.
[0008] Optionally, the memory further stores computer-executable instructions for controlling the one or more processors to: obtain one or more candidate defective devices based on the list of selected process equipment and the list of selected contact equipment that are highly correlated with the defect; wherein each of the one or more candidate defective devices is a device in the list of selected contact equipment and is also a device in the list of selected process equipment.
[0009] In another aspect, the present disclosure provides a computer program product comprising a non-transitory tangible computer-readable medium having computer-readable instructions thereon, the computer-readable instructions being executable by a processor to cause the processor to: calculate a plurality of weight of evidence (WOE) scores for a plurality of process tools, respectively, regarding defects occurring during a manufacturing cycle, a higher WOE score indicating a higher correlation between the defect and the process tool; and sort the plurality of WOE scores to obtain a list of selected process tools that are highly correlated with the defects occurring during the manufacturing cycle, the process tools in the selected list of process tools having a WOE score greater than a first threshold score, wherein each of the plurality of process tools is a respective tool defined by a respective process site at which the respective tool performs a respective operation.
[0010] Optionally, computer-readable instructions may be executed by the processor to further cause the processor to perform: obtaining candidate contact areas for a plurality of contact devices, wherein each of the plurality of contact devices is a respective device that contacts the intermediate product in each chamber during the manufacturing cycle, and each candidate contact area includes a respective theoretical contact area and a respective edge area surrounding each theoretical contact area; respectively obtaining the coordinates of defect points during the manufacturing cycle; selecting a plurality of defective contact areas from the candidate contact areas, wherein each of the plurality of defective contact areas surrounds at least one of the defective point coordinates; and obtaining a list of selected contact devices, wherein each selected contact device is a respective contact device that contacts the intermediate product in each defective contact area.
[0011] Optionally, the computer-readable instructions may be executed by the processor to further cause the processor to perform: obtaining one or more candidate defective devices based on the list of selected process equipment and the list of selected contact equipment that are highly correlated with the defect; wherein each of the one or more candidate defective devices is a device in the list of selected contact equipment and is also a device in the list of selected process equipment.
[0012] On the other hand, the present disclosure provides a defect analysis system, comprising: a distributed computing system, which includes one or more networked computers, the networked computers being configured to execute in parallel to perform at least one common task; one or more computer-readable storage media storing instructions, which, when executed by the distributed computing system, cause the distributed computing system to execute a software module; wherein the software module includes: a data management platform, which is configured to extract, convert or load raw data from multiple data sources into management data, wherein the raw data and the management data include defect information, and the management data is stored in a distributed manner; an analyzer, which is configured to perform defect analysis upon receiving a task request, the analyzer including multiple algorithm servers, the multiple algorithm servers being configured to obtain the management data from the data management platform and perform algorithmic analysis on the management data to derive result data regarding potential causes of defects; and a data visualization and interactive interface, which is configured to generate the task request and display the result data; wherein one or more of the multiple algorithm servers are configured to execute the computer-implemented method described herein.
[0013] On the other hand, the present disclosure provides a computer-implemented method for assessing the likelihood of a defect occurring, comprising: obtaining raw data about defects occurring during a manufacturing cycle; preprocessing the raw data to obtain preprocessed data; extracting features from the preprocessed data to obtain extracted features; selecting main features from the extracted features; inputting the main features into a prediction model; and assessing the likelihood of a defect occurring; wherein extracting features from the preprocessed data includes performing at least one of time domain analysis and frequency domain analysis on the preprocessed data.
[0014] Optionally, time domain analysis extracts statistical information of the preprocessed data.
[0015] Optionally, the frequency domain analysis converts the time domain information obtained in the time domain analysis into frequency domain information.
[0016] Optionally, selecting the principal feature from the extracted features includes performing principal component analysis on the extracted features.
[0017] Optionally, preprocessing the original data includes: excluding a first portion of the original data from the preprocessed data, in which the missing values are greater than or equal to a threshold percentage; and for a second portion of the original data, in which the missing values are less than the threshold percentage, providing interpolation of the missing values.
[0018] Optionally, the prediction model is trained by: obtaining training raw data about defects that occur during a training manufacturing cycle; preprocessing the training raw data to obtain preprocessed training data; extracting features from the preprocessed training data to obtain extracted training features; selecting main training features from the extracted training features; and using the extracted training features to adjust parameters of the initial model to obtain the prediction model for defect prediction, wherein extracting training features from the preprocessed training data includes performing at least one of time domain analysis and frequency domain analysis on the preprocessed training data.
[0019] Optionally, the initial model is an Extreme Gradient Boosting (XGboost) model.
[0020] Optionally, adjusting the parameters of the initial model comprises evaluating the F-measure according to equation (1):
[0021]
[0022] Where Fβ represents the harmonic mean of precision and recall, P represents precision, R represents recall, and β represents a parameter that controls the balance between P and R.
[0023] On the other hand, the present disclosure provides a defect analysis system, comprising: a distributed computing system, which includes one or more networked computers, the networked computers being configured to execute in parallel to perform at least one common task; one or more computer-readable storage media, which store instructions, which, when executed by the distributed computing system, cause the distributed computing system to execute a software module; wherein the software module includes: a data manager, which is configured to store data and extract, convert or load the data; a query engine, which is connected to the data manager and configured to query the data directly from the data manager; an analyzer, which is connected to the query engine and configured to perform defect analysis upon receiving a task request, the analyzer including multiple business servers and multiple algorithm servers, the multiple algorithm servers being configured to query the data directly from the data manager; and a data visualization and interactive interface, which is configured to generate the task request; wherein one or more of the multiple algorithm servers are configured to execute the computer-implemented method described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The following drawings are examples for illustration purposes only, in accordance with various disclosed embodiments, and are not intended to limit the scope of the invention.
[0025] Figure 1 A computer-implemented method for defect analysis in accordance with some embodiments of the present disclosure is shown.
[0026] Figure 2 A computer-implemented method for defect analysis in accordance with some embodiments of the present disclosure is shown.
[0027] Figure 3 A computer-implemented method for defect analysis in accordance with some embodiments of the present disclosure is shown.
[0028] Figure 4A and Figure 4B The conversion process from the individual panel coordinate systems to the unified substrate coordinate system is shown.
[0029] Figure 5 Several candidate contact areas according to some embodiments of the present disclosure are shown.
[0030] Figure 6A is a schematic diagram of the structure of a device according to some embodiments of the present disclosure.
[0031] Figure 6B is a schematic diagram showing the structure of a device in some embodiments of the present disclosure.
[0032] Figure 7 A computer-implemented method for evaluating the likelihood of a defect occurring in accordance with some embodiments of the present disclosure is shown.
[0033] Figure 8 Parameter values of a parameter name "MASK_TEMP" from several glasses in one example of the present disclosure are shown.
[0034] Figure 9 A distributed computing environment according to some embodiments of the present disclosure is shown.
[0035] Figure 10 The software modules in the defect analysis system according to some embodiments of the present disclosure are shown.
[0036] Figure 11 The software modules in the defect analysis system according to some embodiments of the present disclosure are shown.
[0037] Figure 12 A defect analysis method using a defect analysis system according to some embodiments of the present disclosure is shown.
[0038] Figure 13 A defect analysis method using a defect analysis system according to some embodiments of the present disclosure is shown.
[0039] Figure 14 A defect analysis method using a defect analysis system according to some embodiments of the present disclosure is shown.
[0040] Figure 15A defect analysis method using a defect analysis system according to some embodiments of the present disclosure is shown.
[0041] Figure 16 A data management platform according to some embodiments of the present disclosure is shown.
[0042] Figure 17 Depicting multiple sub-tables divided from a data table stored in a general data layer according to some embodiments of the present disclosure.
[0043] Figure 18 A defect analysis method according to some embodiments of the present disclosure is shown.
[0044] Figure 19 A defect analysis method according to some embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0045] The present disclosure will now be described in more detail with reference to the following examples. It should be noted that the following description of some of the embodiments presented herein is for illustration and description purposes only. It is not intended to be exhaustive or limited to the precise forms disclosed.
[0046] The manufacturing of display panels, particularly organic light-emitting diode (OLED) display panels, involves a highly complex and integrated process involving numerous processes, technologies, and equipment. Defects arising from this integrated process are difficult to track. For example, engineers may have to rely on manual data classification to empirically analyze the root cause of defects.
[0047] Therefore, the present disclosure provides, among other things, a computer-implemented method for defect analysis, a computer-implemented method for assessing the likelihood of a defect occurring, an apparatus for defect analysis, a computer storage medium, and a defect analysis system, which substantially eliminate one or more problems due to limitations and disadvantages of the related art. In one aspect, the present disclosure provides a computer-implemented method for defect analysis. In some embodiments, the computer-implemented method for defect analysis includes: calculating, for defects occurring during a manufacturing cycle, a plurality of process equipment, respectively, a plurality of weight of evidence (WOE) scores, wherein a higher WOE score indicates a higher correlation between the defect and the process equipment; and sorting the plurality of WOE scores to obtain a list of selected process equipment that is highly correlated with the defect occurring during the manufacturing cycle, the process equipment in the selected list of process equipment having a WOE score greater than a first threshold score. Optionally, each of the plurality of process equipment is a respective equipment defined by a respective process site at which the respective equipment performs a respective operation.
[0048] Figure 1 A computer-implemented method for defect analysis according to some embodiments of the present disclosure is shown. Figure 1 , in some embodiments, the method includes collecting business data into multiple data sources DS. Examples of data sources include various business systems, including, for example, a yield management system (YMS), a fault detection and classification (FDC) system, a manufacturing data warehouse system (MDW), and a manufacturing execution system (MES). In one example, the method includes obtaining panel defect information during a manufacturing cycle (e.g., a repetitive cycle for automatic detection, or a cycle with a high frequency of defect occurrence). Defect information is detected in a panel cut from a substrate (also referred to as "glass" in this disclosure). The panel and the substrate (or glass) can use different coordinate systems. The panel defect information includes, for example, panel ID information associated with the coordinates of the defect point. Table 1 shows an example of panel ID information obtained from a yield management system.
[0049] Table 1 shows an example of panel ID information obtained from the yield management system.
[0050] Panel ID Y8CP43925109L Y91PA0115109L Y91PA0119109L Y91PA2622109L Y91PA2626109L Y91PA2627109L Y91PB2316109L Y91PB2322109L Y91PB2330109L …
[0051] In another example, the method includes obtaining glass biographical information during the manufacturing cycle. The biographical information is recorded in the glass (or "substrate") that is cut into individual panels. Defects are detected in the panels rather than in the glass. Table 2 shows an example of glass biographical information obtained from a manufacturing data warehouse system.
[0052] Table 2 shows an example of glass history information obtained from the manufacturing data warehouse system.
[0053]
[0054]
[0055] In another example, the method includes obtaining equipment contact information during the manufacturing cycle. The equipment contact information is recorded in the glass (or "substrate") that is cut into individual panels. Table 3 shows an example of equipment contact information obtained from a yield management system.
[0056] Table 3 shows an example of glass history information obtained from the yield management system.
[0057]
[0058]
[0059]
[0060] In some embodiments, the method further includes preprocessing raw data from a plurality of data sources DS. As described above, defect points are detected in panels cut from glass. However, the yield management system only contains panel-level defect information and lacks connection with glass history information. In one example, the panel ID in Table 1 can be preprocessed to obtain the glass ID. For example, the first nine digits in the panel ID represent the glass ID. As shown in Table 4, the panel ID information can be preprocessed to obtain defect point information associated with the glass ID. Therefore, as Figure 1 As shown, the method includes obtaining information about the number of panels having defects classified by glass ID.
[0061] Table 4. Panels with defects and panels without defects classified by glass ID.
[0062]
[0063] The information in Table 2 includes various information, including factories, process sites, and equipment. Table 2 inherently contains information redundancy. Therefore, in some embodiments, the method further includes performing data fusion on equipment and process sites to form fused data information, such as "Process Equipment (Operation Equipment)," which represents the corresponding equipment defined by the corresponding process site and performs the corresponding operation at that process site. The information is then sorted by timestamp and filtered by event ("TrackOut" in Table 2) to obtain process equipment information sorted by glass ID, as shown in Table 5.
[0064] Table 5. Process equipment information classified by glass ID.
[0065] Glass ID Timestamp event factory process equipment Y91PA2626 01 / 11 / 201917:35:16 TrackOut A2 10100_A2P01400 Y91PA2626 01 / 11 / 201918:08:36 TrackOut A2 10103_A2A01A00 Y91PA2626 01 / 12 / 201903:00:50 TrackOut A2 10300_A2P02200 Y91PA2626 01 / 12 / 201904:37:19 TrackOut A2 10303_A2A01600 … … … … …
[0066] As described above, the defect point is detected in the panel coordinate system. On the other hand, the device contact information is recorded in the glass coordinate system. In some embodiments, the method further includes converting the coordinates of the panel defect point in the panel coordinate system into coordinates in the glass coordinate system.
[0067] Figure 2 A computer-implemented method for defect analysis according to some embodiments of the present disclosure is shown. Figure 2The method further includes, after pre-processing the raw data, obtaining defect data associated with a plurality of process tools during a manufacturing cycle, calculating a plurality of weight of evidence (WOE) scores for each of the plurality of process tools, wherein a higher WOE score indicates a higher correlation between the defect and the process tool; and sorting the plurality of WOE scores to obtain a list of selected process tools that are highly correlated with defects occurring during the manufacturing cycle, wherein the process tools in the list of selected process tools have a WOE score greater than a first threshold score. Optionally, each of the plurality of process tools is a tool for performing a corresponding process at a specific process site.
[0068] In one example, Table 4 and Table 5 can be merged based on the glass ID. Specifically, the "number of panels with defects" and "number of panels without defects" information can be reclassified using the merged "process equipment" information. Table 6 shows an example of using the "number of panels with defects" and "number of panels without defects" information by "process equipment."
[0069] Table 6 shows an example of “the number of panels having defects” information and “the number of panels without defects” information using “process equipment”.
[0070]
[0071]
[0072] As shown in Table 6, most of the samples are panels without defects, so a correlation analysis must be performed on the data obtained in Table 6. In one example, the correlation analysis is performed using the weight of evidence method. In another example, the weight of evidence score for each "process equipment" can be calculated according to the following equation:
[0073]
[0074] Among them, i represents the corresponding process equipment, y i represents the number of panels (intermediate products) with defects on the corresponding process equipment, n i represents the number of panels without defects on the corresponding process equipment, y T represents the total number of panels with defects in the data, and n T Then, as shown in Table 7, the process equipment are ranked according to their corresponding WOE scores.
[0075] Table 7 shows a list of process equipment sorted by WOE scores.
[0076] Sorting process equipment WOE 1 19103_A2AOI500 0.76986168 2 162R3_A2AOI100 0.635129086 3 14202_A2PDC300 0.635129086 4 16315_A2CDC300 0.583415711 5 11313_A2AOI600 0.410306273 6 16305_A2CDC100 0.292382166 7 1A3R3_A2AOI600 0.275936306 8 163R3_A2AOI500 0.249466605 9 1A1R3_A2AOI200 0.226866773 10 11305_A2CDC200 0.226866773
[0077] Figure 3 A computer-implemented method for defect analysis according to some embodiments of the present disclosure is shown. Figure 3 In some embodiments, the method further includes obtaining candidate contact areas of a plurality of contact devices, wherein each of the plurality of contact devices is a device that contacts the intermediate substrate in each chamber during the manufacturing cycle, and each candidate contact area includes a theoretical contact area and an edge area surrounding each theoretical contact area; respectively obtaining the coordinates of defect points during the manufacturing cycle; selecting a plurality of defective contact areas from the candidate contact areas, wherein each of the plurality of defective contact areas surrounds at least one of the defective point coordinates; and obtaining a list of selected contact devices, wherein each selected contact device is a contact device that has contact with the intermediate substrate in each defective contact area. In some embodiments, the candidate contact areas include point contacts and line contacts. In some embodiments, each of the plurality of contact devices is in contact with the intermediate substrate in the chamber during the manufacturing cycle. In some embodiments, the intermediate substrate includes glass, half glass, and a panel. As described herein, the "intermediate" of the intermediate substrate refers to an intermediate process.
[0078] Reference Figure 1 In some embodiments, the method includes converting the coordinates in the panel defect coordinate system into the coordinates in the glass defect coordinate system after respectively obtaining the coordinates of the defect points during the manufacturing cycle. Figure 4A and Figure 4B The conversion process from the individual panel coordinate system to the unified substrate coordinate system is shown. Figure 4A , the panel defect points in the plurality of panel defect images (A to N) have panel coordinates assigned according to respective panel coordinate systems (pcs1 to pcs14). For example, the panel defect points in the panel defect image A are assigned panel coordinates according to the first panel coordinate system pcs1, the panel defect points in the panel defect image B are assigned panel coordinates according to the second panel coordinate system pcs2, the panel defect points in the panel defect image C are assigned panel coordinates according to the third panel coordinate system pcs3, and so on. Because the panel is cut from the substrate during the manufacturing process, the coordinates in the respective panel coordinate systems can be converted into a unified substrate coordinate system. In one example, the defect analysis is for the same type of product (e.g., an organic light emitting diode display panel), so the plurality of substrates used to manufacture the panel have the same substrate coordinate system. With reference to Figure 4B , convert the coordinates of the panel defect points in each panel coordinate system into the coordinates of the defect points in the substrate coordinate system scs.
[0079] In one example, the computer-implemented method further includes establishing a mapping relationship between coordinates in the panel coordinate system and coordinates in the substrate coordinate system. In another example, the mapping relationship can be expressed as:
[0080]
[0081] in, which is a rotation matrix; and It is the translation matrix.
[0082] Alternatively, the coordinate transformation can be expressed as:
[0083] Among them, X i and Y i represents the coordinate in the substrate coordinate system; x i and y i Represents coordinates in the panel coordinate system.
[0084] During the manufacturing cycle, various devices have a variety of different contact types with the substrate. Examples of contact types include point contacts, line contacts, circular area contacts, and rectangular area contacts.
[0085] In the actual manufacturing process, (one or more) actual contact positions may deviate from (one or more) theoretical contact positions to a certain extent. Therefore, in some embodiments, the candidate contact areas of the contact device should be expanded accordingly to take into account the deviations in the actual manufacturing process. In some embodiments, each candidate contact area includes each theoretical contact area and each edge area surrounding each theoretical contact area. Figure 5 Several candidate contact areas according to some embodiments of the present disclosure are shown.
[0086] Reference Figure 1 Once the candidate contact regions are calculated, the coordinates of the defect point in the glass coordinate system are matched to the candidate contact regions. Thus, multiple defective contact regions can be selected from the candidate contact regions, each of which encompasses at least one of the defective point coordinates. A list of selected contact devices can be obtained, where each selected contact device is in contact with the intermediate substrate in each defective contact region. Table 8 shows an example of the list of selected contact devices.
[0087] Table 8 shows an example of a list of selected contact devices having contact with the intermediate substrate in the defective contact area.
[0088]
[0089]
[0090] refer to Figure 1 In some embodiments, the method further includes finding common (one or more) devices from the WOE ranked list and the selected contact device list. In some embodiments, the method includes obtaining one or more candidate defective devices based on the selected process equipment list and the selected contact device list that are highly correlated with the defect. Each candidate defective device in the one or more candidate defective devices is a device in the selected contact device list and is also a device in the selected process equipment list.
[0091] On the other hand, the present disclosure provides an apparatus for defect analysis. In some embodiments, the apparatus for defect analysis includes: a memory; and one or more processors. The memory and the one or more processors are connected to each other. In some embodiments, the memory stores computer-executable instructions for controlling the one or more processors to: calculate multiple weight of evidence (WOE) scores for multiple process equipment respectively regarding defects occurring during a manufacturing cycle, a higher WOE score indicates a higher correlation between the defect and the process equipment; and sort the multiple WOE scores to obtain a selected list of process equipment that is highly correlated with the defects occurring during the manufacturing cycle, the process equipment in the selected list of process equipment having a WOE score greater than a first threshold score. Each of the multiple process equipment is a corresponding equipment defined by a corresponding process site at which the corresponding equipment performs a corresponding operation.
[0092] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to: obtain candidate contact areas of a plurality of contact devices, wherein each of the plurality of contact devices is a respective device that contacts the intermediate substrate in each chamber during the manufacturing cycle, and each candidate contact area includes a respective theoretical contact area and a respective edge area surrounding each theoretical contact area; respectively obtain coordinates of defect points during the manufacturing cycle; select a plurality of defective contact areas from the candidate contact areas, wherein each of the plurality of defective contact areas surrounds at least one of the defective point coordinates; and obtain a list of selected contact devices, wherein each selected contact device is a respective contact device that contacts the intermediate substrate in each defective contact area.
[0093] In some embodiments, the memory further stores computer-executable instructions for controlling the one or more processors to obtain one or more candidate defective devices based on the list of selected process tools and the list of selected contact tools that are highly correlated with the defect, wherein each candidate defective device in the one or more candidate defective devices is a device in the list of selected contact tools and a device in the list of selected process tools.
[0094] Figure 6A is a schematic diagram of the structure of the device in some embodiments of the present disclosure. Figure 6A In some embodiments, the device includes a central processing unit (CPU) configured to perform actions according to computer-executable instructions stored in ROM or RAM. Optionally, data and programs required by the computer system are stored in RAM. Optionally, the CPU, ROM, and RAM are electrically connected to each other via a bus. Optionally, an input / output interface is electrically connected to the bus.
[0095] Figure 6B Schematic diagram showing the structure of the device in some embodiments of the present disclosure. Figure 6B In some embodiments, the device includes a display panel DP; an integrated circuit IC connected to the display panel DP; a memory M; and one or more processors P. The memory M and the one or more processors P are connected to each other. In some embodiments, the memory M stores computer-executable instructions for controlling the one or more processors P to perform the method steps described herein.
[0096] In another aspect, the present disclosure provides a computer program product comprising a non-transitory tangible computer-readable medium having computer-readable instructions thereon. In some embodiments, the computer-readable instructions are executable by a processor to cause the processor to perform: calculating a plurality of weights of evidence (WOE) scores for a plurality of process tools, respectively, with respect to defects occurring during a manufacturing cycle, wherein a higher WOE score indicates a higher correlation between the defect and the process tool; and sorting the plurality of WOE scores to obtain a list of selected process tools that are highly correlated with the defect occurring during the manufacturing cycle, wherein the process tools in the selected list of process tools have a WOE score greater than a first threshold score. Each of the plurality of process tools is a respective tool defined by a respective process site at which the respective tool performs a respective operation. Optionally, the weight of evidence represents a variability between a percentage of defects in the respective process tools relative to a percentage of defects in the plurality of process tools as a whole.
[0097] In some embodiments, the computer-readable instructions can also be executed by the processor to cause the processor to perform: obtaining candidate contact areas of multiple contact devices, wherein each of the multiple contact devices is a respective device that contacts the intermediate substrate in each chamber during the manufacturing cycle, and each candidate contact area includes a respective theoretical contact area and a respective edge area surrounding each theoretical contact area; respectively obtaining the coordinates of the defect points during the manufacturing cycle; selecting multiple defective contact areas from the candidate contact areas, wherein each of the multiple defective contact areas surrounds at least one of the defective point coordinates; and obtaining a list of selected contact devices, wherein each selected contact device is a respective contact device that contacts the intermediate substrate in each defective contact area.
[0098] In some embodiments, the computer-readable instructions are further executable by the processor to cause the processor to: obtain one or more candidate defective devices based on the list of selected process equipment and the list of selected contact equipment that are highly correlated with the defect, wherein each candidate defective device in the one or more candidate defective devices is a device in the list of selected contact equipment and is also a device in the list of selected process equipment.
[0099] In another aspect, the present disclosure also provides a computer-implemented method for assessing the likelihood of a defect occurring. In some embodiments, the computer-implemented method for assessing the likelihood of a defect occurring includes: obtaining raw data regarding defects occurring during a manufacturing cycle; preprocessing the raw data to obtain preprocessed data; extracting features from the preprocessed data to obtain extracted features; selecting primary features from the extracted features; using the extracted features to adjust parameters of an initial model to obtain a prediction model for defect assessment; and inputting the primary features into the prediction model to obtain a prediction result for assessing the likelihood of a defect occurring. Optionally, extracting features from the preprocessed data includes performing at least one of time domain analysis and frequency domain analysis on the preprocessed data.
[0100] In one example, the method includes obtaining data stored in multiple data sources, such as a fault detection and classification system. The data is then filtered by, for example, "manufacturing process," "product type," and "product model." The raw data is preprocessed, for example, by line conversion and field concatenation, to generate a table containing parameters during the manufacturing cycle. An example is shown in Table 9.
[0101] Table 9. Panel production parameters.
[0102] glass_id equipment machine chamber Parameter name Parameter value A9B0014A1 ABC01 BCBL BL01 L1_Energy [1017.00,1017.00,1017.00,...,1017.00] A9B0014A1 ABC01 BCBL BL01 L1_Water_Temp [14.90,14.90,14.90,...,14.90] A9B0014A1 ABC01 BCBL BL01 L1_Water_Press [4.90,4.90,4.90,...,4.90] A9B0022A5 ABC02 BCBL BL01 L2_Energy [995.00,995.00,994.00,...,995.00] A9B0022A5 ABC02 BCBL BL01 L2_Water_Temp [14.90,14.90,14.90,...,14.90] A9B0022A5 ABC02 BCBL BL01 L2_Water_Press [4.90,4.90,4.90,...,4.90] … … … … … …
[0103] In another example, the method further includes obtaining defect point information from a yield management system. For example, the data can be classified according to the type of defect. Table 10 shows an example of a classified table.
[0104] Table 10 shows an example of defect information obtained from a data source.
[0105] glass_id Defect Type A9B0014A1 ACTE A9B0014A2 ACTE A9B0014A3 ACTE A9B0014A4 ACTE A9B0014A5 ACTE
[0106] Figure 7 A computer-implemented method for evaluating the likelihood of a defect occurring according to some embodiments of the present disclosure is shown. Figure 7 After obtaining the data, the method further includes at least one of a preprocessing step, a feature extraction step, and a feature selection step.
[0107] In one example, the preprocessing step includes character data preprocessing to convert the character data into a format that can be processed by the various algorithms of the present method. For example, in Table 9, device "ABC01" and device "ABC02" are devices that execute the same process in parallel based on business knowledge. Device "ABC01" and device "ABC02" are associated with the same set of parameter names. Therefore, a code is required for each device (e.g., 0 for ABC01 and 1 for ABC02).
[0108] In another example, machine "BCBL" and machine "HCLN" are machines that perform continuous processes based on business knowledge. Therefore, the parameter name associated with machine "BCBL" is completely different from the parameter name associated with machine "HCLN". Therefore, the "machine" character is not required and can be removed before the subsequent analysis step. Similarly, the "chamber" character is also not required and can be removed before the subsequent analysis step. Basically, based on the "glass ID (glass_id)", "equipment" and "parameter name", a unique value can be determined for each parameter name. Table 11 shows an example of preprocessed data.
[0109] Table 11 shows an example of the preprocessed data.
[0110] glass_id equipment Parameter name Parameter value A9B0014A1 0 L1_Energy [1017.00,1017.00,1017.00,...,1017.00] A9B0014A1 0 L1_Water_Temp [14.90,14.90,14.90,...,14.90] A9B0014A1 0 L1_Water_Press [4.90,4.90,4.90,...,4.90] A9B0022A5 1 L2_Energy [995.00,995.00,994.00,...,995.00] A9B0022A5 1 L2_Water_Temp [14.90,14.90,14.90,...,14.90] A9B0022A5 1 L2_Water_Press [4.90,4.90,4.90,...,4.90] … … … …
[0111] Before feature extraction, the data can be further processed. In some embodiments, the preprocessing step further includes excluding a first portion of the original data from the preprocessed data, in which a value greater than or equal to a threshold percentage is missing; and for a second portion of the original data, in which a value less than a threshold percentage is missing, providing interpolation of the missing values. For example, when the sample data has missing values greater than or equal to 30% (e.g., greater than or equal to 35%, greater than or equal to 40%, or greater than or equal to 45%), it can be considered invalid data and can be excluded from feature extraction and subsequent analysis steps. In another example, the sample data has missing values less than 30% (e.g., less than 35%, less than 40%, or less than 45%) and is therefore retained for subsequent feature extraction and other processes. Missing values can be provided using interpolation calculated, for example, by an average interpolation method.
[0112] Figure 8 Parameter values of parameter name "MASK_TEMP" from several glasses in one example of the present disclosure are shown. Figure 8 , the parameter values recorded for each piece of glass include data slices of different lengths in the time series. When different glasses are processed in the same equipment or chamber (that is, the glasses undergo the same processing process), the number of data points recorded by the sensor is inconsistent, and the actual time when the parameter changes may also differ from the theoretical time. Data with such characteristics are not suitable for direct use as input to machine learning algorithms. Simply truncating or padded the data into slices of equal length does not help either.
[0113] Therefore, reference Figure 7 , the method includes a feature extraction step of extracting features from the preprocessed data. In some embodiments, the feature extraction step includes at least one of a time domain analysis and a frequency domain analysis of the preprocessed data. Optionally, extracting features from the preprocessed data includes performing a time domain analysis on the preprocessed data. Optionally, the time domain analysis extracts statistical information of the preprocessed data, including one or more of count, mean, maximum, minimum, range, variance, deviation, kurtosis and percentile. Optionally, extracting features from the preprocessed data also includes performing a frequency domain analysis on the preprocessed data. Optionally, the frequency domain analysis converts the time domain information obtained in the time domain analysis into frequency domain information, and the frequency domain information includes one or more of a power spectrum, information entropy and a signal-to-noise ratio. Table 12 shows an example of feature extraction results.
[0114] Table 12 shows an example of feature extraction results.
[0115] glass_id L1_Water_Temp_Max L1_Water_Temp_Min … L2_Water_Temp_Mean … A9B0014A1 14.9 14.9 … 14.9 … A9B0014A2 14.9 14.9 … 14.9 … A9B0014A3 14.9 14.8 … 14.8 … A9B0014A4 14.9 14.7 … 14.7 … A9B0014A5 14.9 14.7 … 14.7 … A9B0014B1 14.9 14.9 … 14.9 … A9B0014B2 14.9 14.9 … 14.9 … A9B0014B3 14.9 14.9 … 14.9 … A9B0014B4 14.9 14.8 … 14.8 … A9B0014B5 14.8 14.7 … 14.8 …
[0116] The manufacturing process of display panels is very complex. In one example, at a single process site, the total number of non-repeated parameter names is 491. Therefore, each glass at the process site generates 491 different parameters. After the feature extraction step, the data dimension increases to thousands. Based on business knowledge, equipment in the same space will have more than one sensor of the same type to perform detection simultaneously, resulting in a large number of common linear features and other redundancy problems. Therefore, reference Figure 7 The method further includes a feature selection step to reduce the data dimension. In one example, a principal component analysis method is used to select the main features from the extracted features.
[0117] In some embodiments, a predictive model is trained using a computer-implemented training method. In some embodiments, the computer-implemented training method includes obtaining raw training data regarding defects that occurred during a training manufacturing cycle; preprocessing the raw training data to obtain preprocessed training data; extracting features from the preprocessed training data to obtain extracted training features; selecting primary training features from the extracted training features; and using the extracted training features to adjust parameters of an initial model to obtain a predictive model for defect prediction. Optionally, extracting training features from the preprocessed training data includes performing at least one of time domain analysis and frequency domain analysis on the preprocessed training data, details of which are discussed in the context of methods for assessing the likelihood of defect occurrence.
[0118] Various appropriate models can be used as the initial model for training. Examples of appropriate statistical models include extreme gradient boosting (eXtreme Gradient Boosting, XGBoost) model; random forest (RF) model; gradient boosted machine (GBM) model; generalized linear model (GLM) model. In some embodiments, the plurality of statistical models include one or more of the following: univariate K-fold cross-correlation XGboost model; multivariate XGboost model; univariate K-fold cross-correlation random forest model; multivariate random forest model; univariate K-fold cross-correlation GBM; multivariate GBM model; univariate K-fold cross-correlation GLM model; and univariate non-K-fold cross-correlation GLM model. In some embodiments, the plurality of statistical models include one or more of the following: multivariate, bivariate or univariate K-fold cross-correlation XGboost model; multivariate, bivariate or univariate K-fold cross-correlation random forest model; multivariate, bivariate or univariate K-fold cross-correlation GBM model; multivariate, bivariate or univariate K-fold cross-correlation GLM model; and univariate non-K-fold cross-correlation GLM model. In some embodiments, the multiple statistical models include at least: a univariate K-fold cross-correlation XGboost model; a multivariate XGboost model; a univariate K-fold cross-correlation random forest model; a multivariate random forest model; a univariate K-fold cross-correlation GBM model; a multivariate GBM model; a univariate K-fold cross-correlation GLM model; and a univariate non-K-fold cross-correlation GLM model.
[0119] Optionally, the initial model is an Extreme Gradient Boosting (XGBoost) model.
[0120] In some embodiments, adjusting the parameters of the initial model includes evaluating the F-measure according to equation (1):
[0121]
[0122] Where Fβ represents the harmonic mean of precision and recall, P represents precision, R represents recall, and β represents a parameter that controls the balance between P and R. Optionally, β is in the range of 0.80 to 0.90, for example, 0.85, to achieve higher precision.
[0123] In some embodiments, a grid search method is used to adjust the hyperparameters of a model (eg, an XGboost model). Table 13 shows an example of the adjustment results.
[0124] Table 13 shows an example of adjustment results using the grid search method.
[0125] Hyperparameters value reg_alpha 0.145 Subsample 0.835 max_depth 5 Learning_rate 0.458 n_estimators 75
[0126] In another aspect, the present disclosure also provides a predictive model trained by a computer-implemented method as described herein. Once parameters for one or more steps in manufacturing the glass of the present invention are known, the predictive model can be used to predict the occurrence of defects in the glass of the present invention. The predictive model can be used to determine whether the glass of the present invention is defective and should be discarded.
[0127] In one aspect, the present disclosure provides a defect analysis system. In some embodiments, the defect analysis system includes a distributed computing system comprising one or more networked computers configured to execute in parallel to perform at least one common task; one or more computer-readable storage media storing instructions that, when executed by the distributed computing system, cause the distributed computing system to execute a software module. In some embodiments, the software module includes: a data manager configured to store data and extract, transform, or load the data; a query engine connected to the data manager and configured to query the data directly from the data manager; an analyzer connected to the query engine and configured to perform defect analysis upon receiving a task request, the analyzer comprising a plurality of business servers and a plurality of algorithm servers, the plurality of algorithm servers configured to query the data directly from the data manager; and a data visualization and interactive interface configured to generate the task request. Optionally, the defect analysis system is used for defect analysis in display panel manufacturing. As used herein, the term "distributed computing system" generally refers to an interconnected computer network having a plurality of network nodes that connect a plurality of servers or hosts to each other or to an external network (e.g., the Internet). The term "network node" generally refers to a physical network device. Example network nodes include routers, switches, hubs, bridges, load balancers, security gateways, or firewalls. A "host" generally refers to a physical computing device configured to implement, for example, one or more virtual machines or other suitable virtualized components. For example, a host may include a server having a hypervisor configured to support one or more virtual machines or other suitable types of virtualized components.
[0128] In some embodiments, one or more of the plurality of algorithm servers is configured to perform the computer-implemented method for defect analysis described herein.
[0129] In some embodiments, one or more of the plurality of algorithm servers is configured to perform the computer-implemented method of assessing the likelihood of a defect occurring as described herein.
[0130] Various defects can occur during the manufacture of semiconductor electronic devices. Examples include particles, residues, line defects, voids, spatter, wrinkles, discoloration, and bubbles. Defects that occur during semiconductor electronic device manufacturing are difficult to track. For example, engineers may have to rely on manual data classification to empirically analyze the root cause of defects.
[0131] When manufacturing a liquid crystal display panel, the manufacturing of the display panel includes at least the array stage, the color film (CF) stage, the cell stage and the module stage. In the array stage, a thin film transistor array substrate is manufactured. In one example, in the array stage, a material layer is deposited and the material layer is subjected to photolithography, for example, a photoresist is deposited on the material layer, the photoresist is exposed and then developed. Subsequently, the material layer is etched and the remaining photoresist is removed ("stripping"). In the CF stage, a color filter substrate is manufactured, which involves the following steps, including: coating, exposure and development. In the cell stage, the array substrate and the color filter substrate are assembled to form a cell. The cell stage includes several steps, including coating and rubbing the alignment layer, injecting liquid crystal material, cell sealant coating, cell alignment under vacuum, cutting, grinding and cell inspection. In the module stage, peripheral components and circuits are assembled to the panel. In one example, the module level includes several steps, including assembly of the backlight, assembly of the printed circuit board, attachment of the polarizer, assembly of the chip on the film, assembly of the integrated circuit, aging and final inspection.
[0132] When manufacturing an organic light emitting diode (OLED) display panel, the manufacture of the display panel includes at least four equipment processes, including an array stage, an OLED stage, an EAC2 stage, and a module stage. In the array stage, the backplane of the display panel is manufactured, for example, including the manufacture of multiple thin film transistors. In the OLED stage, multiple light emitting elements (for example, organic light emitting diodes) are manufactured, an encapsulation layer is formed to encapsulate the multiple light emitting elements, and optionally, a protective film is formed on the encapsulation layer. In the EAC2 stage, large glass (glass) is first cut into half glass (hglass), and then further cut into panels (panel). In addition, in the EAC2 stage, inspection equipment is used to inspect the panel to detect defects therein, such as dark spots and bright lines. In the module stage, for example, a flexible printed circuit is bonded to the panel using chip-on-film technology. Cover glass is formed on the surface of the panel. Optionally, further inspections are performed to detect defects in the panel. Data from display panel manufacturing includes biographical information, parameter information, and defect information, which are stored in multiple data sources. Historical information is recorded and uploaded to the database by each processing device from the array stage to the module stage. It includes information such as glass ID, device model, and site information. Parameter information includes data generated by the device during glass processing. Defects may occur at each stage. Inspection information is generated at each of the stages discussed above. Only after the inspection is completed is this information uploaded to the database in real time. Inspection information can include defect type and location.
[0133] In summary, various sensors and inspection equipment are used to obtain historical information, parameter information, and defect information. A defect analysis method or system is used to analyze this historical information, parameter information, and defect information. This defect analysis method or system can quickly identify the equipment, site, and / or stage where the defect occurred, providing key information for subsequent process improvements and equipment repair or maintenance, thereby significantly improving yield.
[0134] Therefore, the present disclosure provides, among other things, a data management platform, a defect analysis system, a defect analysis method, a computer storage medium, and a method for defect analysis thereof, which substantially eliminate one or more problems caused by the limitations and shortcomings of the prior art. The present disclosure provides an improved data management platform with superior functionality. Based on the present data management platform (or other appropriate databases or data management platforms), the inventors of the present disclosure have further developed a novel and unique defect analysis system, defect analysis method, computer storage medium, and method for defect analysis.
[0135] In one aspect, the present disclosure provides a defect analysis system. In some embodiments, the defect analysis system includes a distributed computing system comprising one or more networked computers configured to execute in parallel to perform at least one common task; one or more computer-readable storage media storing instructions that, when executed by the distributed computing system, cause the distributed computing system to execute a software module. In some embodiments, the software module includes: a data management platform configured to store data and extract, convert, or load data, wherein the data includes at least one of historical data information, parameter information, or defect information; an analyzer configured to perform defect analysis upon receiving a task request, the analyzer comprising multiple business servers and multiple algorithm servers, the multiple algorithm servers configured to obtain data directly from the data management platform and perform algorithmic analysis on the data to derive result data regarding the potential cause of the defect; and a data visualization and interactive interface configured to generate the task request. Optionally, the defect analysis system is used for defect analysis in display panel manufacturing. As used herein, the term "distributed computing system" generally refers to an interconnected computer network having multiple network nodes that connect multiple servers or hosts to each other or to an external network (e.g., the Internet). The term "network node" generally refers to a physical network device. Example network nodes include routers, switches, hubs, bridges, load balancers, security gateways, or firewalls. A "host" generally refers to a physical computing device configured to implement, for example, one or more virtual machines or other suitable virtualized components. For example, a host may include a server having a hypervisor configured to support one or more virtual machines or other suitable types of virtualized components.
[0136] Figure 9 FIG2 shows a distributed computing environment according to some embodiments of the present disclosure. Figure 9 In a distributed computing environment, multiple autonomous computers / workstations, called nodes, communicate with each other in a network such as a LAN (local area network) to solve tasks, such as executing applications. Each computer node typically includes its own (one or more) processors, memory, and communication links to other nodes. Computers can be located in a specific location (e.g., a cluster network) or can be connected via a wide area network (LAN) such as the Internet. In such a distributed computing environment, different applications can share information and resources.
[0137] Networks in a distributed computing environment can include local area networks (LANs) and wide area networks (WANs). Networks can include wired technologies (e.g., Ethernet) and wireless technologies (e.g., Code Division Multiple Access (CDMA), Global System for Mobile (GSM), Universal Mobile Telephone Service (UMTS), Bluetooth, wait).
[0138] Multiple computing nodes are configured to join a resource group to provide distributed services. A computing node in a distributed network may include any computing device, such as a computing device or a user device. A computing node may also include a data center. As used herein, a computing node may refer to any computing device or multiple computing devices (i.e., a data center). A software module may be executed on a single computing node (e.g., a server) or distributed across multiple nodes in any suitable manner.
[0139] The distributed computing environment may also include one or more storage nodes for storing information related to the execution of the software modules and / or outputs and / or other functions generated by the execution of the software modules. The one or more storage nodes communicate with each other in the network and with one or more computing nodes in the network.
[0140] Figure 10 FIG2 shows the software modules in the defect analysis system according to some embodiments of the present disclosure. Figure 10 , the defect analysis system includes a distributed computing system, which includes one or more networked computers, which are configured to execute in parallel to perform at least one common task; one or more computer-readable storage media storing instructions, which, when executed by the distributed computing system, cause the distributed computing system to execute a software module. In some embodiments, the software module includes a data management platform DM, which is configured to store data and extract, convert or load data; a query engine QE, which is connected to the data management platform DM and configured to obtain data directly from the data management platform DM; an analyzer AZ, which is connected to the query engine QE and configured to perform defect analysis upon receiving a task request, the analyzer AZ includes multiple business servers BS (similar to back-end servers) and multiple algorithm servers AS, the multiple algorithm servers AS are configured to obtain data directly from the data management platform DM; and a data visualization and interactive interface DI configured to generate task requests. Optionally, the query engine QE is based on Impala TM As used herein, in the context of the present disclosure, the term "connected to" refers to a relationship having a direct flow of information or data from a first component of a system to a second component and / or from a second component of a system to the first component.
[0141] Figure 11 FIG2 shows the software modules in the defect analysis system according to some embodiments of the present disclosure. Figure 11In some embodiments, the data management platform (DM) includes an ETL module (ETLP) configured to extract, transform, or load data from multiple data sources (DS) into a data mart (DMT) and a general data layer (GDL). Upon receiving an assigned task, each of the multiple algorithm servers (AS) is configured to obtain first data directly from the data mart (DMT). When performing defect analysis, each of the multiple algorithm servers (AS) is configured to send second data directly to the general data layer (GDL). It should be understood that the first data is the input data of the algorithm server, and the second data is the result data obtained by the algorithm server's calculations. The multiple algorithm servers (AS) deploy various general algorithms for defect analysis, such as algorithms based on big data analysis. These algorithms can be based on specific machine learning models, such as one or more of decision trees, random forests, GBDT, XGBoost, naive Bayes, support vector machines, Adaboost, neural network models, or other statistical algorithm models, such as WOE&IV and Apriori. The multiple algorithm servers (AS) are configured to analyze data to identify the causes of defects. As used herein, the term "ETL module" refers to computer program logic configured to provide functionality such as extracting, transforming, or loading data. In some embodiments, the ETL module is stored on a storage node, loaded into a memory, and executed by a processor. In some embodiments, the ETL module is stored on one or more storage nodes in a distributed network, loaded into one or more memories in the distributed network, and executed by one or more processors in the distributed network.
[0142] The data management platform DM stores data used in the defect analysis system. For example, the data management platform DM stores data required by multiple algorithm servers AS for algorithm analysis. In another example, the data management platform DM stores the results of the algorithm analysis. In some embodiments, the data management platform DM includes multiple data sources DS (e.g., data stored in an Oracle database), an ETL module ETLP, a data mart DMT (e.g., based on Apache Hbase), and a TM technology) and the general data layer GDL (for example, based on Apache Hive TMData storage technology). For algorithm analysis and interactive display to users, data from multiple data sources DS are cleaned and merged into verification data by the ETL module ETLP. Examples of useful data for defect analysis include tracking history data, data variable (dv) parameter data, mapping defect location data, etc. The amount of data in a typical manufacturing process (for example, the manufacturing process of display panels) is huge. For example, in a typical site, there may be more than 30 million dv parameter data per day. In order to meet the user's demand for defect analysis, it is necessary to increase the speed at which the algorithm server reads production data. In one example, the data required for algorithm analysis is stored in a database based on Apache Hbase. TM In another example, the results of algorithm analysis and other auxiliary data are stored in a data mart based on Apache Hive TM technology's common data layer.
[0143] Apache Hive TM Apache Hive is an open source data warehouse system built on top of Hadoop, which is used to query and analyze structured and semi-structured big data stored in Hadoop files. TM It is mainly used for batch processing, so it is called OLAP. TM Not a database, but a schema model.
[0144] Apache Hbase TM Apache Hbase is a non-relational column-oriented distributed database that runs on top of the Hadoop Distributed File System (HDFS). In addition, it is a NoSQL open source database that stores data in columns. TM It is mainly used for transaction processing, which is called OLTP. However, in Apache Hbase TM In this case, real-time processing is possible. TM It is a NoSQL database and has no schema model.
[0145] In one example, various components of the data management platform (e.g., common data layer, data warehouse, data sources) can be based on, for example, Apache Hadoop. TM and / or Apache Hive TM In the form of distributed data storage.
[0146] Figure 16 FIG. 1 shows a data management platform according to some embodiments of the present disclosure. Figure 16In some embodiments, the data management platform includes a distributed storage system (DFS), such as the Hadoop Distributed File System (HDFS). The data management platform is configured to collect data generated during the factory production process from multiple data sources DS. For example, data generated during the factory production process is stored in a relational database (e.g., Oracle) using RDBMS (relational database management system) grid computing technology. In RDBMS grid computing, a problem requiring a large amount of computer power is divided into many small parts, which are then distributed to many computers for processing. The results of the distributed computing are combined to obtain the final result. For example, in Oracle RAC (Real Application Clusters), all servers have direct access to all data in the database. However, applications based on RDBMS grid computing have limited hardware scalability. When the data volume reaches a certain order of magnitude, the input / output bottleneck of the hard disk makes processing large amounts of data very inefficient. The parallel processing of the distributed file system can meet the challenges posed by the increasing demand for data storage and computing. During the defect analysis process, data from multiple data sources DS is first extracted into the data management platform, greatly accelerating the processing process.
[0147] In some embodiments, the data management platform includes multiple sets of data with different content and / or storage structures. In some embodiments, the ETL module ETLP is configured to extract raw data from multiple data sources DS into the data management platform to form a first data layer (e.g., a data lake DL). The data lake DL is a centralized HDFS or KUDU database configured to store any structured or unstructured data. Optionally, the data lake DL is configured to store the first set of data extracted by the ETL module ETLP from multiple data sources DS. Optionally, the first set of data and the raw data have the same content. The dimensions and attributes of the raw data are saved in the first set of data. In some embodiments, the first set of data stored in the data lake includes dynamically updated data. Optionally, the dynamically updated data includes real-time updated data based on a Kudu-based database, or periodically updated data in the Hadoop distributed file system. In one example, the periodically updated data stored in the Hadoop distributed file system is stored in a database based on Apache Hive. TM In one example, the dynamically updated data includes real-time updated data and periodic updated data. In one example, real-time updates refer to updates at or below the minute level, excluding updates at or above the minute level; and periodic updates refer to updates at or above the minute level, including updates at or above the minute level.
[0148] In some embodiments, the data management platform includes a second data layer, such as a data warehouse DW. The data warehouse DW includes an internal storage system that is configured to provide data in an abstract manner, such as in a table format or a view format, without exposing the file system. The data warehouse DW can be based on Apache Hive TM The ETL module ETLP is configured to extract, clean, transform or load the first set of data to form the second set of data. Optionally, the second set of data is formed by cleaning and standardizing the first set of data.
[0149] In some embodiments, the data management platform includes a third data layer (eg, a general data layer GDL). The general data layer GDL can be based on Apache Hive TM . The ETL module ETLP is configured to perform data fusion on the second set of data to form a third set of data. In one example, the third set of data is data obtained by performing data fusion on the second set of data. Examples of data fusion include cascading based on the same fields in multiple tables. Examples of data fusion also include generating statistical data (e.g., sum and proportion calculation) for the same fields or records. In one example, the generation of statistical data includes counting the number of defective panels in the glass and the proportion of defective panels in multiple panels in the same glass. Optionally, the general data layer GDL is based on Apache Hive TM Optionally, the Generic Data Layer (GDL) is used for data querying.
[0150] In some embodiments, the data management platform includes a fourth data layer (e.g., at least one data mart). In some embodiments, the at least one data mart includes a data mart DMT. Optionally, the data mart DMT is a NoSQL type database that can be used for computing processing. Optionally, the data mart DMT is based on Apache Hbase. TM . Optionally, the data mart DMT is used for computing processing. The ETL module ETLP is configured to transform the third data layer to form a fourth set of data. Optionally, the fourth set of data classifies the data based on different types and / or rules, thereby forming a multi-layer index structure. It will be understood by those skilled in the art that the fourth set of data of the multi-layer index structure may be a plurality of sub-data tables with index relationships. In one example of the present disclosure, classifying data based on different types and / or rules refers to different environmental factors, keys / values, column families, etc. The first index in the multi-layer index structure corresponds to the filtering criteria of the front-end interface, for example, corresponding to the user-defined analysis criteria in the interactive task sub-interface that communicates with the data management platform, thereby facilitating faster data query and calculation processes.
[0151] Those skilled in the art will appreciate that the first set of data, the second set of data, the third set of data, and the fourth set of data may be stored and queried in the form of one or more data tables.
[0152] In some embodiments, the process of converting the third set of data into the fourth set of data can involve importing data from the general data layer (GDL) into the data mart (DMT). In one example, a first table is generated in the data mart (DMT), and a second table (e.g., an external table) is generated in the general data layer (GDL). The first and second tables are configured to be synchronized so that when data is written to the second table, the first table is simultaneously updated to include the corresponding data.
[0153] In another example, the distributed computing processing module can be used to read data written to the general data layer GDL. The MapReduce module in Hadoop can be used as a distributed computing processing module to read data written to the general data layer GDL. The data written to the general data layer GDL can then be written to the data mart DMT. In one example, the data can be written to the data mart DMT using an HBase API. In another example, once the MapReduce module reads the data written to the data mart DMT, it can generate HFile files and load them in bulk to the data mart DMT.
[0154] In some embodiments, this document describes the data flows, data transformations, and data structures between various components of a data management platform. In some embodiments, the raw data collected by multiple data sources (DS) includes at least one of biographical information, parameter information, or defect information. The raw data may optionally include dimensional information (e.g., time, plant, equipment, operator, map, chamber, slot, etc.) and attribute information (e.g., plant location, equipment age, number of bad points, abnormal parameters, energy consumption parameters, processing duration, etc.).
[0155] The history data information includes information about the specific processes that a product (such as a panel or glass) undergoes during manufacturing. Examples of the specific processes that a product undergoes during manufacturing include factories, processes, stations, equipment, chambers, card slots, and operators.
[0156] The parameter information includes information about specific environmental parameters and their variations that a product (such as a panel or glass) experiences during manufacturing. Examples of specific environmental parameters and their variations that a product experiences during manufacturing include ambient particle conditions, equipment temperature, and equipment pressure.
[0157] The defect information includes information about the quality of the product based on the inspection. Example product quality information includes defect type, defect location, and defect size.
[0158] In some embodiments, parameter information includes device parameter information. Optionally, device parameter information includes at least three types of data that can be output from a Generic Model (GEM) interface for communication and control of manufacturing equipment. The first type of data that can be output from the GEM interface is data variables (DVs), which can be collected when an event occurs. Therefore, data variables are only valid in the presence of an event. In one example, the GEM interface can provide an event called PPChanged, which is triggered when a recipe changes, and a data variable called "changed recipe," which is only valid in the presence of a PPChanged event. Polling this value at other times may yield invalid or unexpected data. The second type of data that can be output from the GEM interface is state variables (SVs), which contain device-specific information valid at any time. In one example, the device can be a temperature sensor, and the GEM interface provides temperature state variables for one or more modules. The host can request the value of this state variable at any time and can expect the value to be true. The third type of data that can be output from the GEM interface is device constants (ECs), which contain data items set by the device. Device constants determine the behavior of the device. In one example, the GEM interface provides a device constant named "MaxSimultousTrace" that specifies the maximum number of traces that can be requested from the host simultaneously. The value of the device constant is always guaranteed to be valid and up-to-date.
[0159] In some embodiments, the data lake DL is configured to store a first set of data formed by extracting raw data from multiple data sources by an ETL module ETLP, the first set of data having the same content as the raw data. The ETL module ETLP is configured to extract raw data from multiple data sources DS while maintaining dimensional information (e.g., dimension columns) and attribute information (e.g., attribute columns). The data lake DL is configured to store the extracted data sorted by extraction time. The data can be stored in the data lake DL, which has a new name indicating the "data lake" and / or one or more attributes of each data source, while maintaining the dimensions and attributes of the raw data. The first set of data and the raw data are stored in different forms. The first set of data is stored in a distributed file system, while the raw data is stored in a relational database such as an Oracle database. In one example, the business data collected by the multiple data sources DS includes data from various business systems, such as a yield management system (YMS), a fault detection and classification (FDC) system, and a manufacturing execution system (MES). The data in these business systems has their own signatures, such as product models, production parameters, and equipment model data. The ETL module ETLP uses tools (such as sqoop command, data stack tool, Pentaho tool) to extract the original production data from each business system into the original data format of Hadoop, thereby realizing the integration of data from multiple business systems. The extracted data is stored in the data lake DL. In another example, the data lake DL is based on a database such as Hive TM and Kudu TM The DL data lake contains dimension columns (time, factory, equipment, operator, map, chamber, slot, etc.) and attribute columns (factory location, equipment age, number of bad points, abnormal parameters, energy consumption parameters, processing duration, etc.) involved in the factory automation process.
[0160] In one example, the data management platform integrates various business data (e.g., data related to semiconductor electronic device manufacturing) into multiple data sources DS (e.g., Oracle database). The ETL module ETLP uses, for example, DataStack tools, SQOOP tools, Kettle tools, Pentaho tools, or DataX tools to extract data from multiple data sources DS into the data lake DL. The data is then cleaned, transformed, and loaded into the data warehouse DW and the general data layer GDL. The data warehouse DW, the general data layer GDL, and the data mart DMT utilize tools such as Kudu. TM 、Hive TM and HBase TM Tools to store large amounts of data and analysis results.
[0161] Information generated at various stages of the manufacturing process is acquired by various sensors and inspection equipment and subsequently stored in multiple data sources (DS). The calculation and analysis results generated by the defect analysis system are also stored in multiple data sources (DS). The ETL module (ETLP) enables data synchronization (data flow) between the various components of the data management platform. For example, the ETL module (ETLP) is configured to obtain a parameter configuration template for the synchronization process, including network permissions and database port configuration, inbound database and table names, outbound database and table names, field mappings, task types, scheduling cycles, and so on. The ETL module (ETLP) configures the parameters of the synchronization process based on the parameter configuration template. The ETL module (ETLP) synchronizes data and cleans the synchronized data based on the process configuration template. The ETL module (ETLP) cleans the data using SQL statements to remove null values, outliers, and establish correlations between related tables. Data synchronization tasks include data synchronization between multiple data sources (DS) and the data management platform, as well as data synchronization between the various layers of the data management platform (e.g., data lake (DL), data warehouse (DW), general data layer (GDL), or data mart (DMT).
[0162] In another example, data extraction to the data lake DL can be done in real time or offline. In offline mode, data extraction tasks are scheduled periodically. Optionally, in offline mode, the extracted data can be stored in a storage device based on the Hadoop distributed file system (e.g., based on Hive TM In real-time mode, the data extraction task can be performed by OGG (Oracle GoldenGate) combined with Apache Kafka. Optionally, in real-time mode, the extracted data can be stored in a Kudu-based TM OGG reads log files from multiple data sources (e.g., Oracle database) to obtain add / delete data. In another example, topic information is read by Flink, and Json is selected as the synchronization field type. The data is parsed using the JAR package and the parsed information is sent to the Kudu API to implement the addition / deletion of Kudu table data. In one example, the front-end interface can be based on the data stored in the Kudu TM In another example, the front-end interface can be based on the data stored in the Kudu database. TM Databases, Hadoop distributed file systems (e.g., based on Apache Hive TM Databases) and / or based on Apache Hbase TMIn another example, short-term data (e.g., generated within a few months) is stored in a Kudu-based TM The long-term data (e.g., all data generated in all cycles) is stored in the Hadoop distributed file system (e.g., based on Apache Hive TM In another example, the ETL module ETLP is configured to store TM Extract data from the database to the Hadoop distributed file system (for example, based on Apache Hive TM database).
[0163] By combining data from various business systems (MDW, YMS, MES, FDC, etc.), a data warehouse DW is built based on the data lake DL. The data extracted from the data lake DL is divided according to the task execution time, and the task execution time does not completely match the timestamp in the original data. In addition, there is a possibility of data duplication. Therefore, it is necessary to build a data warehouse DW based on the data lake DL by cleaning and standardizing the data in the data lake DL to meet the needs of upper-level applications for data accuracy and division. The data tables stored in the data warehouse DW are obtained by cleaning and standardizing the data in the data lake DL. Based on user needs, the field format is standardized to ensure that the data tables in the data warehouse DW are completely consistent with the data tables in multiple data sources DS. At the same time, dividing data by date or month, according to time and other fields greatly improves query efficiency and reduces running memory requirements. The data warehouse DW can be based on Kudu TM Database and Apache Hive TM One or any combination of databases.
[0164] In some embodiments, the ETL module ETLP is configured to cleanse the extracted data stored in the data lake into cleaned data, and the data warehouse is configured to store the cleaned data. Examples of cleaning performed by the ETL module ETLP include removing redundant data, removing null value data, removing virtual fields, etc.
[0165] In some embodiments, the ETL module ETLP is further configured to perform standardization (e.g., field standardization and format standardization) on the extracted data stored in the data lake, and the cleansed data is data that has undergone field format standardization (e.g., format standardization of date and time information).
[0166] In some embodiments, at least a portion of the business data from multiple data sources DS is in binary large object (blob) format. After data extraction, at least a portion of the extracted data stored in the data lake DL is in compressed hexadecimal format. Optionally, at least a portion of the cleansed data stored in the data warehouse DW is obtained by decompressing and processing the extracted data. The binary large object data is converted to hexadecimal format when extracted and stored in the data lake; the hexadecimal format data is decompressed when extracted and stored in the data warehouse to form a second set of data. In one example, a business system (e.g., the aforementioned FDC system) is configured to store a large amount of parameter data. Therefore, the data must be compressed into a blob format within the business system. During data extraction (e.g., from an Oracle database to a Hive database), the blob fields are converted to hexadecimal (HEX) strings. To retrieve the parameter data stored in the file, the HEX file is decompressed, and the file contents can then be directly retrieved. The required corresponding data is encoded to form a long string, and specific symbols are used to separate different contents according to output requirements. In order to obtain data in the required format, operations such as cutting according to special characters and converting rows and columns are performed on the long character strings. The processed data is written to the target table together with the original data (for example, the data in the table format stored in the data warehouse DW discussed above).
[0167] In one example, the cleansed data stored in the data warehouse DW maintains the dimension information (e.g., dimension columns) and attribute information (e.g., attribute columns) of the original data in the multiple data sources DS. In another example, the cleansed data stored in the data warehouse DW maintains the same data table names as the data table names in the multiple data sources DS.
[0168] In some embodiments, the ETL module ETLP is further configured to generate a periodically updated dynamic update table. Optionally, as described above, the general data layer GDL is configured to store a dynamic update table including information about high-incidence defects (a type of defect information of concern). In one example of the present disclosure, the high-incidence defect information includes the top five or top ten defect information with a high defect ratio, or other defect information specified by the user. Optionally, the data mart DMT is configured to store a dynamic update table including information about high-incidence defects, as described above.
[0169] The general data layer (GDL) is built based on the data warehouse (DW). In some embodiments, the GDL is configured to store a third set of data formed by merging the second set of data using the ETL module (ETLP). Optionally, data fusion is performed based on different themes. The data in the general data layer (GDL) is highly thematic and aggregated, significantly improving query speed. In one example, tables in the data warehouse (DW) can be used to construct tables with dependencies based on different user needs or different themes, with each table being named according to its respective purpose.
[0170] Various topics may correspond to different data analysis requirements. For example, a topic may correspond to different defect analysis requirements. In one example, a topic may correspond to the analysis of defects attributed to one or more manufacturing node groups (e.g., one or more devices), and data fusion based on the topic may include data fusion of historical information about the manufacturing process and defect information. In another example, a topic may correspond to the analysis of defects attributed to one or more parameter types, and data fusion based on the topic may include data fusion of parameter feature information and defect information. In another example, a topic may correspond to the analysis of defects attributed to one or more equipment operations (e.g., equipment defined by corresponding process sites where corresponding equipment performs corresponding operations), and data fusion based on the topic may include data fusion of at least two types of information among parameter feature information, historical information about the manufacturing process, and defect information. In another example, a topic may correspond to feature extraction of at least one type of parameter information to generate parameter feature information, wherein one or more of the maximum value, minimum value, average value, and median value are extracted for at least one type of parameter information. In one example of the present disclosure, at least one type of parameter information includes data of at least one equipment parameter, such as temperature, humidity, pressure, etc., and also includes data such as environmental granularity.
[0171] In some embodiments, defect analysis includes performing feature extraction on at least one type of parameter information to generate parameter feature information; and performing data fusion on at least two of the parameter feature information, manufacturing process history information, and defect information. Optionally, performing data fusion includes performing data fusion on the parameter feature information and the defect information. Optionally, performing data fusion includes performing data fusion on at least two of the parameter feature information, manufacturing process history information, and defect information. In another example, performing data fusion includes performing data fusion on the manufacturing process parameter feature information and history information to obtain first fused data information; and performing data fusion on the first fused data information and its associated defect information to obtain second fused data information. In one example, the second fused data information includes glass serial number, site information, equipment information, parameter feature information, and defect information. For example, data fusion is performed in the general data layer (GDL) by constructing a table with correlations based on user needs or topics. Optionally, performing data fusion includes performing data fusion on history information and defect information. Optionally, performing data fusion includes performing data fusion on all three of the parameter feature information, manufacturing process history information, and defect information.
[0172] In one example, the CELL_PANEL_MAIN table in the data warehouse DW stores basic historical data of panels in a box-making factory, and the CELL_PANEL_CT table stores detailed data (such as defect information or parameter information) of the CT (cell test) process in the factory. The general data layer GDL is configured to perform relevant operations based on the CELL_PANEL_MAIN table and the CELL_PANEL_CT table to create a wide table YMS_PANEL, thereby realizing data fusion. The basic historical data of the panel and the detailed data of the CT process can be queried in the YMS_PANEL table. The YMS prefix in the table name "YMS_PANEL" represents the subject for defect analysis, and the PANEL prefix represents the specific PANEL information stored in the table. By performing relevant operations on the tables in the data warehouse DW by the general data layer GDL, the data in different tables can be fused and associated.
[0173] According to different business analysis requirements, and based on glass, hglass (halfglass), and panels, the tables in the general data layer GDL can be divided into the following data tags: production records, defect rate, defect MAP, DV, SV, inspection data, and test data.
[0174] The data mart DMT is built based on the data warehouse DW and / or the general data layer GDL. The data mart DMT can be used to provide data required for various reports and analyses, especially highly customized data. In one example, the customized data provided by the data mart DMT includes merged data on defect rates, frequencies of specific defects, etc. In another example, the data in the data lake DL and the general data layer GDL are stored in Hive, and the data in the data mart DMT is stored in a database based on Hbase. Optionally, the table names in the data mart DMT can be kept consistent with those in the general data layer GDL. Optionally, the general data layer GDL is based on Apache Hive TM Technology, and the data mart DMT is based on Apache Hbas TM Technology. The general data layer (GDL) is used for data queries through the user interface. Hive data can be quickly queried in Hive using Impala. The data mart (DMT) provides computational data for the algorithm servers. Leveraging the advantages of HBase's columnar data storage, multiple algorithm servers (AS) can quickly access data in HBase.
[0175] In some embodiments, the data mart DMT is configured as a data table stored in a data table in the general data layer (GDL) divided into multiple sub-tables. In some embodiments, the data stored in the data mart DMT and the data stored in the general data layer (GDL) have the same content. The difference between the data stored in the data mart DMT and the data stored in the general data layer (GDL) is that they are stored in different data models. Depending on the type of NoSQL database used for the data mart DMT, the data in the data mart DMT can be stored in different data models. Examples of data models corresponding to different NoSQL databases include key / value data models, column family data models, versioned document data models, and graph data models. In some embodiments, queries on the data mart DMT can be executed based on specified keys to quickly locate the data to be queried (e.g., values). Therefore, and as discussed in more detail below, the tables stored in the general data layer (GDL) can be divided into at least three sub-tables in the data mart DMT. The first sub-table corresponds to the analysis criteria defined by the user in the interactive task sub-interface. The second sub-table corresponds to the key (e.g., product serial number). The third sub-table corresponds to the value (e.g., the value in the table stored in the general data layer (GDL), including fused data). It is understandable that the product range that the user needs to analyze can be determined through the first sub-table, so that the corresponding data (value) in the third sub-table can be queried based on the serial number (key) of the corresponding product in the second sub-table. In an example, the data mart DMT uses the Apache Hbase TMA NoSQL database of the technology; the specified key in the second sub-table can be a row key; and the fused data in the third sub-table (column family data corresponding to the row key) can be stored in a column family data model. Optionally, the fused data in the third sub-table can be fused data from at least two of parameter feature information, historical information of the manufacturing process, and defect information. In addition, the data mart DMT may include a fourth sub-table. Certain characters in the third sub-table may be stored in codes, for example, due to their length or other reasons. The fourth sub-table includes characters corresponding to these codes stored in the third sub-table (e.g., device names, sites). Indexes or queries between the first sub-table, the second sub-table, and the third sub-table can be based on the codes. The fourth sub-table can be used to replace codes with characters before the results are presented to the user interface.
[0176] In some embodiments, the plurality of sub-tables have an index relationship between at least two of the plurality of sub-tables. Optionally, the data in the plurality of sub-tables is categorized based on type and / or rules. In some embodiments, the plurality of sub-tables includes a first sub-table (e.g., an attribute sub-table) that includes a plurality of environmental factors corresponding to user-defined analysis criteria in an interactive task sub-interface that communicates with the data management platform; a second sub-table that includes a product serial number (e.g., a glass identification number or a batch identification number); and a third sub-table (e.g., a master sub-table) that includes values corresponding to the product serial number in the third set of data. Optionally, the second sub-table may include different designated keys, such as glass identification number or batch identification number, based on different themes (e.g., multiple second sub-tables). Optionally, the values in the third set of data correspond to the glass identification number via an index relationship between the third sub-table and the second sub-table. Optionally, the plurality of sub-tables further includes a fourth sub-table (e.g., a metadata sub-table) that includes values corresponding to the batch identification number in the third set of data. Optionally, the second subtable also includes a batch identification number; the value corresponding to the batch identification number in the third set of data can be obtained through the index relationship between the second subtable and the fourth subtable. Optionally, the multiple subtables also include a fifth subtable (e.g., a code generator subtable) that includes site information and device information. Optionally, the third subtable includes codes or abbreviations for sites and devices, and the site information and device information can be obtained from the fifth subtable through the index relationship between the third subtable and the fifth subtable.
[0177] Figure 17 Depicts multiple sub-tables derived from a data table stored in a common data layer according to some embodiments of the present disclosure. Figure 17In some embodiments, the multiple sub-tables include one or more of the following: an attribute sub-table, which includes multiple environmental factors corresponding to user-defined analysis criteria in an interactive task sub-interface that communicates with a data management platform; a context sub-table, which includes at least a first number of environmental factors and a plurality of manufacturing stage factors among the multiple environmental factors, and a plurality of columns corresponding to a second number of environmental factors among the multiple environmental factors; a metadata sub-table, which includes at least a first manufacturing stage factor among the multiple manufacturing stage factors and a device factor associated with the first manufacturing stage, and a plurality of columns corresponding to parameters generated in the first manufacturing stage; a master sub-table, which includes at least a second manufacturing stage factor among the multiple manufacturing stage factors, and a plurality of columns corresponding to parameters generated in the second manufacturing stage; and a code generator sub-table, which includes at least a third number of environmental factors and device factors among the multiple environmental factors.
[0178] In one example, multiple subtables are subtables stored in a column family database, including one or more of the following: an attribute subtable, which includes a primary key consisting of a data tag, factory information, site information, product model information, product type information, and product serial number. Based on the selection of the interactive task sub-interface, the filtering conditions can be determined; a context subtable, which includes a primary key consisting of the first three digits of the value after the MED5 encryption site, factory information, site information, data tag, manufacturing end time, batch serial number, and glass serial number. The corresponding columns are product model information, product serial number, and product type information. The filtering conditions determined from the attribute subtable locate the corresponding batch serial number and glass serial number based on the context subtable; a metadata subtable, which includes a primary key consisting of the first three digits of the MED5 encrypted batch serial number. The primary key comprises a value, batch serial number, data tag, site information, and device information; the corresponding columns include time and manufacturing parameters; the corresponding values include specific data values; the batch serial number determined from the context subtable can obtain the corresponding data value based on the metadata subtable; a master subtable comprises a primary key comprising the first three digits of the MED5 encrypted batch serial number, the serial number, and the glass serial number; the corresponding columns include time and manufacturing parameters; the corresponding values include specific data values; the glass serial number determined from the context subtable can obtain the corresponding data value based on the metadata subtable; and a code generator subtable comprises a primary key comprising a data tag, site information, and device information; the corresponding column is a serial number; the corresponding site and device data in the code generator subtable can be searched based on the serial number in the master subtable. Optionally, the multiple environmental factors in the attribute subtable include data tags, factory information, site information, product model information, product type information, and product serial number. Optionally, the multiple manufacturing stage factors include batch serial number and glass serial number. Optionally, the device factors include device information.
[0179] Reference Figure 10 and Figure 11In some embodiments, the software module further comprises a load balancer LB connected to the analyzer AZ. Optionally, the load balancer LB (e.g., the first load balancer LB1) is configured to receive a task request and is configured to distribute the task request to one or more of the multiple business servers BS to achieve load balancing between the multiple business servers BS. Optionally, the load balancer LB (e.g., the second load balancer LB2) is configured to distribute tasks from the multiple business servers BS to one or more of the multiple algorithm servers AS to achieve load balancing between the multiple algorithm servers AS. Optionally, the load balancer LB is based on Nginx TM technology load balancer.
[0180] In some embodiments, the defect analysis system is configured to simultaneously meet the needs of many users. By having a load balancer LB (e.g., the first load balancer LB1), the system sends user requests to multiple service servers AS in a balanced manner, thereby maintaining the overall performance of the multiple service servers AS and preventing slow service responses due to excessive pressure on a single server.
[0181] Similarly, by having a load balancer LB (for example, a second load balancer LB2), the system sends tasks to multiple algorithm servers AS in a balanced manner to keep the overall performance of the multiple algorithm servers AS optimal. In some embodiments, when designing a load balancing strategy, not only the number of tasks sent to each of the multiple algorithm servers AS should be considered, but also the amount of computational load required for each task should be considered. In one example, three types of tasks are involved, including defect analysis of type "glass", defect analysis of type "hglass", and defect analysis of type "panel". In another example, the number of defect data items related to type "glass" is an average of 1 million per week, and the number of defect data items related to type "panel" is an average of 30 million per week. Therefore, the computational load required for defect analysis of type "panel" is much greater than the computational load required for defect analysis of type "glass". In another example, load balancing is performed using the formula f(x, y, z) = mx + ny + oz, where x represents the number of defect analysis tasks for the "glass" type; y represents the number of defect analysis tasks for the "hglass" type; z represents the number of defect analysis tasks for the "panel" type; m represents the weight assigned to defect analysis for the "glass" type; n represents the weight assigned to defect analysis for the "hglass" type; and o represents the weight assigned to defect analysis for the "panel" type. Weights are assigned based on the computational load required for each type of defect analysis. Alternatively, m + n + o = 1.
[0182] In some embodiments, the ETL module ETLP is configured to generate a dynamically updated table, which is updated periodically (e.g., daily, hourly, etc.). Optionally, the general data layer GDL is configured to store the dynamically updated table. In one example, the dynamically updated table is generated based on the logic of calculating the defect incidence rate in the factory. In another example, data from multiple tables in the data management platform DM are merged and subjected to various calculations to generate a dynamically updated table. In another example, the dynamically updated table includes information such as job name, defect type, frequency of occurrence of defect type, level of defect type (glass / hglass / panel), factory, product model, date, and other information. The dynamically updated table is updated regularly, and when the production data in the data management platform DM changes, the information in the dynamically updated table will be updated accordingly to ensure that the dynamically updated table can have defect type information for all factories.
[0183] Figure 12 A defect analysis method using a defect analysis system according to some embodiments of the present disclosure is shown. Figure 12 In some embodiments, the data visualization and interaction interface DI is configured to generate a task request; the load balancer LB is configured to receive the task request and is configured to assign the task request to one or more of the multiple business servers to achieve load balancing among the multiple business servers; one or more of the multiple business servers are configured to send a query task request to the query engine QE; the query engine QE is configured to query the dynamically updated table when receiving the query task request from one or more of the multiple business servers to obtain information about high-incidence defects and send the information about high-incidence defects to one or more of the multiple business servers; one or more of the multiple business servers are configured to send the defect analysis task to the load balancer LB to assign the defect analysis task to one or more of the multiple algorithm servers, thereby achieving load balancing among the multiple algorithm servers; when receiving the defect analysis task, one or more of the multiple algorithm servers are configured to obtain data directly from the data mart DMT to perform defect analysis; and when the defect analysis is completed, one or more of the multiple algorithm servers are configured to send the results of the defect analysis to the general data layer GDL.
[0184] The query engine QE is capable of quickly accessing the data management platform DM, for example, quickly reading data from or writing data to the data management platform DM. Compared with direct queries through the general data layer GDL, having a query engine QE is advantageous because it does not require the execution of a MapReduce (MR) program to query the general data layer GDL (for example, a Hive data store). Optionally, the query engine QE can be a distributed query engine that can query the general data layer GDL (HDFS or Hive) in real time, greatly reducing waiting time and improving the responsiveness of the entire system. The query engine QE can be implemented using various appropriate technologies. Examples of technologies for implementing the query engine QE include Impala TM Technology, Kylin TM Technology, Presto TM Technology and Greenpall TM technology.
[0185] In some embodiments, the task request is a recurring task request that defines a recurring period during which the defect analysis is to be performed. Figure 13 A defect analysis method using a defect analysis system according to some embodiments of the present disclosure is shown. Figure 13 In some embodiments, the data visualization and interactive interface DI is configured to generate a periodic task request; the load balancer LB is configured to receive the periodic task request and is configured to distribute the periodic task request to one or more of the multiple business servers to achieve load balancing among the multiple business servers; one or more of the multiple business servers are configured to send a query task request to the query engine QE; the query engine QE is configured to query the dynamically updated table when receiving the query task request from one or more of the multiple business servers to obtain information about high-occurrence defects in a recurring period, and send the information about high-occurrence defects to one or more of the multiple business servers; upon receiving the information about high-occurrence defects in a recurring period When receiving information about defects with a high incidence during a repetitive period, one or more of the multiple business servers are configured to generate a defect analysis task based on the information about defects with a high incidence during a repetitive period; one or more of the multiple business servers are configured to send the defect analysis task to a load balancer LB to assign the defect analysis task to one or more of the multiple algorithm servers, thereby achieving load balancing among the multiple algorithm servers; when receiving the defect analysis task, one or more of the multiple algorithm servers are configured to obtain data directly from the data mart DMT to perform defect analysis; and when completing the defect analysis, one or more of the multiple algorithm servers are configured to send the results of the defect analysis to the general data layer GDL.
[0186] refer to Figure 11In some embodiments, the data visualization and interactive interface DI includes an automatic task sub-interface SUB1, which allows input of a repetition period for defect analysis to be performed. The automatic task sub-interface SUB1 is capable of periodically performing automatic defect analysis on defects with high incidence rates. In automatic task mode, information about defects with high incidence rates is sent to multiple algorithm servers AS to analyze the potential causes of the defects. In one example, the user sets a repetition period for defect analysis to be performed in the automatic task sub-interface SUB1. The query engine QE captures defect information from a dynamically updated table based on system settings at regular intervals, and sends the information to multiple algorithm servers AS for analysis. In this way, the system can automatically monitor defects with high incidence rates, and the corresponding analysis results can be stored in a cache for access and for display in the data visualization and interactive interface DI.
[0187] In some embodiments, the task request is an interactive task request. Figure 14 A defect analysis method using a defect analysis system according to some embodiments of the present disclosure is shown. Figure 14In some embodiments, the data visualization and interactive interface DI is configured to receive user-defined analysis criteria and is configured to generate interactive task requests based on the user-defined analysis criteria; the data visualization and interactive interface DI is configured to generate interactive task requests; the load balancer LB is configured to receive interactive task requests and is configured to distribute the interactive task requests to one or more of the multiple business servers to achieve load balancing among the multiple business servers; one or more of the multiple business servers is configured to send a query task request to the query engine; the query engine QE is configured to, upon receiving a query task request from one or more of the multiple business servers, query a dynamically updated table to obtain information about defects with high incidence rates and send the information about defects with high incidence rates to one or more of the multiple business servers; upon receiving the information about defects with high incidence rates, one or more of the multiple business servers is configured to send the information to the data visualization and interactive interface The data visualization and interactive interface DI is configured to display information about defects with high incidence rates and multiple environmental factors associated with the defects with high incidence rates, and is configured to receive a user-defined selection of one or more environmental factors from the multiple environmental factors, and send the user-defined selection to one or more of the multiple business servers; one or more of the multiple business servers are configured to generate a defect analysis task based on the information and the user-defined selection; one or more of the multiple business servers are configured to send the defect analysis task to the load balancer LB to assign the defect analysis task to one or more of the multiple algorithm servers, thereby achieving load balancing among the multiple algorithm servers; when receiving the defect analysis task, one or more of the multiple algorithm servers are configured to obtain data directly from the data mart DMT to perform defect analysis; and when the defect analysis is completed, one or more of the multiple algorithm servers are configured to send the results of the defect analysis to the general data layer GDL.
[0188] refer to Figure 11In some embodiments, the data visualization and interaction interface DI includes an interactive task sub-interface SUB2 that allows input of user-defined analysis criteria, including user-defined selections of one or more environmental factors. In one example, a user can filter various environmental factors step by step in the interactive task sub-interface SUB2, including data source, factory, site, model, product model, batch, etc. One or more of the multiple business servers BS are configured to generate defect analysis tasks based on information about defects with high incidence rates and user-defined selections of one or more environmental factors. The analyzer AZ continuously interacts with the general data layer GDL and causes the selected one or more environmental factors to be displayed on the interactive task sub-interface SUB2. The interactive task sub-interface SUB2 allows the user to limit the environmental factors to a few, for example, certain selected devices or certain selected parameters, based on the user's experience.
[0189] In some embodiments, the general data layer (GDL) is configured to generate tables based on different topics. In one example, the tables include a history table, which contains information about the sites and equipment that a glass or panel has passed through throughout the manufacturing process. In another example, the tables include a dv table, which contains parameter information uploaded by the equipment. In another example, if the user only wants to analyze equipment dependencies, the user can select the history table for analysis. In another example, if the user only wants to analyze equipment parameters, the user can select the dv table for analysis.
[0190] refer to Figure 11 In some embodiments, the analyzer AZ also includes a cache server CS. The cache server CS is configured to store a portion of the defect analysis task results in cache C. In some embodiments, the data visualization and interaction interface DI also includes a defect visualization sub-interface SUB-3. In one embodiment, the primary function of the defect visualization sub-interface SUB-3 is to allow users to customize queries and display the corresponding defect analysis task results when the user clicks on a defect type. In one example, the user clicks on a defect type, and the system sends a request to one or more of the multiple business servers BS via the load balancer LB. One or more of the multiple business servers BS first queries the result data cached in cache C. If cached result data exists, the system directly displays the cached result data. If result data corresponding to the selected defect type is not currently cached in cache C, the query engine QE is configured to query the general data layer GDL for result data corresponding to the selected defect type. Once queried, the system caches the result data corresponding to the selected defect type in cache C, and this result data can be used for the next query on the same defect type.
[0191] Figure 15A defect analysis method using a defect analysis system according to some embodiments of the present disclosure is shown. Figure 15 In some embodiments, the defect visualization sub-interface DI is configured to receive a user-defined selection of a defect to be analyzed and generate a call request; the load balancer LB is configured to receive the call request and to distribute the call request to one or more of the multiple business servers to achieve load balancing among the multiple business servers; one or more of the multiple business servers is configured to send the call request to the cache server; and the cache server is configured to determine whether information about the defect to be analyzed is stored in the cache. Optionally, when it is determined that the information about the defect to be analyzed is stored in the cache, one or more of the multiple business servers is configured to send the information about the defect to be analyzed to the defect visualization sub-interface for display. Optionally, when it is determined that information about the defect to be analyzed is not stored in the cache, one or more of the multiple business servers is configured to send a query task request to the query engine; the query engine is configured to, when receiving the query task request from one or more of the multiple business servers, query the dynamically updated table to obtain information about the defect to be analyzed, and send the information about the defect to be analyzed to the cache; the cache is configured to store information about the defect to be analyzed; and one or more of the multiple business servers is configured to send the information about the defect to be analyzed to the defect visualization sub-interface for display.
[0192] Optionally, the defect analysis task result includes defect analysis task results based on periodic task requests. Optionally, the defect analysis task result includes defect analysis task results based on periodic task requests; and defect analysis task results obtained based on query task requests.
[0193] By having a cache server CS, high requirements for system response speed (for example, displaying results associated with defect types) can be met. In one example, up to 40 tasks can be generated every half hour through periodic task requests, where each task is associated with up to five different defect types, and each defect type is associated with up to 100 parameter information. If all analysis results are cached, a total of 40*5*100=20,000 queries must be stored in cache C, which will put a lot of pressure on the cluster memory. In one example, a portion of the defect analysis task results is limited to the results associated with the top three highest-ranked defect types, and only this portion is cached.
[0194] Various suitable methods for defect analysis may be implemented by one or more of the plurality of algorithm servers of the defect analysis system described herein. Figure 18 The defect analysis method in some embodiments of the present disclosure is shown. Figure 18 In some embodiments, the method includes obtaining manufacturing data information including defect information; classifying the manufacturing data information into multiple groups of data according to the manufacturing node group, each group of data in the multiple groups of data is associated with each manufacturing node group in the manufacturing node group; calculating the evidence weight of the manufacturing node group to obtain multiple evidence weights, wherein the evidence weight represents the difference between the proportion of defect data in each manufacturing node group relative to the proportion of defect data in all manufacturing node groups; sorting the multiple groups of data based on the multiple evidence weights; obtaining a list of the multiple groups of data sorted based on the multiple evidence weights; and performing defect analysis on one or more selected groups of the multiple groups of data. Optionally, each manufacturing node group includes one or more selected from the group consisting of manufacturing process, equipment, site and process section. Optionally, the manufacturing data information can be obtained from the data mart DMT. Optionally, the manufacturing data information can be obtained from the general data layer GDL.
[0195] Optionally, the method includes processing manufacturing data information including historical data information and defect information to obtain process data; classifying the process data into a plurality of data groups based on the equipment groups, wherein each data group in the plurality of data groups is associated with each equipment group in the equipment groups; calculating evidence weights for the equipment groups to obtain a plurality of evidence weights; ranking the plurality of data groups based on the plurality of evidence weights; and performing defect analysis on one or more groups with the highest ranking in the plurality of data groups. Optionally, the defect analysis is performed at a parameter level.
[0196] In some embodiments, the respective weight of evidence for each device group is calculated according to equation (2):
[0197]
[0198] Wherein, woei represents the respective evidence weights of each device group; P(yi) represents the ratio of the number of positive samples in each device group to the number of positive samples in all manufacturing node groups (e.g., device groups); P(ni) represents the ratio of the number of negative samples in each device group to the number of negative samples in all manufacturing node groups (e.g., device groups); positive samples represent data including defect information associated with each device group; negative samples represent data in which defect information associated with each device group does not exist; #yi represents the number of positive samples in each device group; #yr represents the number of positive samples in all manufacturing node groups (e.g., device groups); #ni represents the number of negative samples in each device group; #nr represents the number of negative samples in all manufacturing node groups (e.g., device groups).
[0199] In some embodiments, the method further comprises processing the manufacturing data information to obtain processing data. Optionally, processing the manufacturing data information comprises performing data fusion on the history data information and the defect information to obtain fused data information.
[0200] In one example, processing manufacturing data information to obtain processed data includes obtaining raw data information of various manufacturing processes of a display panel, including historical data information, parameter information, and defect information; preprocessing the raw data to remove null data, redundant data, and virtual fields, and filtering the data based on preset conditions to obtain verification data; fusing the historical data information and defect information in the verification data to obtain third fused data information; determining whether any defect information in the fused data information contains the same machine-inspected defect information and manually-reviewed defect information, identifying the manually-reviewed defect information (rather than the machine-inspected defect information) as the defect information to be analyzed, thereby generating reviewed data; fusing the review data and the historical data information to obtain fourth fused data information; and removing non-representative data from the fourth fused data information to obtain processed data. For example, data generated during the process of glass passing through a very small number of devices can be eliminated. When the number of devices through which the glass passes only accounts for a small percentage (e.g., 10%) of the total number of devices, non-representative data will bias the analysis, thereby affecting the accuracy of the analysis.
[0201] In one example, the historical data information (used to be fused with the audit data to obtain the fourth fused data information) includes glass data and hglass data (half-glass data, i.e., historical data after the entire glass is cut in half). However, the audited data is panel data. In one example, the glass_id / hglass_id in the fab (manufacturing) stage is fused with the panel_id in the EAC2 stage, where redundant data is removed. The purpose of this step is to ensure that the historical data information in the fab stage is consistent with the defect information in the EAC2 stage. For example, the number of bits in glass_id / hglass_id is different from the number of bits in panel_id. In one example, the number of bits in panel_id is processed to be consistent with the number of bits in glass_id / hglass_id. After data fusion, data with complete information is obtained, including glass_id / hglass_id, site information, equipment information, and defect information. Optionally, the fused data is subjected to additional operations to remove redundant data items.
[0202] In some embodiments, performing defect analysis includes performing feature extraction on at least one type of parameter information to generate parameter feature information, wherein one or more of a maximum value, a minimum value, a mean value, and a median value are extracted for the at least one type of parameter information. Optionally, performing feature extraction includes performing time domain analysis to extract statistical information, wherein the statistical information includes one or more of count, mean value, maximum value, minimum value, range, variance, deviation, kurtosis, and percentile. Optionally, performing feature extraction includes performing frequency domain analysis to convert time domain information obtained in the time domain analysis into frequency domain information including one or more of a power spectrum, information entropy, and a signal-to-noise ratio.
[0203] In one example, feature extraction is performed on a list of multiple sets of data ranked based on multiple weights of evidence. In another example, feature extraction is performed on one or more sets of the multiple sets of data with the highest ranking. In another example, feature extraction is performed on the set of data with the highest ranking.
[0204] In some embodiments, performing defect analysis further includes performing data fusion on at least two of the parameter feature information, the historical information of the manufacturing process, and the defect information. Optionally, performing data fusion includes performing data fusion on the parameter feature information and the defect information. Optionally, performing data fusion includes performing data fusion on the parameter feature information, the historical information of the manufacturing process, and the defect information. In another example, data fusion is performed on the parameter feature information and the historical information of the manufacturing process to obtain first fused data information; data fusion is performed on the first fused data information and the defect information associated with it to obtain second fused data information, the second fused data information including the glass serial number, site information, equipment information, parameter feature information, and defect information. In some embodiments, for example, data fusion is performed in the general data layer GDL by constructing a table having correlations constructed according to user needs or themes as described above.
[0205] In some embodiments, the method further comprises performing a correlation analysis. Figure 19 The defect analysis method in some embodiments of the present disclosure is shown. Figure 19 In some embodiments, the method includes extracting parameter feature information and defect information from the second fused data information; performing a correlation analysis on the parameter feature information and defect information for each type of parameter; generating multiple correlation coefficients for each of the multiple types of parameters; and sorting the absolute values of the multiple correlation coefficients. In one example, the absolute values of the multiple correlation coefficients are arranged from largest to smallest, allowing for visual observation of the relevant parameters that contribute to the occurrence of the defect. Absolute values are used here because correlation coefficients can be positive or negative, i.e., there can be a positive or negative correlation between the parameter and the defect. The larger the absolute value, the stronger the correlation.
[0206] In some embodiments, the plurality of correlation coefficients are a plurality of Pearson correlation coefficients. Optionally, each Pearson correlation coefficient is calculated according to equation (3):
[0207]
[0208] Where x represents the value of the parameter characteristic; y represents the value of the presence or absence of a defect, when a defect exists, y is assigned a value of 1, and when a defect does not exist, y is assigned a value of 0; μ x represents the average value of x; μ y represents the average value of y; σ x σ y denotes the product of the respective standard deviations of x and y; cov(x,y) denotes the covariance of x, y; and ρ(x,y) denotes the respective Pearson correlation coefficients.
[0209] In another aspect, the present disclosure provides a defect analysis method performed by a distributed computing system, the distributed computing system comprising one or more networked computers configured to execute in parallel to perform at least one common task. In some embodiments, the method includes: executing a data management platform configured to store data and extract, transform, or load data; executing a query engine connected to the data management platform and configured to obtain data directly from the data management platform; executing an analyzer connected to the query engine and configured to perform defect analysis upon receiving a task request, the analyzer comprising multiple backend servers and multiple algorithm servers configured to obtain data from the data management platform; and executing a data visualization and interaction interface configured to generate the task request.
[0210] In some embodiments, the data management platform includes an ETL module configured to extract, transform, or load data from multiple data sources into a data mart and a common data layer. In some embodiments, the method further comprises: upon receiving an assigned task by a corresponding one of the plurality of algorithm servers, directly querying the data mart for first data; and, when performing defect analysis, directly sending second data to the common data layer via the corresponding one of the plurality of algorithm servers.
[0211] In some embodiments, the method further includes generating, by the ETL module, a periodically updated dynamically updated table; and storing the dynamically updated table in the universal data layer.
[0212] In some embodiments, the software module further comprises a load balancer connected to the analyzer. In some embodiments, the method further comprises receiving, by the load balancer, a task request, distributing, by the load balancer, the task request to one or more of the plurality of backend servers to achieve load balancing among the plurality of backend servers, and distributing, by the load balancer, the tasks from the plurality of backend servers to one or more of the plurality of algorithm servers to achieve load balancing among the plurality of algorithm servers.
[0213] In some embodiments, the method also includes generating a task request by a data visualization and interactive interface; receiving the task request by a load balancer, and assigning the task request by the load balancer to one or more of a plurality of back-end servers to achieve load balancing between the plurality of back-end servers; sending a query task request by one or more of the plurality of back-end servers to a query engine; when the query engine receives the query task request from one or more of the plurality of back-end servers, the query engine queries a dynamically updated table to obtain information about defects with high incidence rates; the query engine sends the information about defects with high incidence rates to one or more of the plurality of back-end servers; sending a defect analysis task by one or more of the plurality of back-end servers to the load balancer to assign the defect analysis task to one or more of the plurality of algorithm servers to achieve load balancing between the plurality of algorithm servers; when the defect analysis task is received by one or more of the plurality of algorithm servers, one or more of the plurality of algorithm servers directly query data from the data mart to perform defect analysis; and when the defect analysis is completed, one or more of the plurality of algorithm servers send the results of the defect analysis to the common data layer.
[0214] In some embodiments, the method further includes generating a periodic task request. The periodic task request defines a recurring period for which defect analysis is to be performed. Optionally, the method further includes querying a dynamically updated table by a query engine to obtain information about defects with a high incidence rate during the recurring period; and upon receiving information about defects with a high incidence rate during the recurring period, generating a defect analysis task by one or more of the plurality of backend servers based on the information about defects with a high incidence rate during the recurring period. Optionally, the method further includes receiving input of a recurring period for which defect analysis is to be performed, for example, via an automatic task sub-interface of the data visualization and interaction interface.
[0215] In some embodiments, the method further includes generating an interactive task request. Optionally, the method further includes receiving user-defined analysis criteria through the data visualization and interaction interface; generating an interactive task request based on the user-defined analysis criteria by the data visualization and interaction interface; sending information about high-incidence defects to the data visualization and interaction interface by one or more of the multiple backend servers upon receiving information about high-incidence defects; displaying information about high-incidence defects and multiple environmental factors associated with the high-incidence defects through the data visualization and interaction interface; receiving a user-defined selection of one or more environmental factors from the multiple environmental factors through the data visualization and interaction interface; sending the user-defined selection to one or more of the multiple backend servers through the data visualization and interaction interface; and generating a defect analysis task based on the information and the user-defined selection by one or more of the multiple backend servers. Optionally, the method further includes receiving input of user-defined analysis criteria, for example, through an interactive task sub-interface of the data visualization and interaction interface, the user-defined analysis criteria including a user-defined selection of one or more environmental factors.
[0216] In some embodiments, the analyzer further includes a cache server and a cache. The cache is connected to the plurality of backend servers, the cache server, and the query engine. Optionally, the method further includes storing a portion of the defect analysis task results in the cache.
[0217] In some embodiments, the data visualization and interaction interface includes a defect visualization sub-interface. Optionally, the method further includes receiving a user-defined selection of a defect to be analyzed through the defect visualization sub-interface and generating a call request; receiving the call request by a load balancer; distributing the call request to one or more of the multiple back-end servers by the load balancer to achieve load balancing between the multiple back-end servers; sending the call request by one or more of the multiple back-end servers to a cache server; and determining by the cache server whether information about the defect to be analyzed is stored in the cache. Optionally, the method further includes, when it is determined that the information about the defect to be analyzed is stored in the cache, one or more of the multiple back-end servers being configured to send the information about the defect to be analyzed to the defect visualization sub-interface for display. Optionally, the method further includes, when determining that information about the defect to be analyzed is not stored in the cache, sending a query task request by one or more of the plurality of backend servers to the query engine; upon receiving the query task request from the one or more of the plurality of backend servers, querying the dynamically updated table to obtain information about the defect to be analyzed; sending the information about the defect to be analyzed by the query engine to the cache; storing the information about the defect to be analyzed in the cache; and sending the information about the defect to be analyzed by the one or more of the plurality of backend servers to the defect visualization sub-interface for display. Optionally, the portion of the defect analysis task results includes defect analysis task results based on periodic task requests and defect analysis task results obtained based on the query task request.
[0218] In another aspect, the present disclosure provides a computer program product for defect analysis. The computer program product for defect analysis includes a non-transitory tangible computer-readable medium having computer-readable instructions thereon. In some embodiments, the computer-readable instructions are executable by a processor in a distributed computing system so that the processor performs: executing a data management platform configured to store data and extract, convert, or load data; executing a query engine connected to the data management platform and configured to obtain the data directly from the data management platform; executing an analyzer connected to the query engine and configured to perform defect analysis upon receiving a task request, the analyzer including multiple back-end servers and multiple algorithm servers, the multiple algorithm servers configured to obtain the data directly from the data management platform; and executing a data visualization and interactive interface configured to generate a task request, wherein the distributed computing system includes one or more networked computers configured to execute in parallel to perform at least one common task.
[0219] In some embodiments, the data management platform includes an ETL module configured to extract, transform, or load data from multiple data sources into a data mart and a universal data layer. In some embodiments, the computer-readable instructions are further executable by a processor in a distributed computing system to cause the processor to: upon receiving an assigned task from a corresponding one of the plurality of algorithm servers, directly query first data from the data mart by the corresponding one of the plurality of algorithm servers; and, when performing defect analysis, directly send second data to the universal data layer via the corresponding one of the plurality of algorithm servers.
[0220] In some embodiments, the computer-readable instructions can also be executed by a processor in the distributed computing system to cause the processor to generate a periodically updated dynamically updated table by the ETL module; and store the dynamically updated table in the universal data layer.
[0221] In some embodiments, the software module further includes a load balancer connected to the analyzer. In some embodiments, the computer-readable instructions are further executable by a processor in the distributed computing system to cause the processor to perform: receiving a task request by the load balancer, distributing the task request by the load balancer to one or more of the plurality of backend servers to achieve load balancing among the plurality of backend servers, and distributing the tasks from the plurality of backend servers to one or more of the plurality of algorithm servers to achieve load balancing among the plurality of algorithm servers by the load balancer.
[0222] In some embodiments, the computer-readable instructions can further be executed by a processor in a distributed computing system so that the processor executes: generating a task request by a data visualization and interactive interface; receiving the task request by a load balancer, and assigning the task request to one or more of a plurality of back-end servers by the load balancer to achieve load balancing between the plurality of back-end servers; sending a query task request to a query engine by one or more of the plurality of back-end servers; when the query engine receives the query task request from one or more of the plurality of back-end servers, the query engine queries a dynamically updated table to obtain information about defects with high incidence rates; the query engine sends the information about defects with high incidence rates to one or more of the plurality of back-end servers; sending a defect analysis task to the load balancer by one or more of the plurality of back-end servers to assign the defect analysis task to one or more of the plurality of algorithm servers to achieve load balancing between the plurality of algorithm servers; when the defect analysis task is received by one or more of the plurality of algorithm servers, one or more of the plurality of algorithm servers directly query data from the data mart to perform defect analysis; and when the defect analysis is completed, one or more of the plurality of algorithm servers send the results of the defect analysis to a common data layer.
[0223] In some embodiments, the computer-readable instructions are further executable by a processor in a distributed computing system to cause the processor to: generate a periodic task request. The periodic task request defines a recurring period for which defect analysis is to be performed. Optionally, the computer-readable instructions are further executable by a processor in a distributed computing system to cause the processor to: query a dynamically updated table by a query engine to obtain information about defects with a high incidence rate during a recurring period; and upon receiving information about defects with a high incidence rate during a recurring period, generate a defect analysis task by one or more of the multiple back-end servers based on the information about defects with a high incidence rate during a recurring period. Optionally, the computer-readable instructions are further executable by a processor in a distributed computing system to cause the processor to: receive input of a recurring period for which defect analysis is to be performed, for example, through an automatic task sub-interface of a data visualization and interaction interface.
[0224] In some embodiments, the computer-readable instructions are further executable by a processor in a distributed computing system to cause the processor to: generate an interactive task request. Alternatively, the computer-readable instructions are further executable by a processor in the distributed computing system to cause the processor to: receive user-defined analysis criteria via a data visualization and interactive interface; generate the interactive task request based on the user-defined analysis criteria via the data visualization and interactive interface; upon receiving information about high-incidence defects, send information to the data visualization and interactive interface from one or more of the plurality of backend servers; display information about high-incidence defects and a plurality of environmental factors associated with the high-incidence defects via the data visualization and interactive interface; receive a user-defined selection of one or more environmental factors from the plurality of environmental factors via the data visualization and interactive interface; send the user-defined selection to one or more of the plurality of backend servers via the data visualization and interactive interface; and generate a defect analysis task based on the information and the user-defined selection. Alternatively, the computer-readable instructions are further executable by a processor in the distributed computing system to cause the processor to: receive input of user-defined analysis criteria, for example, via an interactive task sub-interface of the data visualization and interactive interface, the user-defined analysis criteria including a user-defined selection of one or more environmental factors.
[0225] In some embodiments, the analyzer further includes a cache server and a cache. The cache is connected to the plurality of backend servers, the cache server, and the query engine. Optionally, the computer-readable instructions can also be executed by a processor in the distributed computing system to cause the processor to perform the following steps: storing a portion of the defect analysis task results in the cache.
[0226] In some embodiments, the data visualization and interaction interface includes a defect visualization sub-interface. Optionally, the computer-readable instructions can also be executed by a processor in the distributed computing system to cause the processor to perform: receiving a user-defined selection of a defect to be analyzed through the defect visualization sub-interface and generating a call request; receiving the call request by a load balancer; distributing the call request to one or more of a plurality of back-end servers by the load balancer to achieve load balancing between the plurality of back-end servers; sending the call request by one or more of the plurality of back-end servers to a cache server; and determining by the cache server whether information about the defect to be analyzed is stored in the cache. Optionally, the computer-readable instructions can also be executed by a processor in the distributed computing system to cause the processor to perform: when it is determined that information about the defect to be analyzed is stored in the cache, one or more of the plurality of back-end servers are configured to send the information about the defect to be analyzed to the defect visualization sub-interface for display. Optionally, the computer-readable instructions can also be executed by a processor in a distributed computing system to cause the processor to perform the following steps: upon determining that information about the defect to be analyzed is not stored in the cache, one or more of the multiple backend servers sending a query task request to a query engine; upon receiving the query task request from one or more of the multiple backend servers, the query engine querying a dynamically updated table to obtain information about the defect to be analyzed; the query engine sending the information about the defect to be analyzed to the cache; storing the information about the defect to be analyzed in the cache; and one or more of the multiple backend servers sending the information about the defect to be analyzed to a defect visualization sub-interface for display. Optionally, the portion of the defect analysis task results includes defect analysis task results based on periodic task requests and defect analysis task results obtained based on the query task request.
[0227] The various illustrative operations described in conjunction with the configurations disclosed herein can be implemented as electronic hardware, computer software, or a combination of the two. These operations can be implemented or executed using a general-purpose processor, a digital signal processor (DSP), an ASIC or ASSP, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to produce the configurations disclosed herein. For example, such a configuration can be implemented at least in part as a hard-wired circuit, a circuit configuration manufactured into an application-specific integrated circuit, or a firmware program loaded into a non-volatile memory, or a software program loaded from or into a data storage medium as a machine-readable code, such code being an instruction executable by an array of logic elements such as a general-purpose processor or other digital signal processing unit. A general-purpose processor can be a microprocessor, but in an alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. The software modules may reside in a non-transitory storage medium, such as RAM (random access memory), ROM (read-only memory), non-volatile RAM (NVRAM), such as flash RAM, erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, a hard disk, a removable disk, or a CD-ROM; or in any other form of storage medium known in the art. The illustrative storage medium is coupled to the processor so that the processor can read information from and write information to the storage medium. In an alternative embodiment, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative embodiment, the processor and the storage medium may reside in a user terminal as discrete components.
[0228] The foregoing description of the embodiments of the present invention has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit the invention to the precise forms or exemplary embodiments disclosed. Therefore, the foregoing description should be considered illustrative rather than restrictive. Obviously, many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described to explain the principles of the invention and its best mode practical application, thereby enabling those skilled in the art to understand the various embodiments of the invention and various modifications as are suited to the particular use or implementation contemplated. The scope of the present invention is intended to be defined by the appended claims and their equivalents, in which all terms are to be used in their broadest reasonable sense unless otherwise indicated. Therefore, the terms "the invention," "the present invention," etc. do not necessarily limit the scope of the claims to specific embodiments, and reference to exemplary embodiments of the present invention is not intended to limit the invention, and no such limitation should be inferred. The present invention is limited solely by the spirit and scope of the appended claims. Furthermore, the claims may use the terms "first," "second," etc., followed by a noun or element. These terms should be understood as nomenclature and should not be construed as limiting the number of elements to which they refer unless a specific number is provided. Any advantages and benefits described may not apply to all embodiments of the present invention. It should be understood that those skilled in the art may make changes to the described embodiments without departing from the scope of the present invention as defined by the appended claims. In addition, no element or component in this disclosure is intended to be dedicated to the public, regardless of whether the element or component is explicitly stated in the appended claims.
Claims
1. A computer-implemented method for defect analysis, comprising: Capture defect data associated with multiple process tools during the manufacturing cycle; For the plurality of process tools, a plurality of weight of evidence (WOE) scores are calculated, wherein a higher WOE score indicates a higher correlation between the defect and the process tool; wherein each WOE score for each process tool is calculated according to equation (1): Wherein, woei represents each WOE score of each process tool; P(yi) represents the ratio of the number of positive samples in each process tool to the number of positive samples in all process tools; P(ni) represents the ratio of the number of negative samples in each process tool to the number of negative samples in all process tools; the positive samples represent data including defect information associated with each process tool; the negative samples represent data in which defect information associated with each process tool does not exist; #yi represents the number of positive samples in each process tool; #yr represents the number of positive samples in all process tools; #ni represents the number of negative samples in each process tool; #nr represents the number of negative samples in all process tools; and sorting the plurality of WOE scores to obtain a list of process tools associated with the defect occurring during the manufacturing cycle, the process tools in the list of process tools having WOE scores greater than a first threshold score, Each of the plurality of process equipments is an equipment for performing a corresponding process at a specific process site.
2. The computer-implemented method of claim 1 , further comprising: Obtain candidate contact regions of a plurality of contact devices, wherein each contact device in the plurality of contact devices is a device that is in direct contact with an intermediate product during the manufacturing cycle, and each candidate contact region includes each theoretical contact region and each edge region surrounding each theoretical contact region; respectively obtaining coordinates of defect points during the manufacturing cycle; selecting a plurality of defective contact regions from the candidate contact regions, each of the plurality of defective contact regions surrounding at least one of the defect point coordinates; A list of selected contact devices is obtained, wherein each selected contact device is a respective contact device that is in contact with the intermediate product in each defective contact region.
3. The computer-implemented method of claim 2, further comprising obtaining one or more candidate defective devices based on the list of process devices associated with the defect and the list of selected contact devices; in, Each candidate defective device of the one or more candidate defective devices is a device in the list of selected contact devices and is also a device in the list of process devices.
4. The computer-implemented method of claim 2, wherein: The candidate contact areas include point contacts and line contacts.
5. The computer-implemented method of claim 2, wherein: The intermediate product is an intermediate substrate.
6. A device for defect analysis, comprising: Memory; one or more processors; wherein the memory and the one or more processors are connected to each other; as well as The memory stores computer-executable instructions for controlling the one or more processors to perform the computer-implemented method according to any one of claims 1 to 5.
7. A computer-readable medium having computer-readable instructions thereon, the computer-readable instructions being executable by a processor to cause the processor to perform the computer-implemented method according to any one of claims 1 to 5.
8. A defect analysis system comprising: A distributed computing system comprising one or more networked computers configured to execute in parallel to perform at least one common task; one or more computer-readable storage media storing instructions that, when executed by the distributed computing system, cause the distributed computing system to execute software modules; The software modules include: a data management platform configured to extract, convert, or load raw data from a plurality of data sources into management data, wherein the raw data and the management data include defect information, and the management data is stored in a distributed manner; an analyzer configured to perform defect analysis upon receiving a task request, the analyzer comprising a plurality of algorithm servers configured to obtain the management data from the data management platform and perform algorithmic analysis on the management data to derive result data regarding potential causes of defects; and a data visualization and interactive interface configured to generate the task request and display the result data; Wherein, one or more of the plurality of algorithm servers are configured to perform the computer-implemented method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Method for predicting wafer defects
CN108122799A
Defect diagnosis with dynamic root cause detection
US20240070371A1