Data processing method and data comparison method

By preprocessing data from semiconductor manufacturing facilities and using ID convolution to segment and transform the data into the same size, the problem of data size mismatch is solved, enabling effective data learning and comparison.

CN115438013BActive Publication Date: 2026-04-10SYSTEM ENGINEERING MEGA SOLUTION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SYSTEM ENGINEERING MEGA SOLUTION CO LTD
Filing Date
2022-06-02
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In semiconductor manufacturing facilities, existing technologies cannot effectively utilize twin networks for data learning and comparison due to data size mismatch.

Method used

By preprocessing the data, ID convolution is used to segment and transform the data into the same size, suitable for the input of the Siamese network, and then the data is assembled into input data suitable for deep learning.

Benefits of technology

This enables effective data learning and comparison within twin networks, improving the efficiency and accuracy of data comparison.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115438013B_ABST
    Figure CN115438013B_ABST
Patent Text Reader

Abstract

The inventive concept provides a method for processing data generated in substrate processing. The method includes segmenting the data according to each process of the substrate processing, and converting the segmented data to a same size. Converting the segmented data to the same size includes converting the segmented data to the same size using ID convolution.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the inventive concept described herein relate to a data processing method and a data comparison method, and more particularly to a data preprocessing method applied to a Siamese network and a data comparison method using the same. BACKGROUND

[0002] Data analysis is an important matter in a semiconductor manufacturing facility because data generated by the semiconductor manufacturing facility can be used for error detection through data analysis, equipment maintenance using the data, etc. In this case, identifying data that has changed is one of the important matters. To this end, it is important to determine whether data in each corresponding facility is identical.

[0003] Figures 1-2 A data comparison method in a conventional method is shown. In the case of the prior art, when a matching rate (i.e., similarity) is calculated for log sample 1 and log sample 2 in the same process, the similarity is calculated in units of I / O data. That is, in the prior art, the identity of data is determined using a method of testing whether data of the entire process (according to each I / O) matches.

[0004] Figure 3 A problem in the data comparison method in the conventional method is shown. The data size of log sample 1 and log sample 2 to be compared in a process does not always match. In most cases, the data size does not match. To compare the identity of data, input 1 and input 2 can be respectively input to a Siamese network. However, if the size of input 1 is changed in Siamese network A, the following problem occurs: Siamese network A becomes a different Siamese network B without preserving what is learned in network A. Therefore, it is necessary to adjust the size of input 1 and input 2 in Siamese network A. SUMMARY

[0005] Embodiments of the inventive concept provide a preprocessing method of data for learning a Siamese network.

[0006] The technical objects of the inventive concept are not limited to the above-mentioned, and other technical objects not mentioned will be readily apparent to those skilled in the art from the following description.

[0007] The inventive concept provides a method for processing data generated in substrate processing. The method includes segmenting data according to each process of substrate processing, and converting the segmented data into the same size.

[0008] In an embodiment, converting the segmented data into the same size includes converting the segmented data into the same size using an ID convolution.

[0009] In an embodiment, the same size is a maximum data among the divided data.

[0010] In an embodiment, the same size is a data size for input to a twin network.

[0011] In an embodiment, the method further comprises assembling the converted data having the same size.

[0012] In an embodiment, a computer-readable recording medium having a program for executing the method is included.

[0013] The inventive concept provides a method for comparing data of a first facility and data of a second facility. The method includes pre-processing first data of the first facility, pre-processing second data of the second facility, and determining whether the pre-processed first data of the first facility and the pre-processed second data of the second facility are identical.

[0014] In an embodiment, pre-processing the first data of the first facility includes dividing the first data of the first facility according to each process of a substrate process by the first facility, and converting the divided first data into a same size.

[0015] In an embodiment, pre-processing the second data of the second facility includes dividing the second data of the second facility according to each process of a substrate process by the second facility, and converting the divided second data into a same size.

[0016] In an embodiment, converting the divided first data and the divided second data into the same size respectively includes converting the divided first data and the divided second data into the same size respectively using ID convolution.

[0017] In an embodiment, the same size is a maximum data among the divided first data and the divided second data respectively.

[0018] In an embodiment, the same size is a data size for input to a twin network.

[0019] In an embodiment, the method further comprises assembling the converted first data.

[0020] In an embodiment, the method further comprises assembling the converted second data.

[0021] In an embodiment, the determining whether the pre-processed first data of the first facility and the pre-processed second data of the second facility are identical includes determining using a twin network.

[0022] In an embodiment, the method includes determining data similarity between the assembled first data and the assembled second data by a twin network.

[0023] In an embodiment, a computer-readable recording medium having a program for executing the method is included.

[0024] According to an embodiment of the inventive concept, by proposing a preprocessing method for learning data of a twin network, it is possible to solve a problem that learning using deep learning cannot be performed because many data are provided in different lengths.

[0025] Effects of the inventive concept are not limited to what has been described above and other effects which are not mentioned will also become apparent to those skilled in the art from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0026] The above and other objects and features will become apparent from the following description, taken in conjunction with the accompanying drawings, wherein:

[0027] Figures 1-2 A data comparison method in a conventional method is shown.

[0028] Figure 3 A problem in the data comparison method in the conventional method is shown.

[0029] Figures 4-5 A data comparison method according to the inventive concept is shown.

[0030] Figure 6 Data comparison by the data comparison method according to the inventive concept is shown.

[0031] Figure 7 An embodiment of ID convolution is shown.

[0032] Figure 8 Data assembly according to the data comparison method according to an embodiment of the inventive concept is shown.

[0033] Figure 9 An intermediate process of Figure 8 is shown.

[0034] Figure 10 Application of assembled data according to the inventive concept is shown.

[0035] Figure 11 is a flowchart showing a data processing method according to an embodiment of the inventive concept.

[0036] Figure 12 is a flowchart showing a data comparison method according to an embodiment of the inventive concept. Detailed Implementation

[0037] The inventive concept can be modified and taken in various forms, and specific embodiments thereof will be shown and described in detail in the accompanying drawings. The embodiments are provided to explain the inventive concept more fully to those skilled in the art. However, the embodiments according to the inventive concept are not intended to limit the specific forms disclosed, and it should be understood that the inventive concept includes all variations, equivalents, and substitutions included within the spirit and scope of the inventive concept. In the description of the inventive concept, detailed descriptions of relevant known technologies may be omitted where this might obscure the essence of the inventive concept. Furthermore, in all the drawings, the same reference numerals are used for parts having similar functions and actions.

[0038] The terminology used herein is for describing only particular embodiments and is not intended to limit the inventive concept. It should also be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of the described features, integrals, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or combinations thereof.

[0039] The absence of a specified quantity does not preclude a majority unless explicitly stated in the context. Additionally, the shape and size of elements in the accompanying drawings may be exaggerated for clarity.

[0040] Unless otherwise defined, all terms used herein (including technical or scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Terms (such as those defined in commonly used dictionaries) should be interpreted in accordance with the context of the relevant art and should not be idealized or over-formalized unless clearly defined herein.

[0041] Figures 4-5 A data comparison method based on the present invention is illustrated.

[0042] refer to Figure 4 This paper discloses an embodiment of comparing two different data samples, Sample 1 and Sample 2. Sample 1 and Sample 2 show different result samples obtained with respect to the same process. According to one embodiment, Sample 1 may be result data processed in a first substrate processing facility. According to another embodiment, Sample 2 may be result data processed in a second substrate processing facility. In other words, Sample 1 and Sample 2 may be samples obtained from the same processing process performed by the first substrate processing facility and the second substrate processing facility, respectively.

[0043] Sample 1 and sample 2 can be time series data of each processing facility, respectively. Since the processing time of each processing facility varies from time to time, the size of data can not match when the data is compared by itself. To solve this problem, in the inventive concept, when it is determined whether sample 1 and sample 2 are the same, the data samples can be cut and compared according to each process step of each data sample. However, referring to Figure 4 , there can be a problem that the length of each corresponding process step also does not match.

[0044] According to the inventive concept, in a method for cutting data according to each corresponding process step and for adjusting the length of data in a substantially equal manner according to each corresponding process step, the data can be converted using ID convolution. This will be described in more detail with reference to Figure 6 and Figure 7 .

[0045] Referring to Figure 4 , sample data 1 and sample data 2 can be cut according to each process step, respectively. According to an embodiment of Figure 4 , sample data 1 and sample data 2 can be cut into three process steps. Referring to Figure 4 , it can be seen that the length of data cut according to each process step for sample data 1 is different from the length of data cut according to each process step for sample data 2.

[0046] In order to adjust the length of data that is not the same according to each step, the length of data in each process step can be converted by ID convolution to adjust the length of data, respectively.

[0047] Conventionally, there can be a problem that learning through a twin network can not be performed when the size of data is not adjusted, but in the inventive concept, data preprocessing can be performed through a length adjustment operation to facilitate learning through a twin network. According to an embodiment, since there are nearly 1000 pieces of data, the more the number of data, the higher the comparison efficiency of data.

[0048] Referring to Figure 5 , data cut from sample data 1 according to a process step and data cut from sample data 2 according to a process step can be selected to be compared, and only data corresponding to the step can be converted in size.

[0049] According to an embodiment, when comparing current data with normal data to find an abnormal I / O, data of a specific I / O can be selected. Referring to Figure 5, the selected process step can be step 2. That is, the transformation can be performed only on the data of the process step for which it is determined whether a change has occurred, and not on all steps. In the following embodiments, it will be assumed that the data size transformation is performed for all process steps and described as an example.

[0050] Referring to Figure 4 and Figure 5 , the size of the data divided according to each process step can not be the same. The condition for determining the reference size for the resizing can be as follows. According to an embodiment, it can be transformed into a size suitable for input into a twin network. That is, it can be transformed into a size matching the input size of a deep learning neural network. According to another embodiment, it can be transformed into a size matching the maximum size of the data of the process step selected from among the process steps.

[0051] In this case, the data transformation can be performed using an ID convolution.

[0052] Figure 6 A data comparison according to the data comparison method according to the inventive concept is shown.

[0053] Referring to Figure 6 , the sample data 1 (top) and the sample data 2 (bottom) can be divided according to each process step. Even within the sample data 1, the size of the data according to each process step can be different, and even within the corresponding process steps of the sample data 1 and the sample data 2, the data size can be different. For their matching, input values having a fixed size can be created by an ID convolution. This can be processed as input values of a twin network.

[0054] Figure 7 An embodiment of an ID convolution is shown.

[0055] Referring to Figure 7 , time series data according to each process step is disclosed. Referring to Figure 7 , an embodiment of transforming the data in step 4 by an ID convolution is disclosed. In this case, the parameters of the ID convolution can be provided as fixed values. In this way, by performing data transformation on the data according to each process step using an ID convolution, a result value having a fixed size can be obtained. In Figure 7 , the data values and the parameter values in the ID convolution are only embodiments, and the parameter values of the ID convolution can be differently applied according to the features to be extracted.

[0056] Figure 8 Data assembly according to the data comparison method according to the embodiment of the inventive concept is shown.

[0057] Referring to Figure 8, with the effect that time series data of a substrate processing facility can be converted into pre-processed image data form to be inputted into a twin network. According to an embodiment, all input sizes can be set to be the same to be converted, and one input data can be set by assembling according to each process step. The same process step performed in different facilities can result in multiple input data. Alternatively, the same process step performed in the same facility with a time difference can also result in multiple input data.

[0058] As shown in Figure 8 , after converting the length of the process steps having different lengths respectively, they can be assembled into one plane to become the optimal shape that can be inputted into deep learning, thus making the application easier.

[0059] Figure 9 The intermediate process of converting from 2D to 3D in Figure 8 is shown. Figure 10 The application of assembled data according to the inventive concept is shown.

[0060] According to the inventive concept, when calculating the matching rate (similarity) for the log data 1 and the log data 2 of the facility, by testing according to the segmented process steps, the similarity reflecting the data characteristics for each process step can be calculated, while the conventional way is to test according to one process unit.

[0061] Referring to Figure 9 , in the conventional art, the similarity is calculated by dividing the log data into I / O data units and then further subdividing according to each process step. When performing the test, the starting point of each process step can be the same. According to the inventive concept, it can be tested whether the data according to the process step matches according to each I / O data. According to the inventive concept, when checking the similarity of the log data of two facilities through deep learning, the similarity of the log data can be divided according to I / O and each process step and checked through data preprocessing.

[0062] Referring to Figure 9 , in the first figure, when the process data is divided according to each process step, the data size of each step is not constant. Referring to Figure 10 , the size of input 1 and input 2 of the twin network is the same and will be constantly provided. According to the inventive concept, after the preprocessing process, the size of the different process step data can become the same as the input size.

[0063] Referring again to Figure 9of FIG. 2, when pre-processing log data of a facility, process step data of various sizes can be converted into an input size for deep learning to be pre-processed. In this case, the input size for pre-processing can be set to a maximum value among sizes of the process step data. Alternatively, pre-processing can be performed in a size suitable for an input size of deep learning to be applied. When converting log data of a facility into an input size, conversion can be performed using ID convolution.

[0064] Referring to Figure 9 of FIG. 3, process step data obtained through first pre-processing (ID convolution) can be merged according to each I / O to generate 2D data. That is, by assembling the process step data, a state for input data can be entered. Referring to Figure 10 , by combining 2D data of the remaining I / Os, 3D data can be finally set as one input data.

[0065] Figure 11 is a flowchart illustrating a data processing method according to an embodiment of the inventive concept.

[0066] A method of processing data generated during a substrate processing process can include a step for dividing data according to each process step, a step for converting data divided according to each process step into the same size through ID convolution, and a step for assembling data divided according to each process step. In this way, data divided according to each process step can be converted into the same size, and by assembling the data, they can be converted into input data suitable for deep learning.

[0067] Figure 12 is a flowchart illustrating a data comparison method according to an embodiment of the inventive concept.

[0068] For convenience, each data of a first facility and a second facility will be described as an example of comparison. In this case, a step for pre-processing first data of the first facility, a step for pre-processing second data of the second facility, and a step for determining whether the pre-processed data of the first facility and the pre-processed data of the second facility are the same can be included.

[0069] The step for pre-processing first data of the first facility can include a step for dividing the first data of the first facility according to each process step, and a step for converting the first data divided according to each process step into the same size.

[0070] The step for pre-processing second data of the second facility can include a step for dividing the second data of the second facility according to each process step, and a step for converting the second data divided according to each process step into the same size.

[0071] In this case, the size conversion can be converted using an ID convolution.

[0072] Each converted data can be assembled according to each of the first data and the second data, and can be provided in the form of 2D data. The assembled data can be input as an input value of a twin network and used to determine similarity.

[0073] According to the present inventive concept, by proposing a preprocessing method for data used to learn a twin network, it is possible to solve a problem that learning cannot be performed using deep learning because many data are provided in different lengths.

[0074] Meanwhile, the data processing method and the data comparison method according to the embodiments of the present inventive concept described above can be implemented in the form of program commands that can be executed by various computer means and recorded in a computer-readable recording medium. In this case, the computer-readable recording medium can include program commands, data files, data structures, etc. alone or a combination thereof. Meanwhile, the program commands recorded on the recording medium can be designed and configured specifically for the present inventive concept, or known and usable to those skilled in the computer software field.

[0075] The computer-readable recording medium can include hardware devices specially configured to store and execute program instructions, such as magnetic media such as a hard disk, a floppy disk, and a magnetic tape, optical media such as a CD-ROM and a DVD, magneto-optical media such as a floptical disk and a ROM, RAM, a flash memory, etc. In addition, the program instructions include machine language codes such as compiler-generated codes, and high-level language codes that can be executed by a computer using an interpreter, etc. The above-described hardware devices can be configured to operate one or more software modules to perform the operations of the present inventive concept.

[0076] Effects of the present inventive concept are not limited to the above-described effects, and those skilled in the art to which the present inventive concept pertains can clearly understand other unmentioned effects from the above description and the attached drawings.

[0077] Although preferred embodiments of the present inventive concept have been illustrated and described with respect to the accompanying drawings, the present inventive concept is not limited to the specific embodiments described above, and it should be noted that those skilled in the art to which the present inventive concept pertains can implement the present inventive concept in various ways without departing from the spirit and scope of the present inventive concept as claimed in the claims, and modifications made should not be interpreted as departing from the technical spirit or prospect of the present inventive concept.

Claims

1. A method for processing data generated in a substrate process, the method comprising: segmenting the data according to each process of the substrate process; converting the segmented data to a same size with ID convolution; and assembling the converted data with the same size.

2. The method of claim 1, wherein the same size is a maximum data among the segmented data.

3. The method of claim 1, wherein the same size is a data size for input to a twin network.

4. A computer-readable recording medium having a program for executing the method according to any one of claims 1 to 3.

5. A method for comparing data of a first facility and data of a second facility, the method comprising: preprocessing first data of the first facility, wherein preprocessing first data of the first facility comprises: segmenting the first data of the first facility according to each process of a substrate process by the first facility; and converting the segmented first data to a same size with ID convolution; preprocessing second data of the second facility, wherein preprocessing second data of the second facility comprises: segmenting the second data of the second facility according to each process of a substrate process by the second facility; and converting the segmented second data to a same size with ID convolution; determining whether the preprocessed first data of the first facility and the preprocessed second data of the second facility are identical; and assembling the converted first data and the converted second data.

6. The method of claim 5, wherein the same size is a maximum data among the segmented first data and the segmented second data, respectively.

7. The method of claim 5, wherein the same size is a data size for input to a twin network. determining with a twin network.

8. The method of claim 5, the determining whether the first data of the first facility after pre-processing is identical to the second data of the second facility after pre-processing comprises: determining data similarity between the assembled first data and the assembled second data by the twin network.

9. The method of claim 8, further comprising:

10. A computer-readable recording medium having a program for executing the method according to any one of claims 5 to 9. ​

Citation Information

Patent Citations

  • Enhanced siamese trackers

    US20180129934A1

  • Data processing method, charged particle beam writing apparatus, and charged particle beam writing system

    US20180366297A1