Semiconductor defect classification model training method and defect classification method

By conducting multi-model fusion training on semiconductor defect pictures, a semiconductor defect classification model suitable for classification of common, common and uncommon defects is obtained, which solves the problem of insufficient classification in the existing technology, and achieves more efficient defect classification.

CN120182741APending Publication Date: 2025-06-20SHANGHAI INTEGRATED CIRCUIT RESEARCH & DEVELOPMENT CENTER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311762729.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-19
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The semiconductor defect classification model in the prior art is only applicable to the classification of common workpiece defects, and it is impossible to effectively classify uncommon workpiece defects, resulting in the classification being incomplete and accurate enough.

Method used

By obtaining multiple defect pictures and category information showing different semiconductor defects, the initial converter model is trained until the preset training conditions are met, and the target semiconductor defect classification model is obtained. This model combines the weights of multiple initial classification models and integrates the advantages of different computational models, and is suitable for classification common, common and uncommon defects.

Benefits of technology

The accurate classification of common, commonly common and uncommon defects of semiconductors is achieved, which improves the accuracy and comprehensiveness of defect classification and reduces the time and cost of manual classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182741A_ABST
    Figure CN120182741A_ABST
Patent Text Reader

Abstract

The invention provides a training method of a semiconductor defect classification model and a defect classification method, and the method comprises the steps: obtaining a plurality of defect pictures which display different defects of a semiconductor, and the defect class information of each defect picture, taking the defect pictures as an input quantity, and taking the defect class information as an output quantity; and training the initial converter model until a first preset training condition is met, and obtaining a trained initial classification model. And carrying out weight training on the trained initial classification model based on a preset loss function until a second preset training condition is met to obtain a target semiconductor defect classification model, and classifying common, generally common or uncommon defects of the semiconductor through the target semiconductor defect classification model. And the accuracy and comprehensiveness of semiconductor defect classification are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of semiconductor production, and in particular, to a method for training a semiconductor defect classification model and a method for classifying defects. Background Art

[0002] In the process of semiconductor manufacturing, various semiconductor defects will occur on the wafer due to factors such as abnormal semiconductor production equipment or dust in the production environment. By detecting and classifying semiconductor defects, the production yield of semiconductors can be effectively improved.

[0003] In the prior art, when classifying semiconductor defects, a defect image of the semiconductor can be obtained through computer vision technology, and then a single classification model, such as a convolutional neural network, an autoencoder, a recurrent neural network, etc., is used to determine the defect category to which the semiconductor belongs according to the defect image.

[0004] However, the classification models in the prior art are often only applicable to the classification of common workpiece defects. Therefore, there is an urgent need for a model and method that can also classify uncommon workpiece defects at the same time. Summary of the Invention

[0005] The present application provides a method for training a semiconductor defect classification model and a method for classifying defects, so as to solve the technical problem that the classification models in the prior art are only applicable to the classification of common workpiece defects and the defect classification is not comprehensive and accurate enough.

[0006] In a first aspect, the present application provides a method for training a semiconductor defect classification model, including:

[0007] Obtaining multiple defect pictures showing different defects of the semiconductor, and defect category information to which each defect picture belongs;

[0008] Using the defect pictures as input quantities and the defect category information as output quantities to train an initial converter model until a first preset training condition is met, and obtaining a trained initial classification model;

[0009] Wherein, the initial classification model includes: a first initial classification model that satisfies a preset first loss function, a second initial classification model that satisfies a preset second loss function, and a third initial classification model that satisfies a preset third loss function;

[0010] Based on a preset fourth loss function and a machine learning algorithm, performing weight training on the trained initial classification model until a second preset training condition is met, and obtaining a target semiconductor defect classification model.

[0011] Optionally, for the method described above, when training the initial converter model with the defective picture as the input quantity and the defective category information as the output quantity until the first preset training condition is met to obtain the trained initial classification model, it includes:

[0012] Taking the defective picture as the input quantity and the defective category information as the output quantity, training the initial converter model based on a preset cross-entropy loss function until the first preset number of training times is reached to obtain the trained first initial classification model;

[0013] Taking the defective picture as the input quantity and the defective category information as the output quantity, training the initial converter model based on a preset class balance loss function until the first preset number of training times is reached to obtain the trained second initial classification model;

[0014] Taking the defective picture as the input quantity and the defective category information as the output quantity, training the initial converter model based on a preset reverse weighted loss function until the first preset number of training times is reached to obtain the trained third initial classification model;

[0015] The trained initial classification model includes: the first initial classification model, the second initial classification model, and the third initial classification model.

[0016] Optionally, for the method described above, the

[0017] Based on a preset fourth loss function and a machine learning algorithm, training the weights of the trained initial classification model until the second preset training condition is met to obtain the target semiconductor defect classification model, including:

[0018] Obtaining the first weight preset for the first initial classification model, the second weight preset for the second initial classification model, and the third weight preset for the third initial classification model;

[0019] According to the first weight, the second weight, the third weight, and the initial classification model, establishing an initial semiconductor defect classification model with the defective picture as the input and the defective category information as the output;

[0020] Obtaining multiple defective pictures showing semiconductor defects;

[0021] Inputting each defective picture into the initial semiconductor defect classification model to obtain the defective category information corresponding to each defective picture;

[0022] Based on the defect category information corresponding to each defect picture, and based on a preset fourth loss function and a machine learning algorithm, weight training is performed on the initial semiconductor defect classification model until the second preset number of training times is reached, and the target semiconductor defect classification model is obtained.

[0023] Optionally, in the method as described above, the obtaining multiple defect pictures showing semiconductor defects includes:

[0024] Obtain M defect pictures showing semiconductor defects, where M is an integer greater than or equal to 1;

[0025] Perform data augmentation processing on the M defect pictures showing semiconductor defects to obtain N defect pictures showing semiconductor defects after data augmentation processing, where N is an integer greater than M.

[0026] Optionally, in the method as described above, the performing weight training on the initial semiconductor defect classification model according to the defect category information corresponding to each defect picture, based on a preset fourth loss function and a machine learning algorithm, includes:

[0027] According to the defect category information corresponding to each defect picture, based on a preset marginal entropy loss function or mutual information loss function, and a preset neural network algorithm or least squares method or maximum likelihood estimation algorithm, perform weight training on the initial semiconductor defect classification model.

[0028] In a second aspect, the present application provides a method for classifying semiconductor defects, including:

[0029] Obtain a defect picture showing a defect of a target semiconductor;

[0030] Input the defect picture into a target semiconductor defect classification model to obtain the target defect category to which the defect picture belongs, where the target semiconductor defect classification model is trained according to the training method of the semiconductor defect classification model according to any item in the first aspect.

[0031] Optionally, in the method as described above, the obtaining a defect picture showing a defect of a target semiconductor includes:

[0032] Obtain original defect pictures of the target semiconductor taken from multiple angles;

[0033] Perform picture processing on the original defect pictures of the target semiconductor taken from multiple angles to obtain a defect picture showing a defect of the target semiconductor.

[0034] In a third aspect, the present application provides a training device for a semiconductor defect classification model, including:

[0035] An acquisition module, configured to acquire multiple defect images showing different defects of a semiconductor, and defect category information to which each defect image belongs;

[0036] A training module, configured to use the defect images as input quantities and the defect category information as output quantities to train an initial converter model until a first preset training condition is met, thereby obtaining a trained initial classification model;

[0037] Wherein, the initial classification model includes: a first initial classification model that satisfies a preset first loss function, a second initial classification model that satisfies a preset second loss function, and a third initial classification model that satisfies a preset third loss function;

[0038] The training module is further configured to perform weight training on the trained initial classification model based on a preset loss function and a machine learning algorithm until a second preset training condition is met, thereby obtaining a target semiconductor defect classification model.

[0039] In a fourth aspect, the present application provides a semiconductor defect classification device, including:

[0040] An acquisition module, configured to acquire a defect image showing a defect of a target semiconductor;

[0041] A processing module, configured to input the defect image into a target semiconductor defect classification model to obtain a target defect category to which the defect image belongs, where the target semiconductor defect classification model is trained according to the training method of the semiconductor defect classification model according to any one of the first aspect.

[0042] In a fifth aspect, the present application provides an electronic device, including:

[0043] A processor and a memory;

[0044] The memory is configured to store executable instructions of the processor;

[0045] Wherein, the processor is configured to execute the training method of the semiconductor defect classification model according to any one of the first aspect and / or execute the classification method of the semiconductor defect according to the second aspect by executing the executable instructions.

[0046] In a sixth aspect, the present application provides a readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it is used to implement the training method of the semiconductor defect classification model according to any one of the first aspect and / or execute the classification method of the semiconductor defect according to the second aspect.

[0047] In a seventh aspect, the present application provides a computer program product, including a computer program which, when executed by a processor, is used to implement the training method of the semiconductor defect classification model according to any one of the first aspect, and / or execute the classification method of the semiconductor defect according to the second aspect.

[0048] The training method of the semiconductor defect classification model and the classification method of the defect provided by the present application obtain multiple defect pictures showing different defects of the semiconductor, as well as the defect category information to which each defect picture belongs. Taking the defect pictures as input quantities and the defect category information as output quantities, the initial converter model is trained until the first preset training condition is met, and the trained initial classification model is obtained. Among them, the initial classification model includes: a first initial classification model that satisfies a preset first loss function, a second initial classification model that satisfies a preset second loss function, and a third initial classification model that satisfies a preset third loss function. And based on a preset fourth loss function and a machine learning algorithm, weight training is performed on the trained initial classification model until the second preset training condition is met, and the target semiconductor defect classification model is obtained. Through this target semiconductor defect classification model, common, generally common or uncommon defects of the semiconductor can be classified, improving the accuracy and comprehensiveness of defect classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0050] Figure 1 It is a schematic flowchart of a training method of a semiconductor defect classification model provided by an embodiment of the present application;

[0051] Figure 2 It is a schematic diagram of a general network framework of a SWin Transformer provided by an embodiment of the present application;

[0052] Figure 3 It is a schematic diagram of training an initial classification model provided by an embodiment of the present application;

[0053] Figure 4 It is a schematic flowchart of a method for obtaining a trained initial classification model provided by an embodiment of the present application;

[0054] Figure 5 It is a schematic flowchart of a method for obtaining a target semiconductor defect classification model provided by an embodiment of the present application;

[0055] Figure 6 It is a schematic diagram of weight training provided by an embodiment of the present application;

[0056] Figure 7 The structural schematic diagram of a training device for a semiconductor defect classification model provided by an embodiment of the present application;

[0057] Figure 8 The structural schematic diagram of a classification device for semiconductor defects provided by an embodiment of the present application;

[0058] Figure 9 The structural schematic diagram of an electronic device provided by an embodiment of the present application.

[0059] Through the above-mentioned drawings, the specific embodiments of the present application have been shown, and more detailed descriptions will be provided later. These drawings and text descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Specific Embodiments

[0060] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0061] In the description of the embodiments of the present application, "first", "second", "third", etc. are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can also be implemented in an order other than those illustrated or described in the present application. Terms such as "inside", "outside", etc., indicating directions or positional relationships, are based on the directions or positional relationships shown in the drawings, which are only for the convenience of description and do not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation of the present application.

[0062] In addition, in the description of the embodiments of the present application, unless otherwise clearly specified and limited, the terms "connected" and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the embodiments of the present application can be understood according to specific circumstances.

[0063] During the semiconductor manufacturing process, due to factors such as abnormalities in the equipment for producing semiconductors or a relatively large amount of fine dust in the workshop environment for semiconductor production, various defects will occur on the semiconductor wafers. By detecting and classifying the defects generated in the semiconductor, the corresponding causes can be found according to the defect categories, thereby improving the production yield of the semiconductor.

[0064] Currently, when classifying semiconductor defects, although there is also a method based on the work experience of staff to determine the defect category to which the semiconductor belongs according to the staff's observation and judgment of the defects. However, due to the large number of semiconductor defect types, this method relies too much on manpower, resulting in a long time-consuming process for determining the defect category to which the semiconductor belongs and a low degree of intelligence, which affects production efficiency.

[0065] In order to reduce the time consumption and cost caused by manual classification of defect categories, there is also a method in the prior art that uses a classification model, that is, uses computer vision technology and combines common classification algorithms, such as convolutional neural networks, autoencoders, recurrent neural networks, etc. to determine the defect category to which the semiconductor belongs. However, in the actual application process, due to the different probabilities of occurrence of different defect categories, the classification models in the prior art are often more suitable for classifying common defects with a higher probability of occurrence, while the classification of general common or uncommon defects is not accurate enough.

[0066] Therefore, in view of the above technical problems in the prior art, the inventors found in the research on defect classification that by setting different calculation models, using different calculation models to classify defects with different probabilities of occurrence, and by setting the weights corresponding to different calculation models to integrate different calculation models, a semiconductor defect classification model can be obtained that can finally output the accurate defect category regardless of the type of defect that occurs in the semiconductor according to the set different operation models. Based on this, the present application proposes a training method for a semiconductor defect classification model and a classification method for defects.

[0067] The following will specifically describe the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems with specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0068] Figure 1 It is a schematic flowchart of a training method for a semiconductor defect classification model provided by an embodiment of the present application. The execution subject of this method can be a semiconductor production device or a computer, server, or server cluster that can interact with the semiconductor production device, etc. The method in this embodiment can be implemented in a software, hardware, or software-hardware combination manner. Such as Figure 1As shown in the figure, the method specifically includes the following steps:

[0069] S101. Obtain multiple defect pictures showing different defects of semiconductors, and defect category information to which each defect picture belongs.

[0070] In this embodiment, taking a computer as an example of the execution subject, the computer can obtain multiple original defect pictures of semiconductors from a database storing production images of semiconductors, and preprocess the original defect pictures.

[0071] A possible way of preprocessing is as follows: Obtain original defect pictures of each semiconductor taken from multiple angles, perform image processing on the original defect pictures of each semiconductor to obtain multiple complete defect pictures showing different defects of the semiconductor, where the image processing includes but is not limited to: splicing processing, enhancement processing, synthesis processing, etc.

[0072] For the same semiconductor, when the image acquisition device acquires images of the defective part, due to different shooting light angles of the image acquisition device, there will be multiple original defect pictures of the same semiconductor with different angles. For example, there will be an original defect magnified picture from the left light perspective, an original defect magnified picture from the top light perspective, and an original defect magnified picture from the right light perspective. By integrating multiple original defect pictures with different angles, a picture showing the complete defect can be obtained.

[0073] In this embodiment, since the original defect pictures have been pre-annotated with the defect categories to which they belong when stored, after preprocessing, defect pictures of each defective semiconductor and defect category information to which the defect pictures belong can be obtained. In this embodiment, the specific categories of defects are not limited, and they include but are not limited to: particle defects, scratch defects, bulge defects, etc.

[0074] S102. Use the defect pictures as input quantities and the defect category information as output quantities to train the initial converter model until the first preset training condition is met, and obtain the trained initial classification model.

[0075] After obtaining multiple defect pictures showing different defects of semiconductors through step 101, they can be scaled by a preset ratio to improve the training efficiency.

[0076] In this embodiment, the initial converter model Transformer can be evolved from the original framework of the SWin Transformer general network based on the moving window self-attention mechanism. The original framework contains a total of 4 transformation stages, namely Transformer Stage. In this embodiment, the first 3 Transformer Stages are used as the shared network, and the last Transformer Stage is copied three times to obtain three expert networks, namely Expert Network 1, Expert Network 2, and Expert Network 3. The schematic diagram of its structure is as Figure 2 shown, Figure 2 This is a schematic diagram of the SWin Transformer general network framework provided by the embodiment of the present application. The shared network is used to transmit training data, and the expert network is used to train models for defect categories with different occurrence probabilities based on the transmitted training data, and finally three initial classification models are obtained.

[0077] Among them, the initial classification models include: a first initial classification model that satisfies a preset first loss function, a second initial classification model that satisfies a preset second loss function, and a third initial classification model that satisfies a preset third loss function. The first initial classification model can be used to classify defect categories with a relatively high occurrence probability, the second initial classification model can be used to classify defect categories with a medium occurrence probability, and the third initial classification model can be used to classify defect categories with a relatively low occurrence probability.

[0078] A defect category with a relatively high occurrence probability indicates that the defect is a common defect, a defect category with a medium probability indicates that the defect is a generally common defect, and a defect category with a relatively low occurrence probability indicates that the defect is an uncommon defect.

[0079] In summary, the schematic diagram as Figure 3 shown, Figure 3 This is a schematic diagram of the training of an initial classification model provided by the embodiment of the present application. For the specific implementation method, please refer to the content in the above steps and will not be repeated here.

[0080] S103. Based on a preset fourth loss function and a machine learning algorithm, perform weight training on the trained initial classification model until the second preset training condition is met, and a target semiconductor defect classification model is obtained.

[0081] In this embodiment, a preset fourth loss function and a machine learning algorithm are used to perform weight training on the trained initial classification model to enhance the processing ability of the finally obtained target semiconductor defect classification model, so that it can classify common, generally common, or uncommon defects of semiconductors at the same time, thereby improving the accuracy of defect classification.

[0082] In the above embodiments of the present application, by obtaining multiple defect pictures showing different defects of the semiconductor and the defect category information to which each defect picture belongs, using the defect pictures as input quantities and the defect category information as output quantities, training the initial converter model r until the first preset training condition is met, a trained initial classification model is obtained, and then based on a preset fourth loss function and machine learning algorithm, weight training is performed on the trained initial classification model until the second preset training condition is met, a target semiconductor defect classification model is obtained, and the accuracy and comprehensiveness of defect classification are improved through this target semiconductor defect classification model.

[0083] Further, on the basis of the above embodiments, the following embodiments illustrate the process of using defect pictures as input quantities and defect category information as output quantities to train the initial converter model until the first preset training condition is met and a trained initial classification model is obtained.

[0084] Figure 4 The flowchart of a method for obtaining a trained initial classification model provided by an embodiment of the present application is shown as Figure 4 shown, and the method includes:

[0085] S401. Use the defect pictures as input quantities and the defect category information as output quantities, and based on a preset cross-entropy loss function, train the initial converter model until the first preset number of training times is reached, and a trained first initial classification model is obtained.

[0086] The preset cross-entropy loss function is shown in formula (1):

[0087]

[0088] where x i represents the i-th defect picture after preprocessing; y i represents the defect category to which the i-th defect picture belongs; f θ is a shared network; Φ1 represents the decision network of the expert network 1, that is, the first initial classification model; represents the cross-entropy function; n s represents the total number of all defect picture training samples.

[0089] This cross-entropy loss function is friendly to defect categories with higher occurrence probabilities. Therefore, the first initial classification model is suitable for classifying common defects with higher occurrence probabilities.

[0090] S402. Use the defect image as the input quantity and the defect category information as the output quantity. Based on the preset class balance loss function, train the initial transformer model until the first preset number of training times is reached, and obtain the trained second initial classification model.

[0091] The preset class balance loss function is shown in formula (2):

[0092]

[0093] Among them, x i represents the i-th defect image after preprocessing; y i represents the defect category to which the i-th defect image belongs; f θ represents the shared network; Φ2 represents the decision network of the expert network 2, that is, the second initial classification model; represents the cross-entropy function; n s represents the total number of training samples of all defect images; n k represents the number of k training samples of the defect category to which the i-th defect image belongs.

[0094] This class balance loss function uses the sample frequency of each category for balancing and is friendly to defect categories with medium occurrence probabilities. Therefore, the second initial classification model is suitable for classifying common defects with medium occurrence probabilities.

[0095] S403. Use the defect image as the input quantity and the defect category information as the output quantity. Based on the preset inverse weighted loss function, train the initial transformer model until the first preset number of training times is reached, and obtain the trained third initial classification model.

[0096] The preset inverse weighted loss function is shown in formula (3):

[0097]

[0098] Among them, x i represents the i-th defect image after preprocessing; y i represents the defect category to which the i-th defect image belongs; f θ represents the shared network; Φ3 represents the decision network of the expert network 3, that is, the third initial classification model; represents the cross-entropy function; n s represents the total number of training samples of all defect images; n k represents the number of k training samples of the category to which the i-th defect image belongs; c represents the total number of categories; n c-k represents the number of training samples of the symmetric category of the category to which the i-th defect image belongs after arranging the training samples from small to large.

[0099] The reverse weighted loss function uses the sample frequency of each class for balancing, and then further uses the reverse sample frequency for reverse adjustment to be friendly to the defect classes with lower occurrence probabilities. Therefore, the third initial classification model is suitable for classifying defects with lower occurrence probabilities.

[0100] In this embodiment, the execution order between steps S401 - S403 is not limited.

[0101] In the above - mentioned embodiment of the present application, by taking the defect pictures as the input quantity, the defect category information as the output quantity, and training the initial transformer model respectively based on the preset cross - entropy loss function, class balance loss function, and reverse weighted loss function, the trained initial classification model is obtained. Among them, the initial classification model includes: the first initial classification model, the second initial classification model, and the third initial classification model. In this embodiment, different initial classification models are used to classify the defect categories to which the defect pictures belong, improving the accuracy of the classification results.

[0102] Furthermore, through Figure 5 the illustrated embodiment describes the process of performing weight training on the trained initial classification model based on the preset fourth loss function and machine learning algorithm until the second preset training condition is met, and then obtaining the target semiconductor defect classification model.

[0103] Figure 5 It is a schematic flowchart of a method for obtaining a target semiconductor defect classification model provided by an embodiment of the present application, as Figure 5 shown:

[0104] S501. Obtain the first weight preset for the first initial classification model, the second weight preset for the second initial classification model, and the third weight preset for the third initial classification model.

[0105] In this embodiment, the first weight w1, the second weight w2, and the third weight w3 can be generated by customizing or randomly. Among them, w1 + w2 + w3 = 1.

[0106] S502. According to the first weight, the second weight, the third weight, and the initial classification model, establish an initial semiconductor defect classification model with defect pictures as the input and defect categories as the output.

[0107] The initial semiconductor defect classification model established according to w1, w2, w3, and the initial classification model is shown in formula (4):

[0108] y = w1·Φ1(f θ (x)) + w2·Φ2(f θ (x)) + w3·Φ3(fθ (x)), i ∈ {1, 2, ..., M} (4)

[0109] Wherein, x represents the input defective image, and y represents the defect category to which the defective image belongs.

[0110] S503. Obtain multiple defective images showing semiconductor defects.

[0111] In this embodiment, M defective images showing semiconductor defects can be selected from the multiple defective images showing different semiconductor defects obtained in step S101, where M is an integer greater than or equal to 1. Alternatively, multiple original defective images of the semiconductor can be obtained from the database storing the production images of the semiconductor, and the original defective images are preprocessed to obtain M defective images showing semiconductor defects. There can also be other methods, which are not limited in this embodiment.

[0112] Perform K times of data augmentation processing on each defective image showing semiconductor defects to obtain N defective images showing semiconductor defects after data augmentation processing, where N is equal to M × K. The data augmentation processing includes, but is not limited to: rotation processing, flipping processing, contrast adjustment, brightness adjustment, adding salt-and-pepper noise processing, adding Gaussian noise processing, etc.

[0113] Perform data augmentation processing on each defective image showing semiconductor defects in one or more of the above ways. The K enhanced defective images showing semiconductor defects obtained are respectively

[0114] S504. Input each defective image into the initial semiconductor defect classification model to obtain the defect category corresponding to each defective image, where the input defective image is the defective image obtained after the data augmentation processing in step S502.

[0115] Input the defective image after data augmentation processing into the above formula (4) to obtain the defect category as shown in formula (5):

[0116]

[0117] Wherein, is the defective image after data augmentation processing input, is the defect category of the defective image after data augmentation processing obtained.

[0118] S505. Based on the defect category information corresponding to each defective image, train the weights of the initial semiconductor defect classification model based on a preset fourth loss function and a machine learning algorithm until the second preset number of training times is reached to obtain the target semiconductor defect classification model.

[0119] Optionally, according to the defect category information corresponding to each defect picture, based on a preset marginal entropy loss function or mutual information loss function, and a preset neural network algorithm or least squares method or maximum likelihood estimation algorithm, weight training is performed on the initial semiconductor defect classification model.

[0120] The preset marginal entropy loss function is shown in formulas (6)-(7):

[0121]

[0122]

[0123] Where K represents K defect pictures after data augmentation processing; i represents the i-th defect picture after data augmentation processing; represents the defect category of the i-th defect picture after data augmentation processing; represents the mean of the defect categories of K defect pictures after data augmentation processing; D t represents the data set containing M defect pictures obtained in step S502; x is any one of the defect pictures in D t in.

[0124] Train the weights according to the above formula, and update w1, w2, and w3. When the second preset training times are reached, substitute the updated w1, w2, and w3 into formula (4) to obtain the final target semiconductor defect classification model.

[0125] In summary, the schematic diagram of this embodiment can be as Figure 6 shown Figure 6 is a schematic diagram of a weight training provided by an embodiment of the present application. For the specific implementation method, please refer to the content in the specific steps of this embodiment and will not be repeated here.

[0126] In the above embodiments of the present application, by obtaining the first weight preset by the first initial classification model, the second weight preset by the second initial classification model, and the third weight preset by the third initial classification model, an initial semiconductor defect classification model with defective pictures as input and defect categories as output is established according to the first weight, the second weight, the third weight, and the initial classification models. Then, multiple defective pictures showing semiconductor defects are obtained, and each defective picture is input into the initial semiconductor defect classification model to obtain the defect category corresponding to each defective picture. Furthermore, based on the preset fourth loss function and machine learning algorithm, the weights of the initial semiconductor defect classification model are trained according to the defect category information corresponding to each defective picture until the second preset number of training times is reached, and a target semiconductor defect classification model is obtained. The method of this embodiment improves the accuracy of the obtained semiconductor defect classification model by training the weights corresponding to the initial classification models, making different initial classification models more suitable for classifying defects with different occurrence probabilities.

[0127] The present application also provides a method for classifying semiconductor defects. By obtaining a defective picture showing a target semiconductor defect, specifically, obtaining the original defective pictures of the target semiconductor taken from multiple angles, and performing image processing on the original defective pictures of the target semiconductor taken from multiple angles to obtain a defective picture showing the defects of the target semiconductor. Then, the defective picture is input into the target semiconductor defect classification model to obtain the target defect category to which the defective picture belongs. By using the trained target semiconductor defect classification model for semiconductor classification, the accuracy of the classification result is improved.

[0128] Among them, the target semiconductor defect classification model is trained according to the training method of the semiconductor defect classification model shown in any of the above embodiments.

[0129] Figure 7 FIG. 10 is a schematic structural diagram of a training device for a semiconductor defect classification model provided by an embodiment of the present application. The device includes: an acquisition module 701 and a training module 702.

[0130] The acquisition module 701 is configured to obtain multiple defective pictures showing different defects of a semiconductor, and the defect category information to which each defective picture belongs.

[0131] The training module 702 is configured to use the defective pictures as input quantities and the defect category information as output quantities to train the initial converter model until the first preset training condition is met, and obtain the trained initial classification model.

[0132] Among them, the initial classification model includes: a first initial classification model that satisfies a preset first loss function, a second initial classification model that satisfies a preset second loss function, and a third initial classification model that satisfies a preset third loss function.

[0133] The training module 702 is further configured to perform weight training on the trained initial classification model based on a preset loss function and a machine learning algorithm until the second preset training condition is met, thereby obtaining the target semiconductor defect classification model.

[0134] A possible implementation is that the training module 702 is specifically configured to:

[0135] Use the defect images as input and the defect category information as output, and train the initial transformer model based on the preset cross-entropy loss function until the first preset number of training times is reached, thereby obtaining the first trained initial classification model.

[0136] Use the defect images as input and the defect category information as output, and train the initial transformer model based on the preset class balance loss function until the first preset number of training times is reached, thereby obtaining the second trained initial classification model.

[0137] Use the defect images as input and the defect category information as output, and train the initial transformer model based on the preset inverse weighted loss function until the first preset number of training times is reached, thereby obtaining the third trained initial classification model.

[0138] The trained initial classification model includes: the first initial classification model, the second initial classification model, and the third initial classification model.

[0139] A possible implementation is that the training module 702 is further specifically configured to:

[0140] Obtain the first weight preset for the first initial classification model, the second weight preset for the second initial classification model, and the third weight preset for the third initial classification model.

[0141] Based on the first weight, the second weight, the third weight, and the initial classification model, establish an initial semiconductor defect classification model with defect images as input and defect categories as output.

[0142] Obtain multiple defect images showing semiconductor defects.

[0143] Input each defect image into the initial semiconductor defect classification model to obtain the defect category corresponding to each defect image.

[0144] Based on the defect category information corresponding to each defect image, perform weight training on the initial semiconductor defect classification model based on a preset fourth loss function and a machine learning algorithm until the second preset number of training times is reached, thereby obtaining the target semiconductor defect classification model.

[0145] One possible implementation is that the training module 702 is further specifically configured to:

[0146] Obtain M defect pictures showing semiconductor defects, where M is an integer greater than or equal to 1.

[0147] Perform data augmentation processing on the M defect pictures showing semiconductor defects to obtain N defect pictures showing semiconductor defects after data augmentation processing, where N is an integer greater than M.

[0148] One possible implementation is that the training module 702 is further specifically configured to:

[0149] Based on the defect category information corresponding to each defect picture, perform weight training on the initial semiconductor defect classification model based on a preset marginal entropy loss function or mutual information loss function, and a preset neural network algorithm or least squares method or maximum likelihood estimation algorithm.

[0150] The device provided in this embodiment is used to execute the foregoing embodiment of the method for training a semiconductor defect classification model, and its implementation principle and technical effects are similar, so details are not described herein again.

[0151] Figure 8 FIG. is a schematic structural diagram of a classification device for semiconductor defects provided in an embodiment of the present application. The device includes: an acquisition module 801 and a processing module 802.

[0152] The acquisition module 801 is configured to acquire a defect picture showing a defect of a target semiconductor.

[0153] The processing module 802 is configured to input the defect picture into a target semiconductor defect classification model to obtain a target defect category to which the defect picture belongs.

[0154] One possible implementation is that the acquisition module 801 is further specifically configured to:

[0155] Acquire original defect pictures of the target semiconductor taken from multiple angles.

[0156] Perform picture processing on the original defect pictures of the target semiconductor taken from multiple angles to obtain a defect picture showing a defect of the target semiconductor.

[0157] The device provided in this embodiment is used to execute the foregoing embodiment of the method for classifying semiconductor defects, and its implementation principle and technical effects are similar, so details are not described herein again.

[0158] It should be noted that the division of the above Figure 7 and Figure 8 The division of each module shown is only for illustration. The present application does not limit the division of each module and the naming of each module.

[0159] Figure 9 The structural schematic diagram of an electronic device provided by an embodiment of this application is shown as Figure 9 shown. The device may include: at least one processor 901 and a memory 902.

[0160] The memory 902 is used to store a program. Specifically, the program may include program code, and the program code includes computer operation instructions.

[0161] The memory 902 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory.

[0162] The processor 901 is used to execute the computer-executable instructions stored in the memory 902 to implement the method described in any of the foregoing embodiments. Among them, the processor 901 may be a central processing unit (CPU for short), or an application specific integrated circuit (ASIC for short), or one or more integrated circuits configured to implement the embodiments of this application.

[0163] Optionally, the electronic device may further include a communication interface 903. In a specific implementation, if the communication interface 903, the memory 902, and the processor 901 are implemented independently, the communication interface 903, the memory 902, and the processor 901 may be interconnected through a bus and communicate with each other. The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc., but it does not mean that there is only one bus or one type of bus.

[0164] Optionally, in a specific implementation, if the communication interface 903, the memory 902, and the processor 901 are implemented integrated on a chip, the communication interface 903, the memory 902, and the processor 901 may communicate through an internal interface.

[0165] The electronic device provided by this embodiment is used to execute the foregoing method for training a semiconductor defect classification model and / or the method for classifying semiconductor defects. Its implementation principle and technical effect are similar to those of the method embodiment, and will not be elaborated here.

[0166] The present application also provides a computer-readable storage medium, which may include: various media capable of storing program codes such as USB flash drives, external hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs. Specifically, the computer-readable storage medium stores computer-executable instructions for the training method of the above semiconductor defect classification model and / or the classification method of semiconductor defects.

[0167] The present application also provides a computer program product, which includes executable instructions stored in a readable storage medium. At least one processor of an electronic device can read the executable instructions from the readable storage medium, and the execution of the executable instructions by at least one processor enables the electronic device to implement the training method of the above semiconductor defect classification model and / or the classification method of semiconductor defects provided by the above various embodiments.

[0168] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.

[0169] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A training method for a semiconductor defect classification model, characterized in that, Including: Obtain multiple defect pictures showing different defects of a semiconductor, and defect category information to which each defect picture belongs; Use the defect pictures as input quantities and the defect category information as output quantities to train an initial converter model until a first preset training condition is met, and obtain a trained initial classification model; Among them, the initial classification model includes: a first initial classification model that satisfies a preset first loss function, a second initial classification model that satisfies a preset second loss function, and a third initial classification model that satisfies a preset third loss function; Based on a preset fourth loss function and a machine learning algorithm, perform weight training on the trained initial classification model until a second preset training condition is met, and obtain a target semiconductor defect classification model.

2. The method according to claim 1, characterized in that, The step of using the defect pictures as input quantities and the defect category information as output quantities to train the initial converter model until a first preset training condition is met and obtaining a trained initial classification model includes: Use the defect pictures as input quantities and the defect category information as output quantities, and based on a preset cross-entropy loss function, train the initial converter model until a first preset number of training times is reached, and obtain the trained first initial classification model; Use the defect pictures as input quantities and the defect category information as output quantities, and based on a preset class balance loss function, train the initial converter model until the first preset number of training times is reached, and obtain the trained second initial classification model; Use the defect pictures as input quantities and the defect category information as output quantities, and based on a preset reverse weighting loss function, train the initial converter model until the first preset number of training times is reached, and obtain the trained third initial classification model; The trained initial classification model includes: the first initial classification model, the second initial classification model, and the third initial classification model.

3. The method according to claim 2, characterized in that, The step of performing weight training on the trained initial classification model based on a preset fourth loss function and a machine learning algorithm until a second preset training condition is met and obtaining a target semiconductor defect classification model includes: Obtain a first weight preset for the first initial classification model, a second weight preset for the second initial classification model, and a third weight preset for the third initial classification model; According to the first weight, the second weight, the third weight, and the initial classification model, establish an initial semiconductor defect classification model with defect pictures as input and defect category information as output; Obtain multiple defect pictures showing semiconductor defects; Input each defect picture into the initial semiconductor defect classification model to obtain defect category information corresponding to each defect picture; Based on the defect category information corresponding to each defect picture, and based on a preset fourth loss function and a machine learning algorithm, perform weight training on the initial semiconductor defect classification model until a second preset number of training times is reached, and obtain the target semiconductor defect classification model.

4. The method according to claim 3, characterized in that, The obtaining of multiple defect pictures showing semiconductor defects includes: Obtaining M defect pictures showing semiconductor defects, where M is an integer greater than or equal to 1; Performing data augmentation processing on the M defect pictures showing semiconductor defects to obtain N defect pictures showing semiconductor defects after data augmentation processing, where N is an integer greater than M.

5. The method according to claim 3, characterized in that, The training of the weights of the initial semiconductor defect classification model based on the defect category information corresponding to each defect picture, a preset fourth loss function, and a machine learning algorithm includes: Based on the defect category information corresponding to each defect picture, training the weights of the initial semiconductor defect classification model based on a preset margin entropy loss function or mutual information loss function, and a preset neural network algorithm or least squares method or maximum likelihood estimation algorithm.

6. A classification method for semiconductor defects, characterized in that, It includes: Obtaining a defect picture showing the defects of a target semiconductor; Inputting the defect picture into a target semiconductor defect classification model to obtain the target defect category to which the defect picture belongs, where the target semiconductor defect classification model is trained according to the training method of the semiconductor defect classification model according to any one of claims 1 to 5.

7. The method according to claim 6, characterized in that, The obtaining of the defect picture showing the defects of a target semiconductor includes: Obtaining original defect pictures of the target semiconductor taken from multiple angles; Performing picture processing on the original defect pictures of the target semiconductor taken from multiple angles to obtain a defect picture showing the defects of the target semiconductor.

8. A training device for a semiconductor defect classification model, characterized in that, It includes: An obtaining module, configured to obtain multiple defect pictures showing different defects of a semiconductor, and the defect category information belonging to each defect picture; A training module, configured to use the defect pictures as input quantities and the defect category information as output quantities to train an initial converter model until a first preset training condition is met, and then obtain a trained initial classification model; Wherein, the initial classification model includes: a first initial classification model that satisfies a preset first loss function, a second initial classification model that satisfies a preset second loss function, and a third initial classification model that satisfies a preset third loss function; The training module is further configured to train the weights of the trained initial classification model based on a preset loss function and a machine learning algorithm until a second preset training condition is met, and then obtain a target semiconductor defect classification model.

9. A semiconductor defect classification device, characterized in that, It includes: An obtaining module, configured to obtain a defect picture showing the defects of a target semiconductor; A processing module, configured to input the defect picture into a target semiconductor defect classification model to obtain the target defect category to which the defect picture belongs, where the target semiconductor defect classification model is trained according to the training method of the semiconductor defect classification model according to any one of claims 1 to 5.

10. An electronic device, characterized in that, It includes: A processor, a memory; The memory is used to store executable instructions of the processor; Wherein, the processor is configured to execute the training method of the semiconductor defect classification model according to any one of claims 1 to 5 and / or execute the classification method of the semiconductor defects according to any one of claims 6 to 7 by executing the executable instructions.

11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which are used to implement the training method of the semiconductor defect classification model according to any one of claims 1 to 5, and / or execute the classification method of the semiconductor defect according to any one of claims 6 to 7 when being executed by a processor.