Supercomputer Model Training and Deployment Method

Through supercomputing model training and deployment system, the coordinated work of central nodes and edge nodes is used to realize data sharing across nodes and cross-production lines and iterative updates of model, solving the problems of high cost of AI model deployment and unsatisfactory accuracy, improving detection efficiency and accuracy, and reducing the risks of resource waste and data leakage.

CN118506128BActive Publication Date: 2025-07-04SHANGHAI GANTU NETWORK TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410679539.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-29
Publication Date
2025-07-04
Estimated Expiration
2044-05-29

AI Technical Summary

Technical Problem

In the prior art, the deployment cost of AI models in semiconductor wafer material detection is high and the defect detection accuracy is not ideal. The detection of wafer materials in different production lines or models varies greatly, which poses a risk of resource waste and data leakage.

Method used

The supercomputer model training and deployment system is adopted, and the model training center nodes and edge nodes work together, and the AI ​​push-up instructions are used to obtain defect samples for data annotation and training. Combined with the detection and verification of the intelligent vision AVS system, data sharing and model iterative updates are realized across nodes and across production lines.

Benefits of technology

It improves model training and deployment efficiency, reduces costs, enhances model accuracy and generalization capabilities, reduces the probability of overkill and missed detection, and ensures data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118506128B_ABST
    Figure CN118506128B_ABST
Patent Text Reader

Abstract

This application discloses a supercomputer model training and deployment method, which relates to the field of model training. The method obtains defect samples uploaded by target production line machines based on AI upward push instructions. The defect samples include defect images captured and intercepted by the production line machines from the material plates and defect types. Data annotation is performed on the defect samples according to the model training metrics and added to the defect image training set. Model training is executed based on the newly added defect image training set, and the trained AI model is deployed to the target production line machines according to the downward push instructions. The AVS system based on the deployed AI model detects and verifies the defect images to be confirmed, obtains true defect images and false defect images, and adds the target defect images screened by rejudgment to the defect image training set for model iterative update training. This solution realizes cross-node data defect sharing and model training, improves the efficiency of model training and deployment, and improves the model accuracy and generalization ability by introducing a rejudgment and correction mechanism for true and false defect images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of model training, and particularly to a supercomputer model training and deployment system. Background Art

[0002] In recent years, the learning ability of AI models has provided key support for automatic optical inspection (AOI) and automatic visual inspection (AVI) devices, significantly improving production efficiency and product quality. By learning a large amount of image data, AI models have achieved highly intelligent defect detection and classification in devices, and can quickly and accurately identify various problems on the product surface. Among them, the training of AI models is one of the key steps to achieve high-performance and intelligent systems.

[0003] In related technologies, for the detection of semiconductor wafer materials, defect detection technologies such as AI models are usually used to identify defective wafers, which requires deploying AI models on each inspection machine on each production line. The iterative update of the machine equipment on the production line usually needs to be deployed locally, and operations such as designing the model structure, collecting images, and model training are carried out according to requirements, and then the trained AI model is deployed to the production line machine. This solution of setting up a model training system separately will cause serious waste of resources and cost investment for enterprises with large factories and off-site branch factories, and there are differences in the detection methods for different production line machines or different types of wafer materials. The versatility and portability of the model are poor, the defect detection accuracy and detection ability of a single model are not ideal, and on-site design update and deployment may also lead to the risk of data leakage. Summary of the Invention

[0004] The embodiments of the present application provide a supercomputer model training and deployment method to solve the problems of high cost for updating and deploying AI detection models and unsatisfactory defect detection accuracy of the models. The method is used for a supercomputer model training and deployment system, which includes a model training center node and several terminal edge nodes in the system, and each edge terminal node includes several production line machines. The method includes:

[0005] Obtaining, based on an AI upload instruction, defect samples uploaded by a target production line machine during the defect detection process of a wafer material board, where the defect samples include defect images and defect types intercepted by the production line machine from the material board;

[0006] Performing data annotation on the uploaded defect samples according to model training metrics and adding them to the defect image training set;

[0007] Performing AI model training based on the newly added defect image training set, and deploying the trained AI model to the target production line machine according to a push-down instruction;

[0008] The intelligent vision AVS system based on the deployed AI model detects and verifies the defect images to be confirmed, obtains true defect images and false defect images, and adds the target defect images screened by rejudgment to the defect image training set for iterative update training of the model; the defect images to be confirmed are the images detected and recognized by other detection models from the material board.

[0009] Specifically, each terminal edge node stores defect samples recognized by each production line machine, and the defect samples are classified according to the type of AI model deployed by the machine; when the model training center node receives the push instruction, it obtains the target defect samples from the corresponding terminal edge node.

[0010] Specifically, the data annotation of the uploaded defect samples according to the model training metrics and adding them to the defect image training set includes:

[0011] Determine the version information of the AI model and the model training parameters; the model training parameters include at least one or more of the model recall rate, precision, learning rate, and gradient parameters;

[0012] Screen and determine the defect types to be trained and the corresponding number of defect images according to the model training parameters, and screen and filter the defect samples;

[0013] Add a calibration data set and import the filtered defect samples;

[0014] Set a label group for the defect images in the calibration data set according to the defect types to be trained by the model;

[0015] Perform defect annotation on all defect images to obtain a defect image training set.

[0016] Specifically, the data annotation types include manual annotation and automatic annotation;

[0017] When the data annotation is automatic annotation, enhance the image quality of the defect images through AE data enhancement operations; identify the defect types of the defect images, and automatically set label groups according to the identified defect types, and annotate all defect images;

[0018] When the data annotation is manual annotation, determine the label group corresponding to the defect type in the preset label library and / or create a new label group according to the model training metrics, and annotate all defect images.

[0019] Specifically, the execution of AI model training based on the newly added defect image training set includes:

[0020] Set mapping labels in the calibration data set, and screen the label groups and the corresponding types of defect images based on the mapping labels to form the defect image training set;

[0021] Iteratively train the initial model using the defective image training set, and output the target AI model according to the model version information.

[0022] Specifically, after completing the iterative training, perform a regression test on the target AI model based on the image validation set and the model training parameters; when the regression test on the target AI model passes, push the target AI model for deployment; when the regression test on the target AI model fails, store the misdetected defective images in the calibration data set for re-image annotation, and reset the mapping labels during training.

[0023] Specifically, when the defective samples uploaded by the target production line machine do not contain the target defect type or the number of defective images of the target defect type is less than the set value, obtain a number of normal images; perform noise addition processing on the normal images based on the target defect type to generate the target number of target defect maps, and import the target defect maps into the newly added calibration data set; and / or, obtain defective images containing the target defect type and the target number from the production line machines of other edge terminal nodes.

[0024] Specifically, when the AVS system executes the true point re-inspection mode, re-judge all the true defective images identified by the AVS system;

[0025] When the re-judgment result indicates that the image to be confirmed is determined to be a defective image with defects, skip it directly;

[0026] When the result indicates that the re-judgment result is a normal image, instruct the AVS system that deploys the target AI model to overkill the recognition, mark the misdetected images, and send them to the defective image training set.

[0027] Specifically, when the AVS system executes the false point re-inspection mode, re-judge all the false defective images identified by the AVS system;

[0028] When the re-judgment result indicates that the image to be confirmed is determined to be a normal image without defects, skip it directly;

[0029] When the result indicates that the re-judgment result is a defective image with defects, instruct the AVS system that deploys the target AI model to miss detecting the defect, mark the misdetected images, and send them to the defective image training set.

[0030] Specifically, when pushing the trained target AI model for deployment, obtain the test machines with the same detection function deployed in all terminal edge nodes, and push the target AI model for deployment.

[0031] The beneficial effects brought by the technical solutions provided in the embodiments of this application at least include: The supercomputer model training and deployment system provided in this application realizes the unified scheduling of all terminal edge nodes by the training center node; the production line machines of each edge node upload defect samples according to the update requirements; and the model training center node can extract defect samples according to the AI push instruction. Especially for the invocation of defect images of different production lines within the system, it can complement and iteratively update the defect type pictures required for model update; after the update and testing are completed, the same push deployment is carried out to achieve cross-production line and cross-node data defect sharing and model training, improving the efficiency of model training and deployment. In order to reduce the probability of over-killing and missed detection of the model, defect images to be confirmed identified by the defect detection model of competitors are specially introduced for cross-verification. The images of over-killing and missed detection are marked and added to the defect image training set to correct the subsequent over-killing and missed detection of the model and optimize the performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 Shows a possible network topology diagram;

[0033] Figure 2 Is a schematic structural diagram of the supercomputer model training and deployment system provided in the embodiments of this application;

[0034] Figure 3 Is a data processing flow chart of the model training center node;

[0035] Figure 4 Is a schematic flow chart of the supercomputer model training and deployment method;

[0036] Figure 5 Shows an algorithm schematic diagram of the supercomputer model training and deployment method. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be further described in detail below in conjunction with the accompanying drawings.

[0038] As used herein, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0039] Network topology diagram: A network topology diagram refers to the physical layout of interconnecting various devices with transmission media, that is, connecting devices such as computers in a network in a specific way. A network topology diagram usually shows the network configuration of devices such as network servers, computer devices, and workstations and their interconnections. Its structures mainly include star structure, ring structure, bus structure, distributed structure, tree structure, mesh structure, honeycomb structure, etc. In this embodiment, the network topology diagram includes a central node and edge nodes. The central node is a computer device or a general server, etc., that issues tasks, and the edge node is the edge device, which is a computer device or a slave server, etc., used to execute tasks.

[0040] Please refer to Figure 1 A possible network topology diagram is shown, including: The central node 110 is the model training central node in the supercomputer model training and deployment system of this embodiment, and the edge node 120 is the terminal edge node. A network topology diagram usually includes at least one central node 110 and multiple edge nodes 120. The central node 110 is a computer device or a network server, etc., responsible for model training and model deployment in the supercomputer system. When the central node 110 is a network server, the central node 110 is at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center, mainly providing background services for the edge nodes 120. In addition, the central node 110 can also be a computer device or a workstation responsible for managing all edge nodes 120. Optionally, the central node 110 can also provide data upload and download services to the edge nodes 120, and feedback the processing process and results of model training, etc.

[0041] The central node 110 establishes a communication connection with the edge node 120 through a wireless network or a wired network. In this application, the supercomputer model training and deployment system is mainly used in the semiconductor field. Large wafer production processing plants usually have many partitions distributed across the country or the world. Each partition has many production lines, and each production line is equipped with multiple inspection stations. For each sub-plant area, an edge node is formed. The edge nodes and the central node belong to the same large network domain, which is uniformly scheduled by the central node and isolated from the outside of the domain to improve the data security of the system.

[0042] Figure 2It is a schematic structural diagram of a supercomputer model training and deployment system provided by an embodiment of the present application, including a model training center node 220 and several terminal edge nodes 210 communicatively connected to the model training center node 220. The normally operating terminal edge nodes 210 run a (to-be-iteratively updated) defect detection model, collect defect images and defect information, and the model training center node performs model training and model deployment based on the defect images and defect information. Each terminal edge node 210 includes a data server 211 and several production line machines 212 running the defect detection model. The production line machines 212 detect the wafer material board based on the defect detection model and collect defect images and confirm defect information; while the data server 211 generates defect samples according to the defect images and defect information, uploads the defect samples, and receives push-down instructions and the trained AI model, and deploys the AI model to the target production line.

[0043] The model training center node 220 includes an image annotation server 221 and a model training server 222. The image annotation server 221 classifies and annotates the uploaded defect samples; the model training server 222 performs iterative training based on the classified and annotated defect samples.

[0044] As Figure 2 In the structure of, a schematic diagram of the interaction between a terminal edge node 210 and the model training center node 220 is exemplarily listed. There are several production line machines set in a factory area. The model parameters of these machines can be the same or different, and the specific items for defect detection are not limited to the same type of materials. Since these machines perform defect identification after deploying the defect detection model, when any production line machine identifies a defect point based on the AI model, it will intercept the defect image from the panoramic image of the material board and determine the defect information, and upload the panoramic image, defect image, and defect information to the data server 211. The data server 211 here can be the computer room located in each factory area, that is, a temporary storage device deployed in the factory area, and the management authority should belong to the system maintenance party. Because the control right of the machines for producing wafers belongs to the wafer production factory, and the device after data collection can only be used for the subsequent iterative update of the model, it is necessary to temporarily store the data inside the factory area. There are usually many production lines set inside a factory area. To ensure the independent maintenance of different production lines, the data server needs to establish an association relationship based on the machine information, panoramic image, defect image, and defect information of each production line machine, and perform associated storage, so as to generate production line-based defect samples, and the pertinence of model iterative training and update is stronger. The defect information mentioned in this application includes at least one of defect type, part number information, machine information, and defect detection time. For example, the defects can be board cracks, dirt, damage, and missing components, etc. The machine itself will identify these defect types, the batch number information of the material board, and the time when the defect is detected, etc. These are very important for model iteration and model transplantation.

[0045] In order to improve the accuracy of inspection items for different machines or different material plates, defect images generated by different production lines or different defect detection models also need to be classified. For example Figure 2 A structured data server 213 and an unstructured data server 214, which are respectively connected to the data server 211, are also provided at the terminal edge node. The data server 211 only sets different defect samples according to the number or function of the machine production line, and constructs a defect sample set according to the defect sample ID. The structured data server 213 caches defect images and panoramic images according to the defect sample ID, while the unstructured data server 214 caches defect information according to the defect sample ID. This scheme of classifying and storing structured and unstructured data is more convenient for management. In particular, the captured defect images and panoramic images belong to the wafer manufacturer. After centralizing and storing such data, it is also convenient for the manufacturer's permission control and to avoid leakage of private data. Correspondingly, the unstructured data is only the data generated by the system manufacturer for defect detection and is stored separately and associated, which is also convenient for the process control and troubleshooting of data upload and training. The data information such as the association information of the two types of data, the defect sample ID, and the defect sample set is stored in the data server 211. In the subsequent process, the data server 211 starts a data upload task based on the push instruction information, extracts defect samples from the structured data server 213 and the unstructured data server 214 according to the saved association information of the defect samples, and uploads them to the model training center node.

[0046] In some embodiments, the model training center node further includes a central server 223. The central server 223 receives the defect samples uploaded by the data server 211, filters and screens the defect samples according to the model training instructions, and sends them to the image annotation server 221 for image annotation. Image annotation is an essential stage in model training and iteration. Especially for the iterative update of an existing AI model, when the detection and recognition accuracy of a certain type of defect needs to be improved, a large number of defect images of this type need to be obtained. At the same time, to avoid misrecognition of the directly uploaded defect images, image annotation is also required. The specific steps may include the following:

[0047] A. Determine the version information of the AI model and the model training parameters; the model training parameters include at least one or more of the model recall rate, accuracy, learning rate, and gradient parameters;

[0048] B. Screen and determine the defect types to be trained and the corresponding number of defect images according to the model training parameters, and filter and screen the defect samples;

[0049] C. Add a calibration data set and import the filtered defect samples;

[0050] D, set label groups for defect images in the calibration dataset according to the defect types trained by the model;

[0051] E. Label all defect images and obtain a defect image training set.

[0052] Because different model training parameters or different indicators need to focus on different image types for training, this step requires first determining the current AI model version information, and then determining the defect types and corresponding number of defect images to be trained. In particular, if the defect types and quantities contained in the uploaded defect samples meet the standards, no other image data needs to be added, that is, directly add the calibration data set and directly import the sample data, and then set the label group according to the defect type, and set defect labels for all defect images.

[0053] For newly added defect types or rules, it is necessary to manually label or call defect images. For the central server, it is necessary to screen and judge based on the uploaded data. For example, the production line under the A edge node has an A detection model that specifically identifies and detects the defect type of "board cracking". The defect samples it collects are mainly defect images of board cracking. When it is necessary to add this defect recognition and detection function to the production line machine under the B edge node, you can directly send instructions to the data server of the A edge node through the central server, retrieve specific defect samples to the cloud, and then filter the defect images of the target defect type for iterative training of the B detection model on the B edge node. Of course, it also includes adding new defect rules. At this time, you need to actively execute the defect image creation task. For example, the defect image generation model is used to generate defect images in a targeted manner, and then the defect images are manually annotated.

[0054] The central server also caches the AI ​​model trained by the model training server, and generates push instructions for deploying the AI ​​model based on the terminal edge nodes or production line distribution. For example, after training the A detection model based on the defect samples of the B edge node, the A detection model of the A edge node is deployed and updated through the push instruction. The deployment and update task first reaches the data server of the node, and then is deployed to the target production line machine by the data server. Similarly, for the same type of production line machines under different edge nodes, that is, machines with the same functions deployed in different factory areas, they can be deployed and updated. This can break through the single production line and cross-regional mobilization of data updates, improve the real-time performance of the system and reduce the cost of training and deployment.

[0055] After obtaining the defective image training set output by the image annotation server, the model training server directly performs model training according to the model training parameters and outputs the target AI model according to the version information. This includes adding new AI models and iterating existing AI models. For example, in a scenario of completely new generation and deployment, the number of defective samples will be larger, while for simple iterative updates, the requirement for the number of samples is relatively small.

[0056] In some possible implementation manners, the model training central node further includes a regression test server 224. The regression test server 224 receives the target AI model output by the model training server 222 and performs regression testing based on the image validation set and the model training parameters. The data storage for regression testing and defective samples requires additional use of an archive server. The storage server stores all the AI models output by the model training server, as well as the defective image training set and the image validation set used for model training and validation. Of course, for convenient management, the archive server can be further divided into a structured archive server 225 and an unstructured archive server 226. The historical AI models and associated data are stored in the unstructured archive server 226. For the image validation set, the temporarily stored defective samples and the source images for generating defective images are stored in the structured archive server 225. When the defective samples do not contain the target defect type or the number of defective images of the target defect type is less than the set value, a number of normal images are obtained, and then the normal images are noise-added based on the target defect type to generate the target number of target defect images, and the target defect images are imported into the newly added calibration data set.

[0057] The purpose of regression testing is to ensure the security of model deployment and avoid potential risks. Figure 3 It is the data processing flow chart of the model training central node; when the regression test of the target AI model passes, the target AI model is sent to the central server; when the regression test of the target AI model fails, the misdetected defective images are sent to the image marking server for re-image annotation, and the model training server performs retraining. The image re-annotation in this process is for the defective images in the validation set, and the re-annotation is audited through the image marking server or manually, so as to specifically expand the training set and improve the recognition ability of the model.

[0058] In summary, the supercomputer model training and deployment system provided by this application realizes the unified scheduling of all terminal edge nodes by the training center node; the data servers of each edge node collect defect samples separately according to the production line machines, classify and store them; upload the defect samples according to requirements or instructions; and the central server can uniformly schedule the model distribution and data allocation, classify and label the defect samples through the image annotation server, and construct a defect image training set. Especially for the invocation of defect images within the system, it can complement and iteratively update the defect type pictures required for model update; after the update and testing are completed, the central server conducts unified allocation and push deployment to realize cross-production line and cross-node data defect sharing and model training, improving the efficiency of model training and deployment.

[0059] This application also provides a supercomputer model training and deployment method, which is specifically used for the above-mentioned supercomputer model training and deployment system. A general control server or a general control computer can be additionally set up to be responsible for globally coordinating the system. The general control computer includes, but is not limited to, monitoring the operating status of each production line machine in the body monitoring terminal edge node and the model training request, and controlling the operating status of each server in the model training center. Figure 4 It is a schematic flow diagram of the supercomputer model training and deployment method, including the following steps:

[0060] Step 401, based on the AI upload instruction, obtain the defect samples uploaded by the target production line machine for the wafer material board during the defect detection process. The defect samples include the defect images intercepted by the production line machine from the material board and the defect types.

[0061] The upload instruction can be issued by a production line machine of a certain edge node, or the general control computer receives the instruction or triggers according to the set conditions to control the normal operation of the central control node and the edge nodes. This step controls the model training center node to obtain the defect samples uploaded by the target production line machine for the wafer material board during the defect detection process. For example, the defect samples detected and uploaded by the sixth production line machine of the first edge node. Because each terminal edge node stores the defect samples identified by each production line machine, and the defect samples are classified according to the type of AI model deployed by the machine.

[0062] When the model training center node receives the upload instruction and determines the target production line machine, it obtains the target defect samples from the corresponding terminal edge node. In particular, when there are multiple production line machines with the same function in the node, the defect samples of the production line machines with the same detection function can be merged to obtain more defect samples, and when uploading the samples, as comprehensive defect images as possible should be covered.

[0063] Step 402, perform data annotation on the uploaded defect samples according to the model training metrics and add them to the defect image training set.

[0064] Figure 5It shows the algorithmic schematic diagram of the supercomputer model training and deployment method. Data annotation can be specifically divided into manual annotation and automatic annotation. For the image annotation server, when the data annotation is automatic annotation, the defective images are enhanced in image quality through AE data enhancement operations, and then the obtained defective images are imported. This step can use image recognition to confirm the defect types of the defective images, or combine the defect types recognized by the defective images themselves, and automatically set the tag groups according to the recognized defect types to annotate all the defective images added to the calibration data set.

[0065] When the data annotation is manual annotation, determine the tag group corresponding to the defect type in the preset tag library and / or create a new tag group according to the model training metrics, and then annotate all the defective images. For defective samples, when the defective samples uploaded by the target production line machine do not contain the target defect type or the number of defective images of the target defect type is less than the set value, obtain a number of normal images; perform noise addition processing on the normal images based on the target defect type to generate the target number of target defect images, and import the target defect images into the newly added calibration data set; and / or, obtain defective images containing the target defect type and the target number from the production line machines of other edge terminal nodes. This process can additionally retrieve the required defective images from other machines to achieve cross-complementation. The specific operations of this step are as described above and will not be elaborated here.

[0066] Step 403 performs AI model training based on the newly added defective image training set and deploys the trained AI model to the target production line machine according to the push-down instruction;

[0067] As mentioned above, the defective samples obtained by the central node are added to the calibration data set. The calibration data set can only contain the defective images required for training the AI model, or can also contain more other types of pictures, because subsequent model verification and iterative updates may require the ability or function to recognize defects of other defect types to be introduced, as well as the purposeful correction of the behaviors of overkill and missed detection of the model. For this reason, the defective images in the calibration data set can be tagged according to the defect types, and for a certain model training task, a mapping tag can be set. For each specific training task, training is selectively performed according to the mapping tag, that is, based on the mapping tag, the tag groups and the corresponding types of defective images are screened to form a defective image training set. The initial model is iteratively trained through the defective image training set, and the target AI model is output according to the model version information.

[0068] Schematically, for example, the calibration dataset contains defect labels such as "plate crack", "blurred handwriting", "component missing", "dirty", and "scratch". However, according to the function, the AI model trained this time only needs to train the defect images with the labels of "plate crack", "dirty", and "scratch" among them. At this time, a mapping label can be established according to the training task, and several types of defect images can be selected from them according to the mapping label, and the training task can be executed according to the composed defect image training set.

[0069] The above method of cross-node and cross-production line calling defect sample training can aggregate the unique defect detection capabilities of each machine tool device. The model trained thereby reduces the model performance fluctuations caused by device differences, making the generated target AI model maintain the unity of the model. This consistent deployment ensures that the same version of the model is used when different devices are running, thereby reducing the performance differences caused by model version differences and improving the stability and reliability of the entire system.

[0070] During the training process, it is also necessary to first determine the model version information, such as changing from AI 2.0 to AI 3.0, then create a new model or continue training on the basis of the original model, and then add training parameters to execute the training process.

[0071] To ensure the accuracy of the trained model, after training is completed, the model can be verified according to the uploaded defect images, that is, a regression test is performed on the target AI model based on the image validation set and the model training parameters. When the regression test on the target AI model passes, the target AI model is pushed down for deployment; when the regression test on the target AI model fails, the misdetected defect images are stored in the calibration dataset for re-image annotation, the mapping label is reset during training, and the defect image training set is selected again. The specific steps are as described above and will not be elaborated here. In addition, it can also be carried out by means of manual verification, that is, select defect images to verify the accuracy manually, and release the model for push-down deployment after confirming that the verification passes.

[0072] Step 404, the intelligent vision AVS system based on the deployed AI model detects and verifies the defect images to be confirmed, obtains true defect images and false defect images, and adds the target defect images selected by re-judgment to the defect image training set for model iterative update training; the defect images to be confirmed are images detected and recognized from the material board by other detection models.

[0073] The intelligent vision AVS system is a device independent of the production line machine. Its purpose is to optimize the performance of the target AI model, specifically by optimizing it with the help of other external defect detection models or defect images to be confirmed identified by external machines. Because in actual wafer fabs and AI model system suppliers are two organizations, large wafer fabs may adopt multiple systems and use them to achieve the effect of mutual promotion among model system suppliers, or in order to make their own AI models achieve better detection results, model system suppliers will compare the target AI model with the defect detection model of friendly companies to make up for their own advantages and disadvantages. At this time, you can run the defect detection model of a friendly company first, and after detecting the wafer material board, you will get a batch of defect images to be confirmed, that is, defect images to be confirmed. The defect images to be confirmed identified by the defect detection model of a friendly company mean that they may contain normal wafer images, but they are mistakenly detected as defect images, or there are problems with the identification of defect categories. At this time, the target AI model trained by this system is deployed in the intelligent vision AVS system, and the defect image to be confirmed is recognized again. When the AVS system recognizes that it is also a defect image, it indicates that the image is a true defect image, that is, AVS recognizes a true point; when the AVS system recognizes that it is not a defect image (that is, a normal image), it indicates that the image is a false defect image, that is, AVS recognizes a false point.

[0074] Whether AVS identifies true points or false points, they are all detected and identified by the machine model. For this kind of identification result, manual intervention or machine detection of the actual defect points to be confirmed can be used to determine which of the two detection results is true or false, that is, to introduce a model correction mechanism in the final stage.

[0075] See also Figure 5 The defect image re-judgment mechanism shown in the figure, when the AVS system executes the true point re-inspection mode, re-judgment is performed on all true defect images identified by the AVS system. When the re-judgment result indicates that the image to be confirmed is a defect image with defects, it is directly skipped and the next image is re-inspected; when the result indicates that the re-judgment result is a normal image, it indicates that the AVS system of the deployed target AI model has over-identified, and the mis-detected images need to be marked and sent to the defect image training set for model correction.

[0076] Correspondingly, when the AVS system executes the false point re-inspection mode, all false defect images identified by the AVS system are re-judged. When the re-judgment result indicates that the image to be confirmed is a normal image without defects, it is directly skipped and the next image is re-inspected; when the result indicates that the re-judgment result is a defective image with defects, the AVS system deploying the target AI model is instructed to miss the defect, and the mis-detected image is marked and sent to the defect image training set for model correction.

[0077] In the above steps, for the cases where the AVS system fails to detect defects, the actual defect images need to be labeled according to the defect types and then added to the defect image training set, that is, sent into the calibration dataset for subsequent model iteration update or iteration after a certain trigger condition is met; for the cases where the AVS system overkills, actually normal labels are set for normal pictures, and the normal labels are added to the calibration dataset. When constructing the defect image training set subsequently, mapping labels are established, that is, normal images are additionally used for training. By feeding the normal images and normal labels to the AI model for training, the recognition capabilities of all aspects of the model can be integrated, which helps to improve the generalization ability of the model. In particular, introducing the defect images to be confirmed recognized by the defect detection model of competitors for cross-validation, marking the overkilled and undetected images and then adding them to the defect image training set to correct the overkilling and undetected situations of the subsequent model and optimize the performance of the model.

[0078] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware or by a program instructing relevant hardware. The described program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a disk or an optical disc, etc.

[0079] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A supercomputer model training and deployment method, characterized in that The method is used for a supercomputer model training and deployment system, which includes a model training center node and several terminal edge nodes. Each edge terminal node includes several production line machines. The method includes: Obtaining, based on an AI push instruction, defect samples uploaded by a target production line machine during the defect detection process of a wafer material board. The defect samples include defect images captured and intercepted by the production line machine from the material board and defect types. Each terminal edge node stores defect samples recognized by each production line machine, and the defect samples are classified according to the type of AI model deployed by the machine. When the model training center node receives a push instruction, it obtains the target defect samples from the corresponding terminal edge node; Performing data annotation on the uploaded defect samples according to model training metrics and adding them to the defect image training set; Performing AI model training based on the newly added defect image training set and deploying the trained AI model to the target production line machine according to a pull instruction; Detecting and verifying the to-be-confirmed defect images based on the intelligent vision AVS system deploying the AI model, obtaining true defect images and false defect images, and adding the target defect images screened by rejudgment to the defect image training set for model iterative update training. The to-be-confirmed defect images are images detected and recognized from the material board by other detection models. Among them, when the AVS system executes the true point recheck mode, all true defect images recognized by the AVS system are rejudged. When the rejudgment result indicates that the to-be-confirmed image is determined to be a defective image with defects, it is directly skipped. When the result indicates that the rejudgment result is a normal image, it indicates that the AVS system deploying the target AI model overkills the recognition, marks the misdetected images, and sends them to the defect image training set; When the AVS system executes the false point recheck mode, all false defect images recognized by the AVS system are rejudged. When the rejudgment result indicates that the to-be-confirmed image is determined to be a normal image without defects, it is directly skipped. When the result indicates that the rejudgment result is a defective image with defects, it indicates that the AVS system deploying the target AI model misses defects, marks the misdetected images, and sends them to the defect image training set.

2. The method according to claim 1, wherein The performing data annotation on the uploaded defect samples according to model training metrics and adding them to the defect image training set includes: Determining the version information of the AI model and model training parameters. The model training parameters include at least one or more of model recall rate, precision, learning rate, and gradient parameters; Screening and determining the defect types to be trained and the corresponding number of defect images according to the model training parameters, and screening and filtering the defect samples; Adding a new calibration data set and importing the filtered defect samples; Setting a tag group for the defect images in the calibration data set according to the defect types to be trained by the model; Performing defect annotation on all defect images to obtain a defect image training set.

3. The method according to claim 2, wherein The data annotation types include manual annotation and automatic annotation; When the data annotation is automatic annotation, enhancing the image quality of the defect images through AE data enhancement operations; identifying the defect types of the defect images, automatically setting tag groups according to the identified defect types, and annotating all defect images; When data is manually annotated, determine the tag group corresponding to the defect type in the preset tag library and / or create a new tag group according to the model training metrics, and annotate all defect images.

4. The method according to claim 2, wherein The AI model training based on the newly added defect image training set includes: Set mapping tags in the calibration data set, and filter the tag group and the defect images of the corresponding type based on the mapping tags to form the defect image training set; Iteratively train the initial model through the defect image training set, and output the target AI model according to the model version information.

5. The method according to claim 4, wherein After the iterative training is completed, perform a regression test on the target AI model based on the image validation set and the model training parameters; when the regression test on the target AI model passes, push the target AI model for downstream deployment; when the regression test on the target AI model fails, store the misdetected defect images in the calibration data set for re-image annotation, and reset the mapping tags during training.

6. The method according to any one of claims 1-5, characterized in that, When the defect samples uploaded by the target production line machine do not contain the target defect type or the number of defect images of the target defect type is less than the set value, obtain a number of normal images; perform noise addition processing on the normal images based on the target defect type to generate the target number of target defect images, and import the target defect images into the newly added calibration data set; and / or, obtain defect images containing the target defect type and the target number from the production line machines of other edge terminal nodes.

7. The method according to claim 5, characterized in that When pushing the trained target AI model for downstream deployment, obtain the test machines with the same detection function deployed in all terminal edge nodes, and push the target AI model for downstream deployment.

Citation Information

Patent Citations

  • Object defect recognition model training method and device, electronic equipment and storage medium

    CN112836724A

  • Auto defect screening using adaptive machine learning in semiconductor device manufacturing flow

    US20180164792A1