Method for detecting abnormal pattern of image in focus exposure matrix
The feature map of FEM images is extracted and fused through the deep network model, predicted bounding boxes are generated, and the model is optimized to improve detection accuracy, which solves the problem of abnormal pattern detection of FEM images in lithography technology, and achieves efficient and robust abnormal pattern detection.
Patent Information
- Application Number
- CN202311787379.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-21
- Publication Date
- 2025-06-24
AI Technical Summary
In lithography technology, the abnormal pattern detection of images in focus exposure matrix (FEM) is difficult, affecting the accuracy of CD measurement and is a fatal problem in the screening process condition window.
Using the deep network model, the feature maps of multiple images in the FEM are extracted, and the fusion feature maps that are independent of the order of multiple images are generated. The prediction bounding box in the target image is generated based on the fusion feature map, and the deep network model is optimized to improve detection accuracy.
It greatly improves the detection accuracy of abnormal patterns in FEM images, and can completely robustly process neighborhood images in different orders, which is suitable for the changing trends of different FEM conditions, and realizes efficient and automated detection of abnormal patterns.
Smart Images

Figure CN120198752A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of semiconductor manufacturing technology, and more particularly, to a lithography process. Embodiments of the present disclosure relate to a method, an apparatus, an electronic device, a computer-readable storage medium, and a computer program product for detecting abnormal patterns in an image in a focus exposure matrix. Background Art
[0002] The lithography process is a very critical step in the semiconductor manufacturing process. Focus and dose are important process parameters in the lithography process. Selecting appropriate process parameters to achieve the maximum process window, that is, maximizing the available range of focus and dose, is an important step in lithography process development, which can effectively improve the product yield and reduce raw material loss.
[0003] A focus exposure matrix (FEM) is a design of experiments method that evaluates the lithography effect under different focus and exposure settings. Through this method, engineers can sample the influence of exposure under different settings and find the optimal exposure and focus settings. Under different focus and exposure conditions, sub-resolution assist features (SRAFs) are often introduced for exposure experiments, and the corresponding critical dimension (CD) values are measured to achieve a larger process window. However, at the same time, SRAFs also introduce abnormal patterns during imaging, which not only affect the accuracy of CD measurement but also are one of the fatal problems in screening the process condition window. How to effectively detect abnormal patterns in the FEM image poses a challenge to the current lithography process. Summary of the Invention
[0004] Embodiments of the present disclosure provide a solution for detecting abnormal patterns in a focus exposure matrix (FEM) image.
[0005] According to a first aspect of the present disclosure, a method for detecting abnormal patterns in FEM images is provided. The method includes: using a deep network model to extract feature maps of multiple images in the FEM, where the multiple images include a target image and neighborhood images of the target image; generating a fusion feature map that is independent of the order of the multiple images based on the feature maps of the multiple images; using the deep network model to generate a predicted bounding box in the target image based on the fusion feature map; and optimizing the deep network model based on the predicted bounding box and an annotated bounding box for the abnormal pattern in the target image. In this way, the optimized deep network model can detect abnormal patterns in the FEM target image based on neighborhood information, greatly improving the detection accuracy, and can robustly handle neighborhood images in different orders, that is, the changing trends of different FEM conditions, so that it can be directly generalized to different FEM detection scenarios.
[0006] In some embodiments of the first aspect, the method further includes: using a graph network model to select one or more images from the neighborhood of the target image as neighborhood images, and the image neural network is configured to determine that the selected one or more images and the target image have positive correlation regarding the abnormal pattern. In this way, beneficial auxiliary samples in the neighborhood can be selectively chosen, reducing model confusion caused by invalid samples, and the number of input images can be reduced during screening, improving the model performance.
[0007] In some embodiments of the first aspect, selecting one or more images from the neighborhood of the target image includes: extracting features of the target image and images in the neighborhood; using the features as nodes and the distances between the target image and the images in the neighborhood in the FEM as edges, and inputting them into the graph network model; and selecting one or more images from the images in the neighborhood based on the correlation degree output by the graph network model being greater than a threshold. In this way, the target image and its neighborhood images can be input into the graph network model in the form of a graph, and the graph network model is used to screen out positively correlated auxiliary samples.
[0008] In some embodiments of the first aspect, generating a fusion feature map that is independent of the order of the multiple images includes: using a permutation-invariant symmetric function to generate the fusion feature map, where the permutation-invariant symmetric function takes the feature maps of the multiple images as inputs, and the output of the permutation-invariant symmetric function is independent of the order of the feature maps of the multiple images. In this way, output features independent of the input order can be constructed from the target image and its neighborhood images, addressing the problem of disorder caused by different focus and exposure settings in neighborhood conditions, and making it completely robust to the order of neighborhood conditions.
[0009] In some embodiments of the first aspect, the permutation-invariant symmetric function includes a pointwise max pooling operation for the feature maps of the multiple images. In this way, a simple and effective permutation-invariant symmetric function is provided.
[0010] In some embodiments of the first aspect, extracting feature maps of multiple images in the FEM includes: for each of the multiple images, using a deep network model to generate multi-scale feature maps of the image. Based on this approach, abnormal patterns of different scales can be detected. For example, large-scale feature maps are used to predict annotation boxes of small areas, and small-scale feature maps are used to predict annotation boxes of large areas.
[0011] In some embodiments of the first aspect, the fused feature maps include fused multi-scale feature maps. Generating predicted bounding boxes in the target image based on the fused feature maps includes: for each scale of feature maps in the fused multi-scale feature maps, using the regression network of the deep network model to generate at least one vector representing the predicted bounding boxes in the target image, where the at least one vector indicates the position of the predicted bounding boxes in the target image, the category of the predicted bounding boxes, and the confidence. Based on this approach, regression prediction of the deep network model is provided, which helps to more accurately detect the position and category of abnormal patterns.
[0012] In some embodiments of the first aspect, generating at least one vector representing the predicted bounding boxes in the target image includes: for each position in the feature maps of each scale, generating a corresponding vector; and determining at least one vector representing the predicted bounding boxes from the generated vectors through non-maximum suppression with respect to the confidence. Based on this approach, the bounding boxes most likely to contain abnormal patterns can be effectively selected while removing redundant other bounding boxes.
[0013] In some embodiments of the first aspect, optimizing the deep network model includes: optimizing the deep network model based on at least one of a position loss, a category loss, and a confidence loss. Based on this approach, an optimization objective of the deep network model is provided, and this optimization objective helps to more accurately detect the position and category of abnormal patterns.
[0014] In some embodiments of the first aspect, the confidence loss is determined based on a comparison between the predicted bounding boxes and the annotated bounding boxes.
[0015] In some embodiments of the first aspect, the multiple images in the FEM are critical dimension scanning electron microscope (CD-SEM) images.
[0016] According to a second aspect of the present disclosure, a method for detecting abnormal patterns in FEM images is provided. The method includes: using an optimized deep network model to extract feature maps of multiple images in the FEM, where the multiple images include a target image and neighborhood images of the target image; generating a fusion feature map independent of the order of the multiple images based on the feature maps of the multiple images; and using the optimized deep network model to generate a predicted bounding box in the target image based on the fusion feature map. In this way, the optimized deep network model can detect abnormal patterns in the FEM target image based on neighborhood information, greatly improving the detection accuracy, and can robustly handle neighborhood images in different orders, that is, the changing trends of different FEM conditions, enabling it to be directly generalized to different FEM detection scenarios.
[0017] In some embodiments of the second aspect, the method further includes: using a graph network model to select one or more images from the neighborhood of the target image as neighborhood images, and the image neural network is configured to determine that the selected one or more images and the target image have a positive correlation regarding abnormal patterns. In this way, beneficial auxiliary samples in the neighborhood can be selectively chosen, reducing model confusion caused by invalid samples, and the number of input images can be reduced during screening, improving the model performance.
[0018] In some embodiments of the second aspect, selecting one or more images from the neighborhood of the target image includes: extracting the features of the target image and the images in the neighborhood; and using the features as nodes of a graph and the distances between the target image and the images in the neighborhood in the FEM as edges of the graph, and inputting them into the graph network model; and selecting one or more images from the images in the neighborhood based on the correlation degree output by the graph network model being greater than a threshold. In this way, the target image and its neighborhood images can be input into the graph network model in the form of a graph, and the graph network model is used to screen out positively correlated auxiliary samples.
[0019] In some embodiments of the second aspect, generating a fusion feature map independent of the order of the multiple images includes: using a permutation-invariant symmetric function to generate the fusion feature map, where the permutation-invariant symmetric function takes the feature maps of the multiple images as inputs, and the output of the permutation-invariant symmetric function is independent of the input order of the feature maps of the multiple images. In this way, output features independent of the input order can be constructed from the target image and its neighborhood images, addressing the problem of disorder caused by different focus and exposure device settings in neighborhood conditions, making it completely robust to the order of neighborhood conditions.
[0020] In some embodiments of the second aspect, the permutation-invariant symmetric function includes a point-wise max pooling operation for the feature maps of the multiple images. In this way, a simple and effective permutation-invariant symmetric function is provided.
[0021] According to a third aspect of the present disclosure, there is provided a method for constructing a graph network model, including: obtaining a target image in a focus exposure matrix (FEM) and images within its neighborhood, where the target image and the images within the neighborhood have annotation information regarding abnormal patterns; extracting features of the target image and the images within the neighborhood; constructing a graph model based on the features and the distances between the target image and the images within the neighborhood, the graph model including a target node corresponding to the target image and neighboring nodes corresponding to the images within the neighborhood; and optimizing the graph network model based on the constructed graph model and the annotation information. In this way, beneficial auxiliary samples can be selectively chosen in the neighborhood of the target image, reducing model confusion caused by invalid samples, reducing the number of input images during screening, and improving model performance.
[0022] In some embodiments of the third aspect, the annotation information indicates the classification regarding abnormal patterns, and optimizing the graph network model includes: optimizing the graph network model based on the classification indicated by the annotation information of the target node and the neighboring nodes. In this way, the graph network model can learn the correlation information between the neighboring nodes and the target node.
[0023] In some embodiments of the third aspect, the classification indicates whether there is an abnormal pattern in the image, and in the case where the target node has an abnormal pattern, the neighboring nodes with abnormal patterns are positively correlated nodes, and in the case where the target node does not include an abnormal pattern, the neighboring nodes without abnormal patterns are positively correlated nodes. In this way, an effective annotation method for correlation information is provided.
[0024] In some embodiments of the third aspect, the graph network model includes a graph convolutional network configured to output the correlation between the nodes of the graph model. In this way, the graph network model can learn the correlation information through the weights of the convolutional network.
[0025] In some embodiments of the third aspect, the method further includes: after the graph network model has been optimized, determining a recommended correlation threshold for selecting neighboring images related to the target image based on the correlation generated by the positively correlated nodes. In this way, it can be ensured that effective neighboring images are comprehensively considered to assist in detecting abnormal patterns.
[0026] According to a third aspect of the present disclosure, there is provided an electronic device, including: a processing unit and a memory, the processing unit executing instructions in the memory such that the electronic device executes the method according to the first aspect or the second aspect or the third aspect of the present disclosure.
[0027] According to a fourth aspect of the present disclosure, there is provided an apparatus for detecting abnormal patterns in FEM images. The apparatus includes: a feature map extraction unit configured to use a deep network model to extract feature maps of a plurality of images in the FEM, the plurality of images including a target image and neighborhood images of the target image; a fusion unit configured to generate a fusion feature map independent of the order of the plurality of images based on the feature maps of the plurality of images; a prediction unit configured to use a deep network model to generate a predicted bounding box in the target image based on the fusion feature map; and an optimization unit configured to optimize the deep network model based on the predicted bounding box and an annotated bounding box for the abnormal pattern in the target image.
[0028] According to a fifth aspect of the present disclosure, there is provided an apparatus for constructing a graph network model. The apparatus includes: an image acquisition unit configured to acquire a target image in the FEM and images within its neighborhood, the target image and the images within the neighborhood having annotation information about abnormal patterns; a feature extraction unit configured to extract features of the target image and the images within the neighborhood; a model construction unit configured to construct a graph model based on the features and the distances between the target image and the images within the neighborhood, the graph model including a target node corresponding to the target image and neighboring nodes corresponding to the images within the neighborhood; and an optimization unit configured to optimize the graph network model based on the constructed graph model and the annotation information.
[0029] According to a sixth aspect of the present disclosure, there is provided a computer-readable storage medium having one or more computer instructions stored thereon, wherein the one or more computer instructions, when executed by a processor, cause the processor to execute the method according to the first aspect or the second aspect or the third aspect of the present disclosure.
[0030] According to a seventh aspect of the present disclosure, there is provided a computer program product including machine-executable instructions that, when executed by a device, cause the device to execute the method according to the first aspect or the second aspect or the third aspect of the present disclosure. Description of the Drawings
[0031] In conjunction with the drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the various embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0032] Figure 1A A schematic diagram showing the process flow of focus exposure matrix (FEM) process condition development is shown;
[0033] Figure 1B A schematic diagram showing the exposure process of introducing sub-resolution assist feature (SRAF) is shown;
[0034] Figure 1CShows a schematic diagram of an image and an abnormal pattern in an exemplary FEM;
[0035] Figure 2 Shows a schematic diagram of an example application scenario capable of implementing the embodiments of the present disclosure;
[0036] Figure 3 Shows a schematic flowchart of a method for constructing a graph network model according to an embodiment of the present disclosure;
[0037] Figure 4 Shows a schematic diagram of a neighborhood information screening process based on a graph network model according to an embodiment of the present disclosure;
[0038] Figure 5 Shows a schematic flowchart of a method for constructing a depth detector according to an embodiment of the present disclosure;
[0039] Figure 6 Shows a schematic diagram of the construction process of a depth network model based on disorder modeling according to an embodiment of the present disclosure;
[0040] Figure 7 Shows a schematic flowchart of a process for detecting abnormal patterns in an FEM image according to an embodiment of the present disclosure;
[0041] Figure 8 Shows a schematic diagram of an inference process based on a graph network model and a depth detection model according to an embodiment of the present disclosure;
[0042] Figure 9 Shows a schematic diagram of the training and optimization process of a model according to an embodiment of the present disclosure;
[0043] Figure 10 Shows a schematic diagram of the model deployment and inference process according to an embodiment of the present disclosure;
[0044] Figure 11A and 11B Shows a curve graph of the optimization process of a depth detection model according to an embodiment of the present disclosure;
[0045] Figure 12 Shows a schematic block diagram of a device for detecting abnormal patterns in an FEM image according to an embodiment of the present disclosure;
[0046] Figure 13 Shows a schematic block diagram of a device for constructing a graph network model according to an embodiment of the present disclosure; and
[0047] Figure 14 Shows a schematic block diagram of an electronic device capable of implementing the embodiments of the present disclosure. Detailed Description of the Invention
[0048] The technical solutions in the present disclosure will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all of the embodiments.
[0049] The technical solutions in the embodiments of the present disclosure will be described below with reference to the accompanying drawings in the embodiments of the present disclosure. Among them, in the description of the embodiments of the present disclosure, unless otherwise specified, " / " means "and / or". For example, A / B may represent A or B, or A and B. The "and / or" herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present disclosure, "a plurality of" or "multiple" means two or more than two.
[0050] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise specified, the meaning of "a plurality of" is two or more than two.
[0051] The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present disclosure. As used in the specification and the appended claims of the present disclosure, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the following embodiments of the present disclosure, "at least one" and "one or more" mean one, two, or more than two. The term "and / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships; for example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0052] References to "one embodiment" or "some embodiments" etc. described in this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of the present disclosure. Thus, statements such as "one embodiment", "some embodiments", "another embodiment", "some other embodiments", etc. that appear in different places in this specification are not necessarily all referring to the same embodiment, but mean "one or more but not all of the embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.
[0053] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. A neural network model is an example of a model based on deep learning. In this document, "model" can also be referred to as "network model", "machine learning model", "learning model", "machine learning network" or "learning network", and these terms are used interchangeably herein.
[0054] A "neural network" is a machine learning network based on deep learning. A neural network can process inputs and provide corresponding outputs, and generally includes an input layer and an output layer and one or more hidden layers between the input layer and the output layer. Neural networks used in deep learning applications generally include many hidden layers, thus increasing the depth of the network. The layers of the neural network are connected in sequence, so that the output of the previous layer is provided as the input of the next layer, where the input layer receives the input of the neural network, and the output of the output layer is the final output of the neural network. Each layer of the neural network includes one or more nodes (also called processing nodes or neurons), and each node processes the input from the previous layer.
[0055] Generally, machine learning can roughly include three stages, namely the training stage (also known as the optimization stage), the testing stage, and the usage stage (also known as the inference stage). In the training stage, a given model can be trained using a large amount of training data, continuously iterating and updating the parameter values until the model can obtain consistent inferences that meet the expected goals from the training data. Through training, the model can be considered to be able to learn the association from input to output (also known as the input-to-output mapping) from the training data. The parameter values of the trained model are determined. In the testing stage, the test input is applied to the trained model to test whether the model can provide the correct output, thereby determining the performance of the model. In some implementations, the testing stage can be omitted. In the usage stage, the model can be used to process the actual input based on the parameter values obtained from training and determine the corresponding output
[0056] The numbers or numerical values used in this specification are illustrative only. Their purpose is only to facilitate the understanding of the technology of the embodiments of the present disclosure and are by no means used to limit the scope of the present disclosure.
[0057] Overview
[0058] The lithography process is a very critical step in the semiconductor manufacturing process. Focus and Dose are the most important process parameters in the lithography process. Selecting appropriate process parameters to achieve the largest process window, that is, maximizing the available Focus and Dose ranges, is an important step in lithography process development, which can effectively improve the product yield and reduce raw material losses. The currently commonly used process condition window development and optimization method is the condition test based on the Focus Exposure Matrix (FEM), and its specific steps are as Figure 1A shown, including the FEM condition sampling and construction stage, the machine exposure stage, and the FEM condition screening stage.
[0059] FEM is an experimental design method that evaluates the lithography effect under different focus and exposure settings. Through this method, engineers can sample the influence of exposure under different settings and find the optimal exposure and focus settings. Under each setting, the process engineer needs to fabricate wafers and test them. The manufacturing process includes the exposure step and subsequent lithography processes such as development and etching. Finally, the process engineer will collect and analyze the relevant measurement data. This mainly includes the Critical Dimension (CD), as well as other important lithography parameters or defect parameters.
[0060] Sub-Resolution Assist Features (SRAF) is a technology widely used in the lithography process. It is used to improve the resolution of patterns and optimize the process window. The development process of introducing SRAF is as follows Figure 1B As shown, SRAF is regularly distributed around the target Graphic Data System (GDS) mask, aiming to make the developed pattern closer to the standard image.
[0061] Sub-Resolution Assist Features are particularly important for achieving CD features beyond the resolution limit of lithography equipment. Specifically, when the feature size to be developed in the process approaches the resolution limit of the lithography equipment, the control of critical dimensions becomes difficult. In this case, the image quality can be improved by adding Sub-Resolution Assist Features. These Sub-Resolution Assist Features affect the exposure process through interference, change the exposure distribution, can assist in achieving smaller critical dimensions, and improve the control of critical dimensions. Therefore, by reasonably designing and using Sub-Resolution Assist Features, smaller critical dimensions can be achieved and controlled beyond the resolution limit of lithography equipment. Currently, high-end lithography processes all consider the strategy of Sub-Resolution Assist Features to achieve more precise development control. This strategy is also applied to the development of FEM process conditions. Specifically, under different focus and exposure conditions, Sub-Resolution Assist Features are introduced for exposure experiments, and the corresponding CD values are measured to achieve a larger process window. However, at the same time, Sub-Resolution Assist Features also introduce abnormal patterns (Print-out Patterns) during imaging. These patterns not only affect the accuracy of CD measurement but are also one of the fatal (Killing Issue) factors for screening the process condition window, which is an extremely serious problem found in the development of the lithography process. It will directly lead to the unusability of the finished product and therefore needs to be excluded from the process window. Therefore, detecting these abnormal patterns in CD-SEM (Scanning Electron Microscope) images in the FEM matrix is a necessary step in the development of process conditions.
[0062] Take Figure 1C as an example. The line width in the target standard GDS is 97nm. The condition of the FEM black background has a CD value within the acceptable range. However, due to the generation of abnormal patterns during the manufacturing process of some of the conditions (black borders), they cannot be confirmed as available process conditions. Therefore, the final feasible process conditions are only the remaining white border areas.
[0063] To more precisely establish the lithography process window, effectively excluding these process conditions with abnormal patterns is an essential step. In current actual production, these abnormal patterns are mainly detected by manual visual inspection. This method consumes a lot of manpower and greatly prolongs the process development cycle. At the same time, the method of manual visual inspection cannot quantitatively analyze the abnormal patterns in the image, which also hinders the development process of automatically establishing the process window in the later stage. Existing detection technologies also cannot meet the requirements of the actual production line due to their poor timeliness and low accuracy. Currently, the abnormal pattern detection technology based on deep learning is one of the feasible solutions, but the deep detection technology based on a single image has obvious limitations in processing FEM images. The reason is that since the abnormal patterns and the GDS target patterns are generated from the same source, for abnormal patterns with strong imaging, a single-frame detector cannot meet the requirement of detection accuracy. Moreover, for abnormally weak patterns, they are basically integrated into the background, so they cannot be accurately detected by the method of single-frame detection.
[0064] In view of this, the present disclosure provides a deep learning-based technical solution that introduces neighborhood information to detect abnormal patterns in FEM images. The embodiments of the present disclosure aim to efficiently and automatically detect abnormal patterns in CD-SEM images under the FEM condition matrix to more reliably determine the appropriate lithography process window. The embodiments of the present disclosure also focus on the correlation of neighborhood images under the FEM matrix and model based on this correlation to achieve more accurate detection of difficult (too strong or too weak) abnormal patterns.
[0065] Exemplary Application Scenarios
[0066] Figure 2 The schematic diagram shows an example application scenario capable of implementing the embodiments of the present disclosure. It should be understood that Figure 2 The application scenario shown is only exemplary and should not constitute any limitation on the functions and scopes of the implementation described in the present disclosure.
[0067] The embodiments of the present disclosure can be applied to the FEM experimental method in the lithography process development stage. Specifically, the embodiments of the present disclosure can automatically detect abnormal patterns in the image by analyzing the CD-SEM images in the FEM matrix, thereby outputting a reliable FEM process window. The example application scenario mainly includes four stages: constructing an annotated data set 210, modeling the neighborhood relationship of the graph network model 220, optimizing the disordered depth detector (also known as the "deep network model") 230, and deploying the model and detection 240, and its composition is as Figure 2As shown. The main system architecture of the embodiments of the present disclosure may include a storage medium (such as a NAS server, a database, etc.) for storing dataset images and their annotations, a CPU processor and memory for processing image data and calculating thresholds, a GPU processor for accelerating deep network calculations, etc.
[0068] As shown in the figure, in the dataset construction stage 210, the user pulls a sufficient amount of images in the CD-SEM machine and stores them in the NAS server, and then examines each image one by one and boxes the annotation boxes of the abnormal patterns in the images. The images can be CD-SEM images or other types of images.
[0069] In the stage of constructing the FEM neighborhood association relationship 220, each CD-SEM image is used as a node (vertex) of the graph, and the distance of the image under the FEM matrix is used as an edge to construct an undirected graph. Then, the graph network model is optimized in the form of graph convolution. The optimized graph network model can output a quantified neighborhood association relationship matrix, and the user can interact with it to screen beneficial neighborhood data. In this article, the graph network model refers to a statistical model that uses a graph to represent the conditional dependence structure between variables.
[0070] In the stage of optimizing the disorder depth detector 230, the target image and the beneficial neighborhood images screened by the graph network model are used as inputs to optimize the depth detector and fit the annotation boxes in the dataset construction stage 210.
[0071] In the model deployment and detection stage 240, the graph network model and the depth detector optimized in the stages 220 and 230 are deployed to the production line to perform detection inference on the input target CD-SEM image and neighborhood images, and output the detection boxes of the abnormal patterns and their confidence levels.
[0072] The following further details the exemplary processes of the graph network neighborhood relationship modeling 220, the depth detector optimization 230, and the model deployment and detection 240 according to the embodiments of the present disclosure.
[0073] Graph network model
[0074] This stage corresponds to Figure 2 the graph network neighborhood relationship modeling 220.
[0075] Figure 3 Fig. shows a schematic flow chart of a method 300 for constructing a graph network model according to an embodiment of the present disclosure. The method 300 aims to construct the correlation between the target image in the target FEM and the images within its FEM neighborhood, with the purpose of screening out neighborhood information beneficial for detecting abnormal patterns in the target image. It can be understood that the method 300 may further include additional actions not shown and / or may omit the shown actions, and the scope of the present disclosure is not limited in this regard.
[0076] As shown Figure 3 in Figure 310, the target image in the Focus Exposure Matrix (FEM) and the images within its neighborhood are obtained. The target image and the images within the neighborhood have annotation information regarding abnormal patterns. The neighborhood can be the surrounding area of the target image, such as a 3*3 square neighborhood, which means the target image can have 8 adjacent images. The neighborhood can be a larger or smaller area and can have different shapes, and the present disclosure does not limit in this regard. The annotation information can include the classification information of the abnormal patterns made by the user and the bounding box. The classification information can be, for example, the category of the abnormal patterns annotated by the user, or the information on whether there is an abnormal image in the image. For example, the annotated bounding box means that there is an abnormal pattern in the image and indicates the corresponding position information, and no bounding box means that there is no abnormal pattern in the image.
[0077] In block 320, the features of the target image and the images within the neighborhood are extracted. The features can be in the form of vectors, matrices, or tensors, and the features of the target image and the images within the neighborhood can be extracted through a deep network or a hand-designed model.
[0078] In block 330, a graphic model is constructed based on the extracted features and the distance between the target image and the images within the neighborhood. The graphic model can include a target node corresponding to the target image and neighboring nodes corresponding to the images within the neighborhood, and the edges of the graphic model can be the distances between two images in the FEM (e.g., Euclidean distance, block distance, etc.). The graphic model can be an undirected graph. In other words, the image features extracted in block 320 are used as the nodes of the undirected graph, and the distances between the target image and the adjacent images in the FEM are used as the edges of the undirected graph.
[0079] At block 340, based on the constructed graph model and annotation information, the graph network model is optimized. In some embodiments, the graph network model may include a graph convolutional network, which may be configured to output the correlation between the nodes of the graph model. As mentioned above, the annotation information may include classification information about abnormal images, indicating whether there are abnormal images in the images. The annotation information indicates the correlation information between the target image and its surrounding images. Therefore, in the case where the target image has an abnormal pattern, the neighborhood nodes corresponding to the images with abnormal patterns are positively correlated nodes, and in the case where the target image does not include an abnormal pattern, the neighborhood nodes corresponding to the images without abnormal patterns are positively correlated nodes. Therefore, using the target image and its surrounding images obtained at block 310 and their annotation information as training data, the graph network model can learn the neighborhood correlation information about abnormal patterns. In some embodiments, in combination with the classification information of abnormal patterns (if any), training can also be performed for specific categories of abnormal patterns, so that the graph network model learns the correlation information for specific categories.
[0080] Figure 4 FIG. shows a schematic diagram of a neighborhood information screening process based on a graph network model. Figure 4 The process shown may be an exemplary specific implementation of method 300. As shown, the graph network model optimized for neighborhood information screening takes as input the target CD-SEM image and the images 401 within its neighborhood, together with its annotation box 409, for judging the positivity or negativity of the correlation therebetween.
[0081] At Figure 4 , the image preprocessing and enhancement module 402 receives the input CD-SEM image 410. This module performs scale normalization on the input image samples in terms of space and image intensity to facilitate model fitting and prevent gradient explosion. In some implementations, the image preprocessing and enhancement module 402 may also perform data augmentation operations, including random image affine transformation, random Gaussian blur, random block erasing, etc., to enrich the sample space. The calculations of this module can be completed by CPU processing in memory.
[0082] The image feature extraction module 403 is used to extract high-semantic features of the input image. This module may be composed of a pre-trained deep model, such as ResNet, Transformer, etc., or may be composed of manually designed features, such as Local Binary Patterns, Scale-Invariant Feature Transform (SIFT), etc. If deep model features are used, this module can be completed by GPU processing; if manual features are used, it can be completed by CPU processing.
[0083] The graph model construction module 404 is used to construct a graph model from the target CD-SEM image and the images within its neighborhood. Specifically, the feature of each CD-SEM image from the image feature extraction module 403 is a node z i (vertex) of this graph model, while the distance between two CD-SEM images in the FEM matrix is the edge α i,j (edge) between the corresponding nodes of the graph model.
[0084] The graph neural network optimization module 405 is used to construct a shallow graph convolutional neural network as the graph network model to be optimized, and apply it to the graph model constructed by the graph model construction module 404, for outputting the positivity and negativity of the association relationships between the nodes of the graph model. In some implementations, the information transmission method between nodes is:
[0085]
[0086] where l represents the l-th graph convolutional layer; σ represents the non-linear function; N i is the set of nodes connected to the node z i ; W l represents the convolutional kernel in the graph convolutional network. The optimization objective of this module 405 is to classify the central target node representing the target image and the neighboring nodes of the FEM image. That is, if the central node contains an abnormal pattern, then the nodes among the neighboring nodes of the image within the neighborhood that also contain abnormal patterns are positive nodes with high probability; conversely, if the central node does not contain an abnormal pattern, then the nodes among the neighboring nodes that do not contain abnormal patterns are positive nodes with high probability.
[0087] In some implementations, the graph neural network optimization module 405 can be optimized with the cross-entropy (BCE loss) as the supervision signal:
[0088]
[0089] where y i represents the true label of the i-th sample, is the predicted probability of the i-th sample, and N is the total number of samples.
[0090] The correlation threshold selection module 406 is used to collect all the high-probability nodes in the training data, and select the minimum value as the recommended threshold for implementing neighborhood CD-SEM image screening during the job inference process. In some implementations, the threshold for screening can also be adjusted manually. The neighborhood nodes with a correlation degree greater than the threshold output by the graph network model can be considered to be positively correlated with the target node, and the corresponding images can be input into the depth detector together with the target image to detect abnormal patterns.
[0091] Depth detector for disorder modeling
[0092] This stage corresponds to Figure 2 the disorder depth detector optimization 230.
[0093] Figure 5 FIG. 500 is a schematic flow chart of a method for constructing a depth detector according to an embodiment of the present disclosure. The method 500 is intended to regress and fit the annotation box in the target image based on the CD-SEM images of the target and its neighborhood in the FEM. The depth detector can be implemented as an optimized deep neural network model, which can be referred to as a deep network, a deep network model, etc. It can be understood that the method 500 may further include additional actions not shown and / or the actions shown may be omitted, and the scope of the present disclosure is not limited in this regard.
[0094] In block 510, using the deep network model, feature maps of multiple images in the FEM are extracted, where the multiple images include the target image and the neighborhood images of the target image. In some implementations, the deep network model may include a convolutional network having multiple convolutional layers for extracting features of the images. The output results from different convolutional layers can be combined together to form multi-scale feature maps, which may have the form of a feature pyramid. At the bottom of the pyramid are the feature maps obtained after shallow convolution, and at the top of the pyramid are the feature maps obtained after deep convolution.
[0095] It should be noted that the neighborhood images input to the deep network model are images that can be screened from the neighborhood of the target image. In some embodiments, a graph network model can be used to select one or more images from the neighborhood of the target image (e.g., a 3*3 region) as neighborhood images. The graph network model can be optimized by referring to Figure 3 and Figure 4 the processes shown, so as to be configured to determine that the one or more selected images and the target image have a positive correlation regarding the abnormal pattern. To this end, the features of the target image and the images within the neighborhood (e.g., the surrounding 8 CD-SEM images) can be extracted, and the extracted features are used as nodes, and the distances between the target image and these images in the FEM are used as edges, and input into the graph network model. The graph network model can output the correlation degree between these images within the neighborhood and the target image. If the correlation degree is greater than a threshold (a manually selected threshold or a recommended threshold), the corresponding image can be selected as the neighborhood image to be input to the depth detector.
[0096] In addition, the target image and its neighborhood images input to the deep network model may have annotation bounding boxes as ground truth. The annotation bounding boxes may have location information, category information, etc. of the abnormal pattern. These information are used to optimize the deep network model so that it has the ability to detect abnormal patterns after being optimized.
[0097] At block 520, a fused feature map independent of the order of multiple images is generated based on the feature maps of the multiple images. The feature maps of the multiple images obtained by the convolutional network can be fused together, and the fused feature map does not depend on the input order of these images. That is to say, it is not necessary to consider the previous spatial relationship of these images, which image is the target image and which image is the neighborhood image. As long as the feature maps of these images are used as the input, the obtained output remains unchanged. This can address the problem of disorder caused by different focus and exposure settings in the neighborhood conditions, making the detection result of abnormal patterns more stable. In some embodiments, the fusion operation can use a permutation-invariant symmetric function, for example, a point-to-point max pooling operation.
[0098] At block 530, using a deep network model, a predicted bounding box in the target image is generated based on the fused feature map. The deep network model can also include a regression network, which predicts the prediction box of the abnormal pattern based on the input fused feature map. The output of the regression network can be one or more vectors, and each vector can include components indicating the position of the predicted bounding box in the target image (e.g., the center coordinates, width, height, etc. of the bounding box). Additionally, if the category of the abnormal pattern is to be determined, the vector can also include components indicating the category of the predicted bounding box. Additionally, the vector can include components indicating the confidence level.
[0099] At block 540, the deep network model is optimized based on the predicted bounding box and the annotated bounding box for the abnormal pattern in the target image.
[0100] Figure 6 A schematic diagram showing the construction process of a deep network model based on disorder modeling according to an embodiment of the present disclosure is shown. Figure 6 The process shown can be an exemplary specific implementation of method 500. It is worth noting that a permutation-invariant symmetric function is introduced in the feature processing stage of this process to address the problem of disorder caused by different focus and exposure settings in the neighborhood conditions, so that the present invention can be completely robust to the order of neighborhood conditions.
[0101] As shown in the figure, the image preprocessing and enhancement module 602 receives the input target CD-SEM image and the filtered neighborhood images 601, and performs unified preprocessing on the received multiple images, normalizing them in terms of spatial scale and intensity. In some embodiments, the image preprocessing and enhancement module 602 can also adopt more complex data enhancement means for the target image, including but not limited to random image affine transformation, random Gaussian blur, random block erasing and other enhancement methods. Additionally, the data enhancement means can also include enhancement methods such as Mixup and CutMix, as shown below
[0102] Mixup: where λ ~ Beta(β, β); (3)
[0103] CutMix:
[0104] where I1 and I2 are the enhanced image and two random original images respectively; in Mixup, γ is a random parameter that satisfies the beta distribution, while in CutMix, M is a masked area with width W and height H.
[0105] The multi-scale feature extraction module 603 uses a backbone network to extract multi-scale features from the input image set. The backbone network can include a convolutional-based ResNet structure, or a Transformer structure based on an attention mechanism, etc. In some implementations, the multi-scale features can come from different stages of the backbone network and can finally output feature maps of different sizes in the form of a feature pyramid.
[0106] The permutation-invariant symmetric function 604 receives the multi-scale feature maps generated by the backbone network for the target image and the neighborhood images, and fuses these feature maps into a fusion feature map that is independent of the order. In some embodiments, the permutation-invariant symmetric function can be, for example, The processing method, which fuses the information from the neighborhood into the target image to assist the process of inferring the abnormal image bounding box. Specifically, the permutation-invariant symmetric function can satisfy:
[0107]
[0108] where h(·) is an arbitrary transformation function, such as the feature extractor used by the multi-scale feature extraction module; γ represents a set of linear transformation factors; ε(·) represents an order perturber, which will perform an order perturbation on the input data set {x i}; max{·} represents a pointwise maximum pooling operation, that is, a maximum pooling operation is performed for each dimension or component. Based on this way, the image features in the set after passing through this module will become order-independent features for fitting the bounding box of the abnormal pattern.
[0109] The multi-scale feature fusion module 605 is used to gradually fuse the multi-scale order-independent features of the neighborhood set images in the scale dimension, and at the same time fit a set of matching bounding boxes for each scale. That is to say, small-scale feature maps can be used to predict large-area annotation boxes, and large-scale feature maps can be used to predict small-area annotation boxes.
[0110] The bounding box regression module 606 may include a regression network for performing regression of annotation box metrics on feature maps at various scales. In some embodiments, for each scale of the feature maps in the fused multi-scale feature maps, the bounding box regression module 606 may generate at least one vector representing the predicted bounding boxes in the target image. In some implementations, the bounding box regression module 606 collapses the spatial dimension into a vector through an adaptive average pooling operation and then changes its length through a fully connected layer. The vector obtained after passing through the fully connected layer may indicate the position of the predicted bounding box in the target image, the category of the predicted bounding box, and the confidence. For example, the vector length may be transformed into (4 + n + 1), where 4 represents the position information of the annotation box as the center coordinate offset (Δx, Δy) and the width and height offset (Δw, Δh); n represents the number of types of abnormal patterns; 1 represents the confidence of the bounding box.
[0111] In some embodiments, for each position in the feature maps of each scale, the bounding box regression module 606 may generate a corresponding vector (including position information, classification, and confidence information). Then, for the predicted bounding boxes with overlap, the vector representing the predicted bounding box may be determined from the generated vectors through non-maximum suppression with respect to the confidence. In this way, the bounding boxes that are most likely to contain abnormal patterns can be effectively selected while removing redundant other bounding boxes. The bounding box regression module 606 may output the corresponding vectors of the selected bounding boxes.
[0112] The deep detector optimization module 607 acts on the output vector of the bounding box regression module 606 and fits it to the true metrics of the annotation box. The optimization objective may include optimization based on the position loss. In some embodiments, the position information may be supervised and fitted using the mean squared error (MSE):
[0113]
[0114] where represents the predicted value of the bounding box regression module.
[0115] Additionally, the optimization objective may further include optimization based on the category loss. For the judgment of the bounding box category and the classification of abnormal patterns, binary cross-entropy (BCE) may be used for supervised fitting:
[0116]
[0117] where y i represents the true label of the i-th sample, is the predicted probability of the i-th sample, and N is the total number of samples.
[0118] Additionally, the optimization objective can also include optimization based on class loss. The regression of confidence can also use BCE as supervision, where the true confidence of the target is defined as the intersection over union (IoU) of the predicted bounding box and the true bounding box, which is expressed as:
[0119]
[0120] where represents the predicted bounding box, and A represents the corresponding true bounding box. For each predicted bounding box operation, such as 3 true boxes and 5 predicted boxes (multi-scale accumulation), 15 IoUs can be calculated, and then the loss is accumulated. In some implementations, each multi-scale pixel can generate multiple regression vectors of length 4 + n + 1.
[0121] In this way, the optimized deep network model can detect abnormal patterns in the FEM target image based on neighborhood information, greatly improving the detection accuracy, and can robustly handle neighborhood images of different orders, that is, the change trends of different FEM conditions, so that it can be directly generalized to different FEM detection scenarios.
[0122] Model Deployment and Inference for Abnormal Pattern Detection
[0123] This stage corresponds to Figure 2 model deployment and detection 240.
[0124] Figure 7 FIG. shows a schematic flow chart of a process 700 for detecting abnormal patterns in an FEM image according to an embodiment of the present disclosure. Method 700 relates to the actual operation of detecting abnormal patterns in an FEM image. It can be understood that method 700 may further include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this regard.
[0125] Method 700 can be implemented by a deep network model, which is a deep detector obtained by referring to the Figure 5 and Figure 6 described process. The inference process of the deep network model conforms to the feed-forward mode of the model.
[0126] In block 1410, using the optimized deep network model, feature maps of multiple images in the FEM are extracted, and the multiple images include the target image and the neighborhood images of the target image. In block 720, based on the feature maps of the multiple images, a fusion feature map independent of the order of the multiple images is generated. In block 730, using the optimized deep network model, predicted bounding boxes in the target image are generated based on the fusion feature map.
[0127] In some embodiments, the images input into the deep network model can be obtained after being screened by the graph network model. Specifically, method 700 uses the graph network model to select one or more images from the neighborhood of the target image as neighborhood images. The image neural network can be, for example, a network model obtained through the optimization process described in Figure 3 and Figure 4 to select neighborhood images that have a positive correlation with the target image to assist in reasoning and detecting abnormal patterns. The graph network model receives the target image and all the images within its neighborhood, and extracts the features of the target image and the images within the neighborhood. Then, the extracted features are used as the nodes of the graph, and the distances between the target image and the images within the neighborhood in the FEM are used as the edges of the graph, which are input into the graph network model. The graph network model outputs relevance information. If the relevance is greater than the threshold, the corresponding image can be selected as an auxiliary sample.
[0128] Method 700 can detect abnormal patterns in the FEM target image based on neighborhood information, which improves the detection accuracy and can robustly handle neighborhood images in different orders, that is, the change trends of different FEM conditions, enabling it to be directly generalized to different FEM detection scenarios.
[0129] Figure 8 FIG. shows a schematic diagram of the reasoning process based on the graph network model and the depth detection model according to an embodiment of the present disclosure. Figure 8 The process shown can be an exemplary specific implementation of method 700. Figure 8 The process shown aims to infer whether the CD-SEM image in the FEM contains abnormal patterns based on the FEM matrix information. If so, the corresponding bounding box is output.
[0130] As shown in the figure, the image preprocessing module 802 is similar to the preprocessing in the previous module, and is used to uniformly process the input multiple images, that is, perform normalization in terms of spatial scale and intensity.
[0131] The graph network model inference module 803 is used to abstract the input image set into a graph model, and output the positive and negative probabilities of the relevance relationship between the target image (i.e., the target node) and the neighborhood images (neighborhood nodes) through the graph convolutional network model.
[0132] The threshold screening module 804 is used to compare the positive relationship probability between the target node and the neighborhood nodes output by the graph network model inference module 803 with the relevance threshold, and screen out the neighborhood nodes that are beneficial to the judgment of abnormal patterns of the target node.
[0133] The depth detector inference module 805 uses the optimized depth detector to infer the abnormal patterns contained in the target CD-SEM image based on the screened CD-SEM image set, and its output can include information such as predicted bounding boxes and confidence levels.
[0134] Test Process and Result Analysis
[0135] The test process is a real business scenario in semiconductor process technology development. During the test process, it is necessary to automatically and real-time process the source files and images output by the CD-SEM machine, and return the corresponding FEM detection results for process development engineers to analyze and judge.
[0136] Specifically, during the actual chip fabrication development process, 20 groups of different types of CD-SEM machine source files and their corresponding CD-SEM images were collected. Among them, 15 groups were used to optimize the deep detector model proposed in this paper, corresponding to approximately 30,000 images; the other 5 groups were used for verification, corresponding to approximately 10,000 images. In the verification data, approximately 2,000 difficult samples were selected to further confirm the advantages of the present invention. The performance indicators and accuracy indicators of the model can be considered simultaneously, that is, the performance needs to match the real-time output of the machine, and the accuracy should successfully detect more than 95% of the abnormal patterns according to the FAB production requirements.
[0137] Figure 9 The schematic diagram of the training and optimization process of the model according to an embodiment of the present disclosure is shown. As Figure 9 shown, the training and optimization of the model includes three stages, namely dataset construction, graph model optimization, and deep detector model optimization.
[0138] In the dataset construction stage, the system pulls the source files and CD-SEM image files output by the machine to the storage medium deployed in this embodiment for backup storage and for subsequent analysis. Secondly, the system will parse the source files output by the machine and structure them into FEM matrix information. Specifically, this module will parse the coordinates of each measurement value in the FEM matrix and associate the corresponding image of this measurement with this coordinate. Therefore, the FEM matrix will contain the image file and its neighborhood structure information. For all the collected CD-SEM images, during the test process, by means of rectangular selection, the position information, size information, and category information of each abnormal pattern on the image are determined, so as to form the sample correspondence between the CD-SEM image and the annotation box.
[0139] In the optimization stage of the graph network model, randomly sample CD-SEM images and image annotation boxes in the FEM matrix, and form a sub-batch to be input into the graph network model for training. In this embodiment, a sub-batch consists of 32 target CD-SEM images, and each target image is accompanied by 8 neighborhood images (i.e., a 3*3 neighborhood). Preprocess the sampled sub-batch images and their neighborhood images to obtain unified grayscale images with a resolution of 224×224 and 8-bit depth. At the same time, perform enhancement operations on the images, including random image affine transformation, random Gaussian blur, random block erasing, etc., to enrich the sample space.
[0140] When the graph network model is in forward propagation, ResNet18 pre-trained on ImageNet22K can be used as a feature extractor to encode each CD-SEM image into a 512-dimensional feature vector as the node descriptor of the graph model. In this embodiment, an information transfer model based on a graph convolutional network is constructed, which includes 8 implicit intermediate graph convolutional layers. Finally, the probability of positive and negative between nodes is output through a fully connected layer.
[0141] Construct the true label of the positive and negative of node connection according to whether there is an annotation box in the target image and neighborhood images. Specifically, neighborhood images of the same category as the central target image will be assigned a label of "1", and different categories will be assigned a label of "0". The BCE loss can be calculated using the loss function of formula (2).
[0142] In the backpropagation stage, calculate the gradient value of each layer's parameters in the graph convolutional network, and update its parameters according to the calculated gradient value. The Stochastic Gradient Descent (SGD) optimizer can be used, with an initial learning rate of 1e-3 as the initial update step size, and the learning rate decays to 0 according to the cosine function in 100 main update rounds.
[0143] Repeat the above process from random sampling to gradient update until the model fitting is completed, and output the optimized graph network model.
[0144] In the optimization stage of the depth detector model, similar to the graph network optimization stage, the depth detector optimization first requires randomly sampling the dataset to construct sample sub-batches. In this step, random and non-replacement sampling is performed, and 16 groups of sample pairs are used to construct sub-batches, with 8 neighborhood images for each group of samples.
[0145] Construct the target image and neighborhood images into a graph model based on the distance relationship under the FEM matrix, and input it into the optimized graph network in the graph network model optimization stage to infer the positive and negative of the neighborhood image and the central target image. According to the comparison between the positive and negative values of the neighborhood image and the threshold, filter out the neighborhood samples with positive correlation.
[0146] Then, preprocess the target image and its selected neighborhood images to obtain a unified grayscale image with a resolution of 512×512 and 8-bit depth in terms of space. Meanwhile, image enhancement operations can be performed on the images, and random geometric transformations and complex sample fusion transformations such as Mixup and CutMix can be adopted. It should be noted that a single target image and its neighborhood images need to be processed using a unified enhancement mode.
[0147] Input the sub-batch data into the depth detector model for forward propagation. The input tensor where n represents the sub-batch size, p represents the total number of a group of target images and neighborhood images, c represents the number of image channels, and w and h are the width and height of the image. Among them, the backbone network extracts features for each image. For parallel acceleration of feature extraction, the dimension of the input tensor x is transformed In this embodiment, the convolutional backbone of YOLOv7 is adopted to extract multi-scale features of each image. From the features of p images at each scale, this embodiment uses the permutation-invariant symmetric function of formula (5) to fuse the features of p images, that is, uses point-to-point maximum pooling operation to obtain the multi-scale features of a single frame and uses them to regress the bounding box parameters. The output feature scales in this embodiment are 16×16, 32×32, and 64×64. Features of different scales are responsible for regressing the annotation boxes of the corresponding scales. Finally, each position on the feature map is encoded as a 6-channel feature vector, where the first 4 channels encode the position parameters (Δx, Δy, Δw, Δh) of the bounding box, the 5th channel is the probability that the bounding box contains an abnormal pattern, and the 6th channel is the confidence of the bounding box.
[0148] Calculate the total loss between the regressed bounding box parameters and the actual annotation boxes according to formulas (6), (7), and (8) respectively, which includes position loss, class loss, and confidence loss.
[0149] Then, perform backpropagation according to the calculated loss, calculate the gradient of each layer in the depth detector network, and update the network parameters according to the obtained gradients. In this embodiment, the Adam optimizer is adopted, with an initial learning rate of 1e-3 as the update step size, and the learning rate decays to 0 according to the cosine function in 500 main update rounds. Repeat the above steps until the model fitting is completed, and output the optimized depth detector model.
[0150] Figure 10 A schematic diagram showing the model deployment and inference process according to an embodiment of the present disclosure is shown. The model deployment and inference process may include a data preprocessing part, a model inference part, and a post-processing part.
[0151] At block 1001, the source file and image data are pulled. When predicting the FEM samples in the actual production line, in this embodiment, the source file and all CD-SEM images are pulled from the metrology tool side and stored on the storage medium of the local server, waiting to be processed.
[0152] At block 1002, the metrology source file is parsed and the FEM is constructed. By processing the source file, the FEM matrix coordinates corresponding to each measurement value, namely the focal length and exposure value, are output, and the result image of this measurement is related to the FEM coordinates to construct the Figure 1C FEM image matrix as shown.
[0153] At block 1003, the images to be detected and their neighboring images in the FEM matrix are preprocessed in the manner required by the graph model. For example, they can be interpolated into a unified grayscale image with a resolution of 224×224 and 8-bit depth, and 512-dimensional features are extracted for each image through ResNet18 pre-trained on ImageNet22K.
[0154] At block 1004, the target image and its neighboring information are constructed into a graph model. Through the optimized graph convolutional network, the probability values of the positive relationships between each neighboring node and the central target node are output. According to the comparison between the threshold in the training stage and the output probability, the positive neighboring samples are selected to assist the inference of the subsequent depth detector.
[0155] At block 1005, the selected target image and neighboring images are preprocessed in the detector stage, including resizing to a unified grayscale image with a resolution of 512×512 and 8-bit depth.
[0156] At block 1006, the preprocessed image set is input into the optimized depth detector for forward propagation to obtain the predicted bounding box encoding. This step outputs a bounding box for each position of the feature map at each scale (there are 3 scales). Therefore, a total of (16×16 + 32×32 + 64×64)×3 = 16,128 bounding boxes are output, and each bounding box has its position information and confidence information.
[0157] At block 1007, non-maximum suppression is performed on the 16,128 bounding boxes to select the bounding boxes that are most likely to contain abnormal patterns and remove the redundant other bounding boxes. Specifically, the bounding boxes can be sorted in descending order according to the confidence of each bounding box; then, the IoU between the bounding boxes is calculated. If the IoU is greater than the preset threshold, the bounding box with a smaller confidence is regarded as redundant and removed; then, the above two steps can be repeated until all the bounding boxes are retained or removed.
[0158] At block 1009, the remaining bounding boxes after non-maximum suppression are decoded into rectangular box coordinates in the actual image coordinate system to identify abnormal patterns.
[0159] Figure 11A A curve graph showing the optimization process of the depth detection model according to the first embodiment is presented. To clearly compare and analyze the beneficial effects of the embodiments of the present disclosure, other detection models and methods are also constructed for comparison and analysis, including basic sliding window detectors, single-step single-frame detectors, and multi-step single-frame detectors. At the same time, the results of manual visual inspection and variants of different backbone networks of the present invention are also compared in this embodiment to corroborate the effectiveness and performance value of the solution proposed in the present disclosure. The specific comparison results are shown in Table 1 below:
[0160] Table 1. Effects of the algorithm of the present invention and comparative methods in the first embodiment.
[0161] Method Feature Extractor AP@0.5 AP@0.75 AP@0.5Hard FPS Manual Visual Inspection - 1.00 1.00 1.0 0.29 Sliding Window Detector Wide-ResNet50 0.14086 0.2433 0.4022 6.27 Faster RCNN Wide-ResNet50 0.7592 0.3017 0.5106 25.11 SSD Wide-ResNet50 0.7145 0.2891 0.4561 35.65 YOLOv7 YOLOv7 0.7934 0.3286 0.5228 45.42 Example 1 Wide-ResNet50 0.8237 0.3613 0.7497 36.44 Example 1 YOLOv7 0.8179 0.3442 0.7483 40.39
[0162] As shown in Table 1, the first embodiment based on the present invention is close to the effect of manual visual inspection in terms of accuracy, and can achieve an average accuracy index of more than 80% under both convolutional feature extractors. This accuracy index is the best solution in automated detection methods. The average accuracy of more than 80% can also meet the requirements of lithography process development. At the same time, in difficult scenarios, that is, when the CD-SEM image contains abnormally strong or weak patterns, the accuracy index of the first embodiment is much higher than that of the comparative methods, further highlighting the advantageous value of the present invention. In terms of comparison of performance, since the present invention considers the modeling and analysis of neighboring images, more image data needs to be processed. The performance of the solution of the present invention reaches 40.39 images processed per second. Although this performance is not as good as that of some single-step single-frame detectors, it is sufficient to match the real-time throughput of the CD-SEM measurement machine. Therefore, the present invention can greatly improve the detection accuracy of abnormal patterns of the depth detector under the FEM matrix, and at the same time can also meet the performance requirements of the actual production line.
[0163] Figure 11B A curve graph showing the optimization process of the depth detection model according to the second embodiment is presented. The second embodiment focuses on alternative solutions for convolutional feature extractors, that is, changing the convolutional-based feature extraction device in the first embodiment to a self-attention-based transformer extractor. Specifically, the backbone network of the depth detector in the second embodiment is constructed using CSWin Transformer. At the same time, to cooperate with the transformer-based feature extraction, the data augmentation method used in the second embodiment also adds erasure for different image block tokens.
[0164] The specific comparison results of the second embodiment are shown in Table 2 below. In this embodiment, other transformer-based methods are compared:
[0165] Table 2. Effects of the algorithm of the present invention and the comparative method in the second embodiment.
[0166] Method Feature Extractor AP@0.5 AP@0.75 AP@0.5Hard FPS Manual Visual Inspection - 1.00 1.00 1.0 0.29 Sliding Window Detector CSWin-Tiny 0.6618 0.2097 0.3473 3.90 DeTR CSWin-Tiny 0.7228 0.3319 0.4181 13.87 YOLOS CSWin-Tiny 0.7003 0.3101 0.4199 20.49 Example 2 CSWin-Tiny 0.7781 0.3511 0.7272 36.44
[0167] It can be seen from this embodiment that the method proposed in the present disclosure can also be applied to the self-attention depth framework based on transformers.
[0168] The above reference Figures 2 to 11B has described the concept and details of the embodiments of the present disclosure, and discussed and analyzed the test results. Compared with the prior art, the improvements of the embodiments of the present disclosure include the following aspects: establishing the relationship between the neighborhood information under the FEM matrix in the form of a graph model, and its advantage lies in that the introduction of neighborhood information greatly improves the detection accuracy of abnormal patterns, especially for the detection effect of overly strong or overly weak abnormal patterns. The embodiments of the present disclosure also use the method of disordered modeling to fuse the information of the neighborhood images for the detection of abnormal patterns, and the advantage lies in that it can completely and robustly process the neighborhood images in different orders, that is, the changes of different FEM conditions, so that it can be directly generalized to different FEM detection scenarios. In some embodiments, the method of graph convolutional neural network can also be used to distinguish the positive and negative of each node in the graph model. Thus, beneficial auxiliary samples can be selectively used, reducing the model confusion caused by invalid samples, and moreover, the number of input images is reduced by screening, improving the model performance.
[0169] Exemplary Devices and Equipment
[0170] FIG. 11 shows a schematic block diagram of a device 1200 for detecting abnormal patterns in a focus exposure matrix (FEM) image according to an embodiment of the present disclosure. The device 1200 can be implemented at any electronic device with computing capabilities.
[0171] The device 1200 is suitable for optimizing or training a deep network model. The device 1200 includes: a feature map extraction unit 1210 configured to extract feature maps of a plurality of images in the FEM using a deep network model, the plurality of images including a target image and neighborhood images of the target image; a fusion unit 1220 configured to generate a fusion feature map independent of the order of the plurality of images based on the feature maps of the plurality of images; a prediction unit 1230 configured to generate a predicted bounding box in the target image using the deep network model based on the fusion feature map; and an optimization unit 1240 configured to optimize the deep network model based on the predicted bounding box and an annotated bounding box for abnormal patterns in the target image.
[0172] In some embodiments, the apparatus 1200 may further include a neighborhood image screening unit configured to select one or more images from the neighborhood of the target image as the neighborhood images using a graph network model, wherein the image neural network is configured to determine that the selected one or more images and the target image have a positive correlation regarding the abnormal pattern.
[0173] In some embodiments, the neighborhood image screening unit may be configured to: extract features of the target image and the images within the neighborhood; and input the features as nodes and the distances between the target image and the images within the neighborhood in the FEM as edges into the graph network model; and select the one or more images from the images within the neighborhood based on that the correlation degree output by the graph network model is greater than a threshold.
[0174] In some embodiments, the fusion unit 1220 may be configured to generate the fused feature map using a permutation-invariant symmetric function, wherein the permutation-invariant symmetric function takes the feature maps of the multiple images as input, and the output of the permutation-invariant symmetric function is independent of the order of the feature maps of the multiple images.
[0175] In some embodiments, the permutation-invariant symmetric function may include a point-wise max pooling operation for the feature maps of the multiple images.
[0176] In some embodiments, the feature map extraction unit 1210 may further be configured to: for each image of the multiple images, generate a multi-scale feature map of the image using the deep network model.
[0177] In some embodiments, the fused feature map may include a fused multi-scale feature map, and the prediction unit 1230 may further be configured to: for each scale of the feature maps in the fused multi-scale feature map, generate at least one vector representing a predicted bounding box in the target image using a regression network of the deep network model, wherein the at least one vector indicates the position of the predicted bounding box in the target image, the category of the predicted bounding box, and the confidence.
[0178] In some embodiments, the prediction unit 1230 may further be configured to: generate a corresponding vector for each position in each scale of the feature maps; and determine the at least one vector representing the predicted bounding box from the generated vectors through non-maximum suppression regarding the confidence.
[0179] In some embodiments, the optimization unit 1240 may further be configured to optimize the deep network model based on at least one of a position loss, a category loss, and a confidence loss.
[0180] In some embodiments, the confidence loss can be determined based on a comparison between the predicted bounding box and the annotated bounding box.
[0181] In some embodiments, the plurality of images in the FEM can be critical dimension scanning electron microscope (CD-SEM) images.
[0182] The apparatus 1200 can also be applicable to the inference of a deep network model. Accordingly, the feature map extraction unit 1210 can be configured to cause a configured deep network model to be used to extract feature maps of a plurality of images in the FEM, the plurality of images including a target image and neighborhood images of the target image. The fusion unit 1220 can be configured to generate a fusion feature map independent of the order of the plurality of images based on the feature maps of the plurality of images. The prediction unit 1230 can be configured to use the deep network model to generate a predicted bounding box in the target image based on the fusion feature map. In the model inference stage, the optimization unit 1240 can be omitted.
[0183] Figure 13 FIG. 10 shows a schematic block diagram of an apparatus 1300 for detecting an abnormal pattern in a focus exposure matrix (FEM) image according to an embodiment of the present disclosure. The apparatus 1300 can be implemented at any electronic device having computing capabilities.
[0184] The apparatus 1300 is used for the construction of a graph network model. The apparatus 1300 includes: an image acquisition unit 1310, configured to acquire a target image in a focus exposure matrix (FEM) and images within its neighborhood, the target image and the images within the neighborhood having annotation information about an abnormal pattern; a feature extraction unit 1320, configured to extract features of the target image and the images within the neighborhood; a model construction unit 1330, configured to construct a graph model based on the features and the distance between the target image and the images within the neighborhood, the graph model including a target node corresponding to the target image and neighboring nodes corresponding to the images within the neighborhood; and an optimization unit 1340, configured to optimize the graph network model based on the constructed graph model and the annotation information.
[0185] In some embodiments, the annotation information can indicate the classification of the abnormal pattern, and the optimization unit 1340 can be configured to: optimize the graph network model based on the classification indicated by the annotation information of the target node and the neighboring nodes.
[0186] In some embodiments, the classification may indicate whether there is an abnormal pattern in the image, and in the case where the target image has an abnormal pattern, the neighborhood nodes corresponding to the image with the abnormal pattern are positively correlated nodes, and in the case where the target image does not include an abnormal pattern, the neighborhood nodes corresponding to the image without the abnormal pattern are positively correlated nodes.
[0187] In some embodiments, the graph network model may include a graph convolutional network configured to output the correlation degree between the nodes of the graph model.
[0188] In some embodiments, the apparatus 1300 may further include a threshold determination unit configured to determine a recommended correlation degree threshold for selecting neighborhood images related to the target image based on the correlation degree generated by the positively correlated nodes after the graph network model has been optimized.
[0189] The present disclosure may be a method, an apparatus, an electronic device, a computing device, a system, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.
[0190] Figure 7 A schematic block diagram of an example device 1400 that may be used to implement embodiments of the present disclosure is shown. For example, the methods 300, 500, and / or 500 for assisted driving according to embodiments of the present disclosure may be implemented by the device 1400. As shown, the device 1400 includes a central processing unit (CPU) 1401 and a graphics processing unit (GPU) 1411, which may perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 1402 or computer program instructions loaded from a storage unit 1408 into a random access memory (RAM) 1403 or the video memory of the GPU 1411. In the RAM 1403 and the video memory of the GPU 1411, various programs and data required for the operation of the device 1400 may also be stored. The CPU 1401, GPU 1411, ROM 1402, and RAM 1403 are connected to each other via a bus 1404. An input / output (I / O) interface 1405 is also connected to the bus 1404.
[0191] Multiple components in device 1400 are connected to I / O interface 1405. The types of I / O interfaces include, but are not limited to, high-speed peripheral component interconnect (PCIe), universal serial bus (USB), high-definition multimedia interface (HDMI), serial attached SCSI (SAS), etc. Components based on I / O interface 1405 may include, but are not limited to: input unit 1406, such as a keyboard, mouse, etc.; output unit 1407, such as various types of displays, speakers, etc.; storage unit 1408, such as a disk, optical disc, etc.; and communication unit 1409, such as a network adapter, modem, wireless communication transceiver, etc. Communication unit 1409 allows device 1400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunications networks.
[0192] Each of the processes and treatments described above, such as methods 300, 500, and / or 700, may be executed by a processing unit in device 1400, such as processing unit 1401, GPU 1411, and / or other processing units (e.g., a microprocessor on the motherboard of device 1400). For example, in some embodiments, method processes 300, 500, and / or 700 may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1408. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 1400 via ROM 1402 and / or communication unit 1409. When the computer program is loaded into RAM 1403 and executed, one or more actions of processes 300, 500, and / or 700 described above may be performed.
[0193] The present disclosure may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure.
[0194] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0195] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0196] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.
[0197] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.
[0198] These computer - readable program instructions can be provided to a processing unit of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data - processing apparatus, a device is created that implements the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions includes a manufacture, which includes instructions for implementing various aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0199] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0200] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0201] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A method for detecting abnormal patterns in a Focus Exposure Matrix (FEM) image, the method comprising: Using a deep network model to extract feature maps of a plurality of images in the FEM, the plurality of images including a target image and neighborhood images of the target image; Generating a fusion feature map independent of the order of the plurality of images based on the feature maps of the plurality of images; Using the deep network model to generate a predicted bounding box in the target image based on the fusion feature map; And Optimizing the deep network model based on the predicted bounding box and a labeled bounding box for the abnormal pattern in the target image.
2. The method according to claim 1, further comprising: Using a graph network model to select one or more images from the neighborhood of the target image as the neighborhood images, Wherein the image neural network is configured to determine that the selected one or more images and the target image have a positive correlation regarding the abnormal pattern.
3. The method according to claim 2, wherein selecting one or more images from the neighborhood of the target image comprises: Extracting features of the target image and the images within the neighborhood; And Inputting the features as nodes and the distances of the target image and the images within the neighborhood in the FEM as edges into the graph network model; And Selecting the one or more images from the images within the neighborhood based on the correlation degree output by the graph network model being greater than a threshold.
4. The method according to claim 1, wherein generating a fusion feature map independent of the order of the plurality of images comprises: Using a permutation-invariant symmetric function to generate the fusion feature map, wherein the permutation-invariant symmetric function takes the feature maps of the plurality of images as inputs, and the output of the permutation-invariant symmetric function is independent of the order of the feature maps of the plurality of images.
5. The method according to claim 4, wherein the permutation-invariant symmetric function comprises a pointwise max pooling operation for the feature maps of the plurality of images.
6. The method according to claim 1, wherein extracting the feature maps of a plurality of images in the FEM comprises: For each image in the plurality of images, using the deep network model to generate a multi-scale feature map of the image.
7. The method according to claim 6, wherein the fusion feature map comprises a fused multi-scale feature map, and generating a predicted bounding box in the target image based on the fusion feature map comprises: For each scale of the feature map in the fused multi-scale feature map, using a regression network of the deep network model to generate at least one vector representing the predicted bounding box in the target image, Wherein the at least one vector indicates the position of the predicted bounding box in the target image, the category of the predicted bounding box, and the confidence.
8. The method according to claim 7, wherein generating at least one vector representing the predicted bounding box in the target image comprises: Generating a corresponding vector for each position in each scale of the feature map; And Determining the at least one vector representing the predicted bounding box from the generated vectors through non-maximum suppression regarding the confidence.
9. The method according to claim 7, wherein optimizing the deep network model comprises: Optimizing the deep network model based on at least one of a position loss, a class loss, and a confidence loss.
10. The method according to claim 9, wherein the confidence loss is determined based on a comparison between the predicted bounding box and the annotated bounding box.
11. The method according to claim 1, wherein the plurality of images in the FEM are critical dimension scanning electron microscope (CD-SEM) images.
12. A method for detecting an abnormal pattern in a focus exposure matrix (FEM) image, the method comprising: Using an optimized deep network model to extract feature maps of a plurality of images in the FEM, the plurality of images including a target image and neighborhood images of the target image; Generating a fused feature map independent of the order of the plurality of images based on the feature maps of the plurality of images; And Using the optimized deep network model to generate a predicted bounding box in the target image based on the fused feature map.
13. The method according to claim 1, further comprising: Using a graph network model to select one or more images from the neighborhood of the target image as the neighborhood images, the image neural network being configured to determine that the selected one or more images and the target image have a positive correlation with respect to the abnormal pattern.
14. The method according to claim 13, wherein selecting one or more images from the neighborhood of the target image comprises: Extracting features of the target image and images within the neighborhood; And Inputting the features as nodes of a graph and the distances between the target image and the images within the neighborhood in the FEM as edges of the graph into the graph network model; And Selecting the one or more images from the images within the neighborhood based on a correlation degree output by the graph network model being greater than a threshold.
15. The method according to claim 12, wherein generating a fused feature map independent of the order of the plurality of images comprises: Using a permutation-invariant symmetric function to generate the fused feature map, wherein the permutation-invariant symmetric function takes the feature maps of the plurality of images as inputs, and the output of the permutation-invariant symmetric function is independent of the input order of the feature maps of the plurality of images.
16. The method according to claim 15, wherein the permutation-invariant symmetric function comprises a pointwise max pooling operation for the feature maps of the plurality of images.
17. A method for constructing a graph network model, comprising: Obtaining a target image in a focus exposure matrix (FEM) and images within its neighborhood, the target image and the images within the neighborhood having annotation information regarding an abnormal pattern; Extracting features of the target image and the images within the neighborhood; Constructing a graph model based on the features and the distances between the target image and the images within the neighborhood, the graph model including a target node corresponding to the target image and neighboring nodes corresponding to the images within the neighborhood; And Optimizing the graph network model based on the constructed graph model and the annotation information.
18. The method according to claim 17, wherein the annotation information indicates a classification of the abnormal pattern, and optimizing the graph network model includes: Optimizing the graph network model based on the classification indicated by the annotation information of the target node and the neighboring nodes.
19. The method according to claim 18, wherein, The classification indicates whether there is an abnormal pattern in the image, and in the case where the target image has an abnormal pattern, the neighborhood nodes corresponding to the image with the abnormal pattern are positive correlation nodes, and in the case where the target image does not include an abnormal pattern, the neighborhood nodes corresponding to the image without the abnormal pattern are positive correlation nodes.
20. The method according to claim 18, wherein the graph network model includes a graph convolutional network configured to output the correlation degree between the nodes of the graph model.
21. The method according to claim 20 further comprises: After the graph network model has been optimized, based on the correlation degree generated by the positive correlation nodes, determine a recommended correlation degree threshold for selecting neighborhood images related to the target image.
22. An electronic device, comprising: A processing unit and a memory, The processing unit executes the instructions in the memory, so that the electronic device executes the method according to any one of claims 1 to 11, any one of claims 12 to 16, or any one of claims 17 to 21.
23. A device for detecting abnormal patterns in a focus exposure matrix (FEM) image, the device comprising, A feature map extraction unit configured to use a deep network model to extract feature maps of a plurality of images in the FEM, the plurality of images including a target image and neighborhood images of the target image; A fusion unit configured to generate a fusion feature map independent of the order of the plurality of images based on the feature maps of the plurality of images; A prediction unit configured to use the deep network model to generate a predicted bounding box in the target image based on the fusion feature map; And An optimization unit configured to optimize the deep network model based on the predicted bounding box and an annotated bounding box for the abnormal pattern in the target image.
24. A device for detecting abnormal patterns in a focus exposure matrix (FEM) image, the device comprising: A feature map extraction unit configured to use an optimized deep network model to extract feature maps of a plurality of images in the FEM, the plurality of images including a target image and neighborhood images of the target image; A fusion unit configured to generate a fusion feature map independent of the order of the plurality of images based on the feature maps of the plurality of images; And A prediction unit configured to use the optimized deep network model to generate a predicted bounding box in the target image based on the fusion feature map.
25. A device for constructing a graph network model, the device comprising: An image acquisition unit configured to acquire a target image in a focus exposure matrix (FEM) and images within its neighborhood, the target image and the images within the neighborhood having annotation information about abnormal patterns; A feature extraction unit configured to extract features of the target image and the images within the neighborhood; A model construction unit, configured to construct a graph model based on the features and the distances between the target image and the images within the neighborhood, where the graph model includes a target node corresponding to the target image and neighboring nodes corresponding to the images within the neighborhood; And An optimization unit, configured to optimize the graph network model based on the constructed graph model and the annotation information.
26. A computer-readable storage medium, having one or more computer instructions stored thereon, where the one or more computer instructions, when executed by a processor, cause the processor to execute the method according to any one of claims 1 to 11, any one of claims 12 to 16, or any one of claims 17 to 21.
27. A computer program product, including machine-executable instructions, where the machine-executable instructions, when executed by a device, cause the device to execute the method according to any one of claims 1 to 11, any one of claims 12 to 16, or any one of claims 17 to 21.