Layout classification method, electronic equipment and computer readable storage medium
By sub-layer segmentation and geometric feature analysis of integrated circuit layouts, pre-classification and clustering, the problem of difficulty in predicting hot spots in lithography is solved, and more efficient clustering and higher yield rates are achieved.
Patent Information
- Application Number
- CN202510191811.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-10
AI Technical Summary
In the process of integrated circuit lithography, it is difficult to efficiently predict and locate possible hot spot areas, resulting in lithography defects and affecting chip yield.
By dividing the layout into multiple sub-layers, the similarity is determined based on the geometric characteristics of the sub-layers, and pre-classification and clustering are performed to determine the type and number of clusters of the sub-layers.
It achieves more efficient clustering and higher clustering accuracy, and can more accurately predict and locate hot spot areas, reduce lithography defects, and improve chip yield.
Smart Images

Figure CN120125864A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure mainly relate to integrated circuits, and more particularly, to a layout classification method, an electronic device, and a computer-readable storage medium. Background Art
[0002] Lithography is an important step in the manufacturing process of integrated circuits. The basic principle of lithography is to use a photoresist, which, after being exposed to light, undergoes a photochemical reaction to etch the pattern on the mask onto the surface to be processed. With the rapid development of very large scale integrated circuit technology, the feature size of transistors has become smaller and smaller, and the circuit design layout has become more and more complex, posing a huge challenge to circuit lithography technology.
[0003] Currently, the wavelength of light has reached the 193nm limit, which is much larger than the existing feature size of transistors. When etching a standard circuit design layout onto a silicon wafer, the diffraction of light causes the circuit pattern on the silicon wafer to change, resulting in defects, also known as hotspots. These hotspots are very likely to cause open or short circuits during the operation of the circuit, burning out the circuit and reducing the yield of the chip, resulting in huge economic losses. Therefore, it is necessary to predict and locate possible hotspot areas before lithography and repair their designs to avoid subsequent lithography defects.
[0004] Generally, hotspots are predicted by lithography simulation or machine learning detection methods. However, due to the large scale of chip problems, in order to reduce the problem scale, it is necessary to cluster the images obtained by clipping the layout, and select representative images for processing for each category. There are deficiencies in the efficiency and accuracy of layout clustering in traditional solutions. Summary of the Invention
[0005] According to an exemplary embodiment of the present disclosure, a layout processing solution is provided to at least partially overcome the above or other potential defects.
[0006] According to one aspect of the present disclosure, a layout classification method is provided. The method includes: determining the similarity of each sub-layout based on the geometric features of the patterns in a plurality of sub-layouts obtained by splitting a layout, where each sub-layout has the same size; and classifying the sub-layouts based on the similarity to determine the number of classifications.
[0007] According to a second aspect of the present disclosure, a layout processing method is provided. The method includes: determining the similarity of each sub-layout based on the geometric features of the graphics in a plurality of sub-layouts obtained by dividing a layout, where each sub-layout has the same size; pre-classifying the sub-layouts based on the similarity to determine the number of classifications; determining geometric feature values based on a feature hierarchy tree, where each layer in the feature hierarchy tree defines corresponding geometric feature information, and the geometric feature values indicate the feature values of each geometric feature obtained using the defined geometric feature information; and clustering the sub-layouts based on the number of classifications and the geometric feature values to determine the type of each sub-layout.
[0008] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes a processor; and a memory coupled to the processor, the memory having instructions stored therein that, when executed by the processor, cause the device to perform operations, the operations including: determining the similarity of each sub-layout based on the geometric features of the graphics in a plurality of sub-layouts obtained by dividing a layout, where each sub-layout has the same size; pre-classifying the sub-layouts based on the similarity to determine the number of classifications; determining geometric feature values based on a feature hierarchy tree, where each layer in the feature hierarchy tree defines corresponding geometric feature information, and the geometric feature values indicate the feature values of each geometric feature obtained based on the defined geometric feature information; and clustering the sub-layouts based on the number of classifications and the geometric feature values to determine the type of each sub-layout.
[0009] In some embodiments, determining the similarity of each sub-layout based on the features of the graphics in a plurality of sub-layouts obtained by dividing a layout includes: dividing each sub-layout into a central region and edge regions located on both sides of the central region; and determining the similarity of two sub-layouts based on the product of the similarity of the central regions of the two sub-layouts and the central similarity weight, and the product of the similarity of the two edge regions and the corresponding edge similarity weights.
[0010] In some embodiments, determining the similarity of two sub-layouts includes: adding the product of the similarity of the central region and the central similarity weight and the product of the similarity of the two edge regions and the corresponding edge similarity weights; and determining the sum value obtained by the addition as the similarity.
[0011] In some embodiments, the ratio of the area of the central region to the total area of the sub-layout is more than fifty percent. In some embodiments, classifying the sub-layouts based on the similarity includes: determining sub-layouts with a similarity greater than a predetermined similarity threshold as the same type.
[0012] In some embodiments, the feature hierarchy tree includes multiple layers, and the geometric feature information in each layer subdivides the corresponding geometric feature information in the previous layer. Determining the geometric feature value based on the feature hierarchy tree includes: classifying the geometric patterns in each sub-layout based on the geometric feature information defined in each layer; and traversing each layer in the feature hierarchy tree to output the feature value.
[0013] In some embodiments, the feature hierarchy tree includes at least one of the following information: the orientation information of the pattern; the quantity information of the pattern; the length information of the pattern; the maximum critical dimension information; the minimum critical dimension information; the distance between different patterns; the distance between adjacent sub-layouts; the geometric information of the pattern being flush or staggered; and the pixel-based feature vector information.
[0014] In some embodiments, the clustering label of each sub-layout is determined by one of the following clustering algorithms: the K-means clustering algorithm; and the density-based clustering algorithm.
[0015] In some embodiments, pre-classifying the sub-layouts based on similarity to determine the number of classifications includes: performing one-hot encoding on the geometric features based on similarity to pre-classify the geometric features; and determining the number of classifications based on the categories of each geometric pattern after pre-classification.
[0016] In some embodiments, it further includes: encoding the sub-layout by an encoder to generate an encoded image; and decoding the encoded image by a decoder to generate a pixel-based feature vector.
[0017] In some embodiments, determining the geometric feature value based on the feature hierarchy tree includes: defining the feature hierarchy tree based on the geometric features and the pixel-based feature vector; and performing binary classification on each geometric feature and the pixel-based feature vector to generate the geometric feature value.
[0018] In a fourth aspect of the present disclosure, another electronic device is provided. The electronic device includes a processor; and a memory coupled to the processor, the memory having instructions stored therein, and the instructions, when executed by the processor, cause the device to perform actions, the actions including: determining the similarity of each sub-layout based on the geometric features of the patterns in the multiple sub-layouts formed by layout segmentation, where each sub-layout has the same size; and classifying the sub-layouts based on the similarity to determine the number of classifications.
[0019] In some embodiments, determining the similarity of each sub-layout based on the geometric features of the patterns in the multiple sub-layouts formed by layout segmentation includes: determining the similarity of each pattern based on the comparison of the geometric features of the patterns in the corresponding regions of each layout; and determining the similarity of each sub-layout based on the similarity of each pattern respectively.
[0020] In some embodiments, determining the similarity of each sub-layout based on the geometric features of the graphics in multiple sub-layouts obtained by layout segmentation includes: dividing each of the sub-layouts into a main region and a secondary region; and determining the main region similarity based on the product of the area ratio of the main regions of two sub-layouts and the main region similarity weight; and determining the secondary region similarity based on the product of the area ratio of the secondary regions of two sub-layouts and the secondary region similarity weight; and determining the sum of the main region similarity and the secondary region similarity as the similarity of the two sub-layouts.
[0021] In some embodiments, the main region is the central region, and the secondary region is the edge region located on both sides of the central region.
[0022] In some embodiments, the ratio of the area of the central region to the total area of the sub-layout is more than fifty percent.
[0023] In some embodiments, the similarity of two sub-layouts is calculated using the following similarity formula:
[0024] S = αS A + βS B + γS C
[0025] where S represents the total similarity, α represents the similarity weight of the central region, β and γ respectively represent the similarity weights of the edge regions, the value ranges of α, β, and γ are all [0, 1], S A represents the similarity of the geometric features within the central region, S B and S C respectively represent the similarity of the geometric features within the edge regions.
[0026] In some embodiments, α = 1, β = γ = 0.5.
[0027] In some embodiments, classifying the sub-layouts based on the similarity includes: determining sub-layouts with a similarity greater than a predetermined similarity threshold as the same type.
[0028] In some embodiments, classifying the sub-layouts based on the similarity to determine the number of classifications includes: performing one-hot encoding on the geometric features based on the similarity to classify the sub-layouts; and determining the number of classifications based on the categories of the classified sub-layouts.
[0029] In the fifth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the methods according to the first and second aspects of the present disclosure are implemented.
[0030] It will be understood from the following description that the technical solutions of the present disclosure can achieve more efficient clustering and higher clustering accuracy.
[0031] The Summary of the Invention section is provided to introduce, in a simplified form, a selection of concepts that will be further described in the Detailed Description below. The Summary of the Invention section is not intended to identify key or essential features of the present disclosure, nor is it intended to limit the scope of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 A schematic diagram showing an example environment in which embodiments of the present disclosure can be implemented;
[0033] Figure 2A A flowchart showing a method for layout classification according to some embodiments of the present disclosure;
[0034] Figure 2B A flowchart showing a method for image processing according to some embodiments of the present disclosure;
[0035] Figure 3 A schematic diagram showing an autoencoder network structure according to some embodiments of the present disclosure;
[0036] Figure 4 A schematic diagram showing the training of an autoencoder network according to some embodiments of the present disclosure;
[0037] Figure 5 A schematic diagram showing a GDS feature vector decision tree according to some embodiments of the present disclosure;
[0038] Figure 6 A schematic diagram showing pattern similarity modeling according to some embodiments of the present disclosure;
[0039] Figure 7 A schematic diagram showing a method for image clustering according to some embodiments of the present disclosure;
[0040] Figure 8 A block diagram of a computing device capable of implementing multiple embodiments of the present disclosure.
[0041] In the respective drawings, the same or corresponding reference numerals indicate the same or corresponding parts. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] The principles of the present disclosure will be described below with reference to various exemplary embodiments shown in the accompanying drawings. It should be understood that the description of these embodiments is only for enabling those skilled in the art to better understand and further implement the present disclosure, and is not intended to limit the scope of the present disclosure in any way. It should be noted that, where feasible, similar or identical reference numerals may be used in the figures, and similar or identical reference numerals may represent similar or identical functions. Those skilled in the art will readily recognize that alternative embodiments of the structures and methods described herein may be employed without departing from the principles of the invention described herein from the following description.
[0043] As used herein, the term "comprising" and its variations mean open-ended inclusion, i.e., "including but not limited to". Unless specifically stated otherwise, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "an example embodiment" and "an embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc. may refer to different or the same objects.
[0044] Hot spot detection is an important step in semiconductor manufacturing to ensure the reliability and performance of integrated circuits (ICs). A hot spot is an area on a chip where overheating or stress may cause defects, thereby reducing the yield and affecting the lifetime and function of the device. Hot spots refer to, for example, line bridging, line breakage, and poor contact hole defects that occur during the lithography manufacturing process. As semiconductor technology nodes continue to shrink, detecting and mitigating these hot spots becomes increasingly important.
[0045] Typically, hot spots are predicted through lithography simulation or machine learning detection methods. However, due to the large scale of chip problems, in order to reduce the problem scale, it is necessary to cluster the images of the layout clips, and select representative images for each category. In this way, the problem is reduced to processing the representative images for each category instead of individual layout samples. Therefore, how to efficiently calculate the layout clustering number and extract layout features in advance before clustering plays a key role in the clustering of the full-chip layout.
[0046] The problem of Graphic Data System (GDS) layout pattern compression (clustering) has been around for a long time. How to efficiently and accurately cluster the GDS patterns of the full chip remains a major challenge in the industry. There are mainly two reasons for this: First, the number of GDS layout patterns of the full chip is on the order of 100 million (in a 4mm * 6mm area, 14nm process, 500nm * 500nm window), and the large amount of data limits the clustering efficiency and commercial viability; Second, the current industry-wide full-process framework for dealing with clustering problems in this sub-field is not yet unified and operable.
[0047] A known clustering method first obtains a circuit layout file, extracts the layout region blocks to be classified from the circuit layout file, then generates a layout region block association graph based on the layout region blocks to be classified, then obtains the complementary graph of the layout region block association graph, and calculates the maximum clique of the complementary graph to obtain the number of layout regions of the maximum clique. Finally, clustering is performed according to the number of layout regions. However, this clustering method only considers the attributes related to the boundaries of the layout, and there are no more geometric features as supplements. For the full-chip layout data, the clustering effect of one-dimensional (1D) patterns or two-dimensional (2D) patterns will not be very ideal and cannot achieve commercial effects.
[0048] In another known solution, by obtaining a sample chip layout and an initial encoder, geometric transformation is performed on the sample chip layout to obtain a reference chip layout; the layout features of the sample chip layout and the layout features of each reference chip layout are extracted through the initial encoder; based on the layout features of the sample chip layout and the layout features of each reference chip layout, the initial encoder is trained to obtain a chip layout encoder. For the chip layout before and after geometric transformation, the chip layout encoder can output similar layout features. However, this method only considers the image pixel features, so there are two drawbacks: First, considering the actual situation, if a new layout appears, a large amount of time is required for encoding and training, which reduces the clustering efficiency of the full-chip layout: Second, the single image data feature cannot fully represent the GDS clustering accuracy.
[0049] There is also a known solution that first customizes a series of feature libraries, then based on the customized feature libraries, saves the feature vectors in a vector database by means of feature extraction, and then clusters the layout using a method that combines supervised and unsupervised learning based on the extracted features. Finally, hotspot prediction is performed through the layout features obtained after clustering. Before using unsupervised learning in this solution, the number of clusters needs to be determined. These numbers of clusters will change in actual situations. How to efficiently determine the number of clusters K is still a great challenge.
[0050] In view of this, the present disclosure provides an improved solution.
[0051] An embodiment of the present disclosure proposes a layout classification method, which includes: determining the similarity of each sub-layout based on the geometric features of the graphics in a plurality of sub-layouts obtained by dividing the layout, where each sub-layout has the same size; and classifying the sub-layouts based on the similarity to determine the number of classifications.
[0052] In an embodiment of the present disclosure, a layout processing method is further proposed. The method includes: determining the similarity of each sub-layout based on the geometric features of the graphics in multiple sub-layouts obtained by dividing the layout, where each sub-layout has the same size; pre-classifying the sub-layouts based on the similarity to determine the number of classifications; determining geometric feature values based on a feature hierarchy tree, where each layer in the feature hierarchy tree defines corresponding geometric feature information, and the geometric feature values indicate the feature values of each geometric feature obtained by using the defined geometric feature information; clustering the sub-layouts based on the number of classifications and the geometric feature values to determine the types of each sub-layout. In the embodiment of the present disclosure, by performing pre-classification to obtain the number of clusters and then implementing a clustering algorithm, more accurate clustering labels can be obtained, that is, more accurate clustering can be achieved.
[0053] Embodiments of the present disclosure will be specifically described below with reference to the accompanying drawings.
[0054] Figure 1 FIG. shows a schematic diagram of an exemplary environment 100 in which embodiments of the present disclosure can be implemented. As Figure 1 shown, the exemplary environment 100 includes a computing device 110 and a client 120.
[0055] In some embodiments, the computing device 110 can interact with the client 120. For example, the computing device 110 can receive an input message from the client 120 and output a feedback message to the client 120. In some embodiments, the input message from the client 120 can be layout data. The computing device 110 can perform corresponding processing on the layout data and output the corresponding operation result to the client 120.
[0056] In some embodiments, the computing device 110 can include, but is not limited to, a personal computer, a server computer, a handheld or laptop device, a mobile device (such as a mobile phone, a personal digital assistant PDA, a media player, etc.), a consumer electronic product, a small computer, a large computer, cloud computing resources, etc.
[0057] It should be understood that describing the structure and function of the exemplary environment 100 only for exemplary purposes is not intended to limit the scope of the subject matter described herein. The subject matter described herein can be implemented in different structures and / or functions. This environment is only illustrative and is not used to limit the application environment of the embodiments of the present disclosure.
[0058] To more clearly explain the principle of the solution of the present disclosure, the following will be described in more detail with reference to Figure 2A and 2B for a more detailed description.
[0059] First, with reference to Figure 2A , Figure 2AShows a flowchart of a layout classification method 200A according to some embodiments of the present disclosure.
[0060] At block 202, the similarity of each sub-layout is determined based on the features of the graphics in the multiple sub-layouts obtained by dividing the layout, where each sub-layout has the same size.
[0061] In some embodiments, the GDS layout can be efficiently divided by a GDS parsing tool and a distributed processing platform (a parallel computing algorithm system based on multi-threading) into a large number of sub-layouts (a sub-layout can also be referred to as a pattern or image herein). That is, a large number of GDS patterns or images are finally clipped. There are multiple graphics (polygons) in each image; on the one hand, the GDS patterns can be stored in the database in the PNG format; on the other hand, the GDS-related point information ( Figure 7 the GDS coordinates shown in) can be saved to the GDS database for later processing, such as geometric feature processing. The GDS-related point information can be extracted from the image by traditional methods.
[0062] In some embodiments, traditional feature extraction methods can be used to extract the geometric features of the graphics from each image, such as the direction, length, critical dimension (CD), etc. of the graphics.
[0063] In some embodiments, determining the similarity of each sub-layout based on the geometric features of the graphics in the multiple sub-layouts obtained by dividing the layout may include: determining the similarity of each graphic based on the comparison of the geometric features of the graphics in the corresponding regions of each layout; and determining the similarity of each sub-layout based on the similarity of each graphic respectively.
[0064] In some embodiments, determining the similarity of each sub-layout based on the geometric features of the graphics in the multiple sub-layouts obtained by dividing the layout may include: dividing each of the sub-layouts into a main region and a secondary region respectively; and determining the main region similarity based on the product of the area ratio of the main regions of two sub-layouts and the main region similarity weight; and determining the secondary region similarity based on the product of the area ratio of the secondary regions of two sub-layouts and the secondary region similarity weight; and determining the sum of the main region similarity and the secondary region similarity as the similarity of the two sub-layouts.
[0065] In some embodiments, determining the similarity of each sub-layout based on the features of the graphics in the multiple sub-layouts obtained by dividing the layout may include: dividing each of the sub-layouts into a central region and edge regions on both sides of the central region; and determining the similarity of the two sub-layouts based on the product of the similarity of the central regions of the two sub-layouts and the central similarity weight, and the product of the similarity of the two edge regions and the corresponding edge similarity weights.
[0066] In some embodiments, determining the similarity of the two sub-layouts includes: adding the product of the similarity of the central region and the central similarity weight to the product of the similarities of the two edge regions and the corresponding edge similarity weights; and determining the sum value obtained from the addition as the similarity.
[0067] In some embodiments, the ratio of the area of the central region to the total area of the sub-layout is more than fifty percent.
[0068] To solve the pattern clustering problem, in some embodiments, patterns (graphical figures) can be modeled. For example, a pattern can be divided into upper, middle, and lower segments. It should be understood that the embodiments of the present disclosure are not limited thereto, and other divisions can be made according to actual needs. Refer to the following Figure 6 . Figure 6 shows a schematic diagram of pattern similarity modeling according to some embodiments of the present disclosure. As Figure 6 shown, a pattern is segmented and divided into three parts, A, B, and C. That is, the respective parts represented by frames 604, 602, and 606 respectively. The specific method of segmentation is: taking the center point of the entire pattern as the center point of the core clustering region, intercepting the region where the area of the symmetric part centered on the center point occupies a certain proportion of the total area, so as to obtain the region of part A; the remaining upper and lower parts can be B and C respectively. The areas of B and C can be determined according to actual needs, and they can be equal or unequal. In addition, in some embodiments, the center of the central region can also deviate from the center point of the entire image within a predetermined range.
[0069] In some embodiments, the similarity can be defined as follows: The similarity between two patterns consists of three parts, namely the similarity of region A, the similarity of region B, and the similarity of region C. Then the total similarity is defined as shown in the following formula (1):
[0070] S = αS A + βS B + γS C (1)
[0071] where α, β, and γ are the similarity weights of each region respectively, and their value ranges are all [0, 1]. S A represents the degree of similarity (i.e., similarity) of geometric features within region A, and is simply referred to as the similarity of part A. Similarly, S BIndicates the similarity of geometric features within part B, and Sc indicates the similarity of geometric features within part B. For the GDS clustering task, generally, α = 1, β = γ = 0.5 can be taken. In this case, that is, the weight of the middle region is the largest, and the weights of the upper and lower regions are relatively small. In this way, the similarity between two sub-layouts can be determined quickly and accurately. It should be understood that the proportions of α and β shown here are only illustrative and can be varied according to actual needs, for example, determined by the user according to actual needs.
[0072] In addition to the above three parts A, B, and C, at least two regions D and E can be respectively divided on both sides of B and C, and weights can be respectively set, and the weights can be lower than those of B and C. The present invention does not make specific limitations.
[0073] For example, to determine the similarity between two sub-layouts, the similarities of the three parts A, B, and C of the two can be respectively determined, which can be determined by comparing the similarity degrees of the geometric features of the graphics in each part. For example, if the geometric graphics in part A of the two are exactly the same, then the similarity of part A of the two is S A = 1. If 50% of the geometric graphics in part A of the two are the same (or similar), then the similarity of part A of the two is S A = 0.5. The same judgment can be made for part B and part C. Substituting the similarities of each part into formula (1), the similarity degree between any two sub-layouts can be determined. In this way, sub-layouts with the same similarity value can be determined as one class. For example, if 10 sub-layouts all have a total similarity of 0.9 (or deviate from 0.9 within a predetermined threshold (such as 0.05)), it can be considered that they have the same similarity, and thus these 10 sub-layouts can be classified into the first group; if 12 sub-layouts all have a total similarity of 0.8 (or deviate from 0.8 within a predetermined threshold range), it can be considered that they have the same similarity, and thus these 12 sub-layouts can be classified into the second group, and so on. In this way, classification of sub-layouts is achieved based on similarity, and then the value of K can be determined.
[0074] In some embodiments, GDS image feature extraction can also be performed to obtain a pixel-based feature vector. The pixel-based feature vector can be combined with geometric features for subsequent clustering processing to obtain higher clustering accuracy.
[0075] For GDS image feature extraction, an unsupervised deep learning network can be mainly used, that is, through an AutoEncoder network, whose main structure includes two main network architectures, an encoder and a decoder (the network structure diagram is as Figure 3As shown in the figure, the input image can be encoded into a representation in a low-dimensional latent space (low-dimensionally encoded image data) through an encoder module, and the low-dimensionally encoded image data can be decoded into the original image through a decoder module. Encoding is for dimensionality reduction to facilitate data processing, and decoding is for restoring the image.
[0076] Refer to the following Figure 3 for further description. Figure 3 FIG. shows a schematic diagram of an autoencoding network structure according to some embodiments of the present disclosure.
[0077] As Figure 3 shown, first, the input image (layout image) 302 is input. The encoder 304 performs encoding processing on the image 302 to generate an encoded image. By performing encoding processing on the image, feature extraction of the image can be achieved, thereby generating a low-dimensional image for easy processing. The low-dimensional image can be referred to as the latent space representation 306. The latent space representation is a method of representing data compressed into a low-dimensional space, which is usually used in machine learning and deep learning. This representation method is called the latent space. The original high-dimensional data can be mapped into a low-dimensional space through an encoder. This process usually involves data compression, using fewer dimensions to represent the original data while retaining important information as much as possible. The concept of the latent space is very important in deep learning. It can capture the essential features of the data while removing noise and redundant information. It can help the model learn the features of the data, simplify the data representation, and thus better discover the patterns in the data. Feature extraction of the image can include performing multiple downsampling processes on the image through the encoder to obtain respective downsampled features. Then, the decoder 308 can perform decoding processing on the low-dimensional image to reconstruct the image to obtain the restored image 310. That is, the image is restored through decoding processing. The decoding processing can include performing upsampling processing on the features to restore the image. In fact, this network structure is a network training model.
[0078] Through Figure 3 the autoencoding network structure shown, a pixel-based feature vector can be obtained. The pixel-based feature vector can be used for subsequent operations such as clustering processing. As is known in the industry, the basic element of an image is a pixel. Both geometric features and pixels can be used as the feature vector of the image.
[0079] Based on the sampled GDS image information, the model of the image encoding network and the decoding network can be trained through the AutoEncoder network structure. Based on the trained network model, any GDS image can be encoded to obtain a pixel-based feature vector. As mentioned before, the AutoEncoder network mainly consists of an encoder and a decoder, and its main function is to reduce the dimension.
[0080] The following is a further description with reference to Figure 4 Figure 2. Figure 4 FIG. 2 shows a schematic diagram of AutoEncoder network training according to some embodiments of the present disclosure. Where x represents the input image, the input image is encoded by the encoder 304 to generate an encoded image c, and the encoded image c is decoded by the decoder 308 to obtain the decoded image The input image is compared with the decoded image (i.e., the restored image), and the square of the difference between the two is denoted as Loss. That is, the loss between the restored image and the original image is calculated. When the loss is greater than a predetermined threshold, iterative processing is performed, that is, the input image is encoded again and the encoded image is decoded, and the difference between the two is calculated. When the difference is less than the predetermined threshold, the iterative processing can be stopped; or when the number of training times reaches the target number (such as 500 times), the iterative processing can be stopped.
[0081] In some embodiments, a GDS feature vector decision tree (referred to as a decision tree, also known as a feature vector hierarchical tree) can be used to extract features from an image. The following is a description of the decision tree with reference to Figure 5 FIG. 3.
[0082] Figure 5 FIG. 3 shows a schematic diagram of a GDS feature vector decision tree according to some embodiments of the present disclosure. Figure 5 The topmost square 502 in FIG. 3 represents the image to be processed.
[0083] First, define as Figure 5The six - layer features F1 - F6 shown: F1 represents the GDS graphic (Polygon) direction information. For example, the left square 504 can represent horizontal, and the right square 504 can represent vertical. F2 represents geometric information such as the number and length of GDS graphics. F3 represents the information of the maximum and minimum CD values within the region (each cropped picture). F4 represents the distances between different graphics within the region and the distances between adjacent two regions (for example, the first group of squares in F4 can represent the distances between different graphics within the region; the second group of squares can represent the distances between adjacent two regions). F5 represents the geometric information of the alignment and staggering of graphics. F6 represents the pixel feature information (i.e., the feature vector based on pixels) extracted by the AutoEncoder network. Secondly, according to the above - defined GDS geometric feature information and the feature extraction information extracted by the auto - encoder network, a classic decision - tree algorithm is used to classify each layer of features. For example, binary classification is performed, that is, each layer is judged until the traversal of the last - layer feature vector is completed, and the decision algorithm ends. The final feature - vector decision tree can output corresponding feature values for subsequent clustering processing. Specifically, for Figure 5 the embodiment shown, the output of the feature - hierarchy tree is a series of geometric feature values corresponding to F1 - F6, that is, the specific numerical values of geometric features, such as the length and width values of rectangles, spacing values, and so on. Through the feature - hierarchy tree, a fine - grained feature analysis of the input layout can be performed to obtain accurate geometric feature values of the graphics in the layout.
[0084] It should be understood that Figure 5 the feature - hierarchy tree shown is only an example. Theoretically, the number of sub - squares split from the squares in the same row is basically the same, but the actual situation may be different from the theory. It can be understood that the actual situation is a special case of the theory, and this special case will also change with different layouts. In other words, which square in the previous level to further divide can vary according to the actual situation.
[0085] Returning to Figure 2A Continue the description. At block 204, the sub - layouts are classified based on the similarity to determine the number of classifications.
[0086] In the traditional solution, the user specifies the clustering number K. In some embodiments of the present disclosure, the number of classifications can be determined by pre - classifying the sub - layouts to be used as the clustering number K for clustering the layout, which can achieve more accurate clustering.
[0087] In some embodiments, the geometric features of the pattern can be one-hot encoded according to the extracted geometric features, such as the length and width of a rectangle, the pitch, etc. Whether they are similar can be determined through one-hot encoding. For example, based on the similarity, one group of sub-layouts can be determined to be of one category, and another group of sub-layouts can be determined to be of a second category, and so on. Based on the full-chip pattern, the total number of clusters K can be initially obtained for use in subsequent clustering algorithms.
[0088] In some embodiments, classifying the sub-layouts based on the similarity includes: determining sub-layouts with a similarity greater than a predetermined similarity threshold as the same type. That is, classification is performed according to the similarity degree of the graphics in the sub-layouts, and those with a similarity degree reaching the predetermined threshold can be determined to be of the same type.
[0089] In some embodiments, classifying the sub-layouts based on similarity to determine the number of classifications may include: one-hot encoding the geometric features based on the similarity for pre-classifying the geometric features; and determining the number of classifications based on the categories of the respective geometric figures after classification.
[0090] The geometric features are related to the total number of clusters. Solving the maximum value of the total number of clusters K requires relying on the specific values of the geometric feature parameters for calculation. It is equivalent to obtaining a preliminary range of K using the geometric features, and then using k-Means for clustering.
[0091] The above combination Figure 2A illustrates a method for layout classification according to some embodiments of the present disclosure. This classification method can be used to determine the number of classifications K, and the value of K can be used for subsequent clustering of the cropped layout images.
[0092] Clustering is a process of classifying and organizing data members that are similar in certain aspects in a dataset. It is a technique for discovering this internal structure, and clustering techniques are often referred to as unsupervised learning.
[0093] The k-means clustering algorithm is the most well-known clustering algorithm. Due to its simplicity and efficiency, it has become the most widely used among all clustering algorithms. Given a set of data points and the number of clusters k required (k can be specified by the user), the k-means clustering algorithm repeatedly divides the data into k clusters according to a certain distance function.
[0094] In some embodiments of the present disclosure, K can be obtained based on a Pre-Clustering algorithm (the method of classifying the sub-layout pairs mentioned above can be called the Pre-Clustering algorithm), and based on a Feature hierarchical Tree, the MiniBatchKMeans algorithm is used to perform fine clustering and optimization on the patterns of the entire chip. Finally, the category corresponding to each pattern will be output, and each category can be output to the database in an encoded manner for subsequent batch clustering processing.
[0095] Different from the traditional method, in some embodiments of the present disclosure, the K value is obtained through calculation instead of being specified by the user, and more accurate clustering results can be obtained compared to the traditional method.
[0096] Refer to the following Figure 2B to describe the clustering method 200B. Figure 2B The boxes 202 and 204 in Figure 2A are the same as the boxes 202 and 204 in Figure 2A respectively, so they will not be described repeatedly. It should be noted that the classification in the Figure 2B description process can be called pre-classification in
[0097] because the subsequent clustering process will further classify.
[0098]
[0099] In some embodiments, the Feature hierarchical Tree may include multiple layers, and the geometric feature information in each layer subdivides the corresponding geometric feature information in the previous layer. Determining the geometric feature value based on the Feature hierarchical Tree may include: classifying the geometric figures in each sub-layout based on the geometric feature information defined in each layer; and traversing each layer in the Feature hierarchical Tree to output the feature value. In some embodiments, the Feature hierarchical Tree may at least include the following information: the direction information of the figure; the quantity information of the figure; the length information of the figure; the maximum critical dimension information; the minimum critical dimension information; the distance between different figures; the distance between adjacent regions (sub-layouts); the geometric information of the figure being flush or staggered; the feature information extracted based on the autoencoder network (pixel-based feature vector information).
[0100] In some embodiments, the method further includes: encoding the sub-layout through an encoder to generate an encoded image; and decoding the encoded image through a decoder to generate a pixel-based feature vector.
[0101] In some embodiments, determining geometric feature values based on a feature hierarchy tree may include: defining a feature hierarchy tree based on geometric features and pixel-based feature vectors; and classifying each geometric feature to generate geometric feature values.
[0102] At block 208, sub-layouts are clustered based on the number of classifications and geometric feature values to determine the type of each sub-layout. Determining the type of each sub-layout is to determine the clustering label of each sub-layout, and the clustering labels respectively indicate the types of the sub-layouts.
[0103] In some embodiments, determining the clustering label of each sub-layout includes clustering the sub-layouts by one of the following clustering algorithms to determine the clustering label of each sub-layout: the K-Means algorithm; and the density-based clustering algorithm.
[0104] The K-Means clustering algorithm is an iterative clustering analysis algorithm. Its steps are as follows: initially divide the data into K groups, then randomly select K objects as the initial clustering centers, and then calculate the distance between each object and each seed clustering center, and assign each object to the clustering center closest to it. The clustering centers and the objects assigned to them represent a cluster. Each time a sample is assigned, the clustering center of the cluster will be recalculated based on the existing objects in the cluster. This process will continue to repeat until a certain termination condition is met. The termination condition can be that no (or the minimum number of) objects are reassigned to different clusters, no (or the minimum number of) clustering centers change anymore, and the sum of squared errors is locally minimized.
[0105] In some embodiments, it further includes: processing the sub-layouts through an autoencoder to obtain pixel-based feature vectors; and generating the geometric feature values based on the geometric features and the pixel-based feature vectors.
[0106] The following is combined with Figure 7 for description.
[0107] Figure 7 shows a schematic diagram of an image clustering method according to some embodiments of the present disclosure. As Figure 7 shown, the GDS layout is parsed (specifically cropped here) to obtain a plurality of sub-layouts. The sub-layouts are input into the GDS database. The images of the obtained sub-layouts are stored in the database. In addition, the GDS coordinate information of the above images is also stored in the GDS database, and this information can be obtained from the images using traditional methods.
[0108] The AutoEncoder module (or AutoEncoder network) can perform encoding and decoding operations on an image to generate a pixel-based feature vector. In some embodiments, the pixel-based feature vector can be combined with the geometric features in each sub-layout obtained by feature extraction of the image, for example, by feature concatenation, to generate a feature hierarchy tree. The geometric features can be presented in the form of a vector. For example, some features can be represented as a 1×12 vector, and the pixel-based feature vector is also presented in the form of a vector, for example, a 1×100 vector. After connecting the two, a 1×112 vector is formed.
[0109] As mentioned above, pre-classification processing can be performed on the geometric features or the combination of geometric features and pixel-based feature vectors. Through the pre-classification processing, the number K of pre-classifications can be obtained. Inputting the number K and the feature hierarchy tree into K-Means for processing can cluster the features, and finally, cluster labels can be output. That is, through the processing of K-means, labels indicating each category of clustering can be obtained. In other words, the cluster labels respectively indicate the types of each sub-layout. The output labels can be a series of numerical values, and these numerical values can be encoded to form a string, and the string can be stored in the GDS database.
[0110] For example, the labels encoded in string form are as follows: label: 1_2_3_4; 1_2(30), 2_3(10), 1_2_3_4(10). Among them, the label 1_2 indicates that there are 30 sub-layouts that meet this type; the label 2_3 indicates that there are 10 sub-layouts that meet this type; the label 1_2_3_4 indicates that there are 10 sub-layouts that meet this type.
[0111] Figure 7 In the K-Means shown, the number of final feature classifications will be a value less than or equal to K. That is to say, this K is only a maximum value, and it depends on the actual situation. Some categories may not exist. Because the number of classifications in the pre-classification may change during the actual clustering process. For example, some features may be found not suitable to be separated into a single category during the actual clustering, or for other reasons.
[0112] The traditional K-Means algorithm requires specifying K. In some embodiments of the present disclosure, K can be automatically calculated through domain knowledge (layout-related geometric features) without specifying the value of K, and then the image can be clustered through a machine learning algorithm (such as k-Means) to achieve a higher accuracy rate than the industry's GDS clustering.
[0113] Some embodiments of the present disclosure provide methods for layout feature processing. It should be noted that the examples given in the above embodiments are only for illustrating the solutions of the embodiments of the present disclosure and do not limit the solutions of the present disclosure.
[0114] In some embodiments of the present disclosure, to solve the full-chip clustering problem, the following algorithm architecture is proposed. First, the full-chip GDS layout is sliced into pictures (or images) of a fixed size (e.g., 500*500 nm). For example, it can be sliced through known parsing tools. Secondly, geometric feature extraction can be performed on the sliced GDS patterns. Then, based on the geometric features, it can be pre-classified by encoding (e.g., one-hot encoding) to obtain the number of clusters and the cluster feature encoding. Finally, clustering can be performed based on the image data (pixel-based feature vectors) to obtain more accurate cluster labels (labels). For example, a more accurate cluster label can be obtained by combining the AutoEncoder network structure and the Kmeans algorithm; and the clustered label is output to the database.
[0115] It should be understood that the embodiments mentioned here are exemplary embodiments of the solutions of the present disclosure, and the embodiments of the present disclosure are not limited thereto. Some steps or features can be omitted, such as the AutoEncoder network structure. In the case of omitting the AutoEncoder network structure, the pixel-based feature vectors are not generated, and correspondingly, F6 mentioned above does not exist in the feature hierarchy tree.
[0116] The full-chip GDS can complete the image clustering task in a short time through this method, achieving commercial purposes. For example, through the GDS clustering algorithm, a compression ratio of about 1:500 of the full-chip pattern can be obtained, greatly reducing the processing time of the full-chip pattern. At the same time, this algorithm can be applied to the pattern clustering of the Hotspot Prediction task to improve its processing efficiency.
[0117] In the above embodiments, the full-chip layout is taken as an example for illustration. It should be understood that obviously, the method of the embodiments of the present disclosure is not limited to the full-chip and is equally applicable to non-full-chips.
[0118] The technical solution of the present disclosure can achieve more efficient clustering and higher clustering accuracy.
[0119] In addition, in the above embodiments, the K-Means algorithm is used to implement clustering. It should be understood that obviously the method of the embodiments of the present disclosure is not limited to this, but other clustering algorithms can be adopted as needed, such as the density-based clustering algorithm (Density-Based Spatial Clustering of Applications with Noise, abbreviated as DBSCAN), and so on.
[0120] In addition, the autoencoder in the above embodiments can also be replaced by other networks with similar functions. For example, a convolutional neural network (Convolutional Neural Networks, CNN) can be used to replace it.
[0121] In addition, in some embodiments, as mentioned above, the autoencoder can be omitted. If it is omitted, a pixel-based feature vector will not be generated. Compared with the above embodiments, the effect will be a little worse. Because the geometric features only consider the geometric features between the graphics in the image, and these features can be distinguished by the human eye. The pixel-based vector represents the higher-order image features of the black and white regions of the entire image that cannot be perceived by the human eye, and can depict the overall features of the image more deeply than the geometric features.
[0122] It should be understood that the embodiments shown in the drawings are only for schematically showing the solutions of some embodiments of the present disclosure and are not used to limit the present disclosure. The embodiments of the present disclosure can also have various other forms.
[0123] An electronic device is also disclosed in the embodiments of the present disclosure. The electronic device includes: a processor; and a memory coupled to the processor, the memory having instructions stored therein, and the instructions, when executed by the processor, cause the device to perform operations, the operations including: determining the similarity of each sub-layout based on the geometric features of the graphics in a plurality of sub-layouts obtained by layout segmentation, where each sub-layout has the same size; pre-classifying the sub-layouts based on the similarity to determine the number of classifications; determining geometric feature values based on a feature hierarchy tree, where each layer in the feature hierarchy tree defines corresponding geometric feature information, and the geometric feature values indicate the feature values of each geometric feature obtained based on the defined geometric feature information; and clustering the sub-layouts based on the number of classifications and the geometric feature values to determine the types of each sub-layout.
[0124] Another electronic device is also disclosed in the embodiments of the present disclosure. The electronic device includes a processor; and a memory coupled to the processor, the memory having instructions stored therein, and the instructions, when executed by the processor, cause the device to perform operations, the operations including: determining the similarity of each sub-layout based on the geometric features of the graphics in a plurality of sub-layouts obtained by layout segmentation, where each sub-layout has the same size; and classifying the sub-layouts based on the similarity to determine the number of classifications.
[0125] Embodiments of the present disclosure also disclose a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the layout classification or layout processing method according to the embodiments of the present disclosure.
[0126] Figure 8 A schematic block diagram of an electronic device according to some exemplary embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0127] As Figure 8 shown, the device 800 includes a CPU 801, which can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 802 or the computer program loaded from the storage unit 808 into the random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the device 800 can also be stored. The CPU 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. The input / output (I / O) interface 805 is also connected to the bus 804.
[0128] Multiple components in the device 800 are connected to the I / O interface 805, and the multiple components include: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0129] Each of the processes and treatments described above, such as methods 200A and 200B, can be executed by CPU 801. For example, in some embodiments, methods 200A and 200B can be implemented as computer software programs tangibly embodied in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by CPU 801, one or more steps of methods 200A and 200B described above can be executed.
[0130] The solutions according to the embodiments of the present disclosure can be a method, an apparatus, a system, and / or a computer program product. The computer program product can include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present disclosure. The computer-readable storage medium can be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable program instructions can be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network.
[0131] The embodiments of the present disclosure have been described above. The above description is exemplary and is only an optional embodiment of the present disclosure, not exhaustive, and is not used to limit the present disclosure. Although the claims in this application have been formulated for specific combinations of features, it should be understood that the scope of the present disclosure also includes any novel feature or any novel combination of features that are explicit or implicit or any generalization thereof disclosed herein, regardless of whether it relates to the same solution in any of the currently claimed claims. The applicant hereby notifies that new claims can be formulated into these features and / or combinations of these features during the examination process of this application or in any further application derived therefrom.
[0132] The selection of the terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other ordinary technicians in the technical field to understand the embodiments disclosed herein. For those skilled in the art, various changes and modifications can be made to the present disclosure. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A layout classification method, comprising: Determining the similarity of each sub-layout based on geometric features of graphics in a plurality of sub-layouts divided from the layout, wherein each sub-layout has the same size; as well as The sub-layouts are classified based on the similarities to determine the number of classifications.
2. The method according to claim 1, wherein determining the similarity of each sub-layout based on the geometric features of the graphics in the plurality of sub-layouts divided from the layout comprises: Determining the similarity of each graphic based on the comparison of geometric features of the graphics in the corresponding areas of each layout; as well as The similarity of each sub-board is determined based on the similarity of each graphic.
3. The method according to claim 2, wherein determining the similarity of each sub-layout based on the geometric features of the graphics in the plurality of sub-layouts divided from the layout comprises: Divide each of the sub-regions into a main region and a sub-region; Determine the similarity of the main region based on the product of the area proportion of the main region of the two sub-patterns and the main region similarity weight; Determine the sub-region similarity based on the product of the area proportion of the sub-regions of the two sub-maps and the sub-region similarity weight; as well as The sum of the main region similarity and the secondary region similarity is determined as the similarity between the two sub-patterns. The method according to claim 3 , wherein the primary region is a central region, and the secondary regions are edge regions located on both sides of the central region.
5. The method according to claim 4, wherein the area of the central region in the sub-layout accounts for more than fifty percent.
6. The method according to claim 4, wherein the similarity between the two sub-layouts is calculated using the following similarity formula: S=αS A +βS B +γS C Where S represents the total similarity, α represents the similarity weight of the central area, β and γ represent the similarity weight of one of the two edge areas respectively, and the value range of α, β and γ are all [0,1]. A represents the similarity of the geometric features in the central area, S B and S C They respectively represent the similarity of the geometric features in the corresponding edge regions. The method according to claim 6 , wherein α=1, β=γ=0.
5.
8. The method according to claim 1, wherein classifying the sub-layouts based on the similarity comprises: Sub-layouts with similarities greater than a predetermined similarity threshold are determined to be of the same type.
9. The method according to claim 1, wherein classifying the sub-layouts based on the similarity to determine the number of classifications comprises: One-hot encoding the geometric features based on the similarity to classify the sub-patterns; as well as The number of classifications is determined based on the categories of each of the classified sub-layouts.
10. An electronic device, comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, the instructions causing the device to perform actions when executed by the processor, the actions comprising: Determining the similarity of each sub-layout based on geometric features of graphics in a plurality of sub-layouts divided from the layout, wherein each sub-layout has the same size; as well as The sub-layouts are classified based on the similarities to determine the number of classifications.
11. A computer-readable storage medium having machine-executable instructions stored thereon, and when the machine-executable instructions are executed by a processor, the processor is enabled to implement the method according to any one of claims 1 to 9.