A method for classifying and integrating hot spots in photolithography processes

By extracting the local features of the lithography hotspot patterns and performing dimensionality reduction and classification of the feature vectors, the problems of large number of hotspot patterns and low matching efficiency in the lithography process hotspot pattern library are solved, and efficient hotspot pattern library integration and matching are achieved.

CN117197508BActive Publication Date: 2025-09-30ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310823739.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-06
Publication Date
2025-09-30
Estimated Expiration
2043-07-06

AI Technical Summary

Technical Problem

In the existing technology, the photolithography process hotspot graphic library has a large number of hotspot graphics, low matching efficiency, high time resource occupancy, a narrow scope of application of clustering effectiveness indicators, and no discussion of the selection method of representative graphics of photolithography process hotspots, resulting in poor classification and integration effect of photolithography process hotspots.

Method used

By extracting the local features of the lithography hotspot graphics, the SIFT algorithm and K-means clustering algorithm are used, combined with the Euclidean distance and similarity threshold, to reduce the dimension and classify the feature vectors. Representative feature vectors are selected as representative hotspot graphics to generate a compressed hotspot graphics library.

Benefits of technology

The calculation speed of the hotspot graphics library for the lithography process is accelerated, the memory usage is reduced, the efficiency of hotspot graphics matching is improved, and the running time and memory usage of the hotspot graphics library are optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117197508B_ABST
    Figure CN117197508B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for classifying and integrating photolithography process hotspots. First, a corresponding graphic area of ​​the same size is intercepted from each photolithography process hotspot in a graphic library to obtain a photolithography hotspot graphic. Then, local features are extracted from the intercepted photolithography hotspot graphic to obtain a feature data set. The local features in the feature data set are then clustered to output a normalized feature vector. Finally, the output feature vectors are reduced in dimension and sorted according to the average similarity value. The graphic mapped by the feature vector with the largest average similarity value is selected as a representative hotspot graphic and stored in a compressed hotspot graphic library. The present invention extracts local features of hotspot graphics and uses feature vectors to judge the similarity between graphics, thereby accelerating the calculation speed and reducing memory usage. In order to better adapt to the photolithography process hotspot clustering in the present invention, the weight coefficient of the clustering effectiveness index function is adjusted to obtain a more comprehensive and accurate clustering effect evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of photolithography technology, and in particular to a photolithography process hotspot classification and integration method. Background Art

[0002] Detecting hotspots in lithography layouts is a crucial technique in design for manufacturability. Typically, a suite of optical models, photoresist models, optical proximity correction programs, and lithography hotspot detection rules is used to perform lithography-friendly checks on design patterns, generating process hotspot markers and silicon wafer simulation patterns. This involves first performing optical proximity correction on the design pattern, then simulating the pattern within the process window, predicting the simulated pattern size within the process window, and then performing rule checks to identify design violations.

[0003] Graphic features are an important application in image processing and can be categorized into global and local features. Using local features instead of the entire image can significantly reduce the information carried in the original image and reduce computational complexity. Image features can be represented by feature descriptors, or feature vectors. The feature vectors of two images can be used to determine the similarity between them.

[0004] Cluster analysis is a statistical analysis method for studying classification problems and is also an important method in data mining. Numerous clustering algorithms have been proposed and applied to various datasets. Among them, the K-means algorithm is a partition-based clustering algorithm. Due to its simple implementation and high accuracy, the K-means algorithm is widely used to solve data partitioning problems in various fields. Clustering algorithms divide the dataset to be analyzed into multiple clusters, ensuring that data within the same cluster has greater similarity and data between different clusters has greater diversity. As an unsupervised learning method, the quality of the results produced by clustering algorithms is often measured or evaluated using clustering effectiveness metrics.

[0005] Existing hotspot graph libraries contain a large number of hotspot graphs, resulting in low hotspot graph matching efficiency and high time resource utilization. They typically determine hotspot graph similarity by moving the boundaries of lithography hotspot graphs. This method has significant limitations, long computation times, and high memory utilization, resulting in poor classification and integration of lithography hotspot graph libraries. Existing clustering effectiveness metrics suffer from a narrow range of applicable dataset types, and methods for selecting representative lithography hotspot graphs are not discussed in the prior art. Consequently, research into the lithography hotspot classification and integration process is somewhat lacking. Summary of the Invention

[0006] The present invention aims to address the deficiencies of the prior art and propose a method for classifying and integrating hotspots in photolithography processes.

[0007] The object of the present invention is achieved through the following technical solution: a method for classifying and integrating hot spots in engraving processes, the method comprising the following steps:

[0008] S1, intercepting the corresponding graphic area of ​​the same size from each photolithography hotspot in the graphic library to obtain a photolithography hotspot graphic;

[0009] S2, extracting local features from the intercepted lithography hotspot pattern to obtain a feature data set;

[0010] S3, clustering the local features in the feature data set and outputting a normalized feature vector;

[0011] S4. Input the output feature vector into the matching network: set the similarity threshold, reduce the dimension of the feature vector and classify and group it. In each group, sort it according to the average similarity. The graph mapped by the feature vector with the largest average similarity is used as the representative hotspot graph of this type, stored in the compressed hotspot graph library and output.

[0012] Furthermore, in the S1 , each photolithography process hotspot intercepts the corresponding graphic area into a square, and the sides of each intercepted square are equal in length.

[0013] Furthermore, the extraction of local features is achieved by using a SIFT algorithm.

[0014] Furthermore, the clustering range for clustering local features is a square area of ​​the same size.

[0015] Furthermore, clustering the local features is specifically as follows:

[0016] S3.1. Calculate the Euclidean distance between each pair of local features, and determine the Euclidean distance threshold based on the maximum and minimum values ​​of the Euclidean distance;

[0017] S3.2. Count the number of data points whose Euclidean distance to each data point is less than a threshold value, and use this as the density information of the data points;

[0018] S3.3. Select the two points with the largest density information as the initial center points;

[0019] S3.4. Cluster the dataset using the K-means clustering algorithm based on the initial center point; calculate the clustering effectiveness index through the global intra-class similarity of the dataset and the global inter-class separation of the dataset; output the clustering result of the feature point dataset when the clustering effectiveness index is optimal.

[0020] Furthermore, the method for determining the threshold ε is ε=(D max +D min ) 2 / (2*C max ), where Cmax is the upper limit of the number of clusters of the pre-set data set to be clustered, and C max Not greater than A positive integer, P is the minimum value of the local feature in all lithography hotspot patterns; D max and D min are the maximum and minimum values ​​of the Euclidean distance between local feature points, respectively.

[0021] Furthermore, the function of the clustering effectiveness index DCVI(K) is:

[0022]

[0023] Where P is the minimum value of the local feature in all lithography hotspot patterns; K is the total number of classes in the dataset; com(K) means the global intra-class similarity of the dataset, and sep(K) means the global inter-class separation of the dataset.

[0024] The clustering result when the clustering effectiveness index function value is the smallest is the best clustering result, and the best cluster number K is onc The method for determining is:

[0025]

[0026] When the clustering effectiveness index is optimal, the clustering result of the feature point dataset is output.

[0027] Furthermore, the dimensionality reduction of the output feature vector specifically includes principal component analysis, linear discriminant analysis, independent component analysis, multidimensional scaling or random forest method.

[0028] Furthermore, the similarity average value is calculated as follows: after obtaining the feature vectors after dimensionality reduction, the Euclidean distance between the feature vectors is calculated, and a similarity threshold is set to classify the feature vectors that meet the threshold range into one category;

[0029] The Euclidean distance is expressed as similarity according to the following formula, with a range of (0,1], and the smaller the distance, the greater the similarity;

[0030]

[0031] The dist(A,B) means two vectors A(a1, a2, ..., a n ) and B(b1,b2,...,b n ) between them;

[0032] After obtaining the classified and grouped eigenvectors, the arithmetic mean of the similarity between each eigenvector and other eigenvectors is calculated and sorted. The eigenvector with the largest average similarity is taken as the representative eigenvector. The arithmetic mean of the similarity of the jth eigenvector in the i-th group is calculated as follows:

[0033]

[0034] Where η = α i -1, α i is the total number of group class eigenvectors of the i-th group, η is the number of eigenvectors of the i-th group excluding the j-th eigenvector, and simωj is the similarity between the ω-th eigenvector and the j-th eigenvector in the i-th group excluding the j-th eigenvector.

[0035] Beneficial effects of the present invention:

[0036] The present invention proposes a method for classifying and integrating hot spots in photolithography processes. By extracting local features of hot spot graphics and judging the similarity between graphics through feature vectors, the calculation speed is accelerated and memory usage is reduced. The classification and integration effect of the hot spot graphics library in photolithography processes is good.

[0037] The present invention modifies the clustering effectiveness index function. In order to be more suitable for the lithography process hotspot clustering in the present invention, the weight coefficients before the global intra-class similarity and global inter-class separation of the data set in the clustering effectiveness index function are adjusted to balance the importance of the two indicators and obtain a more comprehensive and accurate clustering effect evaluation.

[0038] The present invention calculates the mean similarity value, classifies and integrates similar graphics in the lithography process hotspot graphic library into a representative graphic, generates a classified and integrated lithography process hotspot graphic library, reduces the number of similar graphics in the hotspot library, and further optimizes the running time and memory usage in the hotspot graphic matching process.

[0039] The present invention uses a MatchingNet matching network structure, the input of which is the feature vector of the local features of the hotspot graphics of the lithography process. The representative feature graphic selection rule proposed by the present invention is used to reduce the number of similar graphics. The output is a compressed hotspot graphic library, which further improves the efficiency of the hotspot matching of the lithography process and reduces memory usage. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 A flow chart for classifying and integrating lithography hotspots provided in an embodiment of the present invention;

[0041] Figure 2 A flow chart of a clustering algorithm provided by an embodiment of the present invention;

[0042] Figure 3A schematic diagram of a hot spot graph of a photolithography process provided by an embodiment of the present invention;

[0043] Figure 4 Schematic diagram of local feature extraction using the SIFT algorithm provided in an embodiment of the present invention;

[0044] Figure 5 Schematic diagram of local feature K-means cluster analysis provided by an embodiment of the present invention;

[0045] Figure 6 A schematic diagram of a normalized output feature vector provided by an embodiment of the present invention;

[0046] Figure 7 A diagram of the MatchingNet matching network structure provided by an embodiment of the present invention;

[0047] Figure 8 A schematic diagram of the Euclidean distance provided by an embodiment of the present invention;

[0048] Figure 9 A schematic diagram of a similarity matrix provided in an embodiment of the present invention;

[0049] Figure 10 A feature vector grouping diagram provided by an embodiment of the present invention;

[0050] Figure 11 A schematic diagram of selecting representative feature vectors provided in an embodiment of the present invention;

[0051] Figure 12 A schematic diagram of the steps for obtaining a compressed hotspot graphics library provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0052] The specific embodiments of the present invention are further described in detail below with reference to the accompanying drawings.

[0053] like Figure 1 As shown, the present invention provides a method for classifying and integrating hot spots in photolithography processes;

[0054] Provides a graphics library with multiple lithography hotspots.

[0055] Taking the photolithography hotspot as the center, a partial area of ​​the pattern corresponding to each photolithography hotspot is intercepted, and the total number of photolithography hotspot patterns obtained is n. In the present invention, the partial area of ​​each pattern corresponding to the photolithography hotspot is intercepted with the same size, that is, the side lengths of the square areas of each pattern corresponding to the photolithography hotspot are equal, such as Figure 3 As shown, in the embodiment, two square areas of the same size are cut out, namely a first photolithography hotspot area 301 and a second photolithography hotspot area 302 , and the side length of the square is 2r, for example, a side length of 0.8 um to 1.2 um.

[0056] For each of the hotspot patterns in the photolithography process, the SIFT algorithm is used to extract the local features of the hotspot pattern and mark the positions of the local feature points. The number of feature points is 10-16 and is recorded as P. Figure 4 As shown, the local features extracted from the two photolithography process hotspot patterns in the embodiment include the first local feature point 401, the second local feature point 402, the third local feature point 403, the fourth local feature point 404, the fifth local feature point 405, the sixth local feature point 406, the seventh local feature point 407, the eighth local feature point 408, the ninth local feature point 409, the tenth local feature point 410, the eleventh local feature point 411, the twelfth local feature point 412, the thirteenth local feature point 413, the fourteenth local feature point 414, the fifteenth local feature point 415, the sixteenth local feature point 416, the seventeenth local feature point 417, the eighteenth local feature point 418, the nineteenth local feature point 419, and the twentieth local feature point 420, thereby obtaining a local feature point dataset.

[0057] according to Figure 2 The clustering algorithm process in the clustering algorithm is used to cluster the local feature point data set. For the clustering algorithm, the upper limit of the number of clusters to be clustered in the data set is set to Cmax, which is not greater than A positive integer, P is the minimum value of the local feature in all lithography hotspot patterns, calculate the Euclidean distance between local feature points, find the maximum value Dmax and the minimum value Dmin of the Euclidean distance, and determine the threshold value according to Dmax and Dmin. The method for determining the threshold value is ε=(D max +D min ) 2 / (2*C max ), the number of data points with a Euclidean distance less than a threshold from each data point is counted as the density information of the data point, and the two points with the largest density information are selected as the initial cluster centers. The number of local feature points extracted from each hotspot graph may be different, but is within 10-16. For the embodiment of the present invention, the specific value of P is 10. After calculating the C in the embodiment of the present invention, max The value of is 4.

[0058] The clustering range is a square area of ​​the same size, with a side length of 0.3um to 0.45um. The number of more representative feature points obtained by the clustering algorithm is m. Figure 5In the embodiment, among the local feature points of the first photolithography process hotspot area 301, the first local feature point 401, the second local feature point 402, and the third local feature point 403 are a cluster area, the fourth local feature point 404 is a cluster area, the fifth local feature point 405, the sixth local feature point 406, and the seventh local feature point 407 are a cluster area, the eighth local feature point 408 and the ninth local feature point 409 are a cluster area, and among the local feature points of the second photolithography process hotspot area 302, the twelfth local feature point 412 and the twelfth local feature point 413 are cluster areas. The thirteenth local feature point 413 is a cluster area, the fifteenth local feature point 415 and the sixteenth local feature point 416 are a cluster area, the seventeenth local feature point 417 and the eighteenth local feature point 418 are a cluster area, the nineteenth local feature point 419 and the twentieth local feature point 420 are a cluster area, and the cluster feature points of the first photolithography process hot spot area 301 are the first cluster feature point 501, the second cluster feature point 502, the third cluster feature point 503, and the fourth cluster feature point 504, among which the first cluster feature point 501 The first local feature point 401, the second local feature point 402, and the third local feature point 403 are clustered together. The second cluster feature point 502 is clustered together with the eighth local feature point 408 and the ninth local feature point 409. The third cluster feature point 503 is clustered together with the fourth local feature point 404. The fourth cluster feature point 504 is clustered together with the fifth local feature point 405, the sixth local feature point 406, and the seventh local feature point 407. The cluster feature points 505, 506, and 507 of the hot spot area 302 of the second lithography process are clustered together. 06, 507, 508, wherein the fifth cluster feature point 505 is obtained by clustering the local feature point 12 and the 13th local feature point 413, the sixth cluster feature point 506 is obtained by clustering the local feature point 19 and the 20th local feature point 420, the seventh cluster feature point 507 is obtained by clustering the local feature point 15 and the 16th local feature point 416, and the eighth cluster feature point 508 is obtained by clustering the local feature point 17 and the 18th local feature point 417.

[0059] Calculate the Euclidean distance between each data point in the data set and the two initial cluster center points respectively, select the cluster center point within the clustering range of 0.1um as the center point category of the data point, and set the class label of each data point to the category of the cluster center point. For the data in each category, set the virtual center point class label to the in-class data category, and the coordinate information of each dimension of the virtual center point is the arithmetic mean of the coordinate information of each dimension of the in-class data that does not contain density information.

[0060] The clustering effectiveness index is calculated based on the clustering results as follows: divide the data set into K classes C = {C1, C2, ..., CK}, where the number of sample points contained in the k-th class is |Ck|, the cluster center of the class is ck, xki is the i-th data point in the K-th class, and the intra-class similarity of the k-th class is:

[0061]

[0062] The global intra-class similarity of the dataset is:

[0063]

[0064] The global inter-class separation of the dataset is:

[0065]

[0066] si is the i-th cluster center, d(si,sj) is the Euclidean distance between the i-th cluster center and the j-th cluster center, and the cluster validity index function DCVI(K) is:

[0067]

[0068] In order to better adapt to the lithography process hotspot clustering in the present invention, the weight coefficients before the global intra-class similarity and global inter-class separation of the data set in the clustering effectiveness index function are adjusted to balance the importance of the two indicators and obtain a more comprehensive and accurate clustering effect evaluation.

[0069] The clustering result when the clustering effectiveness index function value is the smallest is the best clustering result, and the best cluster number K is onc The method to determine is:

[0070]

[0071] When the clustering effectiveness index is optimal, the clustering result of the feature point dataset is output.

[0072] like Figure 6 As shown, after obtaining the clustering result of the feature point data set, the normalized feature vector of the cluster feature point is output. The normalized feature vector is a matrix vector with 128 rows and m columns. The normalized feature vector of the first photolithography process hot spot area 301 is expressed as The normalized feature vector of the hotspot area 302 of the second lithography process is expressed as Indicates that the element values ​​are between 0 and 1. and It can be expressed by the following formula:

[0073]

[0074]

[0075] like Figure 7 As shown in the figure, the output n 128-row and m-column normalized feature vectors are input into the MatchingNet matching network, and the feature vectors are subjected to dimensionality reduction processing, so that their dimensions are reduced from the original 128 dimensions to 32 dimensions. The dimensionality reduction methods used can be principal component analysis, linear discriminant analysis, independent component analysis, multidimensional scaling, random forest and other methods.

[0076] After obtaining the 32-dimensional feature vector after dimensionality reduction, similarity calculation is performed on the feature vector. Specifically, the similarity between feature vectors is calculated using the Euclidean distance between feature vectors, and the similarity threshold ST is set to 0.8. Feature vectors that meet the threshold range are classified into one category. The total number of group feature vectors of group i is ɑ i .

[0077] like Figure 8 As shown, for two vectors A(a1, a2, ..., a n ) and B(b1,b2,...,b n The Euclidean distance calculation formula between ) is:

[0078]

[0079] The matrix representation method is:

[0080]

[0081] The Euclidean distance is expressed as similarity according to the following formula, with a range of (0,1]. The smaller the distance, the greater the similarity.

[0082]

[0083] After obtaining the classified and grouped eigenvectors, the arithmetic mean of the similarity between each eigenvector and other eigenvectors is calculated and sorted. The eigenvector with the largest average similarity is taken as the representative eigenvector. The arithmetic mean of the similarity of the jth eigenvector in the i-th group is calculated as follows:

[0084]

[0085] Where η = α i -1, α i is the total number of group class eigenvectors of the i-th group, η is the number of eigenvectors of the i-th group excluding the j-th eigenvector, and simωj is the similarity between the ω-th eigenvector and the j-th eigenvector in the i-th group excluding the j-th eigenvector.

[0086] like Figure 9 As shown, after calculating the similarity between each pair of eigenvectors, the similarity matrix between the eigenvectors f1, f2, f3, f4, f5, f6, f7, f8, and f9 is listed, where the similarity between the eigenvectors themselves is 1, the similarity between the eigenvectors f1 and f2 is 0.06, the similarity between the eigenvectors f1 and f3 is 0.95, the similarity between the eigenvectors f1 and f4 is 0.21, the similarity between the eigenvectors f1 and f5 is 0.28, the similarity between the eigenvectors f1 and f6 is 0.97, the similarity between the eigenvectors f1 and f7 is 0.06, and the similarity between the eigenvectors f1 and f The similarity between the feature vectors f1 and f9 is 0.17, the similarity between the feature vectors f2 and f3 is 0.23, the similarity between the feature vectors f2 and f4 is 0.98, the similarity between the feature vectors f2 and f5 is 0.35, the similarity between the feature vectors f2 and f6 is 0.15, the similarity between the feature vectors f2 and f7 is 0.96, the similarity between the feature vectors f2 and f8 is 0.05, the similarity between the feature vectors f2 and f9 is 0.28, the similarity between the feature vectors f3 and f4 is 0.21, and the similarity between the feature vectors f3 and f5 is 0. The similarity between the feature vectors f3 and f6 is 0.99, the similarity between the feature vectors f3 and f7 is 0.12, the similarity between the feature vectors f3 and f8 is 0.28, the similarity between the feature vectors f3 and f9 is 0.11, the similarity between the feature vectors f4 and f5 is 0.06, the similarity between the feature vectors f4 and f6 is 0.17, the similarity between the feature vectors f4 and f7 is 0.94, the similarity between the feature vectors f4 and f8 is 0.13, the similarity between the feature vectors f4 and f9 is 0.29, and the similarity between the feature vectors f5 and f6 is 0.31. The similarity between feature vectors f5 and f7 is 0.28, the similarity between feature vectors f5 and f8 is 0.97, the similarity between feature vectors f5 and f9 is 0.99, the similarity between feature vectors f6 and f7 is 0.04, the similarity between feature vectors f6 and f8 is 0.18, the similarity between feature vectors f6 and f9 is 0.25, the similarity between feature vectors f7 and f8 is 0.12, the similarity between feature vectors f7 and f9 is 0.19, and the similarity between feature vectors f8 and f9 is 0.96. According to the set similarity threshold ST of 0.8, the feature vectors are grouped.

[0087] like Figure 10In the illustrated embodiment, the feature vectors are divided into three groups: Cluster 1, Cluster 2, and Cluster 3. Cluster 1 includes feature vectors f1, f3, and f6, the similarity between f1 and f3 is 0.95, the similarity between f1 and f6 is 0.97, and the similarity between f3 and f6 is 0.99. Cluster 2 includes feature vectors f2, f4, and f7, the similarity between f2 and f4 is 0.98, the similarity between f2 and f7 is 0.97, and the similarity between f4 and f7 is 0.94. Cluster 3 includes feature vectors f5, f8, and f9, the similarity between f5 and f8 is 0.97, the similarity between f5 and f9 is 0.99, and the similarity between f8 and f9 is 0.96.

[0088] like Figure 11 As shown in the figure, the average similarity between each eigenvector and other eigenvectors in the group is calculated, and the similarity averages are sorted in the group. The eigenvector with the largest average similarity is taken as the representative eigenvector of the group. Among them, the representative eigenvector in Cluster 1 is f6, the representative eigenvector in Cluster 2 is f2, and the representative eigenvector in Cluster 3 is f5.

[0089] like Figure 12 As shown, the hotspot graphics mapped by the representative feature vectors of each group are stored as representative hotspot graphics in the compressed hotspot graphics library. The hotspot graphics P1, P2, P3, P4, P5, P6, P7, P8, and P9 in the hotspot graphics library in the embodiment are compressed to obtain a compressed hotspot graphics library that only includes hotspot graphics P2, P5, and P6.

[0090] Through the above method steps, the hotspot graphic library is further integrated and classified to generate a hotspot graphic library with fewer hotspots, thereby improving the efficiency of hotspot graphic matching.

[0091] This method classifies and integrates a total of 9,826 hotspot graphics in the hotspot graphic library, and finally generates a hotspot graphic library with 112 hotspots, which greatly reduces the number of hotspots in the hotspot graphic library, is conducive to checking the design layout, improves the efficiency of hotspot graphic matching, reduces time and resource usage, and accelerates the tape-out process.

[0092] The above embodiments are used to illustrate the present invention rather than to limit the present invention. Any modifications and changes made to the present invention within the spirit of the present invention and the protection scope of the claims shall fall within the protection scope of the present invention.

Claims

1. A method for classifying and integrating hot spots in photolithography processes, characterized in that: The method comprises the following steps: S1, intercepting the corresponding graphic area of ​​the same size from each photolithography hotspot in the graphic library to obtain a photolithography hotspot graphic; S2. Extracting local features from the intercepted lithography hotspot pattern to obtain a feature data set; extracting local features is achieved by using a SIFT algorithm; S3. Clustering the local features in the feature data set and outputting a normalized feature vector. The clustering of the local features in the feature data set is specifically as follows: S3.

1. Calculate the Euclidean distance between each pair of local features and determine the Euclidean distance threshold based on the maximum and minimum values ​​of the Euclidean distance. The method for determining the Euclidean distance threshold ε is ε=(D max +D min ) 2 / (2*C max ), where C max is the upper limit of the number of clusters to be clustered for the pre-set feature data set, and C max Not greater than A positive integer, P is the minimum value of the local feature in all lithography hotspot patterns; D max and D min are the maximum and minimum values ​​of the Euclidean distance between local feature data points respectively; S3.

2. Count the number of data points whose Euclidean distance to each data point is less than a threshold value, and use this as the density information of the data points; S3.

3. Select the two points with the largest density information as the initial center points; S3.

4. Cluster the feature dataset using the K-means clustering algorithm based on the initial center point; calculate the clustering effectiveness index based on the global intra-class similarity and the global inter-class separation of the feature dataset; output the feature dataset clustering result when the clustering effectiveness index is optimal; the function DCVI(K) of the clustering effectiveness index is: Where P is the minimum value of the local feature in all lithography hotspot patterns; K is the total number of classes in the feature dataset; com(K) means the global intra-class similarity of the dataset, and sep(K) means the global inter-class separation of the dataset; The clustering result when the clustering effectiveness index function value is the smallest is the best clustering result, and the best cluster number K is onc The method to determine is: When the clustering effectiveness index is optimal, the clustering result of the feature data set is output; S4. Input the output feature vector into the matching network: set the similarity threshold, reduce the dimension of the feature vector and classify and group it. In each group, sort it according to the average similarity. The graph mapped by the feature vector with the largest average similarity is used as the representative hotspot graph of this type, stored in the compressed hotspot graph library and output.

2. The method for classifying and integrating hot spots in photolithography processes according to claim 1, characterized in that: In the S1 , each photolithography hotspot cuts out a corresponding graphic area into a square.

3. The method for classifying and integrating hot spots in photolithography process according to claim 2, characterized in that: The clustering range for clustering local features in the feature dataset is a square area of ​​the same size.

4. The method for classifying and integrating hot spots in photolithography process according to claim 1, characterized in that: Methods for dimensionality reduction of feature vectors include principal component analysis, linear discriminant analysis, independent component analysis, multidimensional scaling, or random forest methods.

5. The method for classifying and integrating hot spots in photolithography process according to claim 1, characterized in that: The specific method for calculating the average similarity is as follows: after obtaining the feature vectors after dimensionality reduction, calculating the Euclidean distance between the feature vectors, and setting a similarity threshold, and classifying the feature vectors that meet the threshold range into one category; The Euclidean distance is expressed as similarity according to the following formula, with a range of (0,1], and the smaller the distance, the greater the similarity; Where dist(A,B) means two vectors A(a1,a2,...,a n ) and B(b1,b2,...,b n ) between them; After obtaining the classified and grouped eigenvectors, the arithmetic mean of the similarity between each eigenvector and other eigenvectors is calculated and sorted. The eigenvector with the largest average similarity is taken as the representative eigenvector. The arithmetic mean of the similarity of the jth eigenvector in the i-th group is calculated as follows: Where η = α i -1, α i is the total number of group class eigenvectors of the i-th group, η is the number of eigenvectors of the i-th group excluding the j-th eigenvector, and simωj is the similarity between the ω-th eigenvector and the j-th eigenvector in the i-th group excluding the j-th eigenvector.