Space-time coupling virtual metering intelligent labeling method based on feature space geometry
By employing a virtual metrology intelligent annotation method that combines feature space geometry and spatiotemporal coupling, samples with high curvature and scarce regions are identified. An adaptive weight model and multi-objective optimization algorithm are established, solving the problem of low annotation efficiency for extremely imbalanced datasets in semiconductor manufacturing and achieving high-precision sample selection and quality prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-19
- Publication Date
- 2026-04-10
AI Technical Summary
Existing virtual metrology technologies, when processing extremely imbalanced datasets in semiconductor manufacturing, suffer from several drawbacks: reliance on model predictions for sample value assessment, neglect of high-dimensional feature space geometry, lack of adaptive spatiotemporal analysis, and lack of hierarchical optimization strategies. These issues result in low annotation efficiency and make it difficult to achieve high-precision defect detection rates.
A spatiotemporal coupled virtual metrology intelligent annotation method based on feature space geometry is adopted. By identifying samples in high curvature and sparse regions, an adaptive temporal and spatial weight model is established. Combined with a multi-objective optimization algorithm, the global optimal sample combination and intelligent annotation are achieved.
It improves labeling efficiency, increases defect detection rate, meets the requirements of semiconductor manufacturing for high-precision quality prediction, and reduces labeling costs.
Smart Images

Figure CN121542791B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of semiconductor manufacturing quality control, and particularly relates to a time-space coupling virtual metrology intelligent labeling method based on feature space geometry. BACKGROUND
[0002] In the process of semiconductor manufacturing, virtual metrology technology can significantly reduce the cost of physical measurement by analyzing process sensor data to predict product quality status. Taking a 12-inch wafer production line as an example, the typical defective rate is controlled at about 1.8%, forming a highly unbalanced data distribution. This extremely unbalanced data feature poses a severe challenge to the sample labeling strategy of the virtual metrology system. However, existing virtual metrology technologies have significant technical limitations when dealing with such extremely unbalanced data sets:
[0003] Firstly, existing sample value evaluation methods are heavily dependent on model prediction results. Traditional active learning methods such as uncertainty sampling and diversity sampling require the prediction probability of the model to be improved to evaluate the value of the sample, forming a circular dependency technical dilemma: sample value evaluation depends on model prediction, while model improvement depends on sample labeling results. This circular dependency limits the labeling effect to the quality of the initial model, and in the extremely unbalanced data environment of semiconductor manufacturing, the initial model often cannot provide reliable prediction probability, thereby affecting the accuracy of sample selection.
[0004] Secondly, traditional methods ignore the geometric structure characteristics of high-dimensional feature space. Semiconductor manufacturing equipment is usually equipped with 200-300 sensors, and after dimensionality reduction processing such as principal component analysis, 64-dimensional feature vectors are obtained. These high-dimensional feature vectors form complex geometric manifold structures in the feature space. High-value samples are often located at special geometric positions of the feature manifold, such as high-curvature regions near the decision boundary, low-density regions, and regions with sharp gradient changes. However, existing technologies fail to effectively utilize these geometric structure information, resulting in low labeling efficiency.
[0005] Thirdly, existing time-space analysis methods lack adaptability. Semiconductor manufacturing has obvious time sequence characteristics, and the complexity and distribution characteristics of data in different periods differ significantly. Fixed time weight functions cannot adapt to such dynamic changes, and fixed spatial influence radius cannot adapt to the density differences in different regions. This lack of adaptability in time-space analysis methods cannot accurately capture the dynamic change rules of sample value.
[0006] Finally, there is a lack of hierarchical optimization strategy for high-dimensional feature space. Existing sample selection methods mostly use simple greedy strategies, which are prone to local optimization in high-dimensional space and cannot achieve global optimal sample combination. Especially in a 64-dimensional feature space, the similarity measurement and diversity guarantee between samples become more complex, requiring a special hierarchical optimization algorithm.
[0007] These limitations of the prior art result in that the labeling efficiency of the virtual metrology system is low in an extremely unbalanced data environment, and the rejection rate under the same labeling budget can only reach about 82%, which is difficult to meet the strict requirements of semiconductor manufacturing for high-precision quality prediction. SUMMARY
[0008] The purpose of the present application is to provide a feature space geometry-based spatiotemporal coupling virtual metrology intelligent labeling method capable of reducing labeling costs.
[0009] In order to achieve the above purpose, the technical scheme adopted by the present application is as follows:
[0010] The feature space geometry-based spatiotemporal coupling virtual metrology intelligent labeling method comprises the following steps executed in sequence:
[0011] S1: Obtain multi-dimensional sensor time series data of a semiconductor manufacturing equipment as a labeling sample set, obtain two geometric feature indicators of the labeling sample set located in a high-curvature area of a feature manifold and a rare area and a distribution boundary: identify samples located in the high-curvature area of the feature manifold and samples located in the rare area and the distribution boundary in the labeling sample set, and fuse the two geometric feature indicators into a unified manifold value evaluation;
[0012] S2: Establish an adaptive time weight model based on data complexity, establish a dynamic spatial weight model based on local sparsity based on an adaptive kernel function, and combine the adaptive time weight model, the dynamic spatial weight model and the manifold value evaluation to define a spatiotemporal coupling value score;
[0013] S3: Take the spatiotemporal coupling value score as input, establish a three-layer optimization architecture to realize globally optimal sample combination, specifically comprising the following steps:
[0014] S3-1: Establish a multi-objective optimization clustering selection model, determine the optimal cluster number by minimizing information entropy, and divide the samples into a preset number of clusters;
[0015] S3-2: Based on multi-objective optimization theory, comprehensively consider the diversity, value and redundancy of the clusters, obtain cluster scores through weighted calculation, and allocate labeling budgets for each cluster according to the cluster scores;
[0016] S3-3: In each cluster, an adaptive greedy algorithm is used to select the most representative samples through dynamic adjustment of the diversity weight;
[0017] S4: Intelligent labeling execution and result output: merge the selected samples of each cluster to form a final labeling set, perform quality evaluation on the final labeling set, and generate standardized labeling output according to the quality evaluation result.
[0018] Preferably, step S1 comprises the following steps:
[0019] S1-1: Identify the samples in the labeled sample set that are located in the high-curvature region of the feature manifold, the identification process is as follows:
[0020] Based on the theory of differential geometry, the sample is calculated using the following formula: :
[0021] ;
[0022] wherein, is the matrix determinant, is the L2 norm, and the index 3 / 2 is based on the standard form of curvature in differential geometry, is the local density function at the sample ;
[0023] The local feature gradient vector is calculated using the following formula:
[0024] ;
[0025] wherein, is the number of neighbors, is the Gaussian kernel function, is the kernel density estimation bandwidth parameter, is the neighbor set of the sample , is the feature dimension;
[0026] S1-2: Identify the scarce region and distribution boundary in the labeled sample set, the identification process is as follows:
[0027] The density gradient value of the sample is calculated using an improved kernel density gradient estimation: :
[0028] ;
[0029] wherein, is the square of the gradient norm of the density function at , is the local density of the sample , is the global density median, is the exponential function, the ratio is the relative density, and the local density is calculated by Gaussian kernel density estimation: , wherein For the number of nearest neighbors, To estimate the bandwidth parameter for kernel density, For feature dimensions;
[0030] density gradient Calculated using numerical differentiation:
[0031] ;
[0032] in, Let be the differential step size. For the first One standard basis vector;
[0033] S1-3: Integrating two geometric property indices into a unified manifold value assessment:
[0034] ;
[0035] in, For local curvature weights, For density gradient weights, .
[0036] Preferably, step S2 includes the following steps:
[0037] S2-1: Establish an adaptive time weighting model based on data complexity:
[0038] ;
[0039] in, For a moment Time weighting This represents the total number of key time points. For the first The importance weight of each time point satisfies , For the first Key time points, For adaptive time scale parameters, The square of the time distance;
[0040] Adaptive time scale parameters It is expressed by the following formula:
[0041] ;
[0042] in, Based on the basic time scale This is a complexity adjustment factor, used for data complexity scoring. Defined as: ;
[0043] wherein, denotes the feature distribution at the current time t, denotes the historical baseline feature distribution of the past 30 days;
[0044] S2-2: Establish a dynamic spatial weight model based on local sparsity:
[0045] ;
[0046] wherein, is an adaptive Gaussian kernel function, and the adaptive local scale is adjusted according to the local sparsity:
[0047] ;
[0048] wherein, is the local scale parameter of the sample , denotes the median of the distance in the neighborhood set, is a sparsity adjustment coefficient, is the local sparsity, is defined as:
[0049] ;
[0050] wherein, is the volume of a dimensional hypersphere, is the kth nearest neighbor distance of the sample ; The adaptive kernel function
[0051] is:
[0052] ;
[0053] wherein, is the Euclidean distance between samples;
[0054] S2-3: Combine the adaptive time weight model, the dynamic spatial weight model and the manifold value evaluation to define the spatio-temporal coupling value score as:
[0055] .
[0056] Preferably, step S3 comprises the following specific steps:
[0057] S3-1: Establish a multi-objective optimization clustering selection model, and determine the optimal cluster number using the information entropy minimization principle:
[0058] ;
[0059] wherein the cluster information entropy H(C) is defined based on Shannon entropy theory: , is a complexity penalty coefficient, represents the independent variable that makes the objective function minimum;
[0060] The cluster complexity Ω(C) is defined as: ;
[0061] wherein, is a cluster number penalty coefficient, is an intra-cluster compactness weight, is an average intra-cluster distance, and the optimal cluster number C* is obtained by minimizing the above information entropy, and the candidate sample set is divided into C* clusters;
[0062] S3-2: Combining the spatio-temporal coupling value of each sample to perform budget optimization allocation, and the budget allocation formula is:
[0063] ;
[0064] wherein C* is the optimal cluster number, is a softmax temperature parameter, is a label budget allocated to the cluster , satisfying , is a total labeling budget; is a comprehensive score of the cluster ;
[0065] The comprehensive score function is: ;
[0066] wherein the cluster diversity is defined based on geometric diversity theory:
[0067] ;
[0068] wherein, for any two samples in the cluster, there is , the distance set , is the median of the distance set, is the standard deviation;
[0069] The cluster value adopts a weighted combination of the maximum value and the mean value:
[0070] ;
[0071] wherein, maximizing the weight of the maximum value, averaging the weight of the values;
[0072] cluster redundancy based on nearest neighbor distance statistics:
[0073]
[0074] wherein, is a near-neighbor number penalty function, is the number of neighborhood samples having a distance less than a threshold value, is the Euclidean distance between samples;
[0075] S3-3: Cluster-internal sample selection algorithm: For the j-th cluster , sample selection is performed according to the budget allocation using an iterative selection strategy: initialize the set of selected samples , repeatedly perform the selection process until samples have been selected, in each iteration, select the sample from the remaining candidate samples of the cluster that maximizes the objective function:
[0076]
[0077] wherein, is the single sample selected in the current iteration, X is a candidate sample not yet selected from the cluster Cj, S is the set of samples currently selected, the selection objective function is defined as:
[0078]
[0079] wherein, is a diversity enhancement factor,
[0080] When the set S of selected samples is empty, the diversity factor is 1, otherwise, the minimum distance between the candidate sample X and all samples in the set S of selected samples is calculated, the diversity enhancement is quantified by the ratio of this distance to the average distance within the cluster , μ is an adaptive diversity weight, is a base diversity weight, is an adaptive adjustment coefficient.
[0081] Preferably, step S4 comprises the following steps:
[0082] S4-1: Final Sample Set Generation: The samples selected from each cluster are merged to form the final labeled set, which is represented by the following formula:
[0083] ;
[0084] in, For the first The sample set selected for each cluster The final labeled set is a collection of samples from all clusters.
[0085] S4-2: Annotation Quality Assessment and Verification: The final annotation set is assessed for quality, a diversity index is calculated, and the geometric distance distribution between selected samples is measured.
[0086] ;
[0087] Evaluation criteria: Diversity ≥ 1.2 indicates an excellent labeled set, meaning the sample distribution is uniform and the feature space is adequately covered;
[0088] 0.8 ≤ Ddiversity < 1.2 indicates a well-labeled set, meaning the sample distribution is moderate and meets the requirements.
[0089] Ddiversity < 0.8 indicates a labeled dataset that needs improvement, meaning the samples are too concentrated; it is recommended to adjust the algorithm parameters.
[0090] S4-3: Generate standardized annotation output, which includes: feature data of selected samples, spatiotemporal coupling value score, and reasons for sample selection, including cluster classification, geometric characteristics, and value contribution.
[0091] By adopting the aforementioned design scheme, the beneficial effects of the present invention are as follows: This application selects samples worth labeling or verifying from the labeled sample set, establishes a sample value quantification theory based on the geometric structure of the feature space, avoids dependence on model prediction results, derives the labeling value of the sample from the inherent geometric characteristics of the feature distribution, and outputs the manifold value of each sample as a comprehensive quantification result of the sample geometric features. This value integrates local geometric curvature, density gradient and uncertainty features.
[0092] This application establishes an adaptive time weight model based on data complexity and a dynamic spatial weight model based on local sparsity to achieve adaptive adjustment of spatiotemporal parameters, and outputs spatiotemporal coupling value as the final value metric of the sample for hierarchical intelligent selection. This value integrates geometric features, temporal dynamics and spatial correlation. The coupling mechanism of this spatiotemporal coupling value ensures that the sample value assessment can adapt to the dynamic changes in semiconductor manufacturing, and provides an accurate value basis for subsequent intelligent selection.
[0093] The application adopts a three-layer optimization strategy in hierarchical intelligent selection, namely, information entropy clustering determines sample grouping, multi-objective function optimization allocates budgets, and an adaptive greedy algorithm realizes intra-cluster selection, which can effectively solve the local optimal problem of traditional methods. BRIEF DESCRIPTION OF DRAWINGS
[0094] Figure 1 A flowchart of the intelligent labeling method of the application. DETAILED DESCRIPTION
[0095] In order to make the purpose, technical scheme and advantages of the application clearer, the application will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only some of the embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the application.
[0096] The terms "first", "second", "third" and the like in the specification and claims of the application and the above drawings are used to distinguish different objects, not to describe a particular order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.
[0097] The space-time coupling virtual metering intelligent labeling method based on feature space geometry, as shown in Figure 1 includes the following steps executed in sequence:
[0098] S1: Obtain multi-dimensional sensor time series data of a semiconductor manufacturing equipment as a labeling sample set, obtain two geometric characteristic indexes of the labeling sample set located in a high-curvature area of a feature manifold and a rare area and a distribution boundary: identify the samples in the labeling sample set located in the high-curvature area of the feature manifold and the samples located in the rare area and the distribution boundary, fuse the two geometric characteristic indexes into a unified manifold value evaluation and output as a comprehensive quantitative result of the sample geometric characteristics; this step avoids the dependence on the model prediction result by establishing a sample value quantitative theory based on the geometric structure of the feature space, and derives the labeling value of the sample from the inherent geometric characteristics of the feature distribution.
[0099] In this embodiment, the multi-dimensional sensor time series data includes the radio frequency power, chamber pressure, gas flow combination and temperature gradient distribution of the etching device, wherein the temperature gradient distribution includes the temperature difference and temperature uniformity of different regions, and the power matching efficiency of the PECVD device (specifically the power loss, efficiency and waveform stability of each power channel), the reaction gas ratio (specifically including the flow ratio of different gases, such as the ratio of oxygen to fluoride), the heating temperature curve (specifically including the temperature rise rate, steady-state temperature and fluctuation range).
[0100] S1-1: In step S1, samples in the high-curvature region of the feature manifold in the labeled sample set are identified, which are usually close to the decision boundary and have high labeling value. The identification process is as follows:
[0101] Based on the theory of differential geometry, the local curvature value of sample xi is calculated by the following formula: The local curvature matrix on the feature manifold :
[0102] ;
[0103] wherein, is the local curvature value of sample xi, and the value range is ; is the local second-order structure matrix (used to approximate the local curvature / geometric complexity) constructed in the k-neighborhood of sample xi, which has the same dimension as the feature dimension, and is , The outer product mean value of k-neighbor sample difference is calculated as follows: , is the matrix determinant, reflecting the local geometric complexity; is the L2 norm, calculating the Euclidean length of the vector; the power index is 3 / 2, which is used for scale normalization and numerical stabilization of the curvature measure, is the local density function at sample ;
[0104] The local feature gradient vector is calculated by the following formula:
[0105] ;
[0106] wherein, is the number of neighbors, which is determined based on the optimal neighborhood theory of : in a 64-dimensional feature space, the optimal neighborhood size should be close to , considering the balance between calculation efficiency and numerical stability, The local geometric structure features can be effectively captured while avoiding noise accumulation effect in high-dimensional space; wherein, Nk(i) represents a k-neighbor sample index set of the sample xi; xp represents a feature vector of the neighbor sample with index p; p is a neighbor sample index, and p∈Nk(i) and p≠i; is a kernel function, and in the embodiment, is a Gaussian kernel function, ; is a kernel density estimation bandwidth parameter, and is calculated according to a Silverman optimal bandwidth formula: wherein is a standard deviation of the feature data, and for 64-dimensional standardized features, , is a sample number, and is calculated as: ; is a neighbor set of the sample .
[0107] S1-2: In step S1, the rare areas and distribution boundaries in the labeled sample set are identified, and the samples in these areas have important value for improving the model generalization ability. The distribution boundary characteristics are accurately captured through the density gradient, and the identification process is as follows:
[0108] The density gradient value of the sample is calculated by using the improved kernel density gradient estimation:
[0109] ;
[0110] wherein, the value range of , is a gradient norm square of the density function at , reflecting the degree of density change, is a local density of the sample , which is calculated by using the Gaussian kernel density estimation, is a global density median, which is used for normalization processing; is an exponential function, , the ratio is a relative density, reflecting the rarity of the sample in the global distribution, and the local density is calculated by using the Gaussian kernel density estimation: wherein is a neighbor number, , is a kernel density estimation bandwidth parameter, , is a feature dimension, ;
[0111] Density gradient The differential is calculated by numerical differentiation:
[0112] ;
[0113] wherein, is the differential step size, the step size is selected according to: according to the theory of numerical analysis, the optimal step size , wherein is the machine precision of the double-precision floating-point number, is the order of magnitude of the function value of the density function at the current point, for the kernel density function in the present application, the typical value , according to the theoretical formula: , for the present application scenario, considering the engineering balance of calculation efficiency and numerical stability, the actual selection ; is the th standard basis vector.
[0114] S1-3: fuse two geometric indicators into a unified manifold value evaluation:
[0115] ;
[0116] wherein, is the local curvature weight, which reflects the importance of the decision boundary, and is set to the highest weight because high curvature areas are usually close to the classification boundary, and the value of the label is the largest, is the density gradient weight, which reflects the value of scarcity, and the two weights satisfy the normalization constraint: . The weight distribution reflects the importance level of the sample label value, wherein the local curvature weight 0.7 reflects the key role of the proximity of the decision boundary, and the density gradient weight 0.3 reflects the importance of the value of scarcity. The weight configuration ensures that the manifold value can accurately reflect the comprehensive geometric value of the sample in the feature space, and provides a stable and reliable value evaluation basis for the subsequent steps. After step S1 is completed, the manifold value of each sample is output, which comprehensively considers the local geometric curvature, the density gradient and the uncertainty feature, and is transmitted to step S2 as the comprehensive quantitative result of the sample geometric feature.
[0117] S2: establish an adaptive time weight model based on data complexity, calculate a data complexity score, establish a dynamic spatial weight model based on local sparsity, construct an adaptive kernel function, combine the adaptive time weight model, the dynamic spatial weight model and the manifold value evaluation, and define a spatio-temporal coupling value score as: and output the spatio-temporal coupling value Vcoupled(xi, t) as the final value metric of the sample, which integrates the geometric features, temporal dynamics, and spatial correlations; to avoid dimension inconsistency, the above three are normalized to [0, 1] before participating in the product; where xi is the ith sample, and t is its timestamp.
[0118] Based on the manifold value output in step S1 This step establishes an adaptive spatio-temporal coupling model to dynamically modulate the static geometric value according to the production time sequence characteristics and spatial distribution characteristics, and convert it into a dynamic value that adapts to the actual application scenario. Traditional spatio-temporal analysis methods use fixed time weight functions and spatial influence radii, which cannot adapt to the significant differences in data complexity and distribution characteristics at different stages of the semiconductor manufacturing process. This method realizes the adaptive adjustment of spatio-temporal parameters by establishing an adaptive time weight model based on data complexity and a dynamic spatial weight model based on local sparsity. The coupling mechanism of spatio-temporal coupling value ensures that the sample value evaluation can adapt to the dynamic change characteristics of semiconductor manufacturing, providing accurate value basis for intelligent selection in step S3.
[0119] S2-1: According to the change of data complexity at different stages, the decay parameter of time influence is adaptively adjusted.
[0120] An adaptive time weight model based on data complexity is established:
[0121] ;
[0122] wherein, is the time weight at time t, with a value range of [0, 1], is the total number of key time nodes, usually including equipment is the importance weight of the ith time node, satisfying ; is the ith key time point (hour), such as is the equipment maintenance time, is the process switching time, is the adaptive time scale parameter, which is dynamically adjusted according to data complexity, is the square of the time distance, used to calculate the time decay, and the Gaussian kernel form ensures the smooth decay characteristics of the time influence. Adaptive time scale formula is:
[0123] Adaptive time scale formula is:
[0124] ;
[0125] wherein, Using hours as the base timescale, determined based on semiconductor manufacturing batch cycles: Statistical analysis shows that the processing cycle for a single batch of 12-inch wafers is approximately 8-16 hours, with the median of 12 hours taken as the base timescale. This is a complexity adjustment factor, determined based on the characteristics of semiconductor manufacturing processes: when the data complexity score is 1.0 (normal state), the time scale remains at the base value; when the complexity score reaches 2.0 (high complexity), the time scale is increased by 30%. The adaptive adjustment of the time window was ensured to be within a reasonable range; data complexity score. Defined as: ;
[0126] in, This represents the characteristic distribution at the current time t. This represents the historical baseline characteristic distribution over the past 30 days. This time window is determined based on the semiconductor manufacturing process stability cycle and is used to capture long-term process baseline states.
[0127] S2-2: Dynamically adjust the spatial influence radius based on the local density characteristics of the samples in the feature space.
[0128] Establish a dynamic spatial weighting model based on local sparsity:
[0129] ;
[0130] in, An adaptive Gaussian kernel function is used, satisfying the kernel function properties: non-negativity, normalization, and symmetry; and adapting to local scaling. Adjust based on local sparsity:
[0131] ;
[0132] in, For the sample Local scale parameters, Indicates in The median distance in a nearest neighbor set. This is the sparsity adjustment coefficient. For local sparsity, A larger value indicates that the samples in that region are more sparse. The value ranges from [0, +∞), and is usually defined as:
[0133] ;
[0134] in, for The volume of a hypersphere is based on high-dimensional geometry theory. for The volume factor of a 1D hypersphere. For the first neighborhood distance, the formula is used to accurately calculate the local neighborhood volume in high-dimensional feature space.
[0135] Adaptive kernel function For:
[0136] ;
[0137] wherein, is a sparsity adjustment coefficient, used to control the degree of influence of local sparsity on the spatial scale. When the local sparsity is 1.0, the spatial scale increases by 20%, and when the sparsity reaches 2.0, the spatial scale increases by 40%. This parameter ensures that the spatial weight can adapt to the density difference of different regions in the feature space, is the Euclidean distance between samples.
[0138] S2-3: combine the adaptive time weight model, the dynamic spatial weight model and the manifold value evaluation to define the spatio-temporal coupled value score as:
[0139] ;
[0140] The product form embodies the mutual enhancement effect of geometric value, temporal dynamics and spatial correlation. The product form is selected because the spatio-temporal factors have the characteristics of mutual amplification rather than linear superposition in the influence of sample value. Step S2 outputs the spatio-temporal coupled value Vcoupled(xi,t), which integrates geometric features, temporal dynamics and spatial correlation, as the final value measurement of the sample and is passed to step S3 for hierarchical intelligent selection.
[0141] S3: take the spatio-temporal coupled value score of step S2 as input, and establish a three-layer optimization architecture to realize the globally optimal sample combination: first, determine the optimal number of clusters by minimizing the information entropy, divide the samples into a preset number of clusters, and the preset number is set according to actual needs; then, based on the multi-objective optimization theory, comprehensively consider the diversity, value and redundancy of the clusters, obtain the cluster score by weighted calculation, and allocate the labeling budget of each cluster accordingly; finally, use an adaptive greedy algorithm in each cluster to select the most representative sample through dynamic adjustment of the diversity weight μ. This hierarchical strategy effectively avoids the local optimization problem of traditional greedy methods and realizes the global optimization of labeling budget allocation.
[0142] This step adopts a three-layer optimization strategy: information entropy clustering determines sample grouping, multi-objective function optimization allocates budget, and adaptive greedy algorithm realizes intra-cluster selection, effectively solving the local optimization problem of traditional methods.
[0143] S3-1: Clustering optimization driven by information theory, the minimum principle of information entropy is used to determine the optimal number of clusters, and a multi-objective optimization clustering selection model is established:
[0144] ;
[0145] Wherein, the clustering information entropy H(C) is defined based on Shannon entropy theory: , unit: bit, used to measure the uncertainty of clustering, when a certain cluster is empty (|Cc| = 0), 0·log2(0) = 0 is agreed, and the entropy value range is [0, log2(C)]; is the complexity penalty coefficient, which is determined based on the trade-off between model complexity and generalization ability in information theory: too small λ value will lead to over-segmentation, and too large λ value will lead to under-segmentation. λ = 0.05 balances between theoretical optimality and engineering feasibility; represents the independent variable that makes the objective function minimum, if there are multiple parallel minimum values, the smallest C is taken by default.
[0146] The clustering complexity Ω(C) is defined as:
[0147] ;
[0148] Wherein, is the cluster number penalty coefficient, is the intra-cluster compactness weight, is the average intra-cluster distance (reflecting sample similarity), and this complexity term is used to prevent too many clusters and ensure the similarity and interpretability of intra-cluster samples. The optimal number of clusters C* is obtained by minimizing the information entropy, and the candidate sample set is divided into C* clusters. The division results and cluster characteristics of each cluster provide decision basis for the budget allocation of S3-2.
[0149] ;
[0150] ;
[0151] Wherein, N = 50000 is the total number of samples; is the total annotation budget; is the lower bound of the number of clusters, which ensures the basic class separability; is the upper bound of the number of clusters, which prevents over-segmentation; 1000 is the standardization constant, which is determined based on the experience of large-scale data sets; Lower bound estimation based on sample number and budget proportion; Upper bound estimation based on information theory complexity. After obtaining the optimal cluster number C* by minimizing the information entropy, the candidate sample set is divided into C* clusters, each containing samples with similar characteristics. The division results and characteristics of each cluster are directly used for multi-objective budget allocation in S3-2.
[0152] S3-2: Multi-objective optimized budget allocation, based on the optimal cluster number C* obtained in S3-1 and the division of each cluster, combined with the spatio-temporal coupling value of each sample , to perform budget optimization and allocation, the budget allocation formula is:
[0153] ;
[0154] where C* is the total number of clusters, is the temperature parameter, controlling the smoothness of the distribution;
[0155] The comprehensive score function is: ;
[0156] where cluster diversity Based on the definition of geometric diversity theory:
[0157] ;
[0158] To strengthen the contribution of long-distance sample pairs:
[0159] ;
[0160] where for any two samples in the cluster , there is , the distance set , is the median of the distance set, is the standard deviation, used to normalize the distance weight and control the influence of extreme values.
[0161] Cluster value Adopt a weighted combination of maximum and mean value:
[0162] ;
[0163] where is the maximum weight, ensuring that high-value samples are not missed; is the average weight, considering the overall quality of the cluster.
[0164] Cluster redundancy Based on nearest neighbor distance statistics:
[0165] ;
[0166] in, Let the nearest neighbor number penalty function be used. For distance samples Less than the threshold The number of neighborhood samples, The distance between samples is the Euclidean distance.
[0167] Weight configuration: Diversity weights are applied to ensure sample coverage; Assigning the highest weight to a value reflects the dominant role in maximizing value. The weights are used for redundancy to avoid repeated selection of similar samples; the weights satisfy the normalization constraint. .
[0168] The Softmax function implements continuous budget allocation, with temperature parameters... Control the degree of allocation concentration. To ensure that each cluster receives at least one budget, post-processing is employed: ;
[0169] in, To be assigned to clusters The budget for the labeling needs to be met. ; Total budget for annotations; For clusters The overall score reflects the labeling value of the cluster; The softmax temperature parameter controls the smoothness of the distribution. , , Clusters Indicators of diversity, value, and redundancy; The spatiotemporal coupling value calculated in step S2; For clusters The number of samples.
[0170] The budget allocation results for each cluster obtained through the above calculations satisfy the following conditions: This provides a budget constraint for the selection of intra-cluster samples in S3-3.
[0171] S3-3: In-cluster sample selection algorithm, which selects the optimal combination of samples within each cluster, balancing sample value and local diversity. In-cluster sample selection needs to avoid repeated selection of locally similar samples to ensure local diversity. The new method addresses this issue through diversity constraints.
[0172] Based on the submodular function optimization theory, an improved greedy algorithm combined with diversity constraints is adopted. For the j-th cluster Cj, sample selection is performed according to the budget Bj allocated in step S3-2, using an iterative selection strategy: initializing the selected sample set S = , the selection process is repeated until Bj samples are selected. In each iteration, a sample that maximizes the objective function is selected from the remaining candidate samples of cluster Cj:
[0173] ;
[0174] where, is the single sample selected in the current iteration, X is the candidate samples in cluster Cj that have not been selected, S is the current selected sample set, and the selection objective function is defined as:
[0175] ;
[0176] This function synthetically evaluates the selection value of candidate sample xi by two items: spatio-temporal coupling value Diversity enhancement factor ensures the diversity from the selected samples. In the actual selection process, the algorithm calculates the value of all candidate samples in the cluster one by one, selects the sample with the highest score to join the labeled set, then updates the selected set and repeats the process until the budget requirement of the cluster is met.
[0177] Diversity enhancement factor is defined as:
[0178] ;
[0179] where, when the selected set S is empty, the diversity factor is 1; otherwise, the minimum distance between candidate sample xi and all samples in the selected set S is calculated, and the diversity enhancement degree is quantified by the ratio of the distance to the average distance in the cluster, and μ is the adaptive diversity weight. The average distance in the cluster ; is the selected sample set.
[0180] Diversity weight Adopting an adaptive adjustment strategy:
[0181] ;
[0182] where, is the basic diversity weight, is the adaptive adjustment coefficient, which ensures that the diversity requirement is dynamically adjusted with the progress of sample selection.
[0183] Algorithm complexity: the time complexity of selection in a single cluster is where comes from the pairwise distance calculation between samples in the cluster, The budget allocation for the cluster; the overall algorithm time complexity is wherein is the total number of samples, is the total budget, and the complexity has good scalability in large-scale industrial applications.
[0184] At this point, the sample set Sj is selected from each cluster Cj according to the budget Bj, where |Sj| = Bj, that is, the number of samples selected by each cluster is equal to the number of budget allocated to the cluster, providing high-quality candidate samples for subsequent intelligent labeling.
[0185] S4: Intelligent labeling execution and result output, merging the samples selected by each cluster to form a final labeling set, performing quality evaluation on the final labeling set, and generating a standardized labeling output according to the quality evaluation results.
[0186] Based on the aforementioned spatio-temporal coupling value evaluation and hierarchical selection results, the final intelligent labeling is executed and a high-quality labeling sample set is generated. It ensures that the selected high-value samples can be effectively labeled and provides traceable labeling decision basis.
[0187] S4-1: Final sample set generation, merging the samples selected by each cluster to form a final labeling set, then the final labeling set is:
[0188] ;
[0189] wherein, is the sample set selected by the jth cluster; is the final labeling set containing the sample sets of all clusters. S4-2: Labeling quality evaluation and verification, performing quality evaluation on the final labeling set, including the following indicators:
[0190] Diversity index calculation, measuring the geometric distance distribution between selected samples:
[0191]
[0192] ; This index is used to quantify the uniformity of the distribution of selected samples in the feature space, and the larger the value, the better the sample diversity. Evaluation criteria: Excellent (Ddiversity ≥ 1.2): sample distribution is uniform, feature space coverage is sufficient; Good (0.8 ≤ Ddiversity < 1.2): sample distribution is moderate, meeting the requirements; Need to improve (Ddiversity < 0.8): samples are too concentrated, suggesting adjusting algorithm parameters.
[0193]
[0194] To address the lack of diversity, an adaptive adjustment mechanism is established: first, improve the sample distribution by increasing the optimal cluster number C (from 5 to 7-8); second, adjust the diversity weight αdiv in the budget allocation (from 0.3 to 0.4) to give more labeling budget to high diversity clusters; if necessary, adjust the adaptive diversity coefficient β in the sample selection process (from 0.3 to 0.5) to strengthen the diversity constraint in the later selection. Through this hierarchical parameter optimization strategy, the final labeled samples not only maintain high value characteristics, but also have good feature space coverage, meeting the strict requirements of semiconductor manufacturing virtual metrology systems for sample representativeness.
[0195] S4-3: Labeling result output
[0196] The standardized labeling output is generated, including: selected sample feature data, spatio-temporal coupling value score, sample selection reason (cluster classification, geometric characteristics, value contribution), and the output result supports the complete traceability of labeling decisions, facilitating quality control and continuous optimization.
[0197] In summary, the present application selects samples worth labeling or verifying from the labeled sample set, providing accurate value basis for subsequent intelligent selection.
[0198] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A method for spatio-temporal coupled virtual metrology intelligent labeling based on feature space geometry, characterized in that: Comprise the following steps executed in turn: S1: Obtain multi-dimensional sensor time series data of semiconductor manufacturing equipment as a labeled sample set, obtain two geometric feature indicators of the labeled sample set located in the high curvature area of the feature manifold, and the scarce area and the distribution boundary: identify the samples in the labeled sample set located in the high curvature area of the feature manifold and the samples located in the scarce area and the distribution boundary, and fuse the two geometric characteristics indicators into a unified manifold value evaluation, the specific steps are as follows: S1-1: Identify the samples in the labeled sample set located in the high curvature area of the feature manifold, the identification process is as follows: Based on the theory of differential geometry, the following formula is used to calculate the sample Local curvature matrix on the feature manifold : ; wherein, is the matrix determinant, is the L2 norm, based on the standard form of curvature in differential geometry: , is the local density function at the sample point; The local feature gradient vector is calculated by the following formula: ; in, For the number of nearest neighbors, For Gaussian kernel function, To estimate the bandwidth parameter for kernel density, , for sample of nearest neighbor set For feature dimensions; S1-2: Identify the scarce area and the distribution boundary in the labeled sample set, the identification process is as follows: The density gradient values of the sample are calculated using the improved kernel density gradient estimation : ; where, is the squared norm of the gradient of the density function at , is the local density of the sample , is the global density median, is the exponential function, is the ratio is the relative density, the local density is computed by the Gaussian kernel density estimation: where is the number of neighbors, is the kernel density estimation bandwidth parameter, is the feature dimension; Density gradient By numerical differentiation calculation: ; wherein is the differential step size, is the differential step size, is the differential step size, S1-3: Fuse the two geometric characteristics indicators into a unified manifold value evaluation: ; wherein, is a local curvature weight, is a density gradient weight, ; S2: Establish an adaptive time weight model based on data complexity, establish a dynamic spatial weight model based on local sparsity based on an adaptive kernel function, combine the adaptive time weight model, the dynamic spatial weight model and the manifold value evaluation, and define the spatiotemporal coupling value score; S3: Take the spatiotemporal coupling value score as input, establish a three-layer optimization architecture to realize the globally optimal sample combination, which specifically includes the following steps: S3-1: Establish a multi-objective optimization clustering selection model, determine the optimal cluster number by minimizing information entropy, and divide the samples into a predetermined number of clusters; S3-2: Based on multi-objective optimization theory, comprehensively consider the diversity, value and redundancy of the cluster, obtain the cluster score through weighted calculation, and allocate the labeling budget of each cluster according to the cluster score; S3-3: In each cluster, an adaptive greedy algorithm is used to select the most representative samples through dynamic adjustment of the diversity weight; S4: Intelligent labeling execution and result output: combine the selected samples of each cluster to form a final labeling set, perform quality evaluation on the final labeling set, and generate standardized labeling output according to the quality evaluation result.
2. The feature space geometry based spatio-temporal coupled virtual metrology intelligent labeling method of claim 1, wherein: Step S2 includes the following steps: S2-1: Establish an adaptive time weight model based on data complexity: ; wherein, is the time weight of the time point, , is the total number of key time points, is the importance weight of the th time point, satisfying , is the th key time point, is the adaptive time scale parameter, is the square of the time distance; Adaptive time scale parameter Is expressed by the following equation: ; wherein, is a base timescale, is a complexity adjustment coefficient, data complexity score is defined as: ; wherein, represents a feature distribution at the current time t, represents a historical baseline feature distribution for the past 30 days; S2-2: Establish a dynamic spatial weight model based on local sparsity: ; wherein, is an adaptive Gaussian kernel function, adaptive local scale according to local sparsity adjustment: ; wherein, is the local scale parameter of the sample , denotes the median of the distances in the neighborhood set, is the sparsity adjustment coefficient, is the local sparsity, is defined as: ; wherein, is hyper-spherical volume, is a sample of the first nearest neighbor distance; Adaptive kernel function is: ; wherein, is the Euclidean distance between samples; S2-3: Combine the adaptive time weight model, the dynamic spatial weight model and the manifold value evaluation to define the spatiotemporal coupling value score as: 。 3. The feature space geometry based spatio-temporal coupled virtual metrology intelligent labeling method of claim 2, wherein: Step S3 includes the following specific steps: S3-1: Establish a multi-objective optimization clustering selection model, and determine the optimal cluster number by minimizing information entropy: ; wherein the cluster information entropy H(C) is defined based on Shannon entropy theory: , is a complexity penalty coefficient, denotes the independent variable that makes the objective function attain a minimum value; The clustering complexity Ω(C) is defined as: ; wherein, is a cluster number penalty coefficient, is a compactness weight within a cluster, is an average intra-cluster distance, the optimal number of clusters C* is obtained by minimizing the information entropy above, and the candidate sample set is divided into C* clusters; S3-2: combine the spatiotemporal coupling value of each sample , budget optimization allocation, budget allocation formula: ; where C* is the optimal number of clusters, is a softmax temperature parameter, is the label budget assigned to cluster satisfies , is the total label budget; is the aggregate score of cluster . The combined score function is: ; wherein the cluster diversity Based on the geometric diversity theory definition: ; wherein, , for any two samples in the cluster , there is , a set of distances , is the median of the set of distances, is the standard deviation; Cluster value Using a weighted combination of maximum and mean values: ; wherein, is the maximum weight, is the average weight; Cluster redundancy Based on nearest neighbor distance statistics: ; wherein, is a number of neighbors penalty function, is a distance sample is a number of neighborhood samples less than a threshold is a Euclidean distance between samples; S3-3: In-cluster sample selection algorithm: for the j-th cluster , sample selection according to the budget allocation , iterative selection strategy: initialize the selected sample set , repeat the selection process until samples are selected, in each iteration, select the sample that maximizes the objective function from the remaining candidate samples of the cluster : ; wherein, the single sample selected for the current iteration, X is a candidate sample in the cluster Cj that has not yet been selected, S is the current set of selected samples, and the selection objective function is defined as: ; wherein is a diversity enhancer, ; When the selected set S is empty, the diversity factor is 1, otherwise, the minimum distance between the candidate sample X and all samples in the selected set S is calculated, and the diversity enhancement degree is quantified by the ratio of the distance to the average distance within the cluster, , μ is the adaptive diversity weight, , μ is the adaptive diversity weight, , is the basic diversity weight, is the adaptive adjustment coefficient.
4. The feature space geometry based spatio-temporal coupled virtual metrology intelligent labeling method of claim 3, wherein: Step S4 includes the following steps: S4-1: Final sample set generation: combine the selected samples of each cluster to form a final labeling set, which is represented by the following formula: ; wherein, is the final set of annotations, containing the sample sets of all clusters, is the final set of annotations, containing the sample sets of all clusters, is the final set of annotations, containing the sample sets of all clusters, S4-2: Labeling quality evaluation and verification: perform quality evaluation on the final labeling set, calculate the diversity index, and measure the geometric distance distribution between the selected samples: ; Evaluation criteria: Ddiversity ≥ 1.2, excellent labeling set, i.e. uniform sample distribution and sufficient feature space coverage; 0.8 ≤ Ddiversity < 1.2, good labeling set, i.e. moderate sample distribution and meeting the requirements; Ddiversity < 0.8, which means the labeling set needs to be improved, i.e., the samples are too concentrated, and the algorithm parameters need to be adjusted; S4-3: Generate a standardized labeling output, which includes: feature data of selected samples, spatiotemporal coupling value scores, and sample selection reasons, including cluster classification, geometric characteristics, and value contribution.
Citation Information
Patent Citations
Automatic data annotation method of ISP image signal processing visual sensor
CN119723010A
Capacity evaluation method and device based on historical capacity similarity characteristic
US20210365823A1