Space-time coupling virtual metering intelligent labeling method based on feature space geometry

By employing a spatiotemporal coupled virtual metrology intelligent annotation method based on feature space geometry, samples with high curvature and scarce regions are identified. Combined with an adaptive weight model and a three-layer optimization architecture, the method solves the problem of low annotation efficiency for extremely imbalanced datasets in semiconductor manufacturing, and achieves high-precision sample combination and quality prediction.

CN121542791AActive Publication Date: 2026-02-17QUANZHOU INST OF EQUIP MFG +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202610063290.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-19
Publication Date
2026-02-17
Estimated Expiration
2046-01-19

AI Technical Summary

Technical Problem

When faced with extremely imbalanced datasets in semiconductor manufacturing, existing virtual metrology technologies rely on model predictions for sample value assessment, neglecting the geometric characteristics of high-dimensional feature spaces and lacking adaptive and hierarchical optimization strategies. This results in low annotation efficiency, insufficient defect detection rate, and difficulty in meeting the requirements for high-precision quality prediction.

Method used

A spatiotemporal coupled virtual metrology intelligent annotation method based on feature space geometry is adopted. By identifying samples in high curvature regions and scarce regions of the feature manifold, and combining an adaptive temporal and spatial weight model, a three-layer optimization architecture is established to achieve the globally optimal sample combination. This includes multi-objective optimization clustering selection and an adaptive greedy algorithm, and outputs a spatiotemporal coupled value score.

Benefits of technology

It improves labeling efficiency, increases defect detection rate, ensures that sample value assessment adapts to dynamic changes, achieves globally optimal sample combination, and meets the high-precision quality prediction needs of semiconductor manufacturing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121542791A_ABST
    Figure CN121542791A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of semiconductor manufacturing quality control, in particular to a space-time coupling virtual measurement intelligent labeling method based on feature space geometry, which comprises the following steps: S1, acquiring multi-dimensional sensor time sequence data of semiconductor manufacturing equipment as a labeling sample set, and performing manifold value evaluation on the labeling sample set; s2, establishing an adaptive time weight model based on data complexity, establishing a dynamic space weight model based on local sparseness based on an adaptive kernel function, combining the adaptive time weight model, the dynamic space weight model and manifold value evaluation, and defining a time-space coupling value score; s3, taking the time-space coupling value score as an input, and establishing a three-layer optimization architecture to realize a globally optimal sample combination; s4, combining the samples selected by each cluster to form a final label set, performing quality evaluation on the final label set, and generating standardized label output according to a quality evaluation result; and an accurate value basis is provided for subsequent intelligent selection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of semiconductor manufacturing quality control technology, specifically to a spatiotemporal coupled virtual metrology intelligent labeling method based on feature space geometry. Background Technology

[0002] In semiconductor manufacturing, virtual metrology technology predicts product quality by analyzing process sensor data, significantly reducing the cost of expensive physical measurements. Taking a 12-inch wafer production line as an example, the typical defect rate is controlled at around 1.8%, resulting in a highly imbalanced data distribution. This extreme data imbalance poses a significant challenge to the sample labeling strategy of virtual metrology systems. However, existing virtual metrology technologies have significant technical limitations when handling such extremely imbalanced datasets: First, existing sample value assessment methods heavily rely on model prediction results. Traditional active learning methods, such as uncertainty sampling and diversity sampling, depend on the predicted probabilities of the model to be improved to assess sample value, creating a circular dependency: sample value assessment depends on model predictions, while model improvement depends on sample labeling results. This circular dependency means that labeling effectiveness is limited by the quality of the initial model. In the extremely imbalanced data environment of semiconductor manufacturing, the initial model often cannot provide reliable prediction probabilities, thus affecting the accuracy of sample selection.

[0003] Secondly, traditional methods neglect the geometric structure characteristics of high-dimensional feature spaces. Semiconductor manufacturing equipment is typically equipped with 200-300 sensors, which, after dimensionality reduction processing such as principal component analysis, yield 64-dimensional feature vectors. These high-dimensional feature vectors form complex geometric manifold structures in the feature space. High-value samples are often located in special geometric positions within the feature manifold, such as high-curvature regions, low-density regions, and regions with drastic gradient changes near the decision boundary. However, existing technologies fail to effectively utilize this geometric structure information, resulting in low annotation efficiency.

[0004] Furthermore, existing spatiotemporal analysis methods lack adaptability. Semiconductor manufacturing exhibits significant temporal characteristics, with substantial differences in data complexity and distribution across different periods. Fixed time weighting functions cannot adapt to this dynamic change, and fixed spatial influence radii cannot accommodate density differences in different regions. This lack of adaptability in spatiotemporal analysis methods prevents them from accurately capturing the dynamic patterns of sample value changes.

[0005] Finally, there is a lack of hierarchical optimization strategies for high-dimensional feature spaces. Existing sample selection methods often employ simple greedy strategies, which are prone to getting trapped in local optima in high-dimensional spaces and cannot achieve globally optimal sample combinations. Especially in a 64-dimensional feature space, the measurement of similarity and the guarantee of diversity among samples become more complex, requiring specialized hierarchical optimization algorithms.

[0006] These limitations of existing technologies result in low annotation efficiency for virtual metrology systems in extremely unbalanced data environments. Under the same annotation budget, the defect detection rate can only reach about 82%, which is difficult to meet the stringent requirements of semiconductor manufacturing for high-precision quality prediction. Summary of the Invention

[0007] The purpose of this invention is to provide a spatiotemporally coupled virtual metrology intelligent annotation method based on feature space geometry that can reduce annotation costs.

[0008] To achieve the above objectives, the present invention adopts the following technical solution: The spatiotemporal coupled virtual metrology intelligent annotation method based on feature space geometry includes the following steps performed sequentially: S1: Obtain multi-dimensional sensor time-series data of semiconductor manufacturing equipment as a labeled sample set, and obtain two geometric feature indicators of the labeled sample set located in the high curvature region of the characteristic manifold, as well as the scarce region and distribution boundary: identify the samples in the labeled sample set located in the high curvature region of the characteristic manifold and the samples located in the scarce region and distribution boundary, and integrate the two geometric feature indicators into a unified manifold value assessment. S2: Establish an adaptive time weight model based on data complexity, establish a dynamic spatial weight model based on local sparsity based on an adaptive kernel function, combine the adaptive time weight model, the dynamic spatial weight model and the manifold value assessment, and define a spatiotemporal coupled value score. S3: Using the spatiotemporal coupled value score as input, establish a three-layer optimization architecture to achieve the globally optimal sample combination, specifically including the following steps: S3-1: Establish a multi-objective optimization clustering selection model, determine the optimal number of clusters by minimizing information entropy, and divide the samples into a preset number of clusters; S3-2: Based on multi-objective optimization theory, taking into account the diversity, value and redundancy of clusters, a cluster score is obtained through weighted calculation, and the labeling budget for each cluster is allocated according to the cluster score. S3-3: Within each cluster, an adaptive greedy algorithm is used to select the most representative sample by dynamically adjusting the diversity weights; S4: Intelligent annotation execution and result output: Merge the samples selected from each cluster to form the final annotation set, perform quality assessment on the final annotation set, and generate standardized annotation output based on the quality assessment results.

[0009] Preferably, step S1 includes the following steps: S1-1: Identify samples located in the high curvature region of the characteristic manifold within the labeled sample set. The identification process is as follows: Based on differential geometry theory, the sample is calculated using the following formula. Local curvature matrix on the characteristic manifold : ; in, It is the determinant of a matrix. The L2 norm, with an exponent of 3 / 2, is based on the curvature in differential geometry. Standard format For the sample The local density function at that location; The local feature gradient vector is calculated using the following formula: ; in, For the number of nearest neighbors, For Gaussian kernel function, To estimate the bandwidth parameter for kernel density, , for sample of nearest neighbor set For feature dimensions; S1-2: Identify the scarce regions and distribution boundaries in the labeled sample set. The identification process is as follows: The improved kernel density gradient estimation is used to calculate the sample. Density gradient value : ; in, For the density function in The squared gradient norm at point , For the sample Local density, The median of the global density. It is an exponential function. ,ratio Relative density, local density Calculated using Gaussian kernel density estimation: ,in For the number of nearest neighbors, To estimate the bandwidth parameter for kernel density, For feature dimensions; density gradient Calculated using numerical differentiation: ; in, Let be the differential step size. For the first One standard basis vector; S1-3: Integrating two geometric property indices into a unified manifold value assessment: ; in, For local curvature weights, For density gradient weights, . Preferably, step S2 includes the following steps: S2-1: Establish an adaptive time weighting model based on data complexity: ; in, For a moment Time weighting This represents the total number of key time points. For the first The importance weight of each time point satisfies , For the first Key time points, For adaptive time scale parameters, The square of the time distance; Adaptive time scale parameters It is expressed by the following formula: ; in, Based on the basic time scale This is a complexity adjustment factor, used for data complexity scoring. Defined as: ; in, This represents the characteristic distribution at the current time t. This represents the historical baseline characteristic distribution over the past 30 days; S2-2: Establishing a dynamic spatial weighting model based on local sparsity: ; in, For adaptive Gaussian kernel function, adaptive local scaling Adjust based on local sparsity: ; in, For the sample Local scale parameters, Indicates in The median distance in a nearest neighbor set. This is the sparsity adjustment coefficient. For local sparsity, Defined as: ; in, for The volume of the hypersphere For the sample The Nearest neighbor distance; Adaptive kernel function for: ; in, The Euclidean distance between samples; S2-3: Combining the adaptive time-weighted model, the dynamic spatial-weighted model, and the manifold value assessment, the spatiotemporal coupled value score is defined as: .

[0010] Preferably, step S3 includes the following specific steps: S3-1: Establish a multi-objective optimization clustering selection model and determine the optimal number of clusters using the principle of minimizing information entropy: ; Among them, the clustering information entropy H(C) is defined based on Shannon entropy theory: , This is the complexity penalty coefficient. Represents the independent variable that minimizes the objective function; The clustering complexity Ω(C) is defined as: ; in, The cluster number penalty coefficient, The cluster compactness weight. To determine the average intra-cluster distance, the optimal number of clusters C* is obtained by minimizing the information entropy mentioned above, and the candidate sample set is divided into C* clusters. S3-2: Combining the spatiotemporal coupling value of each sample Optimize budget allocation; budget allocation formula: ; Where C* is the optimal cluster number, For softmax temperature parameters, To be assigned to clusters The budget for the labeling needs to be met. , Total budget for annotation; For clusters Overall score; The comprehensive scoring function is: ; Among them, cluster diversity Based on the definition of geometric diversity theory: ; in, For any two samples within the cluster ,have Distance set , The median of this distance set. Standard deviation; Cluster value Use a weighted combination of the maximum and the mean: ; in, Weights are the maximum values. Weighted by average value; Cluster redundancy Based on nearest neighbor distance statistics: ; in, Let the nearest neighbor number penalty function be used. For distance samples Less than the threshold The number of neighborhood samples, The Euclidean distance between samples; S3-3: In-cluster sample selection algorithm: For the j-th cluster According to the budget allocation Sample selection is performed using an iterative selection strategy: the selected sample set is initialized. Repeat the selection process until a selection is made. Each sample, in each iteration, is from the cluster. Choose the sample that maximizes the objective function from the remaining candidate samples: ; in, Let X be a single sample selected in the current iteration, X be a candidate sample in cluster Cj that has not yet been selected, and S be the set of currently selected samples. The selection objective function is... Defined as: ; in, As a diversity enhancement factor, ; When the selected set S is empty, the diversity factor is 1; otherwise, the minimum distance between the candidate sample X and all samples in the selected sample set S is calculated, and this distance is compared with the intra-cluster average distance. The ratio is used to quantify the degree of increased diversity. μ is the adaptive diversity weight. , Based on the weight of diversity, This is the adaptive adjustment coefficient.

[0011] Preferably, step S4 includes the following steps: S4-1: Final Sample Set Generation: The samples selected from each cluster are merged to form the final labeled set, which is represented by the following formula: ; in, For the first The sample set selected for each cluster The final labeled set is a collection of samples from all clusters. S4-2: Annotation Quality Assessment and Verification: The final annotation set is assessed for quality, a diversity index is calculated, and the geometric distance distribution between selected samples is measured. ; Evaluation criteria: Diversity ≥ 1.2 indicates an excellent labeled set, meaning the sample distribution is uniform and the feature space is adequately covered; 0.8 ≤ Ddiversity < 1.2 indicates a well-labeled set, meaning the sample distribution is moderate and meets the requirements. Ddiversity < 0.8 indicates a labeled dataset that needs improvement, meaning the samples are too concentrated; it is recommended to adjust the algorithm parameters. S4-3: Generate standardized annotation output, which includes: feature data of selected samples, spatiotemporal coupling value score, and reasons for sample selection, including cluster classification, geometric characteristics, and value contribution.

[0012] By adopting the aforementioned design scheme, the beneficial effects of the present invention are as follows: This application selects samples worth labeling or verifying from the labeled sample set, establishes a sample value quantification theory based on the geometric structure of the feature space, avoids dependence on model prediction results, derives the labeling value of the sample from the inherent geometric characteristics of the feature distribution, and outputs the manifold value of each sample as a comprehensive quantification result of the sample geometric features. This value integrates local geometric curvature, density gradient and uncertainty features. This application establishes an adaptive time weight model based on data complexity and a dynamic spatial weight model based on local sparsity to achieve adaptive adjustment of spatiotemporal parameters, and outputs spatiotemporal coupling value as the final value metric of the sample for hierarchical intelligent selection. This value integrates geometric features, temporal dynamics and spatial correlation. The coupling mechanism of this spatiotemporal coupling value ensures that the sample value assessment can adapt to the dynamic changes in semiconductor manufacturing, and provides an accurate value basis for subsequent intelligent selection. This application employs a three-layer optimization strategy in hierarchical intelligent selection: information entropy clustering to determine sample grouping, multi-objective function optimization of budget allocation, and adaptive greedy algorithm to achieve intra-cluster selection, which can effectively solve the local optimum problem of traditional methods. Attached Figure Description

[0013] Figure 1 This is a flowchart of the intelligent annotation method of the present invention. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0015] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this invention are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0016] A spatiotemporally coupled virtual metrology intelligent annotation method based on feature space geometry, such as Figure 1 As shown, the steps are executed sequentially as follows: S1: Obtain multi-dimensional sensor time-series data from semiconductor manufacturing equipment as a labeled sample set. Obtain the following two geometric characteristic indicators for the labeled sample set located in the high curvature region of the characteristic manifold, as well as the scarce region and distribution boundary: Identify the samples in the labeled sample set located in the high curvature region of the characteristic manifold and the samples located in the scarce region and distribution boundary. Merge the two geometric characteristic indicators into a unified manifold value assessment and output it as a comprehensive quantitative result of the sample's geometric features. This step avoids dependence on model prediction results by establishing a sample value quantification theory based on the geometric structure of the feature space, and derives the labeled value of the sample from the intrinsic geometric characteristics of the feature distribution.

[0017] In this embodiment, the multidimensional sensor timing data includes the radio frequency power of the etching equipment, chamber pressure, gas flow combination, and temperature gradient distribution. The temperature gradient distribution includes the temperature difference and temperature uniformity in different regions, as well as the power matching efficiency of the PECVD equipment (specifically, the power loss, efficiency, and waveform stability of each power channel), the reaction gas ratio (specifically, the flow ratio of different gases, such as the ratio of oxygen to fluoride), and the heating temperature curve (specifically, the rate of temperature rise, steady-state temperature, and fluctuation range).

[0018] S1-1: In step S1, samples located in the high curvature region of the feature manifold within the labeled sample set are identified. These samples are usually close to the decision boundary and have high labeling value. The identification process is as follows: Based on differential geometry theory, the sample is calculated using the following formula. Local curvature matrix on the characteristic manifold : ; in, Let be the local curvature value of sample xi, with a range of . ; Let be a local second-order structure matrix constructed in the k nearest neighbor neighborhood of sample xi (used to approximate local curvature / geometric complexity), whose dimension is consistent with the feature dimension, and is . , Calculated using the mean of the outer product of the differences between the k nearest neighbors: , It is the matrix determinant, reflecting the local geometric complexity; The L2 norm is used to calculate the Euclidean length of the vector; the exponent is set to 3 / 2, used for scaling and numerical stabilization of the curvature measure. For the sample The local density function at that location; The local feature gradient vector is calculated using the following formula: ; in, For the number of nearest neighbors, based on The optimal neighborhood theory determines that in a 64-dimensional feature space, the size of the optimal neighborhood should be close to... Considering the balance between computational efficiency and numerical stability, It can effectively capture local geometric structure features while avoiding noise accumulation effects in high-dimensional space; where Nk(i) represents the set of k nearest neighbor sample indices of sample xi; xp represents the feature vector of the nearest neighbor sample with index p; p is the nearest neighbor sample index, and p∈Nk(i) and p≠i; The kernel function is used; in this example, it is a Gaussian kernel function. ; The bandwidth parameter for kernel density estimation is calculated using the Silverman optimal bandwidth formula: in Let be the standard deviation of the feature data. For a 64-dimensional standardized feature, , Given the sample size, the following calculations were performed: ; , for sample of Nearest neighbor set.

[0019] S1-2: In step S1, the scarce regions and distribution boundaries in the labeled sample set are identified. Samples in these regions are of great value for improving the model's generalization ability. The distribution boundary characteristics are accurately captured by density gradient. The identification process is as follows: The improved kernel density gradient estimation is used to calculate the sample. Density gradient value : ; in, The range is , For the density function in The squared gradient norm at a given point reflects the degree of drastic change in density. For the sample The local density is calculated using Gaussian kernel density estimation. This is the global density median, used for normalization. It is an exponential function. ,ratio Relative density reflects the scarcity of a sample in the global distribution, while local density... Calculated using Gaussian kernel density estimation: ,in For the number of nearest neighbors, , To estimate the bandwidth parameter for kernel density, , For feature dimension, ; density gradient Calculated using numerical differentiation: ; in, The step size is the differential step size, and the step size selection is based on the numerical analysis theory, specifically the optimal step size. ,in For double-precision floating-point numbers, machine precision. This represents the order of magnitude of the density function at the current point. For the kernel density function in this application, a typical value is... According to the theoretical formula: For this application scenario, considering the engineering balance between computational efficiency and numerical stability, the actual choice is... ; For the first Standard basis vectors.

[0020] S1-3: Integrating two geometric property indices into a unified manifold value assessment: ; in, The local curvature weights reflect the importance of the decision boundary. The highest weight is assigned because high curvature regions are typically close to the classification boundary, maximizing their annotation value. The density gradient weights reflect the value of scarcity, and the two weights satisfy the normalization constraint: The weighting reflects the hierarchy of importance of sample annotations, with a local curvature weight of 0.7 reflecting the crucial role of proximity to the decision boundary, and a density gradient weight of 0.3 reflecting the importance of scarcity value. This weighting configuration ensures the manifold value. This accurately reflects the comprehensive geometric value of the samples in the feature space, providing a stable and reliable basis for value assessment in subsequent steps. After step S1 is completed, the manifold value of each sample is output. This value integrates local geometric curvature, density gradient, and uncertainty features, and is passed to step S2 as a comprehensive quantization result of the sample's geometric features.

[0021] S2: Establish an adaptive time-weighted model based on data complexity, calculate the data complexity score, establish a dynamic spatial weighted model based on local sparsity, construct an adaptive kernel function, combine the adaptive time-weighted model, the dynamic spatial weighted model, and the manifold value assessment, and define the spatiotemporal coupled value score as: The spatiotemporal coupling value Vcoupled(xi,t) is output as the final value measure of the sample and provided to the hierarchical intelligent selection. This value integrates geometric features, temporal dynamics and spatial correlation. To avoid inconsistencies in the units, the above three terms are normalized to [0,1] before participating in the product. Here, xi is the i-th sample and t is its timestamp. Based on the manifold value output in step S1 This step establishes an adaptive spatiotemporal coupling model, dynamically modulating the static geometric value based on production time-series characteristics and spatial distribution features, transforming it into a dynamic value adapted to the actual application scenario. Traditional spatiotemporal analysis methods use fixed time weighting functions and spatial influence radii, which cannot adapt to the significant differences in data complexity and distribution characteristics at different stages of semiconductor manufacturing. This method achieves adaptive adjustment of spatiotemporal parameters by establishing an adaptive time weighting model based on data complexity and a dynamic spatial weighting model based on local sparsity. The coupling mechanism of spatiotemporal coupled value ensures that the sample value assessment can adapt to the dynamic changes in semiconductor manufacturing, providing an accurate value basis for the intelligent selection in step S3.

[0022] S2-1: Adaptively adjust the attenuation parameter of time effect according to the changes in data complexity at different times.

[0023] Establish an adaptive time weighting model based on data complexity: ; in, For a moment The time weight has a range of [0,1]. The total number of key time points, typically including equipment. For the first The importance weight of each time point satisfies ; For the first Key time points (hours), such as For equipment maintenance time, For process changeover time, The time scale parameter is adaptive and dynamically adjusted according to data complexity. The square of the time distance is used to calculate time decay, and the Gaussian kernel form ensures the smooth decay characteristics of the time effect.

[0024] Adaptive time scale The formula is: ; in, Using hours as the base timescale, determined based on semiconductor manufacturing batch cycles: Statistical analysis shows that the processing cycle for a single batch of 12-inch wafers is approximately 8-16 hours, with the median of 12 hours taken as the base timescale. This is a complexity adjustment factor, determined based on the characteristics of semiconductor manufacturing processes: when the data complexity score is 1.0 (normal state), the time scale remains at the base value; when the complexity score reaches 2.0 (high complexity), the time scale is increased by 30%. The adaptive adjustment of the time window was ensured to be within a reasonable range; data complexity score. Defined as: ; in, This represents the characteristic distribution at the current time t. This represents the historical baseline characteristic distribution over the past 30 days. This time window is determined based on the semiconductor manufacturing process stability cycle and is used to capture long-term process baseline states.

[0025] S2-2: Dynamically adjust the spatial influence radius based on the local density characteristics of the samples in the feature space. Establish a dynamic spatial weighting model based on local sparsity: ; in, An adaptive Gaussian kernel function is used, satisfying the kernel function properties: non-negativity, normalization, and symmetry; and adapting to local scaling. Adjust based on local sparsity: ; in, For the sample Local scale parameters, Indicates in The median distance in a nearest neighbor set. This is the sparsity adjustment coefficient. For local sparsity, A larger value indicates that the samples in that region are more sparse. The value ranges from [0, +∞), and is usually defined as: ; in, for The volume of a hypersphere is based on high-dimensional geometry theory. for The volume factor of a 1D hypersphere. For the first Nearest neighbor distance, a formula used to accurately calculate the volume of a local neighborhood in a high-dimensional feature space.

[0026] Adaptive kernel function for: ; in, This is a sparsity adjustment coefficient used to control the degree of influence of local sparsity on the spatial scale. When the local sparsity... When the sparsity is 1.0, the spatial scale increases by 20%; when the sparsity reaches 2.0, the spatial scale increases by 40%. This parameter ensures that the spatial weights can adapt to the density differences in different regions of the feature space. The distance between samples is the Euclidean distance.

[0027] S2-3: Combining the adaptive time-weighted model, the dynamic spatial-weighted model, and the manifold value assessment, the spatiotemporal coupled value score is defined as: ; This product form reflects the mutually reinforcing effect of geometric value, temporal dynamics, and spatial correlation. The product form is chosen because spatiotemporal factors have the characteristic of mutually amplifying rather than nonlinearly superimposing their influence on sample value. Step S2 outputs the spatiotemporal coupled value Vcoupled(xi,t), which integrates geometric features, temporal dynamics, and spatial correlation. This value is passed to step S3 as the final value measure of the sample for hierarchical intelligent selection.

[0028] S3: Scoring based on the spatiotemporal coupling value of step S2 Using this as input, a three-layer optimization architecture is established to achieve the globally optimal sample combination: First, the optimal number of clusters is determined by minimizing information entropy. The samples are divided into a predetermined number of clusters, the number of which is set according to actual needs. Then, based on multi-objective optimization theory, considering the diversity, value, and redundancy of the clusters, a cluster score is obtained through weighted calculation. And allocate the labeling budget for each cluster accordingly. Finally, within each cluster, an adaptive greedy algorithm is employed to select the most representative sample through dynamic adjustment of the diversity weight μ. This hierarchical strategy effectively avoids the local optima problem of traditional greedy methods and achieves globally optimized allocation of the labeling budget.

[0029] This step employs a three-layer optimization strategy: information entropy clustering to determine sample groups, multi-objective function optimization of budget allocation, and adaptive greedy algorithm to achieve cluster selection, effectively solving the local optimum problem of traditional methods.

[0030] S3-1: Information theory-driven clustering optimization, using the principle of minimizing information entropy to determine the optimal number of clusters, and establishing a multi-objective optimization clustering selection model: ; Among them, the clustering information entropy H(C) is defined based on Shannon entropy theory: The unit is bits, used to measure the uncertainty of clustering. When a cluster is empty (|Cc| = 0), it is agreed that 0·log2(0) = 0, and the entropy value range is [0, log2(C)]. λ is the complexity penalty coefficient, determined based on the trade-off between model complexity and generalization ability in information theory: too small a value of λ leads to over-splitting, while too large a value of λ leads to under-splitting. λ = 0.05 strikes a balance between theoretical optimality and engineering feasibility. This represents the independent variable that minimizes the objective function. If there are multiple minimum values, the smallest one, C, is selected by default.

[0031] The clustering complexity Ω(C) is defined as: ; in, The cluster number penalty coefficient, The cluster compactness weight. The average intra-cluster distance (reflecting sample similarity) is used in this complexity term to prevent the generation of too many clusters and to ensure the similarity and interpretability of samples within a cluster. The optimal number of clusters C* is obtained by minimizing the information entropy described above, dividing the candidate sample set into C* clusters. The cluster division results and cluster characteristic analysis provide a basis for decision-making regarding the budget allocation in S3-2. ; ; Where N = 50000 is the total number of samples; Total budget for annotation; This serves as a lower bound for the number of clusters, ensuring basic class separability. 1000 is the upper bound for the number of clusters to prevent over-segmentation; 1000 is a standardization constant, determined empirically based on large-scale datasets. Lower bound estimate based on sample size and budget ratio; An upper bound estimation based on information theory complexity is used. After obtaining the optimal number of clusters C* by minimizing information entropy, the candidate sample set is divided into C* clusters, each containing samples with similar characteristics. The cluster partitioning results and characteristic analysis are directly used for the multi-objective budget allocation in S3-2.

[0032] S3-2: Budget allocation for multi-objective optimization, based on the optimal cluster number C* and cluster partitions obtained in S3-1, combined with the spatiotemporal coupling value of each sample. Optimize budget allocation; budget allocation formula: ; Where C* represents the total number of clusters. Temperature parameter controls the smoothness of the distribution; The comprehensive scoring function is: ; Among them, cluster diversity Based on the definition of geometric diversity theory: ; Used to enhance the contribution of large-distance sample pairs: ; Among them, for any two samples within the cluster ,have Distance set , The median of this distance set. The standard deviation is used to normalize the distance weights and control the influence of extreme values.

[0033] Cluster value Use a weighted combination of the maximum and the mean: ; in, The maximum weight is used to ensure that high-value samples are not missed; The average weight is used to balance the overall quality of the cluster.

[0034] Cluster redundancy Based on nearest neighbor distance statistics: ; in, Let the nearest neighbor number penalty function be used. For distance samples Less than the threshold The number of neighborhood samples, The distance between samples is the Euclidean distance.

[0035] Weight configuration: Diversity weights are applied to ensure sample coverage; Assigning the highest weight to a value reflects the dominant role in maximizing value. The weights are used for redundancy to avoid repeated selection of similar samples; the weights satisfy the normalization constraint. .

[0036] The Softmax function implements continuous budget allocation, with temperature parameters... Control the degree of allocation concentration. To ensure that each cluster receives at least one budget, post-processing is employed: ; in, To be assigned to clusters The budget for the labeling needs to be met. ; Total budget for annotation; For clusters The overall score reflects the labeling value of the cluster; The softmax temperature parameter controls the smoothness of the distribution. , , Clusters Indicators of diversity, value, and redundancy; The spatiotemporal coupling value calculated in step S2; For clusters The number of samples.

[0037] The budget allocation results for each cluster obtained through the above calculations satisfy the following conditions: This provides a budget constraint for the selection of intra-cluster samples in S3-3.

[0038] S3-3: In-cluster sample selection algorithm, which selects the optimal combination of samples within each cluster, balancing sample value and local diversity. In-cluster sample selection needs to avoid repeated selection of locally similar samples to ensure local diversity. The new method addresses this issue through diversity constraints. Based on the submodular function optimization theory, an improved greedy algorithm combined with diversity constraints is adopted. For the j-th cluster Cj, sample selection is performed according to the budget Bj allocated in step S3-2, using an iterative selection strategy: initializing the selected sample set S = Repeat the selection process until Bj samples are selected. In each iteration, select the sample that maximizes the objective function from the remaining candidate samples of cluster Cj. ; in, Let X be a single sample selected in the current iteration, X be a candidate sample in cluster Cj that has not yet been selected, and S be the set of currently selected samples. The selection objective function is... Defined as: ; This function evaluates the selection value of candidate sample xi through a combination of two products: spatiotemporal coupling value. The diversity enhancement factor is derived from the calculation results in step S2. Ensure the difference from the already selected samples. During the actual selection process, the algorithm sequentially calculates the differences between all candidate samples within the cluster. The highest-scoring sample is selected and added to the labeled set. The selected set is then updated and this process is repeated until the budget requirement for that cluster is met.

[0039] Diversity Enhancer Defined as: ; Wherein, when the selected set S is empty, the diversity factor is 1; otherwise, the minimum distance between the candidate sample xi and all samples in the selected sample set S is calculated, and this distance is compared with the average distance within the cluster. The ratio is used to quantify the degree of diversity enhancement, where μ is the adaptive diversity weight. Intra-cluster average distance. ; The selected sample set.

[0040] Diversity weight Adaptive adjustment strategy adopted: ; in, Based on the weight of diversity, The adaptive adjustment coefficient ensures that diversity requirements are dynamically adjusted as the sample selection progresses.

[0041] Algorithm complexity: The time complexity of selection within a single cluster is O(n log n). in It is derived from the pairwise distance calculation between samples within a cluster. Budget allocation for this cluster; the overall algorithm time complexity is O(n log n). ,in The total number of samples, For the total budget, this complexity has good scalability in large-scale industrial applications.

[0042] At this point, a sample set Sj is selected from each cluster Cj according to the budget Bj, where |Sj| = Bj, that is, the number of samples selected for each cluster is equal to the budget allocated to that cluster, providing high-quality candidate samples for subsequent intelligent annotation.

[0043] S4: Intelligent annotation execution and result output, which merges the samples selected from each cluster to form the final annotation set, performs quality assessment on the final annotation set, and generates standardized annotation output based on the quality assessment results.

[0044] Based on the aforementioned spatiotemporal coupling value assessment and hierarchical selection results, the final intelligent annotation is performed, generating a high-quality annotated sample set. This ensures that the selected high-value samples can be effectively annotated and provides traceable annotation decision-making basis.

[0045] S4-1: The final sample set is generated by merging the samples selected from each cluster to form the final labeled set. The final labeled set is as follows: ; in, For the first The sample set selected for each cluster; The final labeled set contains a collection of samples from all clusters.

[0046] S4-2: Annotation Quality Assessment and Verification. The final annotation set is assessed for quality, including the following metrics: Diversity index calculation measures the distribution of geometric distances between selected samples: ; This metric quantifies the uniformity of the distribution of selected samples in the feature space; a higher value indicates better sample diversity. Evaluation criteria: Excellent (Ddiversity ≥ 1.2): Samples are evenly distributed, and the feature space is adequately covered; Good (0.8 ≤ Ddiversity < 1.2): Sample distribution is moderate and meets the requirements; Needs improvement (Ddiversity < 0.8): Samples are too concentrated; it is recommended to adjust the algorithm parameters. To address insufficient diversity, an adaptive adjustment mechanism is established: firstly, the sample distribution is improved by increasing the optimal cluster number C (from 5 to 7-8); secondly, the diversity weight αdiv in the budget allocation is adjusted (from 0.3 to 0.4) to allow highly diverse clusters to receive more labeling budget; and thirdly, the adaptive diversity coefficient β in the sample selection process is adjusted (from 0.3 to 0.5) when necessary to strengthen the diversity constraints in later selections. This hierarchical parameter optimization strategy ensures that the final labeled samples maintain both high-value characteristics and good feature space coverage, meeting the stringent requirements of the semiconductor manufacturing virtual metrology system for sample representativeness.

[0047] S4-3: Output of annotation results Generate standardized annotation output, including: feature data of selected samples, spatiotemporal coupling value score, and reasons for sample selection (cluster classification, geometric characteristics, value contribution). The output results support the complete traceability of annotation decisions, which facilitates quality control and continuous optimization.

[0048] In summary, this application selects samples worth annotating or verifying from the annotated sample set, providing accurate value basis for subsequent intelligent selection.

[0049] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for spatio-temporal coupled virtual metrology intelligent labeling based on feature space geometry, characterized in that: Comprise the following steps executed in turn: S1: Obtain multi-dimensional sensor time series data of semiconductor manufacturing equipment as a labeled sample set, obtain two geometric feature indicators of the labeled sample set located in the high curvature area of the feature manifold, and the scarce area and the distribution boundary: identify the samples in the labeled sample set located in the high curvature area of the feature manifold and the samples located in the scarce area and the distribution boundary, and fuse the two geometric characteristics indicators into a unified manifold value evaluation; S2: Establish an adaptive time weight model based on data complexity, establish a dynamic spatial weight model based on local sparsity based on an adaptive kernel function, combine the adaptive time weight model, the dynamic spatial weight model and the manifold value evaluation, and define a spatiotemporal coupling value score; S3: Take the spatiotemporal coupling value score as input, establish a three-layer optimization architecture to realize globally optimal sample combination, specifically comprising the following steps: S3-1: Establish a multi-objective optimization clustering selection model, determine the optimal cluster number by minimizing information entropy, and divide the samples into a preset number of clusters; S3-2: Based on multi-objective optimization theory, comprehensively consider the diversity, value and redundancy of the cluster, obtain the cluster score through weighted calculation, and allocate the labeling budget of each cluster according to the cluster score; S3-3: In each cluster, an adaptive greedy algorithm is used to select the most representative samples through dynamic adjustment of the diversity weight; S4: Intelligent labeling execution and result output: combine the selected samples of each cluster to form a final labeling set, perform quality evaluation on the final labeling set, and generate standardized labeling output according to the quality evaluation result.

2. The feature space geometry based spatio-temporal coupled virtual metrology intelligent labeling method of claim 1, wherein: Step S1 comprises the following steps: S1-1: Identify the samples in the labeled sample set located in the high curvature area of the feature manifold, and the identification process is as follows: Based on the theory of differential geometry, the following formula is used to calculate the sample Local curvature matrix on the feature manifold : ; wherein is the matrix determinant, is the L2 norm, the exponent 3 / 2 is based on the standard form, is the local density function at the sample point; The local feature gradient vector is calculated by the following formula: ; wherein, is the number of neighbors, is a Gaussian kernel function, is a kernel density estimation bandwidth parameter, is a sample of a set of neighbors, is the feature dimension; S1-2: Identify the scarce area and the distribution boundary in the labeled sample set, and the identification process is as follows: The density gradient values of the sample are calculated using a modified kernel density gradient estimation :​ ; where, is the squared norm of the gradient of the density function at , is the local density of the sample , is the global density median, is the exponential function, is the ratio is the relative density, the local density is computed by the Gaussian kernel density estimation: where is the number of neighbors, is the kernel density estimation bandwidth parameter, is the feature dimensionality; Density gradient By numerical differentiation calculation: ; wherein is the differential step size, is the differential step size, is the differential step size, S1-3: Fuse the two geometric characteristics indicators into a unified manifold value evaluation: ; wherein, is a local curvature weight, is a density gradient weight, .

3. The feature space geometry based spatio-temporal coupled virtual metrology intelligent labeling method of claim 2, wherein: Step S2 comprises the following steps: S2-1: Establish an adaptive time weight model based on data complexity: ; wherein, is the time weight of the time point, is the total number of key time points, is the importance weight of the th time point, satisfying , is the th key time point, is the adaptive time scale parameter, is the square of the time distance;​ Adaptive time scale parameter is expressed by the following equation: ; wherein, is a base timescale, is a complexity adjustment coefficient, data complexity score is defined as: ; wherein, represents the feature distribution at the current time t, represents the historical baseline feature distribution for the past 30 days; S2-2: Establish a dynamic spatial weight model based on local sparsity: ; wherein, is an adaptive Gaussian kernel function, adaptive local scale according to local sparsity adjustment: ; wherein, is the local scale parameter of the sample , is the median of the distances in the neighborhood set, is the sparsity adjustment coefficient, is the local sparsity, is the local sparsity, is defined as: ; wherein, is hyper-spherical volume, is a sample of the first nearest neighbor distance; Adaptive kernel function is: ; wherein, is the Euclidean distance between samples; S2-3: Combine the adaptive time weight model, the dynamic spatial weight model and the manifold value evaluation to define a spatiotemporal coupling value score: 。 4. The feature space geometry based spatio-temporal coupled virtual metrology intelligent labeling method of claim 3, wherein: Step S3 comprises the following specific steps: S3-1: Establish a multi-objective optimization clustering selection model, and determine the optimal cluster number by minimizing information entropy: ; wherein the cluster information entropy H(C) is defined based on Shannon entropy theory: , is a complexity penalty coefficient, denotes the independent variable that makes the objective function attain a minimum value; The clustering complexity Ω(C) is defined as: ; wherein, is a cluster number penalty coefficient, is a compactness weight within a cluster, is an average intra-cluster distance, the optimal number of clusters C* is obtained by minimizing the information entropy above, and the candidate sample set is divided into C* clusters; S3-2: combine the spatiotemporal coupling value of each sample , budget optimization allocation, budget allocation formula: ; where C* is the optimal number of clusters, is a softmax temperature parameter, is the label budget allocated to the cluster satisfying , is the total label budget; is the comprehensive score of the cluster . The combined score function is: ; wherein the cluster diversity Based on the geometric diversity theory definition: ; wherein, , for any two samples in the cluster , there is , a set of distances , , the median of the set of distances, , the standard deviation; Cluster value Using a weighted combination of maximum and mean values: ; wherein, is the maximum weight, is the average weight; Cluster redundancy Based on nearest neighbor distance statistics: ; wherein, is a number of neighbors penalty function, is a distance sample is a number of neighborhood samples less than a threshold is a Euclidean distance between samples; S3-3: In-cluster sample selection algorithm: for the j-th cluster , sample selection according to the budget allocation , iterative selection strategy: initialize the selected sample set , repeat the selection process until samples are selected, in each iteration, select the sample that maximizes the objective function from the remaining candidate samples of the cluster : ; wherein, the single sample selected for the current iteration, X is a candidate sample in the cluster Cj that has not yet been selected, S is the current set of selected samples, and the selection objective function is defined as: ; wherein is a diversity enhancer, ; When the selected set S is empty, the diversity factor is 1, otherwise, the minimum distance between the candidate sample X and all samples in the selected set S is calculated, and the diversity enhancement degree is quantified by the ratio of the distance to the average distance within the cluster, , μ is the adaptive diversity weight, , μ is the adaptive diversity weight, , is the basic diversity weight, is the adaptive adjustment coefficient.

5. The feature space geometry based spatio-temporal coupled virtual metrology intelligent labeling method of claim 4, wherein: Step S4 comprises the following steps: S4-1: Final sample set generation: combine the selected samples of each cluster to form a final labeling set, and the final labeling set is represented by the following formula: ; wherein, is the final set of annotations, containing the sample sets of all clusters, is the sample set selected for the is the final set of annotations, containing the sample sets of all clusters, S4-2: Labeling quality evaluation and verification: perform quality evaluation on the final labeling set, calculate the diversity index, and measure the geometric distance distribution between the selected samples: ; Evaluation criteria: Ddiversity ≥ 1.2, which is an excellent labeling set, i.e. the sample distribution is uniform and the feature space coverage is sufficient; 0.8 ≤ Ddiversity < 1.2, which is a good labeling set, i.e. the sample distribution is moderate and meets the requirements; Ddiversity < 0.8, which means the labeling set needs to be improved, i.e. the samples are too concentrated, and it is suggested to adjust the algorithm parameters; S4-3: Generate standardized labeling output, which includes: feature data of selected samples, spatiotemporal coupling value score and sample selection reason, and the sample selection reason includes cluster classification, geometric characteristics and value contribution.

Citation Information

Patent Citations

  • Network construction energy storage locating and sizing method for supporting transient stability of high-proportion new energy power grid

    CN119675064A

  • Automatic data annotation method of ISP image signal processing visual sensor

    CN119723010A

  • Paleontology diversity analysis method based on dynamic machine learning joint correction model

    CN120561582A

  • Semiconductor anomaly detection method and system based on physical causal relationship modeling

    CN120781269A

  • Image noise mark feature selection method and system, storage medium and computer

    CN120894642A