Etching endpoint online control system and method based on virtual metrology

CN122800518APending Publication Date: 2026-09-22SICHUAN AFARON OPTOELECTRONICS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611053141.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-15
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

随着器件结构向更小线宽和更高深宽比演进,工艺窗口持续收窄,需要提高终点控制的精度,实际量产环境中存在多种干扰,腔体环境恶劣导致传感器信号动态退化、刻蚀反应的非线性物理耦合、晶圆面内负载效应引发的空间速率差异

Benefits of technology

[0037]本申请的有益效果:通过动态信号质量分析并融合,有效抑制数据污染、传感器噪声等引发的信号不佳对终点判断的干扰,提升系统鲁棒性;结合物理分支与数据驱动分支,可以提升模型预测结果的准确性,并提升工艺安全性;通过实时预测与补偿,能够在空间不均匀性扩大前施加干预,有效缩小各区域到达终点的进度差异,减少过刻蚀与欠刻蚀缺陷,提升全片均匀性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122800518A_ABST
    Figure CN122800518A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of etching endpoint online control, and discloses an etching endpoint online control system and method based on virtual measurement, which comprises the following steps: acquiring multi-modal sensor signals collected in real time during etching, weighting and fusing the multi-modal sensor signals to obtain fusion features; inputting the fusion features into a preset deep prediction network to output first etching depths of multiple virtual regions on a wafer surface; based on the first etching depths of the virtual regions and the wafer surface region density, analyzing the etching depth difference between the slowest region and the fastest region of etching progress, and generating a space compensation trigger signal; in response to the space compensation trigger signal, generating a process parameter compensation instruction acting on the slowest region, and issuing an endpoint control instruction when the etching depths of all the virtual regions meet preset endpoint conditions; the application can intervene before the spatial non-uniformity expands, effectively reduce the progress difference of all regions reaching the endpoint, and reduce over-etching and under-etching defects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of online control technology for etching endpoints, and more specifically to an online control system and method for etching endpoints based on virtual measurement. Background Technology

[0002] Currently, in the mass production of etching for semiconductor materials such as silicon carbide, the control precision of the etching endpoint affects device feature size, wafer uniformity, and final yield. As device structures evolve towards smaller linewidths and higher aspect ratios, the process window continues to narrow, necessitating improved endpoint control precision. However, various interferences exist in actual mass production environments, including harsh cavity environments leading to dynamic degradation of sensor signals, nonlinear physical coupling of etching reactions, and spatial velocity differences caused by in-wafer loading effects.

[0003] The existing technology has the following problems: directly using a single signal or fusing multiple signals with fixed weights does not consider the dynamic degradation of signal quality caused by polymer deposition and window contamination in the cavity, which easily leads to endpoint misjudgment or control failure; the partitioned power or back pressure adjustment adopts a control mode of detecting deviation and feedback compensation, which has a time difference from the generation of deviation to the response of the actuator, resulting in further expansion of in-plane non-uniformity and making it difficult to achieve synchronous arrival of the endpoint in each region; in order to solve at least one of the above problems, this application proposes an online control system and method for etching endpoint based on virtual measurement. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the purpose of this application is to provide an online control system and method for etching endpoints based on virtual measurement, which can effectively solve the problems in the background technology. The specific technical solution of this application is as follows:

[0005] An online control method for etching endpoint based on virtual measurement includes:

[0006] The system acquires multimodal sensor signals collected in real time during the etching process, analyzes the signal quality parameters of each sensor, calculates the corresponding dynamic fusion weights based on the signal quality parameters, and performs weighted fusion of the multimodal sensor signals according to the dynamic fusion weights to obtain fusion features.

[0007] The fusion features are input into a preset depth prediction network, which includes a physical branch that calculates the baseline etching rate based on process parameters and a data-driven branch that generates a correction amount based on the fusion features, and outputs the first etching depth of multiple virtual regions on the wafer surface.

[0008] Based on the first etching depth of each virtual region and the density of the wafer surface region, the etching depth difference between the region with the slowest etching progress and the region with the fastest etching progress is analyzed, and a space compensation trigger signal is generated.

[0009] In response to the space compensation trigger signal, a process parameter compensation command is generated for the slowest region based on the preset mapping relationship between process parameters and etching rate. When the etching depth of all virtual regions meets the preset endpoint condition, an endpoint control command is issued to control the etching endpoint.

[0010] Specifically, the multimodal sensor signals include the real-time intensity signal of the optical emission spectrometer at characteristic wavelengths, the radio frequency reflected power signal, the wafer self-bias signal, the cavity pressure signal, and the multi-zone coil current monitoring value. The analysis of the signal quality parameters of each sensor includes:

[0011] For each sensor signal, calculate the ratio of the maximum signal amplitude within a preset time window to the root mean square noise value of the sensor in the idle state of the device, and divide the ratio by a preset constant to obtain the signal-to-noise ratio index.

[0012] Calculate the reciprocal of the signal variance within a preset time window, divide the reciprocal by a preset constant, and obtain the stability index.

[0013] The percentage of each sampling point that deviates from the historical normal range within a preset time window is used as an anomaly ratio indicator.

[0014] The signal-to-noise ratio, stability, and anomaly ratio are weighted and summed according to preset weighting coefficients to obtain the signal quality parameters of the corresponding sensor signals.

[0015] Specifically, the physical branch calculates the baseline etching rate based on process parameters, including:

[0016] Acquire the RF power value, acquire the outer coil current value and the inner coil current value and calculate their current ratio, acquire the wafer self-bias voltage value, cavity pressure value and gas flow rate value;

[0017] The RF power value is raised to a first preset power, the current ratio is raised to a second preset power, the wafer self-bias value is substituted into the exponential decay function, the cavity pressure value and the gas flow rate value are raised to a third preset power and a fourth preset power respectively, and then multiplied to obtain the correction function value. The preset reference rate constant is multiplied with each calculation result to obtain the reference etching rate.

[0018] Specifically, the data-driven branch, based on the fusion feature generation rate correction, includes:

[0019] The fused features are input into a one-dimensional convolutional layer and a long short-term memory network layer. The one-dimensional convolutional layer uses multiple convolutional kernels to slide along the time dimension and extract local temporal feature sequences. The long short-term memory network layer receives the local temporal feature sequences and updates the corresponding hidden states to obtain the long-term dependencies throughout the etching process.

[0020] The hidden states output by the Long Short-Term Memory (LSTM) network layer are mapped to rate correction values ​​corresponding to each virtual region through a fully connected layer.

[0021] Specifically, the analysis of the etching depth difference between the slowest and fastest etching regions and the generation of a spatial compensation trigger signal includes:

[0022] Obtain the first etching depth of each virtual region at the current moment, and calculate the difference between it and the target etching depth to obtain the remaining depth of each virtual region;

[0023] The remaining depth is input into the pre-trained spatial coupling prediction model to obtain the predicted remaining depth of each virtual region at the next prediction time.

[0024] Iterate through the predicted remaining depth of each virtual region, and take the region with the largest predicted remaining depth as the slowest region and the region with the smallest predicted remaining depth as the fastest region.

[0025] Calculate the difference between the predicted remaining depth of the slowest region and the fastest region. If the difference is greater than a preset spatial compensation threshold, generate a spatial compensation trigger signal.

[0026] Specifically, the spatial coupling prediction model includes a cascaded structure of a multi-head self-attention network and a convolutional long short-term memory network. The multi-head self-attention network takes the feature vector composed of the current remaining depth of each virtual region and the pre-acquired graphic density values ​​of each virtual region as input, calculates the influence weights between each pair of virtual regions, and outputs an attention weight matrix. When the convolutional long short-term memory network updates its hidden state step by step, it multiplies the attention weight matrix with the convolutional kernel weights element by element to output the predicted remaining depth of each virtual region at the next prediction time.

[0027] Specifically, the spatial compensation trigger signal includes the location identifier of the slowest region and the corresponding local etching rate increment; the calculation process of the local etching rate increment includes: subtracting the predicted remaining depth of the fastest region from the predicted remaining depth of the slowest region to obtain the maximum remaining depth difference; calculating the average predicted etching rate of all virtual regions, and dividing the predicted remaining depth of the slowest region by the average predicted etching rate to obtain the remaining etching time; dividing the maximum remaining depth difference by the remaining etching time to obtain the local etching rate increment of the slowest region.

[0028] Specifically, the mapping relationship between the preset process parameters and the etching rate includes: performing multi-zone coil power variation analysis on the plasma etching equipment in advance, recording the change in etching rate of each virtual region under different power configurations, and constructing an influence coefficient matrix with virtual regions as rows and independently adjustable power control channels as columns.

[0029] Specifically, the generation of process parameter compensation instructions acting on the slowest region includes:

[0030] Based on the location identifier of the slowest region, extract the row vector of the sensitivity coefficients of the slowest region to each power control channel from the influence coefficient matrix;

[0031] Using the adjustment amount of each power control channel as a variable, and combining the sum of the products of the adjustment amount of each channel and the corresponding sensitivity coefficient, as well as the preset total power change limit, the adjustment amount of each power control channel is obtained, and process parameter compensation instructions are generated.

[0032] The online control system for etching endpoint based on virtual measurement, used to implement the online control method for etching endpoint based on virtual measurement, includes:

[0033] The data analysis module acquires multimodal sensor signals collected in real time during the etching process, analyzes the signal quality parameters of each sensor, calculates the corresponding dynamic fusion weights based on the signal quality parameters, and performs weighted fusion of the multimodal sensor signals according to the dynamic fusion weights to obtain fusion features.

[0034] The etching depth analysis module inputs the fused features into a preset depth prediction network. The depth prediction network includes a physical branch that calculates the baseline etching rate based on process parameters and a data-driven branch that corrects the generation rate of the fused features. It outputs the first etching depth of multiple virtual regions on the wafer surface.

[0035] The etching compensation module analyzes the etching depth difference between the slowest and fastest etching regions based on the first etching depth of each virtual region and the density of the wafer surface region, and generates a space compensation trigger signal.

[0036] The etching endpoint control module, in response to the space compensation trigger signal, generates a process parameter compensation instruction that acts on the slowest region according to the preset mapping relationship between process parameters and etching rate, and issues an endpoint control instruction to control the etching endpoint when the etching depth of all virtual regions meets the preset endpoint conditions.

[0037] The beneficial effects of this application are as follows: By analyzing and fusing dynamic signal quality, the interference of poor signal caused by data pollution and sensor noise on the endpoint judgment is effectively suppressed, thereby improving the robustness of the system; by combining physical branches and data-driven branches, the accuracy of model prediction results can be improved, and process safety can be enhanced; through real-time prediction and compensation, intervention can be applied before the spatial non-uniformity expands, effectively reducing the progress difference of each region in reaching the endpoint, reducing over-etching and under-etching defects, and improving the uniformity of the entire wafer. Attached Figure Description

[0038] Figure 1This is a flowchart illustrating the online control method for etching endpoint based on virtual measurement in the embodiments of this application.

[0039] Figure 2 This is a flowchart illustrating the process of generating the space compensation trigger signal in an embodiment of this application.

[0040] Figure 3 This is a schematic diagram of the online control system for etching endpoint based on virtual measurement in an embodiment of this application. Detailed Implementation

[0041] The present application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0042] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0043] Hereinafter, the terms "first," "second," and other generic terms are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0044] refer to Figure 1 The image shows a specific implementation of the online control method for etching endpoint based on virtual measurement according to this application, including:

[0045] S101. Acquire multimodal sensor signals collected in real time during the etching process, analyze the signal quality parameters of each sensor, calculate the corresponding dynamic fusion weight based on the signal quality parameters, and perform weighted fusion of the multimodal sensor signals according to the dynamic fusion weight to obtain fusion features;

[0046] S102. Input the fusion feature into a preset depth prediction network. The depth prediction network includes a physical branch that calculates the reference etching rate based on process parameters and a data-driven branch that calculates the correction amount of the fusion feature generation rate, and outputs the first etching depth of multiple virtual regions on the wafer surface.

[0047] S103. Based on the first etching depth of each virtual region and the density of the wafer surface region, analyze the etching depth difference between the region with the slowest etching progress and the region with the fastest etching progress, and generate a space compensation trigger signal.

[0048] S104. In response to the space compensation trigger signal, a process parameter compensation command is generated for the slowest region according to the preset mapping relationship between process parameters and etching rate. When the etching depth of all virtual regions meets the preset endpoint condition, an endpoint control command is issued to control the etching endpoint.

[0049] In this embodiment, multimodal sensor signals acquired in real time during the etching process are obtained, the signal quality parameters of each sensor are analyzed, and a dynamic fusion weight is calculated based on these parameters. The multimodal sensor signals are then weighted and fused to obtain fusion features. The acquired multimodal sensor signals specifically include: real-time intensity signals from an optical emission spectrometer at three characteristic wavelengths (silicon fluoride, carbon, and carbon monoxide); the radio frequency reflected power signal acquired at the output of the radio frequency matching unit; the wafer self-bias signal measured at the electrostatic chuck electrode; the cavity pressure signal measured by a capacitive pressure gauge; and the current monitoring value of the multi-zone inductively coupled plasma coil, which includes the inner coil current and the outer coil current. These signals are acquired in real time via the device communication interface at a sampling frequency of 100 Hz.

[0050] For each raw signal, median filtering is first performed, with a filtering window length of 5 sampling points. After filtering, the arithmetic mean of the sampling points within the window is calculated in 0.2-second time windows to obtain the discrete time-series value of each sensor signal within the time window. For each sensor signal, three-dimensional quality indicators are calculated within each time window.

[0051] The first metric is the signal-to-noise ratio (SNR). This is calculated by taking the maximum signal amplitude within the given time window and dividing it by the root mean square (RMS) noise value of the sensor measured beforehand when the equipment is idle (meaning there is no plasma discharge or process gas flow). The resulting ratio is then divided by a preset scaling constant and normalized to between 0 and 1. The second metric is the stability metric. This is calculated by taking the variance of the signal samples within the given time window, reciprocating the variance, dividing this reciprocal by a preset scaling constant, and then applying saturation limiting to normalize it to between 0 and 1. The third metric is the anomaly ratio. Based on historical etching data from the previous ten batches, the mean and standard deviation of each sensor signal are statistically analyzed for each time window to establish a normal range. For the current time window, the number of sampling points whose values ​​exceed the normal range (mean plus or minus three times the standard deviation) is counted. This number is divided by the total number of sampling points within the window to obtain the anomaly ratio. The signal-to-noise ratio, stability, and normal ratio (obtained by subtracting the abnormal ratio from one) are weighted and summed according to preset weighting coefficients. The sum of the three weighting coefficients is one, and the result is the signal quality parameter of the sensor in the current time window, with a value range of 0 to 1.

[0052] As an optional implementation, for optical emission spectrum signals, the ratio of the linear fitting slope over the past 30 seconds to the current window fluctuation amplitude is also calculated. If the ratio is lower than a preset threshold, it indicates that the signal has a non-etch-related trend drift, and the calculated signal quality parameter is multiplied by an attenuation coefficient of 0.7.

[0053] After obtaining the signal quality parameters of each sensor signal, dynamic fusion weights are calculated based on these parameters. This embodiment constructs a gated network to generate the fusion weights. The gated network adopts a three-layer fully connected structure: the number of nodes in the input layer equals the number of sensor signal channels (5 in this embodiment); the hidden layer contains 16 nodes and uses a linear rectified activation function; the output layer has 5 nodes, connected to a flexible maximum normalization layer, outputting 5 weight values ​​whose sum is 1. Training of the gated network is completed offline, collecting multiple batches of etching data under conditions of good cavity condition and high signal-to-noise ratio operation for each sensor. Using the equally weighted fusion results of each signal as a reference target, various signal degradation scenarios are artificially constructed, including applying a gain coefficient that decays over time to the spectral signal, injecting Gaussian noise into the reflected power signal, and applying linear drift to the pressure signal. Under each degradation scenario, the signal quality parameter vector of each sensor is calculated as input, and the loss function is the minimum mean square error between the gated network output fusion result and the reference target. The network parameters are optimized using a backpropagation algorithm.

[0054] During online operation, the five signal quality parameter vectors calculated in real time are input into the trained gating network every 0.5 seconds. After forward inference, a set of dynamic fusion weights are output. The weights are used to sum the five original signal values ​​in the same period to obtain the fusion feature vector.

[0055] It should be noted that by calculating multi-dimensional quality indicators including signal-to-noise ratio, short-term stability, and deviation from historical normal distribution, the real-time reliability of each input signal is quantitatively evaluated. A gating network trained on specific degradation scenarios is used to generate dynamic fusion weights, which can automatically and smoothly reduce the contribution when the signal quality deteriorates. When the signal quality is severely degraded, the weight is reduced to near zero to mask it, ensuring that the fusion features input to the subsequent deep prediction network always have a high signal-to-noise ratio and stability, providing a reliable guarantee for high-precision virtual measurement from the data source.

[0056] Specifically, the fused feature vectors are input into a preset depth prediction network, which includes a physical branch and a data-driven branch, and outputs the first etching depth of multiple virtual regions on the wafer surface. The wafer surface is divided into virtual regions. In this embodiment, a polar coordinate division method is used: it is divided into 8 annular zones along the radial direction and 8 sectors along the circumferential direction, forming a total of 64 virtual regions.

[0057] Specifically, the physical branch of the deep prediction network calculates the baseline etching rate based on process parameters. The input parameters acquired by the physical branch include: RF power, the ratio of the outer coil current to the inner coil current, wafer self-bias voltage, cavity pressure, and gas flow rate. The process of calculating the baseline etching rate by the physical branch is as follows: the RF power is raised to a first preset power, the current ratio is raised to a second preset power, the wafer self-bias voltage is substituted into an exponential decay function, and the cavity pressure and gas flow rate are raised to the third and fourth preset powers respectively, then multiplied to obtain a correction function value. A preset baseline rate constant is multiplied by each of the above calculation results to obtain the baseline etching rate. Each exponent and the baseline rate constant are determined by curve fitting using offline collected experimental data covering different process conditions, aiming to minimize the mean square error between the physical branch's predicted rate and the actual measured rate. The scalar baseline rate output by the physical branch is multiplied by a preset correlation coefficient for each virtual region to obtain the baseline etching rate for each region. The correlation coefficients are obtained through a finite number of experimental designs and calibrations.

[0058] Specifically, the data-driven branch of the deep prediction network generates rate corrections based on fused features. This data-driven branch includes a one-dimensional convolutional layer and a long short-term memory (LSM) network layer. The time series formed by the fused feature vectors in chronological order is input into the one-dimensional convolutional layer. This layer uses multiple convolutional kernels to slide along the time dimension to extract local temporal feature sequences. The LSM network layer receives this local temporal feature sequence and updates the hidden state through internal input gates, forget gates, and output gates to obtain the long-term dependencies throughout the etch process. The hidden state output by the LSM network layer at the final moment is mapped to a vector with the same number of virtual regions through a fully connected layer. Each element of this vector is the rate correction for the corresponding virtual region.

[0059] All parameters of the data-driven branch, including convolutional kernel weights, weights and biases of each gate in the Long Short-Term Memory network, and weights of fully connected layers, were jointly trained end-to-end with some physical branch parameters during the offline pre-training phase. The training data included a small amount of wafer data with offline depth labels measured by scanning electron microscopy, and a large amount of unlabeled mass-produced wafer data. The training loss function included the following terms: calculating the mean square error between the predicted depth and the labeled depth when labeled; penalizing a decrease in predicted depth; penalizing a prediction rate exceeding the rate cap determined based on the physical branch; and penalizing a jump in the etching rate change rate between adjacent time windows. For unlabeled data, only the last three physical constraint terms were used for training.

[0060] The fusion layer adds the baseline etching rate of each virtual region output by the physical branch to the rate correction amount of each virtual region output by the data-driven branch on the corresponding virtual regions to obtain the total predicted etching rate; this rate is then integrated over time in steps of 0.2 seconds to obtain the first etching depth of each virtual region at the current moment.

[0061] Preferably, the physical branch provides a stable and physically interpretable baseline for the prediction model, ensuring that the output always falls within a reasonable range and avoiding deep regression or speed exceeding physical limits that may occur in a purely data-driven model; the data-driven branch automatically extracts complex nonlinear relationships that the physical branch cannot accurately describe from the fused features through one-dimensional convolution and long short-term memory networks, and provides fine compensation to the baseline in the form of corrections; the physical constraint term embedded in the loss function enables the model to perform self-supervised learning using massive amounts of unlabeled mass-produced data, thereby improving model accuracy.

[0062] Specifically, the first etching depth of each virtual region at the current moment is obtained, and the difference between the target etching depth set by the process and this depth is calculated to obtain the remaining depth of each virtual region. The remaining depth is then input into a pre-trained spatial coupling prediction model to obtain the predicted remaining depth of each virtual region at the next prediction time.

[0063] Preferably, the spatial coupling prediction model comprises a cascaded structure of a multi-head self-attention network and a convolutional long short-term memory network. The model input is a feature vector consisting of the current remaining depth of each virtual region and the pre-acquired pattern density value. The pattern density value is extracted offline from the wafer layout design file and is calculated by statistically analyzing the proportion of the total area of ​​the etched opening pattern in each virtual region to the total area of ​​that region. The multi-head self-attention network receives the feature vector and has four parallel attention heads, each with an independent query weight matrix, key weight matrix, and value weight matrix. The calculation process of a single attention head is as follows: the input feature vector of each virtual region is multiplied by the three weight matrices to obtain the query vector, key vector, and value vector; the inner product of the query vector and key vector of any two regions is calculated as a similarity score; flexible maximum normalization is performed on each row of the score matrix to obtain the attention weight matrix under that head; the outputs of the four heads are averaged element-wise to obtain the attention weight matrix. The Convolutional Long Short-Term Memory (LSTM) network receives the current remaining depth vector and an attention weight matrix. When updating its hidden state step-by-step, as the convolutional kernel slides to a central region, it determines the spatial neighborhood it covers. It then extracts a submatrix from the attention weight matrix, corresponding to the row of the central region and the column of the neighborhood region. Each influence weight in this submatrix is ​​element-wise multiplied with the original weights at the corresponding spatial positions in the convolutional kernel. The result is used as the modulated kernel weights for subsequent convolution operations. Through this modulation, neighborhood regions with higher attention weights contribute more significantly to the current hidden state update. The network ultimately outputs the predicted remaining depth of each virtual region at the next prediction time step.

[0064] The training of the spatial coupling prediction model was completed offline. The training data consisted of the remaining depth sequences of multiple historical wafers and the actual depth distribution measured by scanning electron microscopy after etching. Using the mean square error between the model's predicted and measured remaining depths as the loss function, an adaptive moment estimation optimizer was employed with an initial learning rate of 0.001. End-to-end joint optimization was performed on all parameters of the multi-head self-attention network and the convolutional long short-term memory network. The predicted remaining depths of each virtual region were traversed, with the region having the largest predicted remaining depth designated as the slowest region and the region having the smallest predicted remaining depth designated as the fastest region. The difference between the two predicted remaining depths was calculated. If this difference exceeded a preset spatial compensation threshold, a spatial compensation trigger signal was generated. The spatial compensation threshold was set according to the system's etching accuracy requirements; in this embodiment, it was set to 10 nanometers.

[0065] As an optional implementation, the spatial compensation trigger signal also carries the location identifier of the slowest region and the corresponding local etching rate increment. The calculation process of the local etching rate increment is as follows: subtract the predicted remaining depth of the fastest region from the predicted remaining depth of the slowest region to obtain the maximum remaining depth difference; calculate the average value of the current predicted etching rate of all virtual regions, divide the predicted remaining depth of the slowest region by the average rate to obtain the remaining etching time; divide the maximum remaining depth difference by the remaining etching time to obtain the local etching rate increment required for the slowest region.

[0066] Preferably, the complex interaction relationships between dozens of virtual regions on the wafer surface are dynamically modeled by a multi-head self-attention network, which is more accurate than a preset fixed empirical formula. The attention weights are used as modulation factors for the spatial convolution operation of the convolutional long short-term memory network, so that the physical coupling relationship between regions is structurally embedded in the spatiotemporal evolution process, and the prediction results are more in line with the real physical laws. Based on the future prediction of the remaining depth rather than the current remaining depth, the compensation is triggered, which enables the control system to identify risk trends before the spatial non-uniformity actually expands.

[0067] In response to the spatial compensation trigger signal, a process parameter compensation command is generated for the slowest region based on the preset mapping relationship between process parameters and etching rate. An endpoint control command is issued when the etching depth of all virtual regions meets the preset endpoint condition. The preset mapping relationship between process parameters and etching rate is established offline. Multi-zone coil power variation analysis is performed on the plasma etching equipment: under standard process conditions, other parameters are fixed, and the power of the outer coil and inner coil are changed at three levels each, forming nine power configurations. Complete wafer etching is performed on each configuration, and the etching rate of each virtual region is measured. An influence coefficient matrix is ​​constructed based on the experimental data. The rows of this matrix correspond to sixty-four virtual regions, and the columns correspond to two independently adjustable power control channels: the outer and inner coil power. The element in the i-th row and j-th column of the matrix represents the influence coefficient of the unit change in the j-th power channel on the etching rate of the i-th virtual region.

[0068] Upon receiving the space compensation trigger signal, the sensitivity coefficient row vector of the region to each power control channel is extracted from the influence coefficient matrix based on the location identifier of the slowest region. Using the adjustment amount of each power control channel as a variable, and aiming to approximate the required local etching rate increment by the sum of the products of each channel's adjustment amount and its corresponding sensitivity coefficient, the adjustment amount of each channel is solved under the constraint of a preset total power change limit. In this embodiment, the total power change limit is set to 5% of the current total power setting. During the solution process, the power margin is preferentially allocated to the channel with the largest absolute value of the sensitivity coefficient. For example, if the slowest region is located at the wafer edge, the outer ring power sensitivity coefficient is usually significantly greater than the inner ring power, so the outer ring power is increased first. If adjusting the outer ring power within the limit is sufficient to meet the rate increment target, the inner ring power remains unchanged; if the outer ring power is already at the limit but still insufficient, the remaining margin is allocated to the inner ring power, or the inner ring power is slightly reduced while increasing the outer ring power to release total power space. Based on the solution results, a process parameter compensation command containing the adjustment amounts of the outer and inner ring power is generated, and a corresponding back pressure partition adjustment command is generated based on the location of the slowest region. Once the instruction is generated, it is immediately sent out for execution via the device interface.

[0069] In each predictive control cycle, the first etching depth of each virtual region is continuously monitored. When the first etching depth of all virtual regions reaches or exceeds the difference between the target etching depth and the preset endpoint threshold, the endpoint condition is determined to be met; in this embodiment, the endpoint threshold is set to 5 nanometers; an endpoint control command is immediately issued to shut off the RF power supply and process gas valves, terminating the current wafer etching process.

[0070] Preferably, the accuracy and predictability of the compensation action are ensured by using an offline calibrated influence coefficient matrix and online constraint optimization. Power and back pressure compensation are applied directionally to the slowest region to minimize the incidental interference to other regions of the wafer. The compensation command is issued immediately in response to the future deviation prediction signal, forming a complete feedforward control loop, which actively intervenes to cancel the spatial non-uniformity before it is fully manifested. The endpoint determination also depends on all regions meeting the threshold condition to ensure the spatial comprehensiveness of the endpoint determination and effectively prevent over-etching or under-etching defects.

[0071] This application effectively suppresses the interference of poor signal caused by data pollution and sensor noise on the endpoint judgment by dynamic signal quality analysis and fusion, thereby improving the robustness of the system. By combining physical branches and data-driven branches, the accuracy of model prediction results can be improved and process safety can be enhanced. Through real-time prediction and compensation, intervention can be applied before the spatial non-uniformity expands, effectively reducing the progress difference of each region to the endpoint, reducing over-etching and under-etching defects, and improving the uniformity of the whole wafer.

[0072] Furthermore, the multimodal sensor signals include the real-time intensity signal of the optical emission spectrometer at characteristic wavelengths, the radio frequency reflected power signal, the wafer self-bias signal, the cavity pressure signal, and the multi-zone coil current monitoring values. The signal quality parameters of each sensor are analyzed, including:

[0073] S201. For each sensor signal, calculate the ratio of the maximum signal amplitude within a preset time window to the root mean square noise value of the sensor in the idle state of the device, and divide the ratio by a preset constant to obtain the signal-to-noise ratio index.

[0074] S202. Calculate the reciprocal of the signal variance within the preset time window, divide the reciprocal by the preset constant, and obtain the stability index.

[0075] S203. The percentage of each sampling point that deviates from the historical normal range within the preset time window is used as the abnormality ratio indicator.

[0076] S204. The signal-to-noise ratio, stability, and anomaly ratio are weighted and summed according to preset weighting coefficients to obtain the signal quality parameters of the corresponding sensor signals.

[0077] In this embodiment, a signal-to-noise ratio (SNR) is calculated for each sensor signal to quantify the prominence of the current signal strength relative to its background noise. Signal data within a preset time window is acquired. This time window length is set to 0.2 seconds in the overall process. Within this window, all sampling points are traversed to find the maximum signal amplitude. The root mean square (RMS) noise value of the sensor, pre-measured in the equipment idle state, is obtained. The equipment idle state refers to a state where the etching equipment is without plasma discharge and without process gas flow. In this state, the sensor signal is continuously acquired for a period of time, and its RMS value is calculated. This value is the noise floor parameter of the sensor and is pre-stored in the system. The noise floor is a relatively stable equipment characteristic parameter that can be updated and calibrated during regular equipment maintenance.

[0078] The signal-to-noise ratio (SNR) of the sensor within the window is calculated by dividing the maximum signal value by the sensor's noise floor value. A larger SNR indicates a stronger signal and relatively weaker noise. To normalize this ratio to a uniform dimension, the initial SNR is then divided by a preset scaling constant. This scaling constant is chosen because when the signal quality is excellent, the ratio of the maximum signal value to the noise floor value typically reaches a value close to this constant. If the result of the division exceeds one, it is truncated to one. The resulting value is the sensor's SNR within the current window, ranging from 0 to 1.

[0079] Preferably, the signal-to-noise ratio (SNR) is calculated without relying on the absolute amplitude of the signal and can adapt to the sensitivity differences between different sensors. When the effective signal of a sensor is submerged by noise due to window contamination, cable attenuation, or partial discharge interference, the ratio of the maximum value within its window to the noise floor decreases significantly, and the SNR decreases accordingly, providing a key basis for subsequent dynamic adjustment of the fusion weights.

[0080] Specifically, a stability index is calculated for each sensor signal to assess the severity of signal fluctuations within a short time window, reflecting the signal's stability. Within the same preset time window, the variance of all sampled data points is calculated. Variance is a statistic that measures the dispersion of a set of data; the larger the value, the more drastic the signal fluctuations within that window. An index positively correlated with signal quality is constructed by taking the reciprocal of the variance. To prevent excessively large values ​​when the signal is abnormally stable and the variance approaches zero, the variance is multiplied by a preset scaling constant and a minimum constant is added before taking the reciprocal to prevent division by zero errors. The value obtained after taking the reciprocal is compared with a preset saturation upper limit. If it exceeds the upper limit, it is truncated to that upper limit value. The resulting normalized value is the stability index, ranging from 0 to 1.

[0081] Preferably, the stability index enables the system to identify whether there are instantaneous abnormal fluctuations or intermittent noise in the sensor signal. In the etching process, unstable phenomena such as plasma mode jumps, micro-arcs, and airflow disturbances often first manifest as a sudden increase in the variance of the relevant sensor signal. The stability index can sensitively capture such fluctuations, even if the absolute amplitude of the signal does not deviate significantly from the normal range. When the stability index decreases, it indicates that the sensor signal contains more high-frequency disturbance components, and its reliability is reduced.

[0082] Specifically, an anomaly ratio index is calculated for each sensor signal to measure whether the current signal value distribution conforms to the statistical regularity of the equipment under long-term normal operation. A historical normal range benchmark is established; this process is completed offline, collecting all sensor time-series data generated by the etching equipment in the past several batches (e.g., the first ten batches) under stable etching process conditions. For each sensor, data points from all time windows are aggregated, and the overall mean and standard deviation are calculated. The interval formed by adding or subtracting three times the standard deviation from the mean is defined as the historical normal value range of the sensor under that specific process step; this range data is stored in the system as a benchmark parameter. For each sampling point within the current preset time window, it is determined whether its value falls within the aforementioned historical normal value range. The number of sampling points exceeding the range within the window is counted, divided by the total number of sampling points within the window, to obtain the excess ratio. This ratio is the anomaly ratio index, ranging from 0 to 1; the higher the value, the more severe the deviation of the current signal from the historical normal pattern.

[0083] Preferably, the anomaly ratio index introduces a signal quality inspection capability based on statistical distribution laws into the system, which can effectively identify signal degradation situations, including systematic drift of the sensor and the appearance of occasional outliers in the signal, such as zero-point shift causing the overall value to gradually deviate from the historical average. When the anomaly ratio index increases, it means that the current signal characteristics are no longer consistent with the typical characteristics of the equipment in a healthy state, and its reliability as an input to the prediction model is questionable.

[0084] The calculated signal-to-noise ratio (SNR), stability, and anomaly ratio are integrated into a single comprehensive quality parameter. The contribution direction of each indicator to signal quality is clearly defined: SNR and stability are positively correlated with signal quality (higher values ​​indicate better quality), while the anomaly ratio is negatively correlated (higher anomaly ratio indicates worse signal quality). Therefore, "one minus the anomaly ratio" is used as the positive quality contribution component during fusion. Preset weighting coefficients are assigned to each of the three indicators, reflecting the degree of importance given to different quality dimensions. For example, the weight of the SNR can be set to 0.5, the stability to 0.3, and the normality ratio to 0.2; the sum of the three weighting coefficients must be 1. Each indicator is multiplied by its corresponding weighting coefficient, and the products are summed to obtain the final signal quality parameter of the sensor in the current time window. This parameter is also constrained to between 0 and 1; the closer the value is to 1, the higher the overall quality of the sensor and the more reliable its output signal.

[0085] Preferably, by weighted fusion, a comprehensive signal quality parameter is obtained, which unifies the multi-dimensional and independent signal quality evaluation results into a scalar value, thereby realizing a multi-dimensional comprehensive evaluation of the signal quality of different sensors.

[0086] Furthermore, the physics branch calculates the baseline etching rate based on process parameters, including:

[0087] S301. Obtain the RF power value, obtain the outer coil current value and the inner coil current value and calculate their current ratio, obtain the wafer self-bias voltage value, cavity pressure value and gas flow rate value.

[0088] S302. The RF power value is raised to a first preset power, the current ratio is raised to a second preset power, the wafer self-bias value is substituted into the exponential decay function, the cavity pressure value and the gas flow rate value are raised to a third preset power and a fourth preset power respectively, and then multiplied to obtain the correction function value. The preset reference rate constant is multiplied with each calculation result to obtain the reference etching rate.

[0089] In this embodiment, the process parameters and plasma state parameters required for the physical branch are acquired and preprocessed to provide a data foundation for subsequent baseline etching rate calculations. The radio frequency (RF) power value is acquired, derived from the set value or actual output monitoring value of the RF power supply of the etching equipment. In the physical branch, the effective value of the RF power is used as input, directly determining the overall level of electron density and ion energy in the plasma, and is one of the most important driving factors for the etching rate. The current values ​​of the multi-zone coils are acquired and their current ratios are calculated. The current values ​​of the inner and outer coils are read in real time from the equipment sensors. The outer current value is divided by the inner current value, and the resulting ratio is the current ratio. This ratio reflects the distribution of plasma in the radial space of the wafer: a ratio greater than one indicates a higher power proportion in the outer zone, with plasma density concentrated at the edges; a ratio less than one indicates a higher power proportion in the inner zone, with plasma density concentrated at the center, directly related to the spatial differences in etching rate across different regions of the wafer.

[0090] Specifically, the wafer self-bias voltage is obtained, measured from the electrostatic chuck electrode. This is a DC bias voltage, and in inductively coupled plasma etching, the magnitude of the self-bias voltage determines the energy of ions bombarding the wafer surface. The larger the absolute value of the self-bias voltage, the higher the ion bombardment energy, and the stronger the physical breaking effect on the chemical bonds of silicon carbide material, thereby increasing the etching rate. The chamber pressure and gas flow rate are also obtained. The chamber pressure is measured in real time by a capacitive pressure gauge, reflecting the number density and mean free path of neutral particles and ions in the reaction chamber. The gas flow rate comes from the setpoint or feedback value of the mass flow controller, determining the supply rate of the reaction precursor. Both together affect the surface adsorption of the etching reactants and the desorption of the products. All parameters are obtained in real time from the equipment interface or read from the process recipe in each prediction cycle, forming the input parameter set for the physical branch.

[0091] Preferably, by clearly defining the set of input parameters for the physical branch, an input space with clear physical meaning is constructed for subsequent baseline etching rate calculations. The selected RF power, coil current ratio, self-bias voltage, cavity pressure, and gas flow rate correspond to core physical driving factors such as plasma density, spatial distribution, ion bombardment energy, neutral particle transport, and reactant supply, respectively. The parameter selection ensures that the output of the physical branch can respond to changes in process conditions from a first-principles level. By incorporating the coil current ratio and self-bias voltage into the input, the physical branch is able to sense changes in plasma spatial distribution and ion energy.

[0092] The acquired parameters are used to calculate the baseline etching rate through a series of physically meaningful nonlinear transformations and multiplications, based on the ion bombardment-assisted mechanism of silicon carbide etching. A baseline rate constant is preset as the rate base value, and then multiplied by dimensionless correction factors derived from each input parameter. Each correction factor reflects the contribution of the corresponding physical quantity to the etching rate. The RF power value is exponentially calculated, with a preset first exponent, typically between 0.5 and 1. The RF power value is then raised to this power. This process is based on the empirical law in plasma physics that the etching rate increases sublinearly with power; that is, doubling the power usually results in a rate increase of less than double. The current ratio is exponentially calculated, with a preset second exponent, and then raised to this power. This correction reflects the slight influence of changes in plasma spatial distribution on the overall average rate. When the current ratio deviates from the standard value, it indicates a change in plasma spatial distribution, and the baseline rate is adjusted accordingly.

[0093] Specifically, by substituting the wafer self-bias voltage into an exponential decay function, an exponential function with the reciprocal of the self-bias voltage as the independent variable is constructed, and multiplied by the ratio of an equivalent activation energy parameter to the Boltzmann constant as the exponential coefficient. Higher ion bombardment energy more effectively overcomes the surface activation energy barrier of the etching reaction, thus exponentially increasing the etching rate; the larger the absolute value of the self-bias voltage, the larger the value of this exponential term; when the self-bias voltage approaches zero, this exponential term decays sharply, reflecting the characteristic that silicon carbide is almost unetched when ion bombardment is lacking.

[0094] Specifically, the chamber pressure and gas flow rate values ​​are multiplied by exponentiation, with a third power exponent for chamber pressure and a fourth power exponent for gas flow rate. The pressure ratio and flow rate ratio are then multiplied by their respective powers to obtain a combined correction function value for pressure and flow rate. This combined correction function value reflects the synergistic effect of reactant concentration and residence time on the etching rate. Increased pressure increases the neutral particle number density but shortens the mean free path; increased flow rate ensures sufficient reactant supply but shortens residence time. The combined effect of these two factors is reflected in this correction function. Multiplying the preset baseline rate constant by the calculated correction factors yields the baseline etching rate under the current process conditions.

[0095] Preferably, a mechanism-driven modeling of the silicon carbide etching reference rate is achieved through a term-by-term nonlinear transformation and multiplication structure with physical correspondences. The output of the physics branch is highly interpretable; any change in the rate can be traced back to a change in specific physical parameters. The output of the physics branch is strictly limited to the reasonable value range of each physical quantity. Even if the data-driven branch outputs an abnormal correction amount, the reference rate itself will not experience an order-of-magnitude error jump due to small fluctuations in the input parameters, providing a safe baseline anchor for the entire prediction model. By explicitly introducing the self-bias voltage reflecting the spatial distribution of plasma current ratio and ion energy, the reference rate output by the physics branch implicitly contains a preliminary perception of spatial differences, providing a reference value that conforms to physical trends for the subsequent generation of spatial distribution reference rate maps. Each parameter is determined through offline data fitting or joint training, enabling the physical model to be adaptively calibrated for specific etching equipment and process formulations.

[0096] Furthermore, the data-driven branch generates a rate correction based on fused features, including:

[0097] S401. Input the fused features into a one-dimensional convolutional layer and a long short-term memory network layer. The one-dimensional convolutional layer uses multiple convolutional kernels to slide along the time dimension and extract local temporal feature sequences. The long short-term memory network layer receives the local temporal feature sequences and updates the corresponding hidden states to obtain the long-term dependencies of the entire etching process.

[0098] S402. The hidden states output by the Long Short-Term Memory network layer are mapped to rate correction values ​​corresponding to each virtual region through a fully connected layer.

[0099] In this embodiment, the fused features are input into a one-dimensional convolutional layer and a long short-term memory network layer to extract local temporal feature sequences and obtain the long-term dependencies throughout the etching process. During the etching process of each wafer, step S101 outputs a fused feature vector at fixed intervals. This vector is a five-dimensional vector, with each dimension corresponding to the dynamically fused comprehensive features. As time progresses, these vectors are arranged in chronological order to form a time series. This series is the input for this step, with each time step in the series corresponding to a fused feature vector.

[0100] Specifically, a one-dimensional convolutional layer processes the time series. This layer contains multiple learnable convolutional kernels, each consisting of a set of weight parameters with a fixed length. In this embodiment, eight convolutional kernels are used, each covering five consecutive time steps. During computation, each convolutional kernel slides progressively along the time axis, covering a local continuous time window at each step. At each window position, the weights of the convolutional kernel are multiplied element-wise with the fused feature values ​​of the corresponding time steps within the window, and the results are summed to obtain an output value. As the convolutional kernel traverses the entire time series, a new time series is generated, called the local temporal feature sequence extracted by that convolutional kernel.

[0101] The eight local temporal feature sequences output from the one-dimensional convolutional layer are input into a Long Short-Term Memory (LSTM) network layer in their original chronological order. The LSM network is a recurrent neural network that maintains a hidden state vector and a cell state vector at each time step. In this embodiment, the hidden state vector is set to 16 dimensions. At each time step, the network receives the eight input feature sequences of the current moment and combines them with the hidden state and cell state passed from the previous moment. Information is updated through three internal gating structures: the input gate determines how much new information is written into the cell state; the forget gate determines how much information is discarded from the old cell state; and the output gate determines what hidden state to output based on the updated cell state. Through this gating mechanism, the network can selectively remember important long-term trend information, such as the slow rise in cavity temperature, while forgetting irrelevant instantaneous noise throughout the etching process. After processing all time steps of the current batch of wafer etching, the hidden state vector output by the LSM network layer at the final moment is considered to contain the long-term dependencies and dynamic evolution patterns of the entire etching process.

[0102] The model training process includes: all parameters of the one-dimensional convolutional layers and long short-term memory (LSM) network layers, including the weights of each of the eight convolutional kernels, the weight matrices and bias vectors of the three gating units of the LSM network, were determined during the offline pre-training phase through end-to-end joint training with the rest of the depth prediction network. The training dataset consists of two parts: the first part is a small amount of complete time-series data from wafers with offline depth labels measured by scanning electron microscopy, used to provide supervision signals; the second part is a large amount of unlabeled time-series data from mass-produced wafers, used to provide self-supervised signals for physical constraints. During training, data is fed into the network in batches, with a batch size of 32. The network forward propagates to output the predicted depth, calculates the mean squared error between the predicted depth and the labeled depth, and includes penalties for various physical constraints. The gradient of the loss function is backpropagated layer by layer using a time-based backpropagation algorithm, and an adaptive moment estimation optimizer is used to update the parameters of each layer. The initial learning rate is set to 0.001, and the number of training epochs is determined based on the convergence of the validation set error.

[0103] Preferably, the cascaded processing of one-dimensional convolutional layers and long short-term memory network layers provides powerful automatic extraction of temporal features and long-term dependency modeling capabilities for data-driven branches. The one-dimensional convolutional layer, through the sliding scan of multiple convolutional kernels, can decompose various local dynamic patterns with different time scales from the original fused feature sequence. The long short-term memory network layer connects these local features on the time axis, effectively filters instantaneous noise through the gating mechanism, and accurately captures the long-term evolution trend across the entire etching process, laying the temporal information foundation for the subsequent generation of accurate rate corrections.

[0104] Specifically, the hidden state vector output by the Long Short-Term Memory (LSTM) network layer at its final time step is mapped to a rate correction amount corresponding to each virtual region through a fully connected layer. The input is the final hidden state vector, which is 16-dimensional. Each element of this vector is a distributed encoding of the etching process history information. The output is a one-dimensional vector whose length is equal to the number of pre-divided virtual regions on the wafer surface, which is 64 in this embodiment. Each element in the output vector corresponds to the etching rate correction amount required for a specific virtual region, and its unit is consistent with the etching rate, such as nanometers per second.

[0105] Furthermore, mapping is performed through a fully connected layer. The fully connected layer consists of a learnable weight matrix and a bias vector. The weight matrix has 64 rows, corresponding to the dimension of the output vector, and 16 columns, corresponding to the dimension of the input hidden state vector. The bias vector has a length of 64. The specific calculation process is as follows: the input 16-dimensional hidden state vector is multiplied by the 64x16 weight matrix to obtain a 64-dimensional intermediate vector; each element of this intermediate vector is then added to the corresponding bias value in the bias vector to obtain the 64-dimensional output vector. Each element of this output vector represents the rate correction for the corresponding virtual region. This rate correction can be positive or negative: a positive value indicates that the etching rate of this region should be further increased based on the baseline rate output by the physical branch; a negative value indicates that it should be suppressed.

[0106] Regarding the training of the fully connected layer: the weight matrix and bias vector of this layer were also determined end-to-end in the aforementioned offline pre-training stage, together with the one-dimensional convolutional layer and the long short-term memory network layer; during the training process, the gradient of the loss function is backpropagated from the final prediction depth to the fully connected layer, enabling the fully connected layer to learn to decode the rate correction information specific to each virtual region from the abstract hidden state. After training is completed, the parameters are fixed.

[0107] Preferably, a fully connected layer maps temporal features to spatially distributed rate corrections. During training, the weight matrix of the fully connected layer automatically learns the statistical correlation between each dimension in the hidden state and the etching behavior of each virtual region. The data-driven branch can learn complex spatial nonlinear relationships such as load effects and pattern density effects from historical data without any explicit physical equations. The output 64-dimensional rate correction vector serves as a fine-grained compensation for the baseline rate of the physical branch, enabling the overall depth prediction network to make high-precision and differentiated predictions of the etching depth in different regions of the wafer surface, providing a precise decision basis for subsequent spatial uniformity control.

[0108] like Figure 2 As shown, the etching depth difference between the slowest and fastest etching regions is analyzed, and a space compensation trigger signal is generated, including:

[0109] S501. Obtain the first etching depth of each virtual region at the current time, and calculate the difference between it and the target etching depth to obtain the remaining depth of each virtual region.

[0110] S502. Input the remaining depth into the pre-trained spatial coupling prediction model to obtain the predicted remaining depth of each virtual region at the next prediction time.

[0111] S503. Traverse the predicted remaining depth of each virtual region, and take the region with the largest predicted remaining depth as the slowest region and the region with the smallest predicted remaining depth as the fastest region.

[0112] S504. Calculate the difference between the predicted remaining depth of the slowest region and the fastest region. If the difference is greater than the preset spatial compensation threshold, generate a spatial compensation trigger signal.

[0113] In this embodiment, the first etching depth of each virtual region at the current moment is obtained, and the difference between this and the target etching depth is calculated to obtain the remaining depth of each virtual region. The first etching depth of each virtual region at the current moment is then used to obtain the target etching depth of the wafer. This value is preset by the process formula and represents the final depth expected to be achieved by this etching process.

[0114] Specifically, iterate through all virtual regions. For each region, subtract the current first etching depth of the region from the target etching depth. The difference is the remaining etching depth of the region at the current moment. If the difference is positive, it means that the region has not yet reached the target depth and still needs to be etched. If the difference is less than or equal to zero, it means that the region has theoretically been etched. The calculation result is the current remaining depth vector.

[0115] Preferably, by converting the absolute etching depth of each region into the remaining depth relative to a common target, the prediction task is transformed from predicting absolute depth to predicting the distance to the endpoint, providing a unified and intuitive decision-making basis for subsequent spatiotemporal prediction and uniformity control. The absolute depth values ​​can vary greatly due to differences in different process formulations, while the remaining depth normalizes each region to the same target-oriented scale, allowing the spatially coupled prediction model to focus on learning the dynamic behavior at the end of etching and the competitive relationships between regions.

[0116] The calculated remaining depth vectors of each virtual region are input into a pre-trained spatial coupling prediction model to obtain the predicted remaining depth of each virtual region at the next prediction time. For each virtual region, its current remaining depth is concatenated with a pre-acquired static feature to form a two-dimensional feature vector. The static feature is the pattern density value. The pattern density value is extracted offline from the wafer layout design file. Specifically, it is calculated by parsing the pattern geometry information in the layout file, dividing the virtual regions into grids according to the same method as in step S102, and calculating the ratio of the total area of ​​all etched opening patterns in each grid to the total area of ​​that grid. This ratio is the pattern density, with a value ranging from 0 to 1. The two-dimensional feature vectors of all 64 virtual regions are combined to obtain the input feature matrix.

[0117] The input feature matrix is ​​fed into a spatially coupled prediction model, which employs a cascaded structure comprising a multi-head self-attention network module and a convolutional long short-term memory network module. The multi-head self-attention network module calculates the dynamic influence weights between each pair of virtual regions. It contains several parallel attention heads; in this embodiment, four are used. Each attention head has an independent learnable parameter matrix. Taking a single attention head as an example, the operation process is as follows: the two-dimensional input feature vector of each virtual region is mapped to a query vector, a key vector, and a value vector through three different linear transformation matrices. The number of rows and columns of the linear transformation matrix is ​​related to the input feature dimension and the network's preset hidden dimension; in this embodiment, the hidden dimension is set to 32. For any two regions, the inner product of the query vector of the first region and the key vector of the second region is calculated as the original similarity score between them. The original scores of all region pairs form a 64-row, 64-column score matrix. Flexible maximum normalization is performed on each row of this matrix, transforming all values ​​in each row into a probability distribution that sums to one; elements with larger values ​​are closer to one after normalization. The normalized matrix is ​​the attention weight matrix output by the attention head. The above process is executed in parallel by four attention heads, and the attention weight matrices output by each head are averaged element by element to obtain the attention weight matrix; the element in the m-th row and n-th column of the matrix represents the influence weight of the current state of the n-th region on the state of the m-th region at the next time step.

[0118] The Convolutional Long Short-Term Memory (LSTM) network module predicts the remaining depth of each region at the next time step based on the current state and the influence weights between regions. This module receives two inputs: a vector of the remaining depth of each virtual region (length 64) and an attention weight matrix output by the multi-head self-attention network. The network maintains a hidden state tensor whose spatial dimension corresponds to the two-dimensional spatial arrangement of the virtual regions, with 32 channels. During time-step updates to the hidden state, the network performs spatial convolution operations. These operations are modulated by the attention weight matrix. Specifically, when the convolution kernel slides to a central position, it determines the spatial neighborhood it currently covers, for example, the central region and its eight surrounding neighboring regions (nine regions in total). A submatrix is ​​extracted from the attention weight matrix, indexed by the central region (row index) and the nine neighboring regions (column index). This submatrix has one row and nine columns. The nine values ​​in the submatrix represent the influence weights of these neighboring regions on the central region. The original weights in the convolution kernel corresponding to these nine spatial positions are then multiplied element-wise by the corresponding influence weights in the submatrix. The result of the multiplication is used as the modulated convolution kernel weights for subsequent convolution operations. Through this modulation, the contribution of neighborhood regions with higher attention weights to the current hidden state update is amplified. The network finally maps the current hidden state to the remaining prediction depth of each virtual region at the next prediction time through an output fully connected layer, outputting a 64-dimensional vector.

[0119] The training of the spatially coupled prediction model is completed offline, with training data consisting of representative historical etched wafer data. The data for each wafer includes: the remaining depth sequence of each virtual region at each moment recorded during the etching process, and the final remaining depth distribution of each virtual region measured by scanning electron microscopy after the wafer is etched. During training, the historical remaining depth sequence and corresponding pattern density are input into the model, and the model outputs the predicted remaining depth sequence at each moment through forward propagation. The loss function is defined as the mean square error between the model's predicted final remaining depth distribution and the actual remaining depth distribution measured by scanning electron microscopy. A time-based backpropagation algorithm and an adaptive moment estimation optimizer are used to jointly optimize all learnable parameters of the multi-head self-attention network and the convolutional long short-term memory network end-to-end. In the training hyperparameters, the initial learning rate can be set to 0.001, the batch size can be set to 32, and the number of training epochs is determined based on the convergence of the validation set error.

[0120] Preferably, by introducing a spatial coupling prediction model based on multi-head self-attention and modulated convolutional long short-term memory networks, accurate prediction of the future evolution trend of etching progress in multiple regions on the wafer surface is achieved. The multi-head self-attention network can automatically learn the complex and nonlinear dynamic influence relationships between virtual regions from historical data. It uses attention weights as modulation factors for the spatial convolution of the convolutional long short-term memory network, enabling differentiated transmission of information flow between regions according to the learned coupling strength. The prediction results not only depend on the historical time sequence of each region but also integrate the state evolution trend of the spatial neighborhood, improving the spatiotemporal accuracy and physical rationality of the prediction.

[0121] Specifically, the remaining depth of each virtual region is predicted, identifying the slowest and fastest etching progress. The remaining depth prediction vector for the next time step is obtained. Each element in the vector is accessed sequentially, recording the current maximum value and the index of the region where that maximum value was obtained, as well as the current minimum value and the index of the region where that minimum value was obtained. The index can be the grid number of the virtual region, such as an integer from one to sixty-four, or it can be represented in polar coordinates. The virtual region with the largest predicted remaining depth value is identified as the slowest region, indicating that its etching progress lags behind all other regions; the virtual region with the smallest predicted remaining depth value is identified as the fastest region, indicating that its etching progress leads all other regions. The location identifier of the slowest region, the predicted remaining depth value of the slowest region, and the predicted remaining depth value of the fastest region are output.

[0122] Preferably, by locating the slowest region, limited compensation resources can be precisely targeted to the bottleneck area that limits the total etching time of the entire wafer, avoiding the inefficient adjustment of a uniform, scattered approach. Simultaneously, identifying the fastest region provides a benchmark for calculating the time difference between the two. This precise location of the slowest and fastest regions constitutes a crucial information link from front-end spatial prediction to back-end spatial pre-compensation control.

[0123] Specifically, the difference between the predicted remaining depths of the slowest and fastest regions is calculated. The predicted remaining depth of the slowest region is subtracted from the predicted remaining depth of the fastest region to obtain the difference. This difference reflects the remaining depth gap between the fastest and slowest etching positions on the wafer surface at the next prediction time. A larger difference indicates a more significant spatial non-uniformity trend. Without intervention, the fastest region will reach its endpoint before the slowest region, resulting in over-etching in the former and under-etching in the latter. This difference is compared with a preset spatial compensation trigger threshold. This threshold is an empirical parameter set according to process accuracy requirements; in this embodiment, it is set to 10 nanometers.

[0124] Based on the comparison results, if the calculated difference is greater than a preset threshold, it indicates that if the current state continues, the slowest and fastest regions will not be able to reach the endpoint synchronously at approximately the same time. In this case, the system determines that spatial pre-compensation intervention is needed and generates a spatial compensation trigger signal. If the difference is less than or equal to the threshold, it is considered that the current spatial non-uniformity is still within an acceptable range, and compensation is not triggered for the time being.

[0125] As an optional implementation, the spatial compensation trigger signal also carries two key pieces of information: first, the location identifier of the slowest region, used to locate the compensation target; and second, the local etching rate increment required for the slowest region to catch up with the fastest region's progress. This local rate increment is calculated as follows: calculate the average of the current predicted etching rates for all virtual regions; divide the predicted remaining depth of the slowest region by this average rate to estimate the remaining time required to continue etching at the current average rate; divide the difference in predicted remaining depth between the slowest and fastest regions by this remaining time to obtain the required additional etching rate value.

[0126] Preferably, by setting a reasonable trigger threshold and making conditional judgments, necessary decision support is introduced for the spatial pre-compensation mechanism. The compensation instruction generation module only needs to respond to the signal to directly extract the target position and velocity increment for accurate calculation, without having to repeatedly perform spatial traversal and difference analysis, thus ensuring the accuracy, timeliness and system operating efficiency of spatial compensation control.

[0127] Furthermore, the spatially coupled prediction model includes a cascaded structure of a multi-head self-attention network and a convolutional long short-term memory network. The multi-head self-attention network takes the feature vector composed of the current remaining depth of each virtual region and the pre-acquired graphic density values ​​of each virtual region as input, calculates the influence weights between each pair of virtual regions, and outputs an attention weight matrix. When the convolutional long short-term memory network updates its hidden state step by step, it multiplies the attention weight matrix with the convolutional kernel weights element by element to output the predicted remaining depth of each virtual region at the next prediction time.

[0128] In this embodiment, the spatial coupling prediction model adopts a cascaded architecture of a multi-head self-attention network and a convolutional long short-term memory network. The model input is a two-dimensional feature vector formed by concatenating the current remaining depth and corresponding graphic density of each virtual region, totaling 64 virtual regions, forming a 64-row, 2-column input matrix. The multi-head self-attention network internally sets up four parallel attention heads, each with three independent sets of linear transformation matrices, mapping the input features to query vectors, key vectors, and value vectors respectively, with a hidden dimension of 32. Each attention head obtains the original similarity score by calculating the inner product of the query vector and key vector between regions, forming a 64-row, 64-column score matrix. After flexible maximum normalization, the attention weight matrix for that head is obtained. The element-wise average of the output matrices of the four heads is used as the attention weight matrix.

[0129] Specifically, the convolutional long short-term memory network receives the current remaining depth vector and the aforementioned attention weight matrix. Internally, it maintains a hidden state tensor with 32 channels, spatially arranged to correspond to the virtual regions. During time-step updates to the hidden state, for each spatial location, it takes its 3x3 neighborhood (9 regions in total). From the attention weight matrix, it extracts a 9-column submatrix corresponding to the row of the central region and the columns of the neighboring regions. The 9 influencing weights within this submatrix are element-wise multiplied with the original weights corresponding to the 9 spatial locations in the convolution kernel to obtain modulated kernel weights, which are used for the convolution operation at the current central location. The network updates the hidden state and cell states through input gates, forget gates, output gates, and the calculation of candidate cell states. The output fully connected layer then yields a 64-dimensional predicted remaining depth vector for the next time step.

[0130] The training of the spatial coupling prediction model was completed offline. The training data came from the complete etching process of multiple historical wafers, including the remaining depth sequence of each virtual region and the final remaining depth distribution measured by scanning electron microscopy after etching. The loss function was defined as the mean square error between the final remaining depth predicted by the model and the measured remaining depth. The training adopted an adaptive moment estimation optimizer with an initial learning rate of 0.001 and a batch size of 32. All learnable parameters of the multi-head self-attention network and the convolutional long short-term memory network were jointly optimized end-to-end through time-based backpropagation. The number of training epochs was determined based on the convergence of the validation set error.

[0131] Preferably, through the spatial coupling prediction model, the multi-head self-attention network endows the model with the ability to explicitly model the complex interaction relationships between multiple virtual regions on the wafer surface, and the parallel operation of multiple attention heads can simultaneously capture coupling patterns from different representation subspaces; the attention weights are used as modulation factors for the spatial convolution operation of the convolutional long short-term memory network, so that the information flow between regions is transmitted differentially according to the learned dynamic coupling strength, and the neighborhood region with higher attention weights contributes more to the current state update; thus improving the model accuracy and the accuracy of the output results.

[0132] Furthermore, the spatial compensation trigger signal includes the location identifier of the slowest region and the corresponding local etching rate increment; the calculation process of the local etching rate increment includes: subtracting the predicted remaining depth of the fastest region from the predicted remaining depth of the slowest region to obtain the maximum remaining depth difference; calculating the average predicted etching rate of all virtual regions, and dividing the predicted remaining depth of the slowest region by the average predicted etching rate to obtain the remaining etching time; dividing the maximum remaining depth difference by the remaining etching time to obtain the local etching rate increment of the slowest region.

[0133] In this embodiment, during the generation of the spatial compensation trigger signal, the progress difference between the slowest and fastest regions is quantified, and the predicted remaining depth values ​​for the slowest and fastest regions at the next prediction time are extracted. The predicted remaining depth value of the slowest region is subtracted from the predicted remaining depth value of the fastest region. Since the remaining depth of the slowest region is necessarily greater than that of the fastest region, the difference is positive, and this positive value is the maximum remaining depth difference.

[0134] Specifically, the analysis determines how much time is needed for the slowest region to reach the target depth based on the current overall etching capability. The average predicted etching rate of all virtual regions at the current moment is obtained by taking the arithmetic mean of the 64-dimensional total predicted etching rate vector output from the depth prediction network fusion layer in step S102. This average rate reflects the overall etching efficiency of the wafer surface under the current plasma environment and process parameter configuration. The remaining etching time is obtained by dividing the predicted remaining depth of the slowest region by this average predicted etching rate.

[0135] After obtaining the maximum remaining depth difference and the remaining etching time, the maximum remaining depth difference is divided by the remaining etching time. The quotient is the local etching rate increment. The local etching rate increment and the location marker of the slowest region are integrated to generate a spatial compensation trigger signal.

[0136] Preferably, by calculating the local etching rate increment, the physical scale of the progress difference between regions is quantified by the maximum remaining depth difference. The remaining etching time is estimated using the current overall average rate to avoid giving an excessively high rate target that cannot be achieved within the remaining time. The obtained rate increment serves as a clear target value for subsequent constraint optimization solutions, ensuring high accuracy and high efficiency of feedforward compensation.

[0137] Furthermore, the mapping relationship between the preset process parameters and the etching rate includes: performing multi-zone coil power variation analysis on the plasma etching equipment in advance, recording the change in etching rate of each virtual region under different power configurations, and constructing an influence coefficient matrix with virtual regions as rows and independently adjustable power control channels as columns.

[0138] In this embodiment, the mapping relationship between preset process parameters and etching rate is established through systematic experiments in the offline stage, determining independently adjustable power control channels, including at least two channels: outer coil power and inner coil power. Based on the standard silicon carbide etching process formula, other parameters such as total RF power, cavity pressure, gas flow rate, and electrostatic chuck back pressure remain unchanged, with only the outer and inner coil power used as experimental variables. For each power channel, three levels are selected near its current process setting value: the current value, the current value increased by 5%, and the current value decreased by 5%. Two channels each had three levels, and a full factorial experimental design was used to form nine different power configuration combinations. For each power configuration, a complete wafer etching process was performed, with the etching time consistent with the standard formulation. After etching, the depth of 64 virtual regions on the wafer surface was measured using a scanning electron microscope. The measured depth of each region was divided by the etching time to obtain the etching rate value of each virtual region under that power configuration. The correspondence between the nine power configuration vectors and the nine 64-dimensional etching rate vectors was obtained, forming the original dataset for constructing the influence coefficient matrix.

[0139] Specifically, sensitivity analysis was performed on each of the sixty-four virtual regions. Taking a single virtual region as an example, nine etching rate values ​​and corresponding nine sets of outer and inner power values ​​were extracted from nine sets of experiments for that region. The etching rate of the region was used as the dependent variable, and the outer and inner power were used as the two independent variables. A multiple linear fitting method was used to quantify the quantitative relationship between the two. Specifically, a set of coefficients was solved using the least squares method to minimize the sum of squares of the deviations between the rate values ​​predicted by the fitted line and the nine actual rate values. In the fitting results, the coefficient corresponding to the outer power is the sensitivity coefficient of the virtual region to the outer power, and the coefficient corresponding to the inner power is the sensitivity coefficient of the virtual region to the inner power.

[0140] It should be noted that the physical meaning of the sensitivity coefficient is: the average change in the etching rate of the virtual region for every unit change in the corresponding power channel. Arrange the two sensitivity coefficients of each of the 64 virtual regions in rows, with each region in one row. The outer power coefficient is in the first column and the inner power coefficient is in the second column, constructing a 64-row, 2-column influence coefficient matrix.

[0141] Preferably, by constructing an influence coefficient matrix to reflect the differentiated response characteristics of different spatial locations on the wafer surface to the same power adjustment action, during online control, the system can instantly retrieve the corresponding sensitivity coefficient row vector from the matrix based on the location identifier of the slowest region, quickly determine the adjustment efficiency of each power channel, and provide gradient direction for constrained optimization solution.

[0142] Furthermore, process parameter compensation instructions are generated for the slowest region, including:

[0143] S901. Extract the row vector of the sensitivity coefficients of the slowest region to each power control channel from the influence coefficient matrix based on the location identifier of the slowest region.

[0144] S902. Using the adjustment amount of each power control channel as a variable, and combining the sum of the products of the adjustment amount of each channel and the corresponding sensitivity coefficient, as well as the preset total power change limit, the adjustment amount of each power control channel is obtained, and the process parameter compensation command is generated.

[0145] In this embodiment, when the space compensation trigger signal arrives, the slowest region location identifier carried in the parsing signal is analyzed. Based on the location identifier, the influence coefficient matrix pre-stored in the non-volatile memory is accessed. The location identifier is used as the row number, and the values ​​of the first and second columns of the row are read out to form a two-element row vector, which reflects the response sensitivity of the specific region to the adjustment of the outer and inner ring power. The larger the value, the more significant the change in etching rate caused by the unit power change.

[0146] After obtaining the sensitivity coefficient row vector and the required rate increment, the adjustment amount solution stage begins. The known conditions are the outer ring power sensitivity coefficient, the inner ring power sensitivity coefficient, the target value of the required rate increment, and the preset total power change limit, which is usually set to 5% of the current total power setting value. The variables to be solved are the outer ring power adjustment amount and the inner ring power adjustment amount. The solution objective is the comprehensive rate increase after multiplying the two adjustment amounts by their corresponding sensitivity coefficients, which should be as close as possible to the target rate increment. The solution constraint is that the sum of the absolute values ​​of the two adjustment amounts does not exceed the total power change limit.

[0147] The solution strategy employs a priority allocation method: comparing the absolute values ​​of two sensitivity coefficients, priority is given to allocating the available power margin to the channel with the larger sensitivity coefficient. For example, if the outer ring sensitivity coefficient is 0.8, which is greater than the inner ring sensitivity coefficient of 0.2, then the outer ring power is adjusted first. The maximum rate increase that could be achieved by allocating the entire limiting margin to the outer ring power is calculated. If this maximum increase already meets the target, only a portion of the margin needs to be used, and the adjustment amount equals the required rate increment divided by the outer ring sensitivity coefficient. If the maximum increase is still insufficient, the outer ring power is adjusted to the upper limit of the limiting, and the already generated rate increase is calculated. The remaining gap is obtained by subtracting the generated amount from the target value. The remaining gap is divided by the inner ring sensitivity coefficient to obtain the required adjustment amount for the inner ring power, but it is necessary to check whether the total power change exceeds the limit. If it does, an increase-decrease substitution strategy is adopted: while increasing the power of high-sensitivity channels, the power of low-sensitivity channels is appropriately reduced to release total power space. The optimal combination that meets the rate target within the limiting is found through iterative calculation.

[0148] After the solution is completed, the adjustment values ​​of the outer and inner ring power are updated to the current process formula settings. At the same time, the back pressure zones of the electrostatic chuck are determined based on the physical location of the slowest region, and instructions are generated to reduce the back pressure of the zone containing the fastest region and increase the back pressure of the zone containing the fastest region. These adjustment values ​​are packaged into the equipment communication protocol format and immediately sent to the RF power controller and back pressure controller for execution.

[0149] Preferably, by instantly retrieving the influence coefficient matrix, the system can obtain the precise response characteristics of the slowest region to each adjustment channel within milliseconds, and the response speed meets the requirements of real-time control. The priority allocation strategy based on the sensitivity coefficient ensures that the limited power adjustment margin is used for the most efficient channel, exchanging the minimum total power disturbance for the maximum local rate increase, effectively avoiding global fluctuations in plasma state and deterioration of etching quality in other regions caused by over-adjustment. The total power change limit is strictly enforced as a hard constraint, ensuring the overall stability of the process and the safety of equipment operation, and preventing the mismatch of the matching network or the over-limit of reflection power due to excessive amplitude of a single compensation action.

[0150] like Figure 3 As shown, an online control system for etching endpoint based on virtual measurement is used to implement an online control method for etching endpoint based on virtual measurement, including:

[0151] The data analysis module acquires multimodal sensor signals collected in real time during the etching process, analyzes the signal quality parameters of each sensor, calculates the corresponding dynamic fusion weights based on the signal quality parameters, and performs weighted fusion of the multimodal sensor signals according to the dynamic fusion weights to obtain fusion features.

[0152] The etching depth analysis module inputs the fused features into a preset depth prediction network. The depth prediction network includes a physical branch that calculates the baseline etching rate based on process parameters and a data-driven branch that corrects the generation rate of the fused features. It outputs the first etching depth of multiple virtual regions on the wafer surface.

[0153] The etching compensation module analyzes the etching depth difference between the slowest and fastest etching regions based on the first etching depth of each virtual region and the density of the wafer surface region, and generates a space compensation trigger signal.

[0154] The etching endpoint control module, in response to the space compensation trigger signal, generates a process parameter compensation instruction that acts on the slowest region according to the preset mapping relationship between process parameters and etching rate, and issues an endpoint control instruction to control the etching endpoint when the etching depth of all virtual regions meets the preset endpoint conditions.

[0155] The above description is merely a preferred embodiment of this application. The scope of protection of this application is not limited to the above embodiments. All technical solutions falling within the scope of this application's concept are within the scope of protection of this application. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of this application should also be considered within the scope of protection of this application.

Claims

1. A method for online control of etching endpoint based on virtual measurement, characterized in that, include: The system acquires multimodal sensor signals collected in real time during the etching process, analyzes the signal quality parameters of each sensor, calculates the corresponding dynamic fusion weights based on the signal quality parameters, and performs weighted fusion of the multimodal sensor signals according to the dynamic fusion weights to obtain fusion features. The fusion features are input into a preset depth prediction network, which includes a physical branch that calculates the baseline etching rate based on process parameters and a data-driven branch that generates a correction amount based on the fusion features, and outputs the first etching depth of multiple virtual regions on the wafer surface. Based on the first etching depth of each virtual region and the density of the wafer surface region, the etching depth difference between the region with the slowest etching progress and the region with the fastest etching progress is analyzed, and a space compensation trigger signal is generated. In response to the space compensation trigger signal, a process parameter compensation command is generated for the slowest region based on the preset mapping relationship between process parameters and etching rate. When the etching depth of all virtual regions meets the preset endpoint condition, an endpoint control command is issued to control the etching endpoint.

2. The online control method for etching endpoint based on virtual measurement according to claim 1, characterized in that, The multimodal sensor signals include real-time intensity signals from the optical emission spectrometer at characteristic wavelengths, radio frequency reflected power signals, wafer self-bias signals, cavity pressure signals, and multi-zone coil current monitoring values. The analysis of signal quality parameters for each sensor includes: For each sensor signal, calculate the ratio of the maximum signal amplitude within a preset time window to the root mean square noise value of the sensor in the idle state of the device, and divide the ratio by a preset constant to obtain the signal-to-noise ratio index. Calculate the reciprocal of the signal variance within a preset time window, divide the reciprocal by a preset constant, and obtain the stability index. The percentage of each sampling point that deviates from the historical normal range within a preset time window is used as an anomaly ratio indicator. The signal-to-noise ratio, stability, and anomaly ratio are weighted and summed according to preset weighting coefficients to obtain the signal quality parameters of the corresponding sensor signals.

3. The online control method for etching endpoint based on virtual measurement according to claim 2, characterized in that, The physical branch calculates the baseline etching rate based on process parameters, including: Acquire the RF power value, acquire the outer coil current value and the inner coil current value and calculate their current ratio, acquire the wafer self-bias voltage value, cavity pressure value and gas flow rate value; The RF power value is raised to a first preset power, the current ratio is raised to a second preset power, the wafer self-bias value is substituted into the exponential decay function, the cavity pressure value and the gas flow rate value are raised to a third preset power and a fourth preset power respectively, and then multiplied to obtain the correction function value. The preset reference rate constant is multiplied with each calculation result to obtain the reference etching rate.

4. The online control method for etching endpoint based on virtual measurement according to claim 1, characterized in that, The data-driven branch generates a rate correction based on fused features, including: The fused features are input into a one-dimensional convolutional layer and a long short-term memory network layer. The one-dimensional convolutional layer uses multiple convolutional kernels to slide along the time dimension and extract local temporal feature sequences. The long short-term memory network layer receives the local temporal feature sequences and updates the corresponding hidden states to obtain the long-term dependencies throughout the etching process. The hidden states output by the Long Short-Term Memory (LSTM) network layer are mapped to rate correction values ​​corresponding to each virtual region through a fully connected layer.

5. The online control method for etching endpoint based on virtual measurement according to claim 1, characterized in that, The analysis of the etching depth difference between the slowest and fastest etching regions and the generation of a spatial compensation trigger signal includes: Obtain the first etching depth of each virtual region at the current moment, and calculate the difference between it and the target etching depth to obtain the remaining depth of each virtual region; The remaining depth is input into the pre-trained spatial coupling prediction model to obtain the predicted remaining depth of each virtual region at the next prediction time. Iterate through the predicted remaining depth of each virtual region, and take the region with the largest predicted remaining depth as the slowest region and the region with the smallest predicted remaining depth as the fastest region. Calculate the difference between the predicted remaining depth of the slowest region and the fastest region. If the difference is greater than a preset spatial compensation threshold, generate a spatial compensation trigger signal.

6. The online control method for etching endpoint based on virtual measurement according to claim 5, characterized in that, The spatially coupled prediction model comprises a cascaded structure of a multi-head self-attention network and a convolutional long short-term memory network. The multi-head self-attention network takes the feature vector composed of the current remaining depth of each virtual region and the pre-acquired graphic density values ​​of each virtual region as input, calculates the influence weights between each pair of virtual regions, and outputs an attention weight matrix. When the convolutional long short-term memory network updates its hidden state step by step, it multiplies the attention weight matrix with the convolutional kernel weights element by element to output the predicted remaining depth of each virtual region at the next prediction time.

7. The online control method for etching endpoint based on virtual measurement according to claim 5, characterized in that, The spatial compensation trigger signal includes the location marker of the slowest region and the corresponding local etching rate increment. The calculation process of the local etching rate increment includes: subtracting the predicted remaining depth of the fastest region from the predicted remaining depth of the slowest region to obtain the maximum remaining depth difference; Calculate the average predicted etching rate for all virtual regions, and divide the predicted remaining depth of the slowest region by the average predicted etching rate to obtain the remaining etching time; divide the maximum remaining depth difference by the remaining etching time to obtain the local etching rate increment of the slowest region.

8. The online control method for etching endpoint based on virtual measurement according to claim 1, characterized in that, The preset mapping relationship between process parameters and etching rate includes: performing multi-zone coil power variation analysis on the plasma etching equipment in advance, recording the change in etching rate of each virtual region under different power configurations, and constructing an influence coefficient matrix with virtual regions as rows and independently adjustable power control channels as columns.

9. The online control method for etching endpoint based on virtual measurement according to claim 8, characterized in that, The process parameter compensation instructions generated for the slowest region include: Based on the location identifier of the slowest region, extract the row vector of the sensitivity coefficients of the slowest region to each power control channel from the influence coefficient matrix; Using the adjustment amount of each power control channel as a variable, and combining the sum of the products of the adjustment amount of each channel and the corresponding sensitivity coefficient, as well as the preset total power change limit, the adjustment amount of each power control channel is obtained, and process parameter compensation instructions are generated.

10. An online control system for etching endpoint based on virtual measurement, characterized in that, The method for online control of etching endpoint based on virtual measurement as described in any one of claims 1 to 9 includes: The data analysis module acquires multimodal sensor signals collected in real time during the etching process, analyzes the signal quality parameters of each sensor, calculates the corresponding dynamic fusion weights based on the signal quality parameters, and performs weighted fusion of the multimodal sensor signals according to the dynamic fusion weights to obtain fusion features. The etching depth analysis module inputs the fused features into a preset depth prediction network. The depth prediction network includes a physical branch that calculates the baseline etching rate based on process parameters and a data-driven branch that corrects the generation rate of the fused features. It outputs the first etching depth of multiple virtual regions on the wafer surface. The etching compensation module analyzes the etching depth difference between the slowest and fastest etching regions based on the first etching depth of each virtual region and the density of the wafer surface region, and generates a space compensation trigger signal. The etching endpoint control module, in response to the space compensation trigger signal, generates a process parameter compensation instruction that acts on the slowest region according to the preset mapping relationship between process parameters and etching rate, and issues an endpoint control instruction to control the etching endpoint when the etching depth of all virtual regions meets the preset endpoint conditions.