Exposure parameter intelligent recommendation and adaptive feedback method of an exposure machine

CN122815792APending Publication Date: 2026-09-25梅州市鸿利线路板有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611125973.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-28
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]本申请提供一种曝光机的曝光参数智能推荐与自适应反馈方法,旨在解决上述背景技术中提到的感知延迟较大,模型检测和响应设备状态变化的时效性不足的问题

Benefits of technology

[0015]本方案通过构建“状态感知—任务识别—参数生成”三级协同架构,提出一种无需反向传播与标注反馈的在线自适应机制,显著提升了推荐系统在动态工业场景下的鲁棒性与敏捷响应能力。具体而言,通过同步采集曝光周期内多源非成像物理信号——如光源驱动电流频谱、调焦电机反电动势包络、掩模台位置残差序列及腔体温度梯度张量,并经统一归一化后输入轻量级图卷积编码器,有效融合异构时序信息并抑制噪声干扰,生成具备拓扑鲁棒性的128维工艺状态指纹,实现了对设备隐性退化过程的可观测化表征;该指纹不仅保留了关键状态演变特征,还具备跨批次、跨时段的可比性,为后续动态适应提供了可靠依据。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122815792A_ABST
    Figure CN122815792A_ABST
Patent Text Reader

Abstract

The present application relates to a kind of exposure parameter intelligent recommendation and adaptive feedback method of exposure machine, its core scheme is: through the real-time acquisition and fusion of multiple-source sensor and the multiple-dimensional weak physical signal such as exposure light source, high-speed current, focusing motor, mask table and temperature field, using normalization and isomorphic data synchronous processing, constructs the topological graph structure reflecting process state, extracts process state fingerprint by lightweight graph convolution network, then similar drift task set is obtained by comparing with historical meta-task library, drives graph neural network model to generate lightweight adaptive parameter for recommendation model output layer, realizes local re-calibration based on current equipment state and adaptive exposure parameter recommendation.This method improves the robustness and adaptability to process variation, effectively enhances the intelligent control level of exposure process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent parameter recommendation technology for semiconductor manufacturing equipment, and in particular to an intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine. Background Technology

[0002] In the field of semiconductor manufacturing equipment, especially in the automatic parameter recommendation technology of exposure machines, a new stage of intelligent and data-driven development has emerged. The industry widely adopts parameter recommendation models centered on deep learning. By analyzing historical exposure data, process formulations, and finished product quality feedback, a parameter prediction network is trained to achieve automatic recommendation and correction of process parameters. Some high-end equipment is beginning to integrate environmental perception and equipment health diagnosis modules to dynamically capture key signals during equipment operation for model optimization. Furthermore, combining adaptive learning and online fine-tuning technologies, some systems are attempting to improve the recommendation model's ability to respond to changes in production line conditions by continuously collecting new operating data and employing online incremental training.

[0003] Currently, there are several typical technical solutions for this type of intelligent parameter recommendation system: One mainstream solution is based on historical data labels, constructs a deep parameter prediction network through supervised learning, and retrains the model periodically based on production line feedback data; another solution introduces online reinforcement learning or labeled self-feedback loops, and fine-tunes the model weights in real time according to process deviations to achieve adaptive parameter adjustment within a limited range.

[0004] However, in actual production, exposure equipment suffers from significant long-term process drift and gradual degradation of hardware performance. These factors include thermal deformation of optical components, micro-wear of mechanical transmission parts, and slow fluctuations in temperature and humidity within cleanrooms. These factors often manifest as weak, heterogeneous physical signals at multiple sensing nodes within the equipment, and these signals do not directly map to process labels or significant parameter drift. Existing methods typically rely heavily on historical process labels, requiring periodic collection of standardized finished product quality criteria and extensive retraining or fine-tuning of the model's backbone structure or weight parameters. This leads to two major problems: first, significant sensing latency, resulting in insufficient timeliness in model detection and response to changes in equipment status, often requiring reactive adjustments only after significant production line anomalies (such as decreased yield); second, high training and maintenance costs, necessitating manual intervention or model reconstruction whenever equipment status transitions or new types of drift occur, significantly increasing operational burden and system uncertainty. Summary of the Invention

[0005] This application provides an intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine, aiming to solve the problems mentioned in the background art, such as large perception delay and insufficient timeliness of model detection and response to changes in device state.

[0006] This application provides a method for intelligent recommendation and adaptive feedback of exposure parameters for an exposure machine, specifically including:

[0007] S1: Acquire multi-source weak state signals during the exposure execution cycle. The multi-source weak state signals include at least three of the following: high-frequency current spectrum characteristics of the exposure light source driving circuit, back electromotive force fluctuation envelope of the objective lens focusing motor, position feedback residual sequence of the mask stage motion trajectory, and spatial gradient tensor of the distributed temperature sensor in the cavity.

[0008] S2: The multi-source weak state signals are uniformly sampled and normalized to generate a set of normalized heterogeneous time-series signals;

[0009] S3: Input the normalized heterogeneous time-series signal set into a lightweight graph convolutional encoder for topology mapping processing to extract low-dimensional embedding vectors as process state fingerprints.

[0010] S4: Perform cosine similarity calculation based on the process state fingerprint and the prototype vector in the offline meta-task library to select the Top-3 nearest neighbor tasks and form a dynamic task set.

[0011] S5: Input the process state fingerprint and the dynamic task set into the graph neural network model, and output lightweight adapter parameters that only affect the output layer attention head and the feedforward network bias term;

[0012] S6: Process the lightweight adapter parameters to obtain the local remapping inference path, perform local coordinate system recalibration based on the local remapping inference path and generate recommended values ​​for adaptive exposure parameters;

[0013] S7: Based on the recommended values ​​of the adaptive exposure parameters, control the exposure machine to perform exposure operations, and repeatedly collect new multi-source weak state signals in the next execution cycle to trigger real-time updates of the process state fingerprint.

[0014] The intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine provided in this application has the following beneficial effects:

[0015] This solution constructs a three-level collaborative architecture of "state awareness—task recognition—parameter generation" and proposes an online adaptive mechanism that does not require backpropagation or annotation feedback, significantly improving the robustness and agile response capability of the recommendation system in dynamic industrial scenarios. Specifically, by synchronously collecting multi-source non-imaging physical signals during the exposure cycle—such as the spectrum of the light source driving current, the envelope of the back electromotive force of the focusing motor, the mask stage position residual sequence, and the cavity temperature gradient tensor—and inputting them into a lightweight graph convolutional encoder after unified normalization, heterogeneous temporal information is effectively fused and noise interference is suppressed to generate a 128-dimensional process state fingerprint with topological robustness, realizing the observable characterization of the equipment's implicit degradation process. This fingerprint not only retains the key state evolution characteristics but also has comparability across batches and time periods, providing a reliable basis for subsequent dynamic adaptation. Attached Figure Description

[0016] Figure 1 This is the main flowchart of an intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine.

[0017] Figure 2 This is a sub-flowchart of an intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine.

[0018] Figure 3 This is another sub-flowchart of an intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine. Detailed Implementation

[0019] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0020] The following disclosure provides many different embodiments or examples for implementing different structures of the invention. To simplify the disclosure, specific examples of components and arrangements are described below. Of course, these are merely examples and are not intended to limit the invention. Furthermore, reference numerals and / or letters may be repeated in different examples; such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.

[0021] like Figure 1 As shown, this application provides a method for intelligent recommendation and adaptive feedback of exposure parameters for an exposure machine, specifically including:

[0022] S1: Acquire multi-source weak state signals during the exposure execution cycle. The multi-source weak state signals include at least three of the following: high-frequency current spectrum characteristics of the exposure light source driving circuit, back electromotive force fluctuation envelope of the objective lens focusing motor, position feedback residual sequence of the mask stage motion trajectory, and spatial gradient tensor of the distributed temperature sensor in the cavity.

[0023] S2: The multi-source weak state signals are uniformly sampled and normalized to generate a set of normalized heterogeneous time-series signals;

[0024] S3: Input the normalized heterogeneous time-series signal set into a lightweight graph convolutional encoder for topology mapping processing to extract low-dimensional embedding vectors as process state fingerprints.

[0025] S4: Perform cosine similarity calculation based on the process state fingerprint and the prototype vector in the offline meta-task library to select the Top-3 nearest neighbor tasks and form a dynamic task set.

[0026] S5: Input the process state fingerprint and the dynamic task set into the graph neural network model, and output lightweight adapter parameters that only affect the output layer attention head and the feedforward network bias term;

[0027] S6: Process the lightweight adapter parameters to obtain the local remapping inference path, perform local coordinate system recalibration based on the local remapping inference path and generate recommended values ​​for adaptive exposure parameters;

[0028] S7: Based on the recommended values ​​of the adaptive exposure parameters, control the exposure machine to perform exposure operations, and repeatedly collect new multi-source weak state signals in the next execution cycle to trigger real-time updates of the process state fingerprint.

[0029] Step S1: Acquire multi-source weak state signals during the exposure execution cycle. These multi-source weak state signals include at least three of the following: the high-frequency current spectrum characteristics of the exposure light source driving circuit, the back electromotive force fluctuation envelope of the objective lens focusing motor, the position feedback residual sequence of the mask stage motion trajectory, and the spatial gradient tensor of the distributed temperature sensors within the cavity. Specifically, they include:

[0030] S1.1: The real-time output current of the exposure light source driving circuit is sampled at high speed and processed by fast Fourier transform to extract high-frequency current spectrum features that characterize the stability of the light source from the time-domain current waveform, and the high-frequency current spectrum features are output as the first type of original state signal.

[0031] Sub-step S1.1 aims to extract high-frequency spectral features characterizing the stability of the light source from the real-time current signal of the exposure light source driving circuit, serving as the first type of raw input for constructing the process state fingerprint. This step directly operates on the output of the exposure machine's light source driving module, acquiring the time-domain current waveform through a high-speed data acquisition card.

[0032] A high sampling rate analog-to-digital converter is used to synchronously sample the real-time output current of the light source driving circuit. The sampling frequency is set to 200kHz to ensure that microsecond-level current transient fluctuations can be captured, generating an original time-domain current sequence containing a mixture of noise and valid signals.

[0033] The original time-domain current sequence is detrended by using a moving average filter to remove the DC component and low-frequency power frequency interference, while retaining the AC component that reflects the high-frequency jitter characteristics of the light source, thus generating a preprocessed current signal with zero mean.

[0034] The preprocessed current signal is windowed using the Hanning window function to suppress spectral leakage, and the truncated signal segment is converted into a finite-length sequence required for frequency domain analysis, generating a windowed time-domain data frame.

[0035] Perform a Fast Fourier Transform (FFT) operation on the windowed time-domain data frame to map the time-domain signal to the frequency domain space, calculate the complex amplitude of each frequency component, and generate an initial spectral distribution containing amplitude and phase information.

[0036] The following formula is used to calculate the one-sided power spectral density to quantify the energy distribution at each frequency point:

[0037]

[0038] in, For frequency Power spectral density at that point The complex spectrum after FFT transformation. The number of sampling points. The sampling frequency.

[0039] The calculated power spectral density is subjected to frequency band integration processing. The high-frequency range of 5kHz to 50kHz is selected as the feature extraction window. The spectral line energies in this range are accumulated to generate a scalar index characterizing the total energy of high-frequency noise.

[0040] The top five significant peak frequencies and their corresponding amplitudes are extracted from the power spectral density curve to construct a sparse spectral feature vector, which reflects the drift of a specific resonant point in the light source driving circuit.

[0041] Through the above-mentioned high-speed sampling, frequency domain transformation and feature extraction processing methods, the unstructured time-domain current waveform is transformed into a high-frequency current spectrum feature vector with physical meaning, realizing the digital representation of the microscopic stability state of the light source, and providing a highly discriminative first-class original state signal for subsequent multi-source signal fusion.

[0042] S1.2: Based on the acquisition timing of the first type of original state signal, perform differential operation on the terminal voltage and speed data of the objective lens focusing motor in the closed-loop control process to calculate the back EMF fluctuation envelope reflecting the change of mechanical load, and output the back EMF fluctuation envelope as the second type of original state signal.

[0043] The real-time terminal voltage sampling sequence and angular velocity feedback data of the objective lens focusing motor during the closed-loop control cycle are acquired as the raw input signal reflecting the dynamic characteristics of the mechanical load.

[0044] The sampling sequence of the opposite end voltage is subjected to low-pass filtering to remove high-frequency switching noise interference and retain the low-frequency fundamental component that characterizes the electromagnetic state of the motor.

[0045] Numerical differentiation is performed on the angular velocity feedback data to calculate the rate of change of angular velocity at adjacent sampling times, generating an angular acceleration time series.

[0046] Based on the physical model of a DC motor, a back electromotive force calculation equation is constructed, incorporating terminal voltage, armature current, and resistance voltage drop into a unified calculation framework.

[0047] Calculate the instantaneous back electromotive force value, input the calculated instantaneous back electromotive force sequence into the sliding window envelope extractor, and set the window length to half a mechanical vibration cycle.

[0048] Search for local maxima and minima within a sliding window, and construct upper and lower envelopes.

[0049] Calculate the mean trajectory of the upper and lower envelopes, eliminate instantaneous fluctuations caused by random noise, and generate a smooth back EMF fluctuation envelope.

[0050] The fluctuation envelope is subjected to trend term separation processing to remove the DC component caused by the linear change of rotational speed, and retain the AC fluctuation component that reflects the abnormal mechanical resistance.

[0051] By using the above processing method, the time-series electrical signal from the previous step is converted into a back electromotive force fluctuation envelope that characterizes the change in mechanical transmission load, thereby achieving sensitive capture of the micro-wear state of the objective lens focusing mechanism and providing highly discriminative mechanical feature basis for the subsequent construction of process state fingerprint.

[0052] S1.3: Using the synchronous triggering time of the second type of original state signal, the position encoder feedback value of the mask stage motion control system is differentially compared with the theoretical command trajectory to calculate the position feedback residual sequence characterizing the mechanical transmission accuracy, and the position feedback residual sequence is output as the third type of original state signal.

[0053] The system receives a global clock synchronization signal from the main control unit of the exposure machine, locks the precise timestamp corresponding to the peak point of the back electromotive force fluctuation envelope in the second type of original state signal, and establishes the starting reference point for mask stage motion data acquisition.

[0054] The actual position feedback data sequence acquired by the high-resolution grating ruler or laser interferometer is read in real time from the closed-loop control register of the mask stage linear motor driver. This sequence reflects the actual physical displacement trajectory of the mask stage during the scanning exposure process.

[0055] Simultaneously acquire the theoretical motion command trajectory data preset in the exposure process formula. This data includes the expected position coordinates of the mask stage at each time point calculated based on an ideal rigid body model, which serves as a standard reference system for measuring motion accuracy.

[0056] Perform time axis alignment on the actual position feedback data sequence and the theoretical command trajectory data to ensure that the two are comparable at the same time sampling point and eliminate phase deviation caused by data transmission delay.

[0057] By employing the point-by-point difference method, the deviation between the actual position coordinates and the theoretical position coordinates at the same moment is calculated, and an initial position error sequence is constructed. This sequence intuitively reflects the instantaneous positioning error in the mechanical transmission system.

[0058] A sliding window filtering mechanism is introduced to denoise the initial position error sequence, eliminating non-systematic jitter caused by high-frequency noise from the sensor, and retaining the effective residual components that reflect low-frequency drift characteristics such as mechanical wear and guide rail deformation.

[0059] The filtered residual sequence and the calculated root mean square error index are encapsulated in a structured manner to generate a position feedback residual sequence containing time-domain waveform features and statistical features.

[0060] By using differential comparison and filtering, the synchronous trigger signal from the previous step is transformed into position feedback residual sequence data that characterizes the accuracy of mechanical transmission, thus achieving precise quantification and feature extraction of the microscopic motion drift of the mask stage.

[0061] For example, in a 28nm node chip exposure scenario using a stepper lithography machine, the mask stage scanning speed is set to 500mm / s, and the acceleration is 2g. The system synchronously acquires the actual position data fed back by the laser interferometer and the theoretical sinusoidal acceleration / deceleration command trajectory issued by the main controller at a frequency of 10kHz. Within a 200ms scanning window, a total of 2000 data points are acquired. Calculations show that the maximum instantaneous deviation between the actual and theoretical trajectories is 15nm, and the root mean square error (RMS) is 4.2nm. Spectral analysis of the residual sequence reveals a periodic fluctuation component with a frequency of 120Hz, corresponding to mechanical resonance caused by a tiny scratch on the guide rail. The system outputs the residual sequence containing this 120Hz characteristic peak as a third type of raw state signal for subsequent process state fingerprinting, thereby accurately identifying the early wear state of the mask stage guide rail and significantly improving the predictive ability for exposure overlay errors caused by mechanical aging.

[0062] S1.4: Based on the timestamp index of the third type of original state signal, tensor reconstruction processing is performed on the readings of the multi-point temperature sensors that are spatially distributed in the cavity to generate a spatial gradient tensor characterizing the thermal field non-uniformity, and the spatial gradient tensor is output as the fourth type of original state signal.

[0063] The real-time reading set of multiple temperature sensors spatially distributed within the cavity at the timestamp index of the third type of original state signal is obtained. The reading set includes the instantaneous temperature values ​​of multiple discrete temperature monitoring points located near the exposure light source, around the objective lens assembly, and on the mask stage guide side.

[0064] Spatial coordinate mapping processing is performed on the real-time reading set. Based on the physical installation position of the sensor in the exposure machine cavity, the discrete temperature values ​​are mapped to the corresponding nodes in the three-dimensional Cartesian coordinate system to construct sparse grid data of the temperature field with clear geometric topological relationships.

[0065] Based on the sparse grid data of the temperature field, a linear interpolation method is used to fill the blank areas between adjacent sensor nodes to generate a high-resolution continuous temperature distribution matrix covering the entire exposure area, thereby eliminating the spatial information loss caused by the sparse sensor layout.

[0066] The finite difference method is used to perform partial derivative calculations on the continuous temperature distribution matrix along the X-axis, Y-axis and Z-axis to generate a spatial gradient tensor characterizing the intensity and direction of thermal field inhomogeneity.

[0067] The spatial gradient tensor is normalized to eliminate the influence of absolute temperature level on the gradient amplitude, highlight the relative change trend of local thermal deformation, and form a standardized fourth type of original state signal.

[0068] Through the tensor reconstruction and gradient calculation processing methods described above, the timestamp index of the previous step is transformed into spatial gradient tensor data that characterizes the degree of thermal field distortion in the cavity. This enables the quantitative extraction of process drift characteristics caused by thermal effects, providing a key thermodynamic dimension input for the subsequent construction of process state fingerprints.

[0069] S1.5: Integrate the first type of original state signal, the second type of original state signal, the third type of original state signal, and the fourth type of original state signal, and perform multi-channel time alignment and data packet encapsulation processing to generate a set of multi-source weak state signals containing complete exposure cycle physical state information as the final output.

[0070] Step S2: The multi-source weak state signals are uniformly sampled and normalized to generate a normalized heterogeneous time-series signal set. Specifically, this includes:

[0071] S2.1: Perform timestamp alignment processing on the high-frequency current spectrum characteristics, back EMF fluctuation envelope, position feedback residual sequence, and spatial gradient tensor. Resample each asynchronous signal to a unified target sampling rate using a global clock reference based on the exposure execution cycle, thereby generating a time-synchronized multi-source state signal sequence.

[0072] The system receives a set of weak state signals from multiple sources output by S1.5, and analyzes the high-frequency current spectrum characteristics, back EMF fluctuation envelope, position feedback residual sequence, and original timestamp index of the spatial gradient tensor contained therein. It extracts the global clock reference signal for the exposure execution cycle, establishes a unified time reference coordinate system, and eliminates asynchronous deviations caused by hardware trigger delays in various sensors.

[0073] For high-sampling-rate signals such as high-frequency current spectrum characteristics and back EMF fluctuation envelope, a linear interpolation method is used for resampling processing, mapping data points to a preset unified target sampling rate grid to ensure accurate alignment of high-frequency details on the time axis.

[0074] For low-sampling-rate signals such as spatial gradient tensors, zero-order hold or spline interpolation methods are applied to fill in the missing time points, so that their time resolution is consistent with that of high-frequency signals, and a multi-channel data matrix aligned to the time dimension is constructed.

[0075] Cross-correlation analysis is used to calculate the phase lag of each channel signal relative to the global clock reference. Transmission delay is compensated by time-domain shift operation to achieve strict synchronization of multi-source signals with microsecond-level precision.

[0076] Outlier detection is performed on the time-aligned signal sequence to remove invalid data points caused by communication packet loss, and a nearby valid value filling strategy is used to repair data breakpoints, generating a continuous and complete time-synchronized multi-source state signal sequence.

[0077] By using the timestamp alignment and resampling methods described above, the heterogeneous and asynchronous multi-source weak state signals in the previous step are transformed into a time-synchronized multi-source state signal sequence with a unified time base and consistent sampling rate. This achieves the expected technical effect of eliminating timing misalignment interference and providing high-quality data input for subsequent sliding window truncation.

[0078] For example, the global clock reference frequency is set to 10kHz, and the target uniform sampling rate is 1kHz. The original sampling rate of the high-frequency current spectrum feature is 50kHz, which is aligned to a 1kHz grid through 50x downsampling and linear interpolation; the original sampling rate of the spatial gradient tensor is 100Hz, which is padded to 1kHz through 10x upsampling and cubic spline interpolation. A 2ms transmission delay is detected in the mask stage position signal, which is compensated for by shifting the time domain left by 20 data points. The final output is a synchronization signal matrix with a length of 1000 points and 4 channels, with the time deviation of each channel controlled within ±0.1ms, significantly improving the timing consistency of subsequent feature extraction.

[0079] S2.2: Perform sliding window truncation processing on the time-synchronized multi-source state signal sequence to extract continuous and equal-length data segments based on the preset signal analysis frame length, thereby generating a fixed-length heterogeneous signal segment set.

[0080] It receives a time-synchronized multi-source state signal sequence that has been timestamped and aligned. This sequence contains four heterogeneous data channels: high-frequency current spectrum, back EMF fluctuation envelope, position feedback residual, and spatial gradient tensor.

[0081] Based on the preset signal analysis frame length parameters, a sliding window truncation mechanism is constructed, and the window length is set to correspond to the duration of the complete physical process of the exposure execution cycle, ensuring that the truncated data segments cover the full dynamic characteristics of the device from startup acceleration to stable exposure.

[0082] An overlapping sliding strategy is used to slice the continuous signal stream, with the step size set to half the window length, in order to preserve the temporal continuity information between adjacent segments and prevent the loss of transient features due to hard truncation, thereby generating a preliminary set of candidate signal segments.

[0083] Boundary integrity checks are performed on each candidate signal segment to eliminate abnormal segments with data loss or incomplete sampling at the start or end time, ensuring that the data input to the subsequent normalization module has complete physical meaning and statistical representativeness.

[0084] The verified signal segments are structurally reorganized according to the channel order of the light source, motor, mask stage and temperature sensor to form a set of fixed-length heterogeneous signal segments with uniform dimensions and fixed length.

[0085] By using a sliding window truncation method, the time-synchronized multi-source state signal sequence of the previous step is transformed into a set of fixed-length heterogeneous signal segments, which realizes the discretization of the continuous process drift process, provides a unified dimension of input basis for subsequent range standardization, and effectively eliminates the calculation deviation caused by the different signal lengths.

[0086] S2.3: Perform range normalization calculation on the set of fixed-length heterogeneous signal segments to map the original physical quantities to a dimensionless interval of zero to one using the maximum and minimum values ​​in each signal channel, thereby generating a dimensionless signal matrix.

[0087] The system receives a set of fixed-length heterogeneous signal segments generated in the previous steps. This set contains equal-length data sequences from four channels: high-frequency current spectrum, back electromotive force fluctuation envelope, position feedback residual, and spatial gradient tensor. It iterates through each independent channel in the signal set, extracting the global maximum and global minimum values ​​within the data segment of the currently processed channel to construct the dynamic numerical boundary interval for that channel. Based on the extracted maximum and minimum values, it calculates the numerical range of the current channel (the difference between the maximum and minimum values), which serves as the denominator reference parameter for subsequent linear mapping. For each original sampling point within the current channel, it performs a subtraction operation, subtracting the global minimum value of that channel from the original physical quantity to achieve zero-point alignment and eliminate the absolute offset effects caused by installation references or zero-point drift from different sensors. Finally, it divides the aligned data by the pre-calculated numerical range and performs a linear scaling transformation to compress the data distribution to a preset dimensionless interval.

[0088] Repeat the above processing flow until all channels' data are mapped. Reassemble the processed data from each channel according to their original topological order to generate a dimensionless signal matrix. Through range normalization, the fixed-length heterogeneous signal segments from the previous step are transformed into a dimensionless signal matrix with values ​​strictly limited to zero and one. This eliminates dimensional differences and order-of-magnitude disparities between multi-source signals, achieving the expected technical effect of unified weighted processing of heterogeneous features by the subsequent graph convolutional encoder.

[0089] For example, for the high-frequency current spectrum channel, the sliding window length is set to 1024 points. The maximum value of the current harmonic amplitude within this window is 5.2A, the minimum value is 4.8A, and the range is 0.4A. For a sampled value of 5.0A at a certain moment within the window, substituting it into the formula, we get (5.0-4.8) / 0.4=0.5. For the position feedback residual channel, the maximum residual is 20nm, the minimum residual is -15nm, and the range is 35nm. For a residual value of 5nm at a certain moment, we get (5-(-15)) / 35≈0.571. After processing, the values ​​of all channel data converge to the [0,1] interval, eliminating the dimensional gap between the current ampere level and the displacement nanometer level, and significantly improving the numerical stability and convergence speed of subsequent feature extraction.

[0090] S2.4: Perform mean-variance normalization transformation on the dimensionless signal matrix to adjust the data distribution to a standard normal distribution with zero mean and unit variance by subtracting the channel mean and dividing by the channel standard deviation, thereby generating a standardized heterogeneous time-series signal vector.

[0091] For example, a dimensionless signal matrix with 4 channels and a length of 1024 points is processed. The mean of the light source current channel is 0.52, and the standard deviation is 0.15; the mean of the motor back EMF channel is 0.48, and the standard deviation is 0.12. A transformation is performed on the first point (0.67) of the light source current channel: (0.67-0.52) / 0.15=1.0. A transformation is performed on the first point (0.36) of the motor channel: (0.36-0.48) / 0.12=-1.0. After this processing, the mean of all channel data is close to 0, and the standard deviation is close to 1, significantly improving the sensitivity of the subsequent model to subtle process drift characteristics.

[0092] S2.5: Perform multidimensional tensor stacking encapsulation processing on the standardized heterogeneous timing signal vector to integrate all channel data according to the preset topological order of the light source, motor, mask stage and temperature sensor, thereby generating a normalized heterogeneous timing signal set.

[0093] like Figure 2 As shown, step S3 involves inputting the normalized heterogeneous time-series signal set into a lightweight graph convolutional encoder for topology mapping processing to extract low-dimensional embedding vectors as process state fingerprints. Specifically, this includes:

[0094] S3.1: Based on the physical distribution relationship of sensors in the normalized heterogeneous time-series signal set, construct a multi-source signal topology graph structure including light source driving nodes, focusing motor component nodes, mask stage motion nodes and temperature sensing nodes, so as to map discrete time-series data into graph structure data with clear connection relationships and generate a multi-source signal topology graph structure.

[0095] The system receives a normalized heterogeneous time-series signal set generated in step S2. This set contains standardized data fragments from four dimensions: the light source, the motor, the mask stage, and the temperature sensor. The system analyzes the actual installation coordinates and electrical connections of each sensor in the physical space of the exposure machine, defining the light source drive node as the root node and the focusing motor assembly node, the mask stage motion node, and the temperature sensor node as leaf nodes. An undirected graph topology is constructed, setting the edge weights between the light source and temperature nodes based on the optical path thermal coupling effect, and setting the edge weights between the motor and mask stage nodes based on the mechanical transmission chain, generating an initial adjacency matrix. The normalized four-dimensional time-series signals are mapped to the initial feature vectors of the corresponding graph nodes, ensuring that the feature dimension of each node matches the number of signal channels. A sparse matrix storage format is used to optimize the adjacency matrix, reducing the computational complexity of subsequent graph convolution operations, forming a complete multi-source signal topology graph structure containing node feature matrices and edge connections. By using the above-mentioned topological mapping processing method, discrete and isolated time-series signals are transformed into graph structure data with clear physical relationships, realizing a structured representation of the coupling relationship of multiple physical fields inside the equipment, and laying a data foundation for the subsequent extraction of process state fingerprints with topological robustness.

[0096] S3.2: The graph attention mechanism is used to perform weighted aggregation processing on the features of adjacent nodes in the multi-source signal topology graph structure to eliminate dimensional interference between heterogeneous signals and enhance the propagation weight of key fault features, thereby generating a local spatial enhancement feature map.

[0097] The multi-source signal topology structure generated in the previous steps is received. This structure includes four physical nodes: light source driver, focusing motor, mask stage, and temperature sensor, as well as their adjacency matrix.

[0098] A linear transformation mapping is performed on the feature vector of each node in the topology graph. The original heterogeneous signals are projected onto a unified hidden layer feature space through a learnable weight matrix, eliminating the interference of differences in the dimensions of different sensors on subsequent attention calculations.

[0099] Based on the hidden layer features, the attention coefficient between any two connected nodes is calculated. The LeakyReLU activation function is used to perform nonlinear processing on the spliced ​​node features to quantify the information contribution of the source node to the target node.

[0100] The calculated attention coefficients are normalized using Softmax to ensure that the sum of the attention weights of all neighboring nodes of each node is 1, thereby generating a standardized attention distribution matrix that highlights the propagation path of key fault characteristics.

[0101] By using normalized attention weights to perform weighted summation and aggregation on the features of neighboring nodes, the semantic information of neighboring nodes is fused into the central node, generating an intermediate node representation that includes local topological context information.

[0102] Nonlinear activation processing is performed on the aggregated intermediate node representations. The ELU activation function is introduced to alleviate the gradient vanishing problem, enhance the model's sensitivity to weak process drift signals, and generate local spatial enhancement feature maps.

[0103] By using the weighted aggregation processing of the graph attention mechanism, the multi-source signal topology graph structure of the previous step is transformed into a local spatial enhancement feature graph with semantic consistency, thereby realizing dimensional decoupling between heterogeneous signals and adaptive enhancement of key fault features.

[0104] For example, the hidden layer dimension is set to 64 dimensions, and the number of attention heads is 4. For the connection between the light source node and the temperature node, the initial attention coefficient is calculated to be 0.75. After Softmax normalization, the weight is adjusted to 0.82, which is significantly higher than other non-critical connections. After aggregation, 82% of the temperature field gradient information is incorporated into the light source node features, effectively suppressing misjudgments of light source current harmonics caused by thermal deformation, significantly improving the signal-to-noise ratio of the feature map, and providing highly discriminative input for subsequent temporal dependency capture.

[0105] S3.3: Based on the local spatial enhancement feature map, a temporal convolutional network is used to perform multi-layer causal convolution operations on the node feature sequence to capture the long-term and short-term temporal dependencies in the process drift and extract the dynamic evolution pattern, generating a spatiotemporal fusion feature tensor.

[0106] Receive the local spatial enhancement feature map generated in the previous step, which contains spatiotemporal feature data of multi-source sensor nodes aggregated by the graph attention mechanism.

[0107] For the feature sequences of each node in the local spatial enhancement feature map, a one-dimensional causal convolution channel is constructed according to the time dimension to ensure that the feature extraction at the current moment depends only on the history and the current input, and to avoid prediction deviation caused by future information leakage.

[0108] A multi-layer dilated causal convolutional structure is used to perform deep temporal modeling of node feature sequences. The kernel size of the first layer is set to 3, and the dilation rate is set to 1 to capture short-term fluctuation details between adjacent time steps.

[0109] The second convolutional kernel size remains at 3, and the expansion rate is adjusted to 2, thereby expanding the receptive field to capture the evolution trend of process states over a medium time span.

[0110] The kernel size of the third convolutional layer remains at 3, but the dilation rate is further increased to 4, aiming to cover the long-cycle drift pattern throughout the entire exposure execution cycle and form a multi-scale temporal feature fusion.

[0111] After each causal convolution operation, a gated linear unit activation function is applied, and the effective temporal features are selected and noise interference is suppressed through the sigmoid gating mechanism.

[0112] By using residual connection structures to add input features to convolutional output features element-wise, the gradient vanishing problem in deep network training is alleviated, and the integrity of key information in the original signal is maintained.

[0113] The channel dimension concatenation operation is performed on the feature tensor after multi-layer causal convolution to integrate the multi-scale temporal features extracted under different dilation rates into a unified high-dimensional spatiotemporal representation.

[0114] Through the above processing method, the local spatial enhancement feature map of the previous step is transformed into a spatiotemporal fusion feature tensor containing long and short-term temporal dependencies, thereby realizing the accurate extraction of the process drift dynamic evolution mode.

[0115] S3.4: Perform global average pooling and nonlinear dimensionality reduction mapping on the spatiotemporal fusion feature tensor to compress redundant information and project high-dimensional features to a preset low-dimensional latent space to generate an initial low-dimensional embedding vector.

[0116] Receive the spatiotemporal fusion feature tensor generated in the previous step. This tensor contains spatiotemporal dependency information of multi-source signals that has been aggregated by graph attention mechanism and extracted by temporal convolutional network, and serves as the input data object for this step.

[0117] A global average pooling operation is performed on the spatiotemporal fusion feature tensor to calculate the arithmetic mean along the time dimension and the spatial node dimension, compressing the high-dimensional dynamic features into a static feature vector of fixed length, thus eliminating the dimension mismatch problem caused by the difference in temporal length.

[0118] A fully connected dimensionality reduction mapping layer is constructed, with the output dimension set to 128 dimensions to match the spatial dimension requirements of the prototype vectors in the subsequent meta-task library, ensuring the consistency of feature embedding.

[0119] The pooled static feature vector is projected using a linear transformation matrix, and the initial low-dimensional embedding vector is calculated using the following formula:

[0120]

[0121] in, The initial low-dimensional embedding vector, For non-linear activation functions such as ReLU, The weight matrix is ​​a learnable, dimensionality-reduced matrix. The feature vector after global average pooling. This is a bias term.

[0122] The L2 norm normalization process is performed on the projected feature vectors to map the vectors to the surface of a unit hypersphere, thereby eliminating the interference of feature magnitude on cosine similarity calculation and enhancing the discriminative power of feature direction.

[0123] By using global average pooling and nonlinear dimensionality reduction mapping, the high-dimensional spatiotemporal fusion feature tensor from the previous step is transformed into a topologically robust 128-dimensional initial low-dimensional embedding vector, achieving efficient compression from high-dimensional redundant data to low-dimensional semantic fingerprints, and providing a standardized feature basis for subsequent comparative learning optimization.

[0124] S3.5: Based on the contrastive learning loss function, perform cluster center constraint optimization on the initial low-dimensional embedding vector to bring the feature distance of similar process drift states closer and push away the feature distance of dissimilar states, and finally output a process state fingerprint with topological robustness.

[0125] Receive the initial low-dimensional embedding vectors output from S3.4 and construct a cluster center constraint optimization mechanism based on contrastive learning. Define positive sample pairs as fingerprints of different time windows under the same process drift mode, and negative sample pairs as fingerprint combinations between different drift modes. Calculate the Euclidean distance between positive sample pairs in the feature space as a measure of intra-class compactness. Calculate the cosine similarity between negative sample pairs in the feature space as a measure of inter-class separation. Construct a contrastive loss function aimed at minimizing the positive sample distance and maximizing the negative sample distance. Employ a variant of the InfoNCE loss function. The optimization objective is as follows:

[0126]

[0127] in, and For positive sample pairs, For negative samples, The temperature coefficient is used. The gradient of the loss function with respect to the parameters mapped to the embedded vector is calculated using backpropagation. The weights of the last layer projection matrix of the graph convolutional encoder are updated using stochastic gradient descent. Forward propagation and parameter updates are iteratively performed until the loss function converges to a preset threshold. L2 normalization is performed on the converged embedded vectors to eliminate the interference of vector magnitude differences on subsequent similarity calculations. A topologically robust process state fingerprint is generated, ensuring that similar drift states are tightly clustered in the vector space and dissimilar states are significantly separated. Cluster center constraint optimization is performed on the initial low-dimensional embedded vectors using a contrastive learning loss function, transforming the result of the previous step into highly discriminative process state fingerprint data, achieving the expected technical effect of sensitive perception and accurate differentiation of minor process drifts in the recommendation model.

[0128] like Figure 3As shown, step S4 involves performing cosine similarity calculation based on the process state fingerprint and the prototype vector in the offline meta-task library to select the Top-3 nearest neighbor tasks and form a dynamic task set. Specifically, this includes:

[0129] S4.1: Obtain the process status fingerprint generated in the previous steps and the pre-built offline meta-task library. The offline meta-task library contains multiple prototype vectors covering typical aging stages and common drift modes. Perform vector space alignment processing on the process status fingerprint and the multiple prototype vectors to eliminate feature distribution offset and generate an aligned set of process status fingerprints to be matched and standardized prototype vectors.

[0130] Receive a topologically robust 128-dimensional process state fingerprint vector output from the preceding steps. This vector represents the overall physical degradation state of the current exposure equipment within a specific execution cycle.

[0131] The pre-built offline meta-task library is invoked. This library stores N standardized prototype vectors covering typical operating conditions such as light source aging, mechanical wear, and thermal drift. Each prototype vector corresponds to a specific process drift mode and its optimal adaptation parameter context.

[0132] The process state fingerprint and the prototype vector in the meta-task library are subjected to L2 norm normalization to map all vectors to the unit hypersphere space, so as to eliminate the feature distribution offset caused by signal amplitude fluctuations and ensure the geometric consistency of subsequent similarity calculation.

[0133] The normalized fingerprint vector is decorated using a whitening transformation matrix. Based on the feature covariance inverse matrix statistically obtained during the offline training phase, the linear correlation between dimensions is eliminated, and a decorated process state fingerprint to be matched is generated.

[0134] The same whitening transformation matrix is ​​applied to the prototype vectors in the meta-task library to ensure that the query vector and the vectors in the library are in the same orthogonal feature subspace, thus generating a set of standardized prototype vectors.

[0135] The Euclidean distance distribution deviation between the fingerprint to be matched and each standardized prototype vector is calculated. If the deviation exceeds a preset threshold, an online mean sliding update mechanism is activated to fine-tune the center position of the fingerprint vector to compensate for the instantaneous shift caused by short-term environmental noise.

[0136] By using the vector space alignment processing method described above, the original process state fingerprint generated in the previous step is transformed into an aligned process state fingerprint and a standardized prototype vector set that are consistent with the feature distribution of the offline meta-task library. This achieves the expected technical effect of eliminating cross-domain feature distribution offset and improving the matching accuracy of meta-tasks.

[0137] S4.2: Perform vector-wise cosine similarity calculation based on the aligned process state fingerprint to be matched and the standardized prototype vector set to quantify the semantic closeness between the current equipment operating state and each historical drift pattern and generate an initial similarity score sequence containing multiple similarity values.

[0138] The system receives the process state fingerprint vector to be matched (after vector space alignment) and the standardized prototype vector set from the offline meta-task library as input data sources for similarity calculation. It iterates through each prototype vector in the standardized prototype vector set, extracts the prototype vector component at the current index position, and reads it element-by-elementally with the process state fingerprint vector to be matched. It calculates the dot product of the process state fingerprint vector to be matched and the current prototype vector to obtain the projection overlap metric of the two vectors in the multidimensional feature space. It calculates the L2 norm magnitude of both the process state fingerprint vector to be matched and the current prototype vector to quantify their geometric lengths in the feature space. It divides the dot product result by the product of the L2 norm magnitudes of the two vectors, performs normalized division to eliminate the influence of vector magnitude differences on the directional consistency metric, and generates the cosine similarity scalar value corresponding to the current prototype vector. It stores the calculated cosine similarity scalar value in a temporary cache queue and marks the unique identifier ID of the prototype vector to which the value belongs. Repeat the above steps of dot product, norm calculation, and normalized division until all standardized prototype vectors in the offline meta-task library have been traversed. Summarize all calculated cosine similarity scalar values ​​and their corresponding prototype vector IDs, arranging them according to their original entry order or calculation sequence to construct an initial similarity score sequence containing multiple similarity values. Through vector-wise cosine similarity calculation, the vector distance relationship in the high-dimensional feature space is transformed into a dimensionless semantic proximity index, generating the initial similarity score sequence. This achieves a quantitative representation of the correlation strength between the current device operating state and various historical drift patterns.

[0139] S4.3: Perform descending sorting on the initial similarity score sequence to establish the association priority between each prototype vector and the current process state fingerprint based on the similarity value and generate a sorted task candidate list with a clear order relationship.

[0140] Receive the initial similarity score sequence generated by step S4.2. This sequence contains the cosine similarity values ​​between the current process status fingerprint and all prototype vectors in the offline meta-task library, as well as the corresponding task index identifiers.

[0141] Construct a mapping table based on key-value pairs, using each similarity value as the sorting key and the corresponding prototype vector index and historical drift pattern label as the associated value, to ensure that the integrity of the data index is not lost during subsequent reordering.

[0142] The key-value pairs in the mapping table are sorted in descending order using either quicksort or heapsort algorithms. The task candidate list is then reorganized according to the logical order of similarity values ​​from largest to smallest, and the semantic proximity of each historical drift pattern to the current real-time state is prioritized.

[0143] A stability constraint mechanism is introduced during the sorting process. When the difference in similarity between two different prototype vectors is less than a preset floating-point error threshold, their original storage order in the meta-task library is preserved to avoid randomness in task selection caused by numerical jitter.

[0144] Traverse the sorted mapping table, extract the index information of the top M high-priority tasks and their associated prototype vector features, and generate a sorted candidate list of tasks with clear hierarchical relationships, where the first item in the list represents the historical drift pattern that best matches the current device state.

[0145] By using descending sorting and stability constraints, the initial similarity score sequence is transformed into a sorted list of candidate tasks with clear priorities. This enables precise quantification and hierarchical division of the correlation between various historical drift patterns, providing an ordered data foundation for subsequent Top-3 optimal task selection.

[0146] S4.4: Perform a top truncation operation based on the sorted task candidate list to select the top three highly similar tasks as the optimal matching objects and generate a Top-3 nearest neighbor task subset consisting of three prototype vectors.

[0147] Obtain the task candidate list after descending sorting. This list contains the cosine similarity scores of all prototype vectors in the offline meta-task library with the current process state fingerprint and their corresponding index identifiers.

[0148] Set a preset threshold parameter, which is fixed at 3, to limit the maximum number of context prototypes in the dynamic task set in order to balance computational complexity and feature coverage.

[0149] Traverse the first three positions of the sorted task candidate list, and extract the prototype vectors of the first, second and third ranked tasks and their associated similarity confidence scores in turn.

[0150] The three extracted high-similarity prototype vectors are subjected to dimension consistency verification to ensure that their vector dimensions are strictly aligned with the 128-dimensional embedding space of the process state fingerprint, thus preventing dimension mismatch errors in subsequent splicing operations.

[0151] The three prototype vectors that passed the verification were arranged in a structured manner according to the order of similarity from high to low, and an ordered vector group containing the main matching task, the secondary matching task, and the auxiliary matching task was constructed.

[0152] The ordered vector group is encapsulated in a continuous memory manner to generate a fixed-length Top-3 nearest neighbor task subset, which serves as the multi-task context input benchmark for the graph neural network model.

[0153] By performing a top-truncation operation, the full ranking results generated in the previous step are transformed into a simplified task subset containing only the most relevant historical drift patterns. This achieves sparsity constraints on the meta-learning search space, significantly reduces the GPU memory usage during graph neural network inference, and improves the real-time response speed of parameter generation.

[0154] S4.5: Perform joint encapsulation processing using the Top-3 nearest neighbor task subset and the aligned process state fingerprint to be matched, so as to construct a structured data unit containing current real-time state features and historical similar context information and generate a dynamic task set for driving the graph neural network model.

[0155] Step S5: Input the process state fingerprint and the dynamic task set into the graph neural network model, and output lightweight adapter parameters that only affect the output layer attention head and the feedforward network bias term. Specifically, this includes:

[0156] S5.1: Perform splicing and fusion processing on the process state fingerprint and the prototype vector in the dynamic task set to construct a joint representation vector containing the current drift features and the historical similar task context, as the initial input data of the graph neural network model.

[0157] The system receives the aligned process state fingerprint vector to be matched from the output of step S4, as well as a set of normalized prototype vectors composed of the Top-3 nearest neighbors. The process state fingerprint is a 128-dimensional floating-point vector representing the physical degradation characteristics of the equipment in the current exposure cycle; the set of normalized prototype vectors contains three 128-dimensional vectors, each corresponding to one of the three typical drift patterns that are semantically closest in the historical database.

[0158] A weighted average aggregation operation is performed on the three standardized prototype vectors to generate a global task prototype vector representing historically similar contexts. The weight coefficients are normalized based on the cosine similarity score calculated in step S4.2 to ensure that tasks with high similarity contribute more to the global prototype, thereby constructing a statistically significant historical drift trend benchmark.

[0159] The 128-dimensional process state fingerprint vector to be matched is concatenated with the generated global task prototype vector in terms of feature dimensions. Through tensor join operation, the two independent one-dimensional vectors are merged into a joint feature vector of length 256, which simultaneously retains the micro-fluctuation information of the current real-time state and the prior knowledge of the historical macro-drift pattern.

[0160] A linear projection transformation is performed on the 256-dimensional joint feature vector to map it to the hidden layer dimension space of the graph neural network model input. A predefined weight matrix is ​​then used to linearly combine the concatenated features, eliminating the feature distribution imbalance caused by direct concatenation and generating an initial input embedding vector with uniform dimensions and distribution characteristics.

[0161] The initial input embedding vector is subjected to layer normalization to stabilize the numerical distribution during the forward propagation of the graph neural network model. By subtracting the batch mean and dividing by the standard deviation, the internal covariate bias between different batches of data is eliminated, ensuring the stability of the gradient flow in subsequent nonlinear mapping layers. The final output is the joint representation vector that serves as the entry point for the graph neural network encoder.

[0162] Through the above splicing, fusion, and normalization processing, the discrete fingerprints and task prototypes from the previous step are transformed into joint representation vector data with compact structure and rich semantics. This achieves deep fusion of the current process drift state and historical evolution context, providing a highly discriminative input basis for the graph neural network model to accurately generate lightweight adapter parameters.

[0163] S5.2: Based on the joint representation vector, input to the pre-trained graph neural network encoder to perform nonlinear mapping transformation, so as to extract the implicit semantic features of process degradation mode and generate high-dimensional hidden layer state code, providing deep semantic basis for subsequent parameter generation.

[0164] The system receives the joint representation vector generated in step S5.1, which integrates the semantic information of the current process state fingerprint and the Top-3 nearest neighbor task prototype vectors. This joint representation vector is input into the input layer of a pre-trained graph neural network encoder, where linear transformation and bias addition are performed to initially map it to a high-dimensional latent space. The ReLU activation function is used to perform nonlinear rectification on the linearly transformed features, suppressing negative noise and enhancing sparsity, generating the first-level hidden layer feature map. A Dropout mechanism is introduced to randomly discard some neuron connections, preventing the graph neural network from overfitting to specific drift patterns and improving generalization robustness. A multi-layer fully connected perceptron structure is used to deeply abstract the first-level hidden layer feature map, extracting the implicit semantic features of process degradation patterns layer by layer. Layer normalization is used to standardize the output of each layer, accelerating convergence and stabilizing gradient propagation, generating a high-dimensional hidden layer state code. Through the above nonlinear mapping transformation, the joint representation vector from the previous step is transformed into a high-dimensional hidden layer state code containing deep semantics, achieving the expected technical effect of providing accurate semantic basis for subsequent parameter generation.

[0165] For example, the joint representation vector has a dimension of 192. The graph neural network encoder contains three fully connected layers with 256, 128, and 64 hidden units, respectively. The first layer receives a 192-dimensional vector as input, with a weight matrix of 192x256 and a bias vector of 256 dimensions. After a linear transformation and ReLU activation, it outputs 256-dimensional features. The second layer has a weight matrix of 256x128, outputting 128-dimensional features, with a dropout rate of 0.2. The third layer has a weight matrix of 128x64, outputting a 64-dimensional high-dimensional hidden state encoding. This encoding captures the coupled degradation semantics of light source aging and mechanical wear, serving as direct input to the decoder to generate adapter parameters, significantly improving the relevance and accuracy of parameter generation.

[0166] S5.3: The high-dimensional hidden layer state encoding is used to perform linear projection operation through the first branch of the graph neural network decoder to generate an attention adapter parameter set specifically used to correct the attention head weight matrix of the backbone recommendation model output layer.

[0167] Receive the high-dimensional hidden state code generated in the previous steps. This code contains deep semantic association information between the current process drift mode and the context of historically similar tasks.

[0168] The architecture parameters of the multi-head attention mechanism in the output layer of the backbone recommendation model are analyzed to determine the number of attention heads to be corrected and the dimension of the projection matrix corresponding to each attention head. The mapping index relationship between the graph neural network output space and the backbone model parameter space is established.

[0169] The first branch of the graph neural network decoder, the linear projection layer, is constructed. This layer consists of a set of learnable weight matrices and bias vectors. Its input dimension is consistent with the dimension of the high-dimensional hidden layer state encoding, and its output dimension strictly matches the total number of elements in all attention head projection matrices to be corrected.

[0170] The high-dimensional hidden layer state encoding is input into the first branch linear projection layer, and matrix multiplication and bias addition operations are performed to map the hidden semantic features into the original adapter parameter vector.

[0171] The generated original adapter parameter vector is reshaped and restructured into multiple independent two-dimensional weight correction matrices according to the preset attention head structure. Each correction matrix corresponds to a specific attention head.

[0172] A scaling factor gating mechanism is introduced to scale the recombined weight correction matrix element by element, limiting the numerical range of the adapter parameters and preventing catastrophic shifts in the output distribution of the backbone model due to drastic parameter fluctuations.

[0173] By using linear projection and structured reshaping, the high-dimensional hidden layer state encoding from the previous step is transformed into an attention adapter parameter set specifically used to correct the attention head weight matrix of the backbone recommendation model output layer. This achieves fine-grained dynamic calibration of the interaction weights of the attention mechanism features, significantly improving the model's sensitivity to subtle process drift.

[0174] S5.4: Based on the high-dimensional hidden layer state encoding, the bias offset is calculated through the second branch of the graph neural network decoder to generate a bias adapter parameter set specifically used to adjust the activation threshold of the feedforward network neurons in the backbone recommendation model.

[0175] It receives high-dimensional hidden state encodings from a graph neural network encoder, which contain deep semantic association features between the current process drift pattern and the context of historically similar tasks.

[0176] A fully connected linear mapping process is performed on the high-dimensional hidden layer state encoding. The high-dimensional semantic space is projected onto a low-dimensional parameter space that matches the dimension of the feedforward network bias term through a preset second branch weight matrix, generating an initial bias offset vector.

[0177] The Tanh activation function is used to perform nonlinear constraint processing on the initial bias offset vector, limiting the numerical range to a preset small perturbation range to prevent drastic distortion of the backbone model output distribution due to excessive offset.

[0178] A dynamic scaling factor is introduced to adjust the magnitude of the constrained bias offset vector. This factor is adaptively adjusted according to the confidence score of the process state fingerprint to ensure that the parameter correction intensity is reduced when the state recognition is uncertain.

[0179] The adjusted bias offset vector is reshaped into a tensor structure that is completely consistent with the bias term of the feedforward network in the output layer of the backbone recommendation model, forming a bias adapter parameter set specifically used to adjust the activation threshold of neurons.

[0180] Through the above chain-like derivation process, the high-dimensional hidden layer state encoding of the previous step is transformed into a set of bias adapter parameters with physical meaning and controlled amplitude, thereby realizing refined dynamic compensation of the activation threshold of the feedforward network of the backbone model and significantly improving the robustness of the recommendation model to nonlinear process drift.

[0181] S5.5: Perform structured encapsulation processing on the attention adapter parameter group and the bias adapter parameter group to form the final lightweight adapter parameter package, which serves as the direct control variable for local coordinate system recalibration before freezing the output layer of the backbone recommendation model.

[0182] Step S6: Process the lightweight adapter parameters to obtain a local remapping inference path, perform local coordinate system recalibration based on the local remapping inference path, and generate recommended adaptive exposure parameter values. Specifically, this includes:

[0183] S6.1: Based on the lightweight adapter parameters generated in the previous steps and the output layer architecture definition of the frozen backbone recommendation model, the weight matrix of the multi-head attention mechanism in the output layer is decomposed to separate the attention head projection matrix to be corrected and generate the set of attention heads to be recalibrated.

[0184] Receive the lightweight adapter parameter packet output by S5.5, parse the weight correction tensor structure for the attention mechanism, and obtain the dimension index and numerical mapping relationship to be corrected.

[0185] Read the architecture definition file of the frozen backbone recommendation model output layer to locate the physical storage address and dimension specifications of the query, key, and value projection matrices in the multi-head attention mechanism.

[0186] The multi-head attention weight matrix in the output layer is subjected to structured decomposition, and the global weight matrix is ​​decomposed into multiple independent subspace projection matrices according to the number of attention heads.

[0187] Identify the feature channels with the highest sensitivity to process drift in the projection matrix of each attention head, mark them as regions to be recalibrated, and form a set of attention heads to be recalibrated.

[0188] Establish an index mapping table between adapter parameters and the set of attention heads to be recalibrated to ensure address alignment and dimension matching during subsequent parameter loading.

[0189] Through the above decomposition and mapping process, the lightweight adapter parameters generated in the previous step are accurately associated with the specific computational units of the backbone model, realizing the preparatory work for local coordinate system recalibration and ensuring the accuracy and real-time performance of parameter injection.

[0190] For example, assume the backbone model output layer contains 8 attention heads, each with a 64x64 projection matrix. Parse the adapter parameter package to obtain the bias correction vectors corresponding to attention heads 2, 5, and 7, all with a dimension of 64. Read the model architecture to locate the memory addresses of the Query projection matrices for these three heads. Split the global weight matrix by head index, extracting three 64x64 sub-matrices as objects to be recalibrated. Establish a mapping table: adapter parameter index 0 corresponds to head 2, index 1 corresponds to head 5, and index 2 corresponds to head 7. This process takes less than 0.1ms, ensuring parameter preparation is completed within a millisecond-level response window, providing precise operation objects for subsequent addition and fusion operations, and significantly improving the model's sensitivity to minute process drifts.

[0191] S6.2: Using the bias correction vector in the lightweight adapter parameters, perform an additive fusion operation on the projection matrices of each attention head in the set of attention heads to be recalibrated, so as to eliminate the feature space offset caused by process drift and generate a recalibrated attention head matrix group.

[0192] Obtain the original projection matrix and the bias correction vector from the lightweight adapter parameters in the set of attention heads to be recalibrated, and use them as input objects for the addition fusion operation.

[0193] The original projection matrix is ​​analyzed for dimensions to confirm that its row and column dimensions correspond to the input feature dimension and output mapping dimension of the attention head, respectively, thus ensuring the integrity of the matrix structure.

[0194] The numerical sequence in the bias correction vector is extracted. This sequence is dynamically generated by the graph neural network decoder based on the current process state fingerprint and represents the feature space offset caused by equipment aging.

[0195] The bias correction vector is broadcast and extended to the same dimensional space as the original projection matrix, and an offset matrix with the same shape as the projection matrix is ​​constructed to accommodate the matrix addition operation requirements.

[0196] Perform element-wise addition to add each weight value in the original projection matrix to the offset value at the corresponding position, thereby achieving a local translation transformation of the weight space.

[0197] The recalibrated attention head projection matrix is ​​calculated using the following formula:

[0198]

[0199] in, The attention head projection matrix after recalibration. The original projection matrix, This is the bias correction matrix after broadcast expansion.

[0200] The numerical stability of the calculation results is checked to eliminate non-finite values ​​caused by floating-point operation errors, ensuring the numerical validity of the recalibrated matrix in subsequent inference.

[0201] The processed recalibrated attention head projection matrix is ​​encapsulated into a recalibrated attention head matrix group, which serves as the core parameter unit for local coordinate system recalibration.

[0202] By using an additive fusion processing method, the bias correction vector generated in the previous step is transformed into the recalibrated attention head matrix data, thereby achieving the expected technical effect of eliminating feature space offset caused by process drift and improving the adaptability of the recommendation model to changes in equipment status.

[0203] S6.3: Based on the feedforward network scaling factor in the lightweight adapter parameters, perform element-wise multiplication mapping on the feedforward neural network bias term before the frozen backbone recommendation model output layer to adjust the response threshold of the nonlinear activation function and generate a dynamically compensated feedforward network bias tensor.

[0204] Obtain the feedforward network scaling coefficient vector from the lightweight adapter parameter package. This vector is generated by the second branch of the graph neural network decoder, and its dimension is strictly consistent with the dimension of the bias term of the feedforward network in the output layer of the backbone recommendation model.

[0205] Read the static bias tensor of the feedforward neural network layer in the frozen backbone recommendation model, extract its original numerical distribution as benchmark reference data, and ensure that subsequent mapping operations are based on stable model structure features.

[0206] An element-wise multiplication mapping operator is constructed to perform point-to-point multiplication of each component of the feedforward network scaling coefficient vector with the corresponding bias value in the feedforward neural network bias tensor, forming an intermediate variable after dynamic compensation.

[0207] The bias tensor of the feedforward network after dynamic compensation is calculated using the following formula:

[0208]

[0209] in, This is the bias tensor of the feedforward network after dynamic compensation. This is the original static bias tensor. is the scaling coefficient vector of the feedforward network, and ⊙ represents the Hadamard product, i.e., element-wise multiplication.

[0210] The nonlinear activation function response threshold calibration is performed on the intermediate variables after dynamic compensation. By adjusting the value of the bias term, the shift of the neuron activation curve is changed, thereby correcting the nonlinear distortion of the feature space caused by process drift.

[0211] A dynamically compensated feedforward network bias tensor containing calibrated values ​​is generated. This tensor directly replaces the static bias term in the original model and is used for feature transformation calculation in the subsequent inference stage.

[0212] By using an element-wise multiplication mapping method, the scaling coefficients generated in the previous step are transformed into dynamically compensated feedforward network bias tensor data, which enables adaptive adjustment of the response threshold of the nonlinear activation function and significantly improves the feature representation robustness of the recommendation model in process drift scenarios.

[0213] S6.4: Load the recalibrated attention head matrix group and the dynamically compensated feedforward network bias tensor into the inference computation graph of the frozen backbone recommendation model to replace the corresponding static parameters of the original output layer and construct a local remapping inference path with process state awareness.

[0214] The recalibrated attention head matrix generated in S6.2 and the dynamically compensated feedforward network bias tensor generated in S6.3 are obtained as the parameter input sources for the local remapping inference path.

[0215] We analyze the computational graph topology of the frozen backbone recommendation model's output layer and locate the memory addresses of the Query, Key, and Value projection layers of the multi-head attention mechanism and the linear transformation layer of the feedforward neural network.

[0216] After recalibration, the data of each submatrix in the attention head matrix group are directly mapped to the weight storage area of ​​the corresponding attention head in the backbone model through pointer reference, thus overwriting the original static weight values.

[0217] The numerical sequence in the bias tensor of the dynamically compensated feedforward network is injected into the bias registers of each layer of the feedforward network in the order of neuron indices, replacing the original fixed bias terms.

[0218] Perform a computation graph parameter consistency check to confirm that the dimensions of all replaced parameters are fully matched with the default architecture of the backbone model, preventing inference crashes caused by dimension misalignment.

[0219] Activate the local remapping inference path containing new parameters and establish a real-time data flow channel from deep process feature vectors to the exposure parameter output.

[0220] By using the above-mentioned parameter hot replacement and pathway construction processing methods, static model parameters are transformed into dynamic inference configurations with process status awareness capabilities, achieving millisecond-level adaptive recommendation response.

[0221] For example, the backbone model output layer contains 8 attention heads, each with a 128x128 weight matrix. The attention adapter parameter set generated by the graph neural network contains 8 128x128 correction matrices. The system loads these 8 correction matrices into the weight storage area of ​​the 1st to 8th attention heads, respectively. The feedforward network contains 2048 neurons, and the dynamically compensated bias tensor is a vector of length 2048, directly overwriting the original bias terms. The verification process confirms that all matrices are 128x128 in dimension and the vector length is 2048, with no dimension conflicts. After loading, the inference path takes effect immediately without restarting the model service, ensuring that the exposure parameter recommendation latency is controlled within 5ms.

[0222] S6.5: Perform a single forward propagation calculation on the deep process feature vector extracted from the frozen backbone recommendation model through the local remapping inference path, so as to complete the local coordinate system recalibration within a millisecond delay and output the adaptive exposure parameter recommendation value that adapts to the current equipment aging state.

[0223] It receives the deep process feature vectors output by the frozen backbone recommendation model, the recalibrated attention head matrix group generated by the preceding steps, and the dynamically compensated feedforward network bias tensor.

[0224] The deep process feature vectors are input into the multi-head attention mechanism layer in the local remapping inference path where the static parameters have been replaced.

[0225] The input feature vector is subjected to a linear projection transformation using the recalibrated attention head matrix group in order to capture the key feature dependencies under the current process state.

[0226] The scaling dot product attention mechanism is used to calculate the similarity score between the query vector and the key vector, and the normalized attention weight distribution is generated by the softmax function.

[0227] The value vector is weighted and summed based on the attention weight distribution to generate an intermediate feature representation that includes global context information.

[0228] The intermediate feature representations are input into the feedforward neural network layer, and the dynamically compensated feedforward network bias tensor is loaded to adjust the neuron activation threshold.

[0229] The bias tensor is fused into the linear transformation result of the feedforward network by element-wise addition to correct the nonlinear response shift caused by equipment aging.

[0230] Layer normalization eliminates internal covariate bias and stabilizes the numerical distribution range of deep features.

[0231] The processed feature vectors are passed to the linear projection module of the output layer to perform the final mapping from the hidden space to the exposure parameter space.

[0232] The recommended values ​​for adaptive exposure parameters are calculated using the following formula:

[0233]

[0234] in, Recommended values ​​for adaptive exposure parameters. The output layer projection weight matrix. This is the deep feature vector after local recalibration. This is the output layer bias term.

[0235] Physical constraint verification is performed on the calculated original recommended values ​​to ensure that parameters such as dose and focal length are within the safe operating range of the exposure machine.

[0236] By performing a single forward propagation calculation through the local remapping inference path, deep process features are transformed into adaptive exposure parameter recommendations that fit the current equipment aging state, enabling real-time compensation for process drift and high-precision parameter recommendations within millisecond-level latency.

[0237] Step S7: Based on the recommended values ​​of the adaptive exposure parameters, control the exposure machine to perform the exposure operation, and repeatedly collect new multi-source weak state signals in the next execution cycle to trigger real-time updates of the process state fingerprint. Specifically, this includes:

[0238] S7.1: Obtain the recommended values ​​of adaptive exposure parameters and the process recipe information of the wafer to be processed. Use the industrial fieldbus protocol to map and convert the recommended values ​​of adaptive exposure parameters to convert the abstract recommended values ​​into a digital instruction set that can be recognized by the underlying motion control unit and the light source drive unit of the exposure machine, thereby generating a standardized execution instruction package containing dose, focal length and scanning speed information.

[0239] The system receives the adaptive exposure parameter recommendation vector output from S6.5, which includes dose correction coefficients, focal length offset, and scan speed adjustment factors. It reads the process recipe ID of the current wafer to be processed and retrieves the corresponding reference exposure parameter set from the local process database, including reference energy density, reference focal plane position, and reference stage movement speed. It performs element-wise algebraic addition on the adaptive exposure parameter recommendation values ​​and the reference exposure parameter set to calculate the absolute physical control target value for the current wafer. Based on the data frame structure definition of the industrial fieldbus protocol, it constructs a standard communication data packet template containing a frame header identifier, function code field, data payload area, and checksum tail. The calculated absolute physical control target value is converted into a 32-bit floating-point binary sequence conforming to the IEEE 754 standard. According to the preset byte order rules, the binary sequence is mapped to the specified offset address in the data payload area and filled into the corresponding data bits of the dose control register, focal length servo register, and speed command register. A cyclic redundancy check is performed on the complete data packet to generate a 16-bit checksum, which is then filled into the frame tail, forming a standardized execution command packet with integrity protection mechanisms. Through the above mapping and transformation process, the abstract model recommendation values ​​are transformed into a set of digital instructions that can be parsed by the underlying hardware, achieving a seamless connection between the recommendation results and the physical execution mechanism.

[0240] For example, the recommended adaptive exposure parameters are: dose increase of 0.5 mJ / cm², focal length increase of 2 μm, and scan speed decrease of 0.1 mm / s. The baseline formulation is: dose 20 mJ / cm², focal length 0 μm, and scan speed 500 mm / s. The calculated absolute target values ​​are: dose 20.5 mJ / cm², focal length 2 μm, and scan speed 499.9 mm / s. An EtherCAT protocol data frame is constructed, with the function code set to 0x02 to indicate parameter writing. 20.5 is converted to hexadecimal 41A40000H, 2.0 is converted to 40000000H, and 499.9 is converted to 43F9E666H. Bytes 4-7, 8-11, and 12-15 of the data payload are filled sequentially. The calculated CRC checksum is 0xB3A4. The final 18-byte instruction packet is transmitted via fiber optic cable to the light source controller and workpiece stage driver, with an instruction parsing delay of less than 50 μs, ensuring strict synchronization between the exposure action and the recommended parameters.

[0241] S7.2: Based on the standardized execution instruction package, the precision workpiece stage displacement mechanism and the high-energy ultraviolet light source module of the exposure machine are coordinated and adjusted to establish an optical imaging field between the mask and the silicon wafer surface that meets the recommended values ​​of the adaptive exposure parameters, and a single complete exposure scan is performed to produce an exposed wafer with predetermined pattern features.

[0242] The dose setting, focal length compensation, and scanning speed parameters in the standardized execution instruction package are analyzed and mapped to discrete control words in the exposure machine's underlying controller. The target step pulse count for the objective lens focusing motor is calculated based on the focal length compensation, generating a Z-axis displacement drive signal sequence containing direction position and pulse count values. Based on the scanning speed parameters and the grating ruler feedback resolution, the speed feedforward gain and position loop PID parameters of the workpiece stage X / Y axis motion controller are calculated, constructing a dynamic motion trajectory planning curve. The real-time power feedback value of the light source drive module is read and compared with the target exposure dose in the instruction package; the light source current correction coefficient is calculated using the proportional-integral-differential method. The current correction coefficient is injected into the digital-to-analog converter register of the light source high-voltage power supply to adjust the pump energy of the ultraviolet laser to match the instantaneous scanning speed change. The global clock signals of the workpiece stage motion controller and the light source modulation unit are synchronously triggered to ensure strict spatiotemporal alignment between the mask pattern projection and the photoresist exposure on the silicon wafer surface. The exposure shutter mechanism is activated, executing a line-by-line scanning exposure operation on the wafer surface according to the preset acceleration-uniform speed-deceleration phases.

[0243] By coordinating the precise mechanical displacement and high-energy light source output, abstract recommended parameters are transformed into physical exposure actions, achieving high-precision pattern transfer that adapts to process drift.

[0244] For example, the instruction package includes focal length compensation of +0.5μm, a scanning speed of 200mm / s, and a target dose of 30mJ / cm². The objective lens focusing motor receives 1000 step pulses to complete Z-axis fine-tuning. The stage controller adjusts the speed feedforward gain to 1.2 and the position loop proportional gain to 500N / m. The light source power feedback is 98% of the rated value, and the PID algorithm calculates a current correction coefficient of 1.02, which is injected into the DAC register to increase the pump energy. The global clock synchronization error is controlled within 5ns. After the shutter opens, the stage scans at a uniform speed of 200mm / s, and the light source stably outputs a dose of 30mJ / cm², completing the exposure of one layer of circuit pattern with significantly improved linewidth uniformity.

[0245] S7.3: Monitor the end flag signal of the exposure scanning operation. Once the operation is detected to be completed, immediately send a synchronous acquisition trigger pulse to the distributed sensor network to activate multiple physical sensor probes installed in the exposure light source drive circuit, objective lens focusing motor, mask stage guide rail and exposure cavity to enter high-frequency sampling mode.

[0246] Monitor the state machine flags in the exposure scanning operation control logic to capture in real time the end signals that represent the stopping of the wafer stage movement and the shut-off of the light source shutter.

[0247] The level transition edge of the end flag is analyzed to confirm that the physical action of the current exposure cycle has been completely terminated and the system has entered an idle waiting state.

[0248] Based on the confirmed idle state, a high-priority hardware interrupt request is generated and directly mapped to the synchronization trigger controller of the distributed sensor network.

[0249] Nanosecond-level synchronization pulses are broadcast via fiber optic bus to the light source drive circuit, objective lens focusing motor, mask stage guide rail, and sensor probe inside the exposure cavity.

[0250] After receiving the synchronization pulse, each sensor immediately switches the phase of its internal sampling clock, instantly transitioning from a low-power sleep mode to a high-frequency data acquisition mode.

[0251] Lock the initial sampling times of light source current harmonics, motor back electromotive force, position residuals, and temperature gradients to ensure that multi-source signals are strictly aligned within a microsecond-level time window.

[0252] Through the aforementioned synchronous triggering mechanism, discretely distributed physical sensors are activated into a data acquisition array with a unified timing reference, enabling zero-time-difference parallel acquisition of weak state signals from multiple sources.

[0253] For example, in a 193nm ArF immersion lithography machine scenario, when the workpiece stage decelerates to zero and the shutter closure signal is high, the trigger controller emits a TTL synchronization pulse within 50ns. The light source current sensor captures high-frequency harmonics at a 10MHz sampling rate, the focusing motor encoder records back EMF fluctuations at a 5MHz sampling rate, the laser interferometer acquires position residuals at a 20MHz sampling rate, and the 48 temperature sensors in the cavity synchronously refresh their readings. This mechanism ensures that the time deviation of the four types of heterogeneous signals is less than 1μs, eliminates the misalignment of process state fingerprint features caused by asynchronous acquisition, and significantly improves the extraction accuracy of spatiotemporal coupling features by the subsequent graph convolutional encoder.

[0254] S7.4: Using the high-frequency sampling mode, a new round of raw physical response data is captured in parallel from the activated physical sensing probe. The raw physical response data includes the transient waveform of the light source current, the back electromotive force voltage fluctuation of the motor, the position reading deviation of the laser interferometer, and the temperature distribution of the thermistor array, so as to construct an initial multi-source heterogeneous signal stream covering the current latest equipment operating status.

[0255] In response to the synchronous acquisition trigger pulse issued by S7.3, the high-speed data acquisition card in the distributed sensor network immediately enters the interrupt service routine, locks the end time of the current exposure operation as the zero point, and starts the multi-channel parallel sampling sequence.

[0256] For the exposure light source driving circuit, a high-precision current probe with a bandwidth of not less than 100MHz is configured to continuously capture transient current waveforms with a duration of 20ms at a sampling rate of 500MS / s. High-frequency switching noise is filtered out by a hardware low-pass filter, while retaining the fundamental and harmonic components that characterize plasma stability, thereby generating the original timing data of the light source current.

[0257] The terminal voltage feedback signal of the objective lens focusing motor driver and the encoder speed pulse are read synchronously. The voltage values ​​of adjacent sampling points are differentiated using a differential amplifier. Combined with the motor back EMF constant, the back EMF fluctuation envelope reflecting the small changes in mechanical load is calculated to form the original time sequence data of motor voltage fluctuation.

[0258] The laser interferometer system is triggered to read the real-time position coordinates of the mask stage in the X / Y axis direction, compare them point by point with the theoretical command trajectory, calculate the deviation between the two, and extract the position feedback residual sequence with nanometer precision as the original time-series data of position deviation characterizing the rigidity of the mechanical transmission chain and the wear degree of the guide rail.

[0259] The four high-precision platinum resistance temperature sensors arranged in a tetrahedral pattern within the cavity are polled to obtain their respective thermistor resistance values ​​and convert them into absolute temperature readings. This constructs a three-dimensional temperature field discrete point set containing spatial coordinate information, forming the original spatiotemporal data of the thermal distribution.

[0260] The above four types of raw data are each stamped with a high-precision timestamp based on a global clock synchronization protocol, and packaged according to the physical topology order of light source, motor, mask stage, and temperature to construct an initial multi-source heterogeneous signal stream covering the latest equipment operating status.

[0261] By using parallel capture and timestamp marking processing, the trigger signal from the previous step is transformed into an initial multi-source heterogeneous signal stream containing multi-dimensional information such as current, voltage, position, and temperature. This enables holographic recording of the instantaneous physical state of the exposure machine, providing a high-fidelity data foundation for real-time updates of the subsequent process state fingerprint.

[0262] For example, after the 5000th wafer was exposed using a 193nm ArF exposure machine, the system captured a 20ms source current waveform at a sampling rate of 500MS / s, with a data volume of approximately 10... 7 Simultaneously, the voltage fluctuation of the focusing motor was recorded at a sampling rate of 1 MS / s, the mask stage position residual was read at a frequency of 10 kHz, and the data of four temperature points were read at a frequency of 100 Hz. All data were timestamped at the nanosecond level and encapsulated into an initial multi-source heterogeneous signal stream containing four channels. The total data volume was controlled within 5 MB, ensuring that the acquisition and encapsulation were completed within 1 ms, which significantly improved the real-time performance and completeness of state awareness.

[0263] S7.5: Perform timestamp alignment and data frame encapsulation processing on the initial multi-source heterogeneous signal stream to integrate the discretely acquired current harmonic characteristics, back EMF fluctuation envelope, position feedback residual sequence and spatial gradient tensor into a set of multi-source weak state signals for the next execution cycle with a unified timing reference, thereby completing the data input preparation required for process state fingerprint update.

[0264] For those skilled in the art, various other corresponding changes and modifications can be made based on the technical solutions and concepts described above, and all such changes and modifications should fall within the protection scope of the claims of this invention.

[0265] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this application pertains. The terms “first,” “second,” “third,” and similar terms used in this patent application specification and claims do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “comprising” or “including” and similar terms mean that the elements or objects preceding “comprising” or “including” encompass the elements or objects listed following “comprising” or “including” and their equivalents, and do not exclude other elements or objects. The “multiple” mentioned in the embodiments of this application refers to two or more. A and / or B indicate three possibilities: A; B; and A and B.

[0266] The above description is merely an exemplary embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for intelligent recommendation and adaptive feedback of exposure parameters for an exposure machine, characterized in that, Specifically, it includes: S1: Acquire multi-source weak state signals during the exposure execution cycle. The multi-source weak state signals include at least three of the following: high-frequency current spectrum characteristics of the exposure light source driving circuit, back electromotive force fluctuation envelope of the objective lens focusing motor, position feedback residual sequence of the mask stage motion trajectory, and spatial gradient tensor of the distributed temperature sensor in the cavity. S2: The multi-source weak state signals are uniformly sampled and normalized to generate a set of normalized heterogeneous time-series signals; S3: Input the normalized heterogeneous time-series signal set into a lightweight graph convolutional encoder for topology mapping processing to extract low-dimensional embedding vectors as process state fingerprints. S4: Perform cosine similarity calculation based on the process state fingerprint and the prototype vector in the offline meta-task library to select the Top-3 nearest neighbor tasks and form a dynamic task set. S5: Input the process state fingerprint and the dynamic task set into the graph neural network model, and output lightweight adapter parameters that only affect the output layer attention head and the feedforward network bias term; S6: Process the lightweight adapter parameters to obtain the local remapping inference path, perform local coordinate system recalibration based on the local remapping inference path and generate recommended values ​​for adaptive exposure parameters; S7: Based on the recommended values ​​of the adaptive exposure parameters, control the exposure machine to perform exposure operations, and repeatedly collect new multi-source weak state signals in the next execution cycle to trigger real-time updates of the process state fingerprint.

2. The intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine according to claim 1, characterized in that, The graph convolutional encoder includes a graph attention mechanism and a temporal convolutional network.

3. The intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine according to claim 1, characterized in that, S2 specifically includes: The high-frequency current spectrum characteristics, back EMF fluctuation envelope, position feedback residual sequence and spatial gradient tensor are subjected to timestamp alignment processing. The asynchronous signals are resampled to a unified target sampling rate using a global clock reference based on the exposure execution cycle, thereby generating a time-synchronized multi-source state signal sequence. A sliding window truncation process is performed on the time-synchronized multi-source state signal sequence to extract continuous and equal-length data segments based on a preset signal analysis frame length, thereby generating a fixed-length heterogeneous signal segment set. The range normalization calculation is performed on the set of fixed-length heterogeneous signal segments to map the original physical quantities to a dimensionless range of zero to one using the maximum and minimum values ​​in each signal channel, thereby generating a dimensionless signal matrix. A standardized heterogeneous time-series signal vector is obtained by performing mean-variance normalization transformation on the dimensionless signal matrix. The standardized heterogeneous timing signal vector is subjected to multidimensional tensor stacking encapsulation processing to integrate all channel data according to the preset topological order of light source, motor, mask stage and temperature sensor, thereby generating a normalized heterogeneous timing signal set.

4. The intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine according to claim 3, characterized in that, The step of performing mean-variance normalization transformation on the dimensionless signal matrix to obtain a standardized heterogeneous time-series signal vector includes: The dimensionless signal matrix is ​​subjected to mean-variance normalization transformation to adjust the data distribution to a standard normal distribution with zero mean and unit variance by subtracting the channel mean and dividing by the channel standard deviation, thereby generating a standardized heterogeneous time-series signal vector.

5. The intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine according to claim 1, characterized in that, S3 specifically includes: Based on the physical distribution relationship of sensors in the normalized heterogeneous time series signal set, a multi-source signal topology graph structure is constructed to map discrete time series data into graph structure data with clear connection relationships, thereby generating a multi-source signal topology graph structure. A graph attention mechanism is used to perform weighted aggregation processing on the features of adjacent nodes in the multi-source signal topology graph structure to generate a local spatial enhancement feature map. Based on the local spatial enhanced feature map, a temporal convolutional network is used to perform multi-layer causal convolution operations on the node feature sequence to capture the long-term and short-term temporal dependencies in the process drift and extract the dynamic evolution pattern, generating a spatiotemporal fusion feature tensor. Global average pooling and nonlinear dimensionality reduction mapping are performed on the spatiotemporal fusion feature tensor to compress redundant information and project high-dimensional features onto a preset low-dimensional latent space to generate an initial low-dimensional embedding vector. The process state fingerprint is obtained by processing the initial low-dimensional embedding vector.

6. The intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine according to claim 5, characterized in that, The multi-source signal topology includes a light source driving node, a focusing motor assembly node, a mask stage motion node, and a temperature sensing node.

7. The intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine according to claim 5, characterized in that, The process of processing the initial low-dimensional embedding vector to obtain the process state fingerprint includes: Based on the contrastive learning loss function, cluster center constraint optimization is performed on the initial low-dimensional embedding vector to bring the feature distance of similar process drift states closer and push away the feature distance of dissimilar states, and finally output a process state fingerprint with topological robustness.

8. The intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine according to claim 1, characterized in that, S4 specifically includes: An offline meta-task library is pre-built, which contains multiple prototype vectors covering typical aging stages and common drift modes. Vector space alignment processing is performed on the process state fingerprint and the multiple prototype vectors to eliminate feature distribution offset and generate an aligned set of process state fingerprints to be matched and standardized prototype vectors. Based on the aligned process state fingerprint to be matched and the standardized prototype vector set, perform vector-wise cosine similarity calculation to generate an initial similarity score sequence containing multiple similarity values. The initial similarity score sequence is sorted in descending order. The association priority between each prototype vector and the current process state fingerprint is determined based on the similarity value, and a sorted list of candidate tasks with a clear order relationship is generated. Based on the sorted task candidate list, a top truncation operation is performed to select the top three highly similar tasks as the optimal matching objects and generate a Top-3 nearest neighbor task subset consisting of three prototype vectors. The Top-3 nearest neighbor task subset and the aligned process state fingerprint to be matched are used to perform joint encapsulation processing to construct a structured data unit containing current real-time state features and historical similar context information and generate a dynamic task set.

9. The intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine according to claim 1, characterized in that, The high-frequency current spectrum characteristics of the exposure light source driving circuit are obtained by high-speed sampling of the output current, using a moving average filter to remove DC components and power frequency interference, applying a Hanning window function to window the circuit, and then performing a fast Fourier transform to collect high-frequency energy and the first five significant peak frequencies and amplitudes.

10. The intelligent recommendation and adaptive feedback method for exposure parameters of an exposure machine according to claim 1, characterized in that, The method for obtaining the back electromotive force fluctuation envelope of the objective lens focusing motor is to perform low-pass filtering and numerical differentiation on the feedback data of the terminal voltage and angular velocity, collect high-frequency data, extract the envelope by sliding a window according to the length of the mechanical vibration period, construct the back electromotive force calculation equation, remove the linear trend term and retain the AC fluctuation component.