Metrology model training method, metrology method, device, apparatus, medium and product
By extracting features from sensor and measurement data using a multimodal measurement model, the problem of incomplete wafer quality monitoring in existing technologies is solved, enabling full-point measurement information prediction and improving the accuracy and efficiency of semiconductor virtual measurement.
Patent Information
- Application Number
- CN202411993635.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2026-07-03
AI Technical Summary
In existing semiconductor measurement methods, traditional physical measurement methods result in incomplete wafer quality monitoring and time delays. Virtual measurement methods fail to fully combine instrument sensor parameters and actual measurement status parameters, and the measurement accuracy needs to be improved.
A multimodal measurement model is adopted, which extracts text features from sensor data and image features from measurement data, and combines the multimodal model for modeling to achieve prediction of measurement information at all locations.
It improves the accuracy and efficiency of semiconductor virtual measurement, enabling timely detection of problems in the production process and reducing manpower and material costs.
Smart Images

Figure CN122335644A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of semiconductor measurement, and more particularly to a measurement model training method, measurement method, apparatus, equipment, medium, and product. Background Technology
[0002] Semiconductor metrology is a crucial part of the semiconductor manufacturing process, designed to ensure that microelectronic devices manufactured on wafers meet design specifications and quality standards. With the continuous advancement of semiconductor technology, the feature sizes of devices are shrinking, and manufacturing processes are becoming increasingly complex, placing higher demands on metrology techniques.
[0003] Traditional metrology methods rely on physical measurement techniques, which, due to sampling, result in incomplete wafer quality monitoring and a time lag between the occurrence of a problem and its detection through actual measurement. To address this, machine learning technology has been introduced into the semiconductor metrology field, employing virtual measurement methods to improve the efficiency and accuracy of semiconductor metrology.
[0004] However, current virtual measurement methods mostly use a single model and cannot fully combine and effectively utilize the parameters of the machine's sensors and the existing actual measurement state parameters, so the measurement accuracy needs to be improved. Summary of the Invention
[0005] This application provides a measurement model training method, measurement method, apparatus, equipment, medium, and product to improve the accuracy of semiconductor virtual measurement.
[0006] In a first aspect, embodiments of this application provide a measurement model training method, comprising: acquiring multiple training samples, wherein each training sample includes sensor data of adjacent training wafers on the current machine, measurement data of adjacent training wafers after passing the current machine, and measurement data of the next training wafer in the adjacent training wafers after passing the previous machine; for adjacent training wafers in each training sample, performing a first normalization process on the sensor data of adjacent training wafers on the current machine to obtain text features corresponding to the training sample; performing a second normalization process on the measurement data of the next training wafer in the adjacent training wafers after passing the previous machine to obtain a first image feature corresponding to the training sample, performing a second normalization process on the measurement data of the previous training wafer in the adjacent training wafers after passing the current machine to obtain a second image feature corresponding to the training sample, and performing a second normalization process on the measurement data of the next training wafer in the adjacent training wafers after passing the current machine to obtain a third image feature corresponding to the training sample; and training the measurement model based on the text features, first image features, second image features, and third image features corresponding to the multiple training samples until the measurement model training is completed.
[0007] In one possible implementation, the training process includes: inputting the text features, first image features, and second image features corresponding to the training samples into the measurement model to obtain the first wafer image output by the measurement model; feeding back the first loss between the first wafer image and the third image features corresponding to the training samples to the measurement model until the first loss is less than a first threshold, and then determining that the measurement model training is complete.
[0008] In one possible implementation, the method further includes: acquiring multiple verification samples, wherein each verification sample includes sensor data of adjacent verification wafers on the current machine, measurement data of adjacent verification wafers after passing the current machine, and measurement data of the next verification wafer in the adjacent verification samples after passing the previous machine; for adjacent verification wafers in each verification sample, performing a first normalization process on the sensor data of the adjacent verification wafers on the current machine to obtain text features of the adjacent verification wafers, which are used as text features corresponding to the verification sample; and performing a second normalization process on the measurement data of the next verification wafer in the adjacent verification samples after passing the previous machine to obtain the verification sample pair. The first image feature is obtained by performing a second normalization process on the measurement data of the previous verification wafer after passing through the current machine in adjacent verification wafers to obtain the second image feature corresponding to the verification sample. The second normalization process is also performed on the measurement data of the next verification wafer after passing through the current machine in adjacent verification wafers to obtain the third image feature corresponding to the verification sample. After each training process, the text feature, the first image feature, and the second image feature corresponding to the verification sample are input into the measurement model to obtain the second wafer image output by the measurement model. The second loss between the second wafer image and the third image feature corresponding to the verification sample is calculated, and the model with the smallest second loss is selected as the measurement model.
[0009] In one possible implementation, the method further includes: removing samples from multiple training samples and multiple validation samples where the sensor data is constant or where the sensor data is missing greater than a preset threshold.
[0010] In one possible implementation, the first normalization process includes: calculating the average, maximum, minimum, and sum of sensor data for the training wafer at each process of the current machine, based on the number of processes on the current machine; filling missing values in the average of the sensor data for the training wafer at each process of the current machine using the result of averaging the average values of the training wafer at all processes; filling missing values in the maximum values of the sensor data for the training wafer at each process of the current machine using the average of the maximum values at all processes; filling missing values in the minimum values of the sensor data for the training wafer at each process of the current machine using the average of the minimum values at all processes; filling missing values in the sum of the sensor data for the training wafer at each process of the current machine using the average of the sums of the sums of the sensor data at each process of the current machine; and normalizing the filled average, maximum, minimum, and sum of the sensor data for each process of the current machine to obtain the text features of the training wafer.
[0011] In one possible implementation, the measurement data includes measurement values and measurement coordinates. The second normalization process includes: mapping the measurement data of the subsequent training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement coordinates of the subsequent training wafer after passing the current machine to a preset range to obtain processed measurement coordinates; and correspondingly mapping the processed measurement coordinates to the measurement values in the measurement data of the subsequent training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the subsequent training wafer after passing the current machine. For the missing values in the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine in adjacent training wafers, the average value of each is used to fill the gaps. Based on the number of measurement items, the maximum and minimum values of the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine in adjacent training wafers are calculated and normalized to obtain the first image feature, the second image feature, and the third image feature corresponding to the training sample.
[0012] In one possible implementation, the measurement model includes a text encoder, an image encoder, and a feature decoder; the text encoder is used to extract features from text features to obtain a text vector; the image encoder is used to extract features from image features to obtain an image vector; and the feature decoder is used to generate a wafer image based on the text vector and the image vector.
[0013] Secondly, embodiments of this application provide a measurement method, employing a measurement model obtained using the aforementioned measurement model training method. The measurement method includes: acquiring sensor data of the wafer under test and its adjacent preceding wafer at the current equipment station; measurement data of the wafer under test after passing the previous equipment station; and measurement data of the adjacent preceding wafer after passing the current equipment station; performing a first normalization process on the sensor data of the wafer under test and its adjacent preceding wafer at the current equipment station to obtain text features corresponding to the wafer under test; and ... The measurement data of the wafer after passing the previous machine and the measurement data of the adjacent wafer after passing the current machine are subjected to a second normalization process to obtain the image features corresponding to the wafer under test. The text features and image features corresponding to the wafer under test are input into the measurement model corresponding to the current machine, and the output wafer image is used as the measurement result. The measurement model corresponding to the current machine is a pre-trained multimodal model, which is used to predict the wafer image after passing the machine. The wafer image is wafer process data.
[0014] Thirdly, embodiments of this application provide a measurement device that uses a measurement model obtained by the aforementioned measurement model training method. The measurement device includes: a first acquisition module, used to acquire sensor data of the wafer under test and the adjacent previous wafer at the current machine, measurement data of the wafer under test after passing the previous machine, and measurement data of the adjacent previous wafer after passing the current machine; a first processing module, used to perform a first normalization process on the sensor data of the wafer under test and the adjacent previous wafer at the current machine to obtain text features corresponding to the wafer under test; The measurement data of the wafer under test after passing through the previous machine and the measurement data of the adjacent wafer after passing through the current machine are subjected to a second normalization process to obtain the image features corresponding to the wafer under test; the first input module is used to input the text features and image features corresponding to the wafer under test into the measurement model corresponding to the current machine, and use the output wafer image as the measurement result; wherein, the measurement model corresponding to the current machine is a pre-trained multimodal model, and the measurement model corresponding to the current machine is used to predict the wafer image after passing through the machine; wherein, the wafer image is wafer process data.
[0015] Fourthly, embodiments of this application provide an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.
[0016] Fifthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.
[0017] In a sixth aspect, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.
[0018] The measurement model training method, measurement method, apparatus, device, medium, and product provided in this application first acquire sensor data of the wafer under test and the adjacent wafer before the wafer under test at the current machine, measurement data of the wafer under test after passing the previous machine, and measurement data of the adjacent wafer before the wafer under test after passing the current machine. Then, the sensor data are subjected to a first normalization process to obtain text features corresponding to the wafer under test. The measurement data are subjected to a second normalization process to obtain image features corresponding to the wafer under test. Finally, the text features and image features corresponding to the wafer under test are input into the measurement model corresponding to the current machine, and the output wafer image is used as the measurement result. The solution of this application processes the acquired sensor data as text data to obtain text features, processes the acquired measurement data as image data to obtain image features, and inputs the text features and image features into a multimodal measurement model to obtain the wafer image output by the model. A multimodal image generation method is used to model the machine, enabling full-point measurement information prediction and improving the accuracy of semiconductor virtual measurement. Attached Figure Description
[0019] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0020] Figure 1 The diagram below illustrates a flow chart of the measurement method provided in Embodiment 1 of this application.
[0021] Figure 2 This is a schematic diagram of the wafer fabrication process.
[0022] Figure 3 The diagram above illustrates a flow chart of the measurement method provided in Embodiment 2 of this application.
[0023] Figure 4 This is a schematic diagram of the overall structure of the encoder;
[0024] Figure 5 This is a schematic diagram of the overall structure of the residual block;
[0025] Figure 6 The diagram below illustrates a schematic representation of the measuring device provided in Embodiment 4 of this application.
[0026] Figure 7 This is a schematic diagram of the structure of the electronic device provided in Embodiment 5 of this application.
[0027] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0028] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0029] The terms "comprising" and "having" in this application are used to indicate an open-ended inclusion, meaning that additional elements / components / etc. may exist besides the listed elements / components / etc.; the terms "first" and "second," etc., are used only as markings or distinctions and are not intended to limit the order or quantity of the objects. Furthermore, the different elements and areas in the accompanying drawings are only schematic and are therefore not limited to the dimensions or distances shown in the drawings. The technical solutions will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0030] In semiconductor manufacturing, metrology is a crucial way to ensure wafer quality. However, in actual metrology processes, measurements are only performed at certain points on certain wafers and at certain process steps, which cannot achieve comprehensive wafer quality monitoring. Furthermore, actual metrology using equipment consumes manpower and material resources, increasing production time costs.
[0031] Virtual measurement is a key technology that utilizes existing production data and employs data analysis and mining techniques to build models, predicting process parameters such as device dimensions and material thickness, and optimizing the production process. Using virtual measurement technology can reduce actual physical measurements, saving manpower and material costs; at the same time, it can promptly detect potential problems in the production process, improving production efficiency and effectiveness.
[0032] However, current virtual measurement methods typically rely on a single model, which has limitations in integrating and utilizing machine sensor parameters and existing actual measurement state parameters, resulting in insufficient accuracy. Furthermore, machine sensor parameters are usually time-series data, while existing data analysis methods, such as principal component analysis, do not fully consider the characteristics of both time-series and non-time-series data. This deficiency leads to a failure to fully capture the dynamic changes and complex relationships in the data during feature modeling and feature mining, limiting the model's performance and predictive capabilities.
[0033] The technical content provided in this application aims to solve the aforementioned technical problems in related technologies. In the embodiments of this application, sensor data of the wafer under test and its adjacent preceding wafer at the current testing station, measurement data of the wafer under test after passing the preceding testing station, and measurement data of the adjacent preceding wafer after passing the current testing station are first acquired. Then, the sensor data undergoes a first normalization process to obtain text features corresponding to the wafer under test; the measurement data undergoes a second normalization process to obtain image features corresponding to the wafer under test; finally, the text features and image features corresponding to the wafer under test are input into the measurement model corresponding to the current testing station, and the output wafer image is used as the measurement result. The solution of this application processes the acquired sensor data as text data to obtain text features, processes the acquired measurement data as image data to obtain image features, and inputs the text features and image features into a multimodal measurement model to obtain the wafer image output by the model. By employing a multimodal image generation method to model the testing station, full-point measurement information prediction is achieved, improving the accuracy of semiconductor virtual measurement.
[0034] Some aspects of this application's examples involve the above considerations. The following examples illustrate the proposed solutions.
[0035] Example 1
[0036] Figure 1 The diagram above illustrates a flow chart of the measurement method provided in Embodiment 1 of this application. The executing entity in this embodiment can be a measurement device, such as... Figure 1 As shown, the method includes:
[0037] Step 101: Acquire sensor data of the wafer under test and the previous wafer adjacent to the wafer under test on the current machine, measurement data of the wafer under test after passing the previous machine, and measurement data of the previous wafer adjacent to the wafer under test after passing the current machine.
[0038] Step 102: Perform a first normalization process on the sensor data of the wafer to be tested and the adjacent wafer of the wafer to be tested at the current machine to obtain the text features corresponding to the wafer to be tested; perform a second normalization process on the measurement data of the wafer to be tested after passing the previous machine and the measurement data of the adjacent wafer of the wafer to be tested after passing the current machine to obtain the image features corresponding to the wafer to be tested.
[0039] Step 103: Input the text features and image features corresponding to the wafer to be tested into the measurement model corresponding to the current machine, and use the output wafer image as the measurement result; wherein, the measurement model corresponding to the current machine is a pre-trained multi-modal model, and the measurement model corresponding to the current machine is used to predict the wafer image after the wafer passes through the machine; wherein, the wafer image is wafer process data.
[0040] In practical applications, the execution subject of this method can be a measuring device. There are various ways to implement the measuring device. For example, it can be implemented through a computer program, such as application software; or it can be implemented as a medium storing relevant computer programs, such as a USB flash drive or cloud drive; or it can be implemented through a physical device that integrates or installs relevant computer programs, such as a chip.
[0041] Specifically, semiconductor manufacturing is a complex and precise process involving multiple key stages aimed at transforming wafers into fully functional integrated circuits. The process begins by fabricating high-purity silicon into wafers, which are then processed through a series of equipment that perform steps such as photolithography, etching, doping, and thin-film deposition to build the basic structure of the circuit. After each processing step, the wafer is inspected by metrology equipment to ensure that parameters meet design specifications. These metrology steps are crucial for ensuring product quality because they can promptly detect and correct any deviations that may occur during manufacturing.
[0042] In one example, the wafer to be tested is set as wafer B, the adjacent wafer to the wafer to be tested is the wafer that passed through the process station and measurement station before the wafer to be tested, i.e., wafer A, the process station is process station 1 and process station 2 (the current station), and the measurement station is measurement station M. A Measuring machine M B . Figure 2 This is a schematic diagram of the wafer fabrication process. (Example:) Figure 2 As shown, wafer A passes through process equipment 1 and measurement equipment M. A—Process Equipment 2—Measuring Equipment M B Next, wafer B passes through process equipment 1 and measurement equipment M. A —Process Equipment 2—Measuring Equipment M B In this example, the sensor data of the wafer under test (DUT) and the preceding wafer adjacent to it on the current machine are the sensor data of wafer B and wafer A on process machine 2. This sensor data is the FDC (Failure Data Collection) data of wafers B and A on process machine 2, which can be directly obtained from process machine 2. The measurement data of the DUT after passing the preceding machine and the measurement data of the preceding wafer adjacent to it after passing the current machine (i.e., wafer B after passing process machine 1 and then measurement machine M) are also included. A The obtained measurement data and wafer A are then processed by the measurement machine M after passing through the process equipment 2. B The obtained measurement data can be obtained from the measurement machine M. A Measuring machine M B The measurement data includes wafer map data obtained by mapping measurement coordinates and measurement values.
[0043] Based on the previous example, firstly, acquire the sensor data of the wafer under test (wafer B) and the adjacent wafer before the wafer under test on the current machine (process machine 2) (FDC data of wafer B on process machine 2 and FDC data of wafer A on process machine 2), the measurement data of the wafer under test after passing the previous machine, and the measurement data of the adjacent wafer before the wafer under test after passing the current machine (wafer B after passing process machine 1 and then measurement machine M). A The obtained wafer map data and wafer A, after passing through process equipment 2, pass through measurement equipment M. B The obtained wafer map data. Then, the FDC data of wafer B on process equipment 2 is subjected to the first normalization process to obtain the first text feature corresponding to wafer B, and the FDC data of wafer A on process equipment 2 is subjected to the first normalization process to obtain the second text feature corresponding to wafer B; the data of wafer B after passing through process equipment 1 and then through measurement equipment M. A The obtained wafer map data undergoes a second normalization process to obtain the first image features corresponding to wafer B. Wafer A passes through process equipment 2 and then through measurement equipment M. BThe obtained wafer map data undergoes a second normalization process to obtain the second image features corresponding to wafer B. Finally, the first and second text features are used as text input, and the first and second image features are used as image input, fed into the pre-trained measurement model corresponding to process equipment 2. The output is the corresponding wafer map as the measurement result. The measurement model corresponding to process equipment 2 is a multimodal model used to predict the wafer image after passing through the equipment, where the wafer image is wafer process data. The multimodal model can simultaneously process sensor data from the process equipment (such as temperature, pressure, gas flow, etc.) and image data from the measurement equipment (such as wafer surface patterns and defects). By combining data from these different modalities, the model can identify complex patterns and relationships that might be undetectable by a single data source. This integration capability gives the multimodal model higher accuracy and robustness in prediction and analysis.
[0044] The measurement method provided in this embodiment first acquires sensor data of the wafer under test and the adjacent wafer before it on the current machine, measurement data of the wafer under test after passing the previous machine, and measurement data of the adjacent wafer before it after passing the current machine. Then, the sensor data is subjected to a first normalization process to obtain text features corresponding to the wafer under test. The measurement data is then subjected to a second normalization process to obtain image features corresponding to the wafer under test. Finally, the text features and image features corresponding to the wafer under test are input into the measurement model corresponding to the current machine, and the output wafer image is used as the measurement result. The solution of this application processes the acquired sensor data as text data to obtain text features, processes the acquired measurement data as image data to obtain image features, and inputs the text features and image features into a multimodal measurement model to obtain the wafer image output by the model. By employing a multimodal image generation method to model the machine, full-point measurement information prediction is achieved, improving the accuracy of semiconductor virtual measurement.
[0045] Example 2
[0046] Figure 3 The diagram above illustrates a flow chart of the measurement method provided in Embodiment 2 of this application. Figure 3 As shown, based on Embodiment 1, the method further includes:
[0047] Step 301: Obtain multiple training samples, wherein each training sample includes sensor data of adjacent training wafers on the current machine, measurement data of adjacent training wafers after passing the current machine, and measurement data of the next training wafer after passing the previous machine.
[0048] Step 302: For each training sample, perform a first normalization process on the sensor data of the adjacent training wafers on the current machine to obtain the text features corresponding to the training sample; perform a second normalization process on the measurement data of the next training wafer after passing the previous machine to obtain the first image features corresponding to the training sample; perform a second normalization process on the measurement data of the previous training wafer after passing the current machine to obtain the second image features corresponding to the training sample; and perform a second normalization process on the measurement data of the next training wafer after passing the current machine to obtain the third image features corresponding to the training sample.
[0049] Step 303: Based on the text features, first image features, second image features and third image features corresponding to multiple training samples, train the measurement model until the measurement model training is completed.
[0050] In this example, all wafers within the analysis period are selected from the current machine (the machine used for a specific process, i.e., the machine to be modeled). The analysis period can be set according to specific circumstances; for example, it can be a few months or a quarter of a year. After selecting all wafers within the analysis period, the wafers are sorted by processing time, and the top 80% (i.e., training wafers) are selected as training samples. Then, using the wafer ID as an index, sensor data from adjacent training wafers (usually two adjacent training wafers) are collected from the current machine, resulting in {'ID':{'Si':{t1:vi1,t2:vi2,t3:vi3,…,tj:vij}}}, where ID represents the current wafer number, Si represents sensor number i, tj represents the data acquisition time, and vij represents the data collected by sensor Si at time point tj. Measurement data is collected from the measurement stations that the adjacent training wafer passes through after the current machine and from the measurement stations that the next training wafer passes through after the previous machine (the process machine before the current process machine) to obtain {'ID':{'Pp':{vp1,vp2,vp3,…,vpn}}}, where ID represents the current wafer number, Pp represents the p-th measurement item, and vpn represents the measurement data of the n-th measurement point of the p-th measurement item. The measurement data includes measurement values and measurement coordinates. The measurement coordinates are {'ID':{'Pp':{(xp1,yp1),(xp2,yp2),(xp3,yp3),…,(xpn,ypn)}}}, where ID represents the current wafer number, Pp represents the p-th measurement item, and (xpn,ypn) represents the measurement coordinates of the n-th measurement point of the p-th measurement item.
[0051] Accordingly, after acquiring multiple training samples, for each sample's adjacent training wafers, the sensor data of the adjacent training wafers on the current machine undergoes a first normalization process to standardize the data range and eliminate dimensional differences. Through this processing step, the sensor data of adjacent training wafers are transformed into text features, which are then used as the corresponding text features for the training samples.
[0052] In one example, the first normalization process includes: calculating the average, maximum, minimum, and sum of the sensor data of the training wafer at each process of the current machine based on the number of processes of the current machine;
[0053] For missing values in the average value of sensor data for each process of the training wafer at the current machine, fill them with the result of averaging the average values of the training wafer at all processes; for missing values in the maximum value of sensor data for each process of the training wafer at the current machine, fill them with the average of the maximum values at all processes; for missing values in the minimum value of sensor data for each process of the training wafer at the current machine, fill them with the average of the minimum values at all processes; for missing values in the sum of sensor data for each process of the training wafer at the current machine, fill them with the average of the sums at all processes.
[0054] The average, maximum, minimum, and sum of the sensor data for each process of the current machine tool after filling are normalized to obtain the text features of the training wafer.
[0055] In this example, for a machine with i sensors and N process steps, the sampling frequency and sampling time of each sensor are different. The data is aggregated in four ways: average value (vmean), maximum value (vmax), minimum value (vmin), and summation (vsum) for each process step. Here, vmean = {vi1, vi2, ..., viN}, where viN represents the average value of the i-th sensor in the N-th process step; vmax = {vi1, vi2, ..., viN}; and vmax = {vi1, vi2, ..., viN}. ’ vi2 ’ ,…,viN ’},viN ’ This represents the maximum value of the i-th sensor in the N-th process; vmin = {vi1} ” vi2 ” ,…,viN ”},viN ” This represents the minimum value of the i-th sensor in the N-th process step; vsum = {vi1} ”’ vi2 ”’ ,…,viN ”’},viN ”’This represents the sum of values from the i-th sensor in the N-th process step.
[0056] Optionally, for missing values in the average value of a certain sensor in the above aggregated data, the missing values are filled using the average of the average values of all other processes of the training wafer under that sensor; for missing values in the maximum values in the above data, the missing values are filled using the average of the maximum values of all other processes of the sensor; for missing values in the minimum values in the above data, the missing values are filled using the average of the minimum values of all other processes of the sensor; for missing values in the summation values in the above data, the missing values are filled using the average of the summation values of all other processes of the sensor, making the data of equal length in the time dimension, with a length of N steps. For example, the training wafer C calculates the average value 1, maximum value 1, minimum value 1, and sum value 1 in the first sensor data; calculates the average value 2, maximum value 2, minimum value 2, and sum value 2 in the second sensor data; calculates the average value 3, maximum value 3, minimum value 3, and sum value 3 in the third sensor data… Due to process differences or sensor malfunctions, the data corresponding to the 5th step of the second sensor may be missing. Therefore, the average value, maximum value, minimum value, and sum value of the training wafer C in the 5th step of the second sensor data are missing. Then: fill the missing values for the training wafer C in the 5th step… The average value of the fifth step in the data from the two sensors is the result of averaging the average value of the other N-1 steps of the second sensor. The maximum value of the fifth step in the data from the second sensor for the training wafer C is the average value of the maximum value of the other N-1 steps of the second sensor. The minimum value of the fifth step in the data from the second sensor for the training wafer C is the average value of the minimum value of the other N-1 steps of the second sensor. The sum of the values of the fifth step in the data from the second sensor for the training wafer C is the average value of the sum of the values of the other N-1 steps of the second sensor.
[0057] Correspondingly, for the mean vmean, maximum vmax, minimum vmin, and sum vsum, the maximum and minimum values are normalized according to the formula (viN1-min(viN)) / (max(viN),min(viN)) to obtain the text features T∈R. 4×i×N Where max(viN) and min(viN) represent the maximum and minimum values of the i-th sensor data after aggregation by process step, respectively, in the average / maximum / minimum / summary values; viN1 represents the average / maximum / minimum / summary value of the i-th sensor in the N-th process step. After the sensor data of adjacent training wafers on the current machine are processed by the above steps, the first text feature T1∈R is obtained respectively. 4×i×N Second text feature T2∈R4×i×N Where i represents the i-th sensor; N represents the N-th process step.
[0058] In the example above, the first normalization process provides the model with more stable and higher quality input data, thereby improving the model's prediction accuracy and reliability.
[0059] Accordingly, after acquiring multiple training samples, for each adjacent training wafer in the sample, the measurement data of the subsequent training wafer after passing through the previous machine is subjected to a second normalization process to obtain the first image feature corresponding to the training sample. The measurement data of the previous training wafer after passing through the current machine is subjected to a second normalization process to obtain the second image feature corresponding to the training sample. The measurement data of the subsequent training wafer after passing through the current machine is subjected to a second normalization process to obtain the third image feature corresponding to the training sample. This process standardizes the range of the data and eliminates dimensional differences. Through this processing step, the measurement data of adjacent training wafers are transformed into image features, which are then used as the corresponding image features of the training samples.
[0060] In one example, the measurement data includes measurement values and measurement coordinates. The second normalization process includes:
[0061] The measurement coordinates in the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine are mapped to a preset range to obtain processed measurement coordinates. The processed measurement coordinates are then matched with the measurement values in the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine, respectively.
[0062] For the missing values in the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine, the average value of each is used to fill them.
[0063] Based on the number of measurement items, the maximum and minimum values of the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine are calculated and normalized to obtain the first image feature, second image feature, and third image feature corresponding to the training sample.
[0064] In this example, the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine all include measurement values and measurement coordinates. The measurement points of the p measurement items on the wafer correspond to different coordinates. These measurement coordinates are uniformly mapped to a fixed range of 0 to s. Then, based on the measurement points, the measurement values and measurement coordinates are mapped to construct a wafer map, resulting in {'ID':{v}. (xp1,yp1) ,v (xp2,yp2) ,…,v (xpn,ypn)}}, where ID represents the current wafer number, the subscript (xpn, ypn) represents the coordinates of the nth measurement point of the p-th measurement item, and v (xpn,ypn) This represents the measurement value of the p-th measurement item at coordinates (xpn, ypn).
[0065] Optionally, for missing values in the measurement data of the next training wafer after passing the previous machine in an adjacent training wafer, the average value is used to fill the missing values, or radial basis fill is used to fill the missing values; for missing values in the measurement data of the previous training wafer after passing the current machine in an adjacent training wafer, the average value is used to fill the missing values, or radial basis fill is used to fill the missing values; for missing values in the measurement data of the next training wafer after passing the current machine in an adjacent training wafer, the average value is used to fill the missing values, or radial basis fill is used to fill the missing values.
[0066] Accordingly, for p measurement items, according to the formula (v (pxn,pyn) -min(v (pxn,pyn) )) / (max(v (pxn,pyn) ),min(v (pxn,pyn) Normalize the maximum and minimum values respectively to obtain the image features M∈R. p×s×s , where max(v (pxn,pyn) ),min(v (pxn,pyn) ) represent the maximum and minimum values of the p-th measurement item in the dataset at coordinates (xpn, ypn), respectively. (xpn,ypn) Let represent the measurement value of the p-th measurement item at coordinates (xpn, ypn), and s represent the maximum value after the measurement coordinates are mapped. The measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine, after being processed by the above steps, respectively, yield the first image feature M1∈R. p×s×s Second image feature M2∈R p×s×s Third image feature M3∈R p×s×sWhere M1 and M2 are the input image features, and M3 is the output image feature.
[0067] In the example above, the second normalization process provides the model with more stable and higher quality input data, thereby improving the model's prediction accuracy and reliability.
[0068] Based on the previous example, after obtaining the text features, first image features, second image features, and third image features corresponding to multiple training samples, the measurement model is trained until the training is complete. In one example, the training process includes:
[0069] Input the text features, first image features, and second image features corresponding to the training samples into the measurement model to obtain the first wafer image output by the measurement model.
[0070] The first loss between the first wafer image and the third image features corresponding to the training samples is fed back to the measurement model until the first loss is less than the first threshold, at which point the measurement model is considered to have completed training.
[0071] In this example, the text features, first image features, and second image features corresponding to the training samples are input into the measurement model. The text features, derived from the normalization of sensor data, represent key parameters of the wafer at different processing stages. The first and second image features, derived from different image processing steps, provide visual information about the wafer surface state. By inputting these multimodal features into the measurement model, the model can comprehensively analyze different types of data to generate a predicted first wafer image. This process leverages the advantages of multimodal data fusion, enabling the model to capture complex patterns and relationships that might be missed by a single data source, thereby improving prediction accuracy.
[0072] Correspondingly, the first wafer image output by the measurement model is compared with the corresponding third image features in the training samples to calculate the first loss. The first loss uses the mean squared error (MSE) loss function to measure the difference between the predicted wafer image and the actual wafer image. The MSE loss provides a standard for measuring prediction accuracy by calculating the squared difference between the predicted and actual values for each pixel and taking the average. The smaller the loss value, the closer the model's prediction is to the reality. The calculated MSE loss is fed back to the measurement model until the first loss is less than a first threshold, at which point the measurement model is considered to have completed training. The size of the first threshold is set according to actual conditions and is not limited here. The model parameters are adjusted using the backpropagation algorithm. The loss calculation is based only on points with measurement values, and the loss is fed back to the measurement model. This process continuously optimizes the model, enabling it to more accurately reproduce the actual state of the wafer in subsequent predictions. Through this iterative optimization, the model gradually improves its predictive ability, providing important support for quality control and process optimization in semiconductor manufacturing.
[0073] In one example, the measurement model includes a text encoder, an image encoder, and a feature decoder;
[0074] The text encoder is used to extract features from text to obtain text vectors; the image encoder is used to extract features from image to obtain image vectors; and the feature decoder is used to generate wafer images based on the text vectors and image vectors.
[0075] Specifically, the measurement model includes two text encoders to extract text features, two image encoders to extract image features, and one feature decoder to generate a prediction wafer map. The text encoders, for the aforementioned text features T1 and T2, use a residual network as their basic architecture, employing one-dimensional convolution to extract features, performing non-linear mapping via ReLU activation, and then downsampling via average pooling before outputting the result. The specific structure includes a cascaded input layer, a residual block, and an average pooling layer to build the text encoder, resulting in vector T. v1 = [v1, v2, ..., v N ]、T v2 = [v1, v2, ..., v N ], where v N The text encoder represents the high-dimensional features extracted. The image encoder, for the aforementioned image features M1 and M2, uses a residual network as its basic architecture, employs two-dimensional convolution to extract features, performs non-linear mapping using the ReLU activation function, and outputs the features after downsampling via average pooling. The specific structure includes a cascaded input layer, residual block, and average pooling layer to build the image encoder, obtaining vector M. v1 =[v1 ’ v2 ’,…,v N ’ ]、M v2 =[v1 ’ v2 ’ ,…,v N ’ ], where v N ’ The high-dimensional features extracted by the image encoder represent the output vector T of the text encoder. v1 and T v2 By concatenating them together, we obtain the text concatenation vector T. v The output vector M of the image encoder v1 M v2 By concatenating the images together, we obtain the image concatenation vector M. v Then T v and M v By splicing them together, we obtain the fused feature T. Mv For the fusion feature T Mv Based on a residual network architecture, a feature decoder is constructed by recovering features layer by layer using two-dimensional convolutions and controlling the output range between 0 and 1 using the sigmoid function. Figure 4 This is a schematic diagram of the overall structure of the encoder. Figure 5 This is a schematic diagram of the overall structure of the residual block.
[0076] In the example above, the measurement model integrates a text encoder, an image encoder, and a feature decoder to achieve efficient processing and fusion of multimodal data, enabling prediction of measurement information at all locations.
[0077] Based on the aforementioned example, the method further includes: acquiring multiple verification samples, wherein each verification sample includes sensor data of adjacent verification wafers on the current machine, measurement data of adjacent verification wafers after passing the current machine, and measurement data of the next verification wafer in the adjacent verification wafers after passing the previous machine.
[0078] For each verification sample, the sensor data of the adjacent verification wafers on the current machine are subjected to a first normalization process to obtain the text features of the adjacent verification wafers, which are used as the text features corresponding to the verification sample. The measurement data of the next verification wafer after passing through the previous machine is subjected to a second normalization process to obtain the first image features corresponding to the verification sample. The measurement data of the previous verification wafer after passing through the current machine is subjected to a second normalization process to obtain the second image features corresponding to the verification sample. The measurement data of the next verification wafer after passing through the current machine is subjected to a second normalization process to obtain the third image features corresponding to the verification sample.
[0079] After each training process, the text features, first image features, and second image features corresponding to the verification sample are input into the measurement model to obtain the second wafer image output by the measurement model.
[0080] Calculate the second loss between the features of the second wafer image and the third image corresponding to the verification sample, and select the model with the smallest second loss as the measurement model.
[0081] In this example, after selecting all wafers within the analysis period, the wafers are sorted according to their processing time, and 5% of the wafers (i.e., verification wafers) are selected as the verification sample. Then, using the wafer ID as an index, sensor data of adjacent verification wafers (usually two adjacent verification wafers) are collected in the current machine, resulting in {'ID”:{'Si”:{t1 ’ :vi1 ’ ,t2 ’ :vi2 ’ ,t3 ’ :vi3 ’ ,…,tj ’ :vij ’}}}, where ID ’ Represents the current wafer number, Si ’ Representing sensor number i, tj ’ Represents the time of data collection, vij ’ The representative sensor Si in tj ’ Data collected at specific time points. Measurement data is collected from the measurement stations that adjacent verification wafers pass through after the current machine and from the measurement stations that the next verification wafer passes through after the previous machine (the process machine preceding the current process machine), resulting in {'ID”:{'Pp”:{vp1 ’ ,vp2 ’ vp3 ’ VPN ’}}}, where ID ’ Pp represents the current wafer number. ’ Represents the p-th measurement item, VPN ’ This represents the measurement data of the nth measurement point of the p-th measurement item. The measurement data includes the measurement value and measurement coordinates, where the measurement coordinates are {'ID”:{'Pp”:{(xp1)}. ’ yp1 ’ ),(xp2 ’ yp2 ’ (xp3) ’ yp3 ’ ),…,(xpn ’ ypn ’ )}}}, where ID ’Pp represents the current wafer number. ’ Represents the p-th measurement term, (xpn) ’ ypn ’ ) represents the measurement coordinates of the nth measurement point of the pth measurement item.
[0082] Optionally, after acquiring multiple verification samples, for the sensor data of adjacent verification wafers in each verification sample on the current machine, the aforementioned first normalization process is performed to obtain the text features of the adjacent verification wafers, which are used as the text features corresponding to the verification sample. For the measurement data of the next verification wafer after passing the previous machine in the adjacent verification wafers, the aforementioned second normalization process is performed to obtain the first image features corresponding to the verification sample. For the measurement data of the previous verification wafer after passing the current machine in the adjacent verification wafers, the aforementioned second normalization process is performed to obtain the second image features corresponding to the verification sample. For the measurement data of the next verification wafer after passing the current machine in the adjacent verification wafers, the aforementioned second normalization process is performed to obtain the third image features corresponding to the verification sample, so as to standardize the range of data and eliminate dimensional differences.
[0083] Accordingly, after every 100 iterations, the text features, first image features, and second image features corresponding to the validation sample are input into the measurement model to obtain the second wafer image output by the measurement model. The second loss between the second wafer image and the third image features corresponding to the validation sample is calculated, and the model with the smallest second loss is selected as the measurement model. In the above example, by setting validation samples to validate the measurement model, the prediction accuracy and reliability of the model are improved.
[0084] Building upon the aforementioned example, the method further includes removing samples from multiple training and validation samples where sensor data is constant or where sensor data is missing beyond a preset threshold. Specifically, for data in the training and validation samples where sensor data is constant, where the proportion of missing sensor data due to sensor malfunction exceeds 80%, and for anomalous data, all data from that wafer (including FDC data and wafer measurement data) is removed. Samples with constant sensor data typically lack useful information because constant values do not provide any meaningful variation information about process changes or equipment status, which may prevent the model from learning important features and patterns in the data. A sample with more than a preset threshold of missing sensor data may result in incomplete data, affecting the model's training and validation process. This missing data may introduce noise and bias, making it difficult for the model to accurately capture the true relationships within the data.
[0085] In the example above, by removing these low-quality samples, the quality of the dataset was significantly improved, enabling the model to focus on learning from more representative and informative samples, thus improving the model's robustness and stability.
[0086] In the measurement method provided in this embodiment, the measurement model is trained using training samples, and then validated using validation samples after training, until the measurement model training is completed, thereby improving the prediction accuracy and reliability of the measurement model.
[0087] Example 3
[0088] The measurement method provided in this application will be described in detail below with a specific embodiment.
[0089] 1. Data Collection:
[0090] Data from a specific process equipment for a certain product from February 2024 to September 2024 were selected, including the FDC data of the wafer under test and the adjacent wafer before the wafer under test on the current process equipment, the measurement data of the wafer under test after passing through the previous process equipment, and the measurement data of the adjacent wafer before the wafer under test after passing through the current process equipment; wherein, the measurement data includes measurement values and measurement coordinates.
[0091] 2. Data processing and feature construction:
[0092] (1) The FDC data consists of 225 sensors and 9 process steps. After deleting samples with more than 80% missing values, 140 sensors are retained. Each sensor has a different sampling frequency and sampling time. Following the first normalization method in the previous example, the data is aggregated by step, with 4 aggregation types. The missing values are filled with the mean, and the maximum and minimum values are normalized to obtain the first text feature T1∈R. 4×140×9 Second text feature T2∈R 4×140×9 ;
[0093] (2) The wafer measurement data consists of 5 measurement items. Following the second normalization method in the previous example, the coordinates are mapped to the range of 0 to 63 and corresponded to the measurement values to construct a wafer map. The missing values are filled using the radial basis method, and then the maximum and minimum values are normalized to obtain the first image feature M1∈R. 5×64×64 Second image feature M2∈R 5×64×64 Third image feature M3∈R 5×64×64 .
[0094] 3. Model training and prediction:
[0095] (1) Sort the 8013 samples according to the wafer processing time, and take the first 6410 samples as training samples to input into the measurement model.
[0096] (2) Take the first 400 of the remaining samples as validation samples. Input the validation samples into the current measurement model every 100 iterations to verify the model effect.
[0097] (3) Select the model with the minimum loss as the trained measurement model and save it. In the experiment, the minimum loss was about 0.0023.
[0098] (4) The remaining 1203 samples were input into the trained measurement model to obtain the true value, predicted value and r2 score of each measurement item. The r2 scores of the five measurement items were all greater than 0.79, and the regression results were highly reliable.
[0099] The specific measurement methods can be found in the foregoing embodiments. In summary, the measurement method provided in this example processes the acquired sensor data as text data to obtain text features, processes the acquired measurement data as image data to obtain image features, and inputs the text features and image features into a multimodal measurement model to obtain the wafer image output by the model. By employing a multimodal image generation method to model the equipment, full-point measurement information prediction is achieved, improving the accuracy of semiconductor virtual measurement.
[0100] Example 4
[0101] Figure 6 The diagram below illustrates the structure of the measuring device provided in Embodiment 4 of this application. Figure 6 As shown, the device includes:
[0102] The first acquisition module 61 is used to acquire sensor data of the wafer under test and the previous wafer adjacent to the wafer under test at the current machine, measurement data of the wafer under test after passing the previous machine, and measurement data of the previous wafer adjacent to the wafer under test after passing the current machine.
[0103] The first processing module 62 is used to perform a first normalization process on the sensor data of the wafer under test and the adjacent wafer before the wafer under test at the current machine to obtain the text features corresponding to the wafer under test; and to perform a second normalization process on the measurement data of the wafer under test after passing the previous machine and the measurement data of the adjacent wafer before the wafer under test after passing the current machine to obtain the image features corresponding to the wafer under test.
[0104] The first input module 63 is used to input the text features and image features corresponding to the wafer to be tested into the measurement model corresponding to the current machine, and to output the wafer image as the measurement result; wherein, the measurement model corresponding to the current machine is a pre-trained multi-modal model, and the measurement model corresponding to the current machine is used to predict the wafer image after the wafer passes through the machine; wherein, the wafer image is wafer process data.
[0105] In practical applications, this measuring device can be implemented in various ways. For example, it can be implemented through computer programs, such as application software; or it can be implemented as a medium storing relevant computer programs, such as a USB flash drive or cloud drive; or it can be implemented through a physical device that integrates or installs relevant computer programs, such as a chip.
[0106] Specifically, semiconductor manufacturing is a complex and precise process involving multiple key stages aimed at transforming wafers into fully functional integrated circuits. The process begins by fabricating high-purity silicon into wafers, which are then processed through a series of equipment that perform steps such as photolithography, etching, doping, and thin-film deposition to build the basic structure of the circuit. After each processing step, the wafer is inspected by metrology equipment to ensure that parameters meet design specifications. These metrology steps are crucial for ensuring product quality because they can promptly detect and correct any deviations that may occur during manufacturing.
[0107] In one example, the wafer to be tested is set as wafer B, the adjacent wafer to the wafer to be tested is the wafer that passed through the process station and measurement station before the wafer to be tested, i.e., wafer A, the process station is process station 1 and process station 2 (the current station), and the measurement station is measurement station M. A Measuring machine M B . Figure 2 This is a schematic diagram of the wafer fabrication process. (Example:) Figure 2 As shown, wafer A passes through process equipment 1 and measurement equipment M. A —Process Equipment 2—Measuring Equipment M B Next, wafer B passes through process equipment 1 and measurement equipment M. A —Process Equipment 2—Measuring Equipment M B In this example, the sensor data of the wafer under test (DUT) and the preceding wafer adjacent to it on the current machine are the sensor data of wafer B and wafer A on process machine 2. This sensor data is the FDC (Failure Data Collection) data of wafers B and A on process machine 2, which can be directly obtained from process machine 2. The measurement data of the DUT after passing the preceding machine and the measurement data of the preceding wafer adjacent to it after passing the current machine (i.e., wafer B after passing process machine 1 and then measurement machine M) are also included. A The obtained measurement data and wafer A are then processed by the measurement machine M after passing through the process equipment 2. B The obtained measurement data can be obtained from the measurement machine M. A Measuring machine M B The measurement data includes wafer map data obtained by mapping measurement coordinates and measurement values.
[0108] Based on the previous example, firstly, acquire the sensor data of the wafer under test (wafer B) and the adjacent wafer before the wafer under test on the current machine (process machine 2) (FDC data of wafer B on process machine 2 and FDC data of wafer A on process machine 2), the measurement data of the wafer under test after passing the previous machine, and the measurement data of the adjacent wafer before the wafer under test after passing the current machine (wafer B after passing process machine 1 and then measurement machine M). A The obtained wafer map data and wafer A, after passing through process equipment 2, pass through measurement equipment M. B The obtained wafer map data. Then, the FDC data of wafer B on process equipment 2 is subjected to the first normalization process to obtain the first text feature corresponding to wafer B, and the FDC data of wafer A on process equipment 2 is subjected to the first normalization process to obtain the second text feature corresponding to wafer B; the data of wafer B after passing through process equipment 1 and then through measurement equipment M. A The obtained wafer map data undergoes a second normalization process to obtain the first image features corresponding to wafer B. Wafer A passes through process equipment 2 and then through measurement equipment M. B The obtained wafer map data undergoes a second normalization process to obtain the second image features corresponding to wafer B. Finally, the first and second text features are used as text input, and the first and second image features are used as image input, fed into the pre-trained measurement model corresponding to process equipment 2. The output is the corresponding wafer map as the measurement result. The measurement model corresponding to process equipment 2 is a multimodal model used to predict the wafer image after passing through the equipment, where the wafer image is wafer process data. The multimodal model can simultaneously process sensor data from the process equipment (such as temperature, pressure, gas flow, etc.) and image data from the measurement equipment (such as wafer surface patterns and defects). By combining data from these different modalities, the model can identify complex patterns and relationships that might be undetectable by a single data source. This integration capability gives the multimodal model higher accuracy and robustness in prediction and analysis.
[0109] Based on the aforementioned example, the device further includes:
[0110] The second acquisition module is used to acquire multiple training samples, wherein each training sample includes sensor data of adjacent training wafers on the current machine, measurement data of adjacent training wafers after passing the current machine, and measurement data of the next training wafer after passing the previous machine.
[0111] The second processing module is used to perform a first normalization process on the sensor data of adjacent training wafers in the current machine to obtain the text features corresponding to the training sample for each training sample; perform a second normalization process on the measurement data of the next training wafer after passing the previous machine to obtain the first image features corresponding to the training sample; perform a second normalization process on the measurement data of the previous training wafer after passing the current machine to obtain the second image features corresponding to the training sample; and perform a second normalization process on the measurement data of the next training wafer after passing the current machine to obtain the third image features corresponding to the training sample.
[0112] The training module is used to train the measurement model based on the text features, first image features, second image features, and third image features corresponding to multiple training samples until the measurement model training is completed.
[0113] In this example, all wafers within the analysis period are selected from the current machine (the machine used for a specific process, i.e., the machine to be modeled). The analysis period can be set according to specific circumstances; for example, it can be a few months or a quarter of a year. After selecting all wafers within the analysis period, the wafers are sorted by processing time, and the top 80% (i.e., training wafers) are selected as training samples. Then, using the wafer ID as an index, sensor data from adjacent training wafers (usually two adjacent training wafers) are collected from the current machine, resulting in {'ID':{'Si':{t1:vi1,t2:vi2,t3:vi3,…,tj:vij}}}, where ID represents the current wafer number, Si represents sensor number i, tj represents the data acquisition time, and vij represents the data collected by sensor Si at time point tj. Measurement data is collected from the measurement stations that the adjacent training wafer passes through after the current machine and from the measurement stations that the next training wafer passes through after the previous machine (the process machine before the current process machine) to obtain {'ID':{'Pp':{vp1,vp2,vp3,…,vpn}}}, where ID represents the current wafer number, Pp represents the p-th measurement item, and vpn represents the measurement data of the n-th measurement point of the p-th measurement item. The measurement data includes measurement values and measurement coordinates. The measurement coordinates are {'ID':{'Pp':{(xp1,yp1),(xp2,yp2),(xp3,yp3),…,(xpn,ypn)}}}, where ID represents the current wafer number, Pp represents the p-th measurement item, and (xpn,ypn) represents the measurement coordinates of the n-th measurement point of the p-th measurement item.
[0114] Accordingly, after acquiring multiple training samples, for each sample's adjacent training wafers, the sensor data of the adjacent training wafers on the current machine undergoes a first normalization process to standardize the data range and eliminate dimensional differences. Through this processing step, the sensor data of adjacent training wafers are transformed into text features, which are then used as the corresponding text features for the training samples.
[0115] In one example, the second processing module is specifically used to: calculate the average, maximum, minimum, and sum of the sensor data of the training wafer under each process of the current machine based on the number of processes of the current machine;
[0116] For missing values in the average value of sensor data for each process of the training wafer at the current machine, fill them with the result of averaging the average values of the training wafer at all processes; for missing values in the maximum value of sensor data for each process of the training wafer at the current machine, fill them with the average of the maximum values at all processes; for missing values in the minimum value of sensor data for each process of the training wafer at the current machine, fill them with the average of the minimum values at all processes; for missing values in the sum of sensor data for each process of the training wafer at the current machine, fill them with the average of the sums at all processes.
[0117] The average, maximum, minimum, and sum of the sensor data for each process of the current machine tool after filling are normalized to obtain the text features of the training wafer.
[0118] In this example, for a machine with i sensors and N process steps, the sampling frequency and sampling time of each sensor are different. The data is aggregated in four ways: average value (vmean), maximum value (vmax), minimum value (vmin), and summation (vsum) for each process step. Here, vmean = {vi1, vi2, ..., viN}, where viN represents the average value of the i-th sensor in the N-th process step; vmax = {vi1, vi2, ..., viN}; and vmax = {vi1, vi2, ..., viN}. ’ vi2 ’ ,…,viN ’},viN ’ This represents the maximum value of the i-th sensor in the N-th process; vmin = {vi1} ” vi2 ” ,…,viN ”},viN ” This represents the minimum value of the i-th sensor in the N-th process step; vsum = {vi1} ”’ vi2 ”’ ,…,viN ”’},viN ”’This represents the sum of values from the i-th sensor in the N-th process step.
[0119] Optionally, for missing values in the average values of the above data, fill them with the average of the average values of the training wafer under all processes; for missing values in the maximum values of the above data, fill them with the average of the maximum values under all processes; for missing values in the minimum values of the above data, fill them with the average of the minimum values under all processes; for missing values in the summation values of the above data, fill them with the average of the summation values under all processes, making the data of equal length in the time dimension, with a length of N steps.
[0120] Correspondingly, for the mean vmean, maximum vmax, minimum vmin, and sum vsum, the maximum and minimum values are normalized according to the formula (viN1-min(viN)) / (max(viN),min(viN)) to obtain the text features T∈R. 4×i×N Where max(viN) and min(viN) represent the maximum and minimum values of the i-th sensor data in the dataset after aggregation by process step, respectively; viN1 represents the average / maximum / minimum / sum value of the i-th sensor in the N-th process. After the sensor data of adjacent training wafers on the current machine are processed by the above steps, the first text feature T1∈R is obtained respectively. 4 ×i×N Second text feature T2∈R 4×i×N Where i represents the i-th sensor; N represents the N-th process step.
[0121] In the example above, the first normalization process provides the model with more stable and higher quality input data, thereby improving the model's prediction accuracy and reliability.
[0122] Accordingly, after acquiring multiple training samples, for each adjacent training wafer in a sample, the measurement data of the subsequent training wafer after passing through the previous machine is subjected to a second normalization process to obtain the first image feature corresponding to the training sample; the measurement data of the previous training wafer after passing through the current machine is subjected to a second normalization process to obtain the second image feature corresponding to the training sample; and the measurement data of the subsequent training wafer after passing through the current machine is subjected to a second normalization process to obtain the third image feature corresponding to the training sample. This is done to standardize the range of the data and eliminate dimensional differences. By converting the measurement data of adjacent training wafers into image features, these image features are subsequently used as the corresponding image features of the training samples.
[0123] In one example, the measurement data includes measurement values and measurement coordinates. The second processing module is specifically used for:
[0124] The measurement coordinates in the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine are mapped to a preset range to obtain processed measurement coordinates. The processed measurement coordinates are then matched with the measurement values in the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine, respectively.
[0125] For the missing values in the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine, the average value of each is used to fill them.
[0126] Based on the number of measurement items, the maximum and minimum values of the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine are calculated and normalized to obtain the first image feature, second image feature, and third image feature corresponding to the training sample.
[0127] In this example, the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine all include measurement values and measurement coordinates. The measurement points of the p measurement items on the wafer correspond to different coordinates. These measurement coordinates are uniformly mapped to a fixed range of 0 to s. Then, based on the measurement points, the measurement values and measurement coordinates are mapped to construct a wafer map, resulting in {'ID':{v}. (xp1,yp1) ,v (xp2,yp2) ,…,v (xpn,ypn)}}, where ID represents the current wafer number, the subscript (xpn, ypn) represents the coordinates of the nth measurement point of the p-th measurement item, and v (xpn,ypn) This represents the measurement value of the p-th measurement item at coordinates (xpn, ypn).
[0128] Optionally, for missing values in the measurement data of the next training wafer after passing the previous machine in an adjacent training wafer, the average value is used to fill the missing values, or radial basis fill is used to fill the missing values; for missing values in the measurement data of the previous training wafer after passing the current machine in an adjacent training wafer, the average value is used to fill the missing values, or radial basis fill is used to fill the missing values; for missing values in the measurement data of the next training wafer after passing the current machine in an adjacent training wafer, the average value is used to fill the missing values, or radial basis fill is used to fill the missing values.
[0129] Accordingly, for p measurement items, according to the formula (v (pxn,pyn) -min(v (pxn,pyn) )) / (max(v (pxn,pyn) ),min(v (pxn,pyn) Normalize the maximum and minimum values respectively to obtain the image features M∈R. p×s×s , where max(v (pxn,pyn) ),min(v (pxn,pyn) ) represent the maximum and minimum values of the p-th measurement item in the dataset at coordinates (xpn, ypn), respectively. (xpn,ypn) Let represent the measurement value of the p-th measurement item at coordinates (xpn, ypn), and s represent the maximum value after the measurement coordinates are mapped. The measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine, after being processed by the above steps, respectively, yield the first image feature M1∈R. p×s×s Second image feature M2∈R p×s×s Third image feature M3∈R p×s×s Where M1 and M2 are the input image features, and M3 is the output image feature.
[0130] In the example above, the second normalization process provides the model with more stable and higher quality input data, thereby improving the model's prediction accuracy and reliability.
[0131] Based on the previous example, after obtaining the text features, first image features, second image features, and third image features corresponding to multiple training samples, the measurement model is trained until training is complete. In one example, the training module is specifically used for:
[0132] Input the text features, first image features, and second image features corresponding to the training samples into the measurement model to obtain the first wafer image output by the measurement model.
[0133] The first loss between the first wafer image and the third image features corresponding to the training samples is fed back to the measurement model until the first loss is less than the first threshold, at which point the measurement model is considered to have completed training.
[0134] In this example, the text features, first image features, and second image features corresponding to the training samples are input into the measurement model. The text features, derived from the normalization of sensor data, represent key parameters of the wafer at different processing stages. The first and second image features, derived from different image processing steps, provide visual information about the wafer surface state. By inputting these multimodal features into the measurement model, the model can comprehensively analyze different types of data to generate a predicted first wafer image. This process leverages the advantages of multimodal data fusion, enabling the model to capture complex patterns and relationships that might be missed by a single data source, thereby improving prediction accuracy.
[0135] Correspondingly, the first wafer image output by the measurement model is compared with the corresponding third image features in the training samples to calculate the first loss. The first loss uses the mean squared error (MSE) loss function to measure the difference between the predicted wafer image and the actual wafer image. The MSE loss provides a standard for measuring prediction accuracy by calculating the squared difference between the predicted and actual values for each pixel and taking the average. The smaller the loss value, the closer the model's prediction is to the reality. The calculated MSE loss is fed back to the measurement model until the first loss is less than a first threshold, at which point the measurement model is considered to have completed training. The size of the first threshold is set according to actual conditions and is not limited here. The model parameters are adjusted using the backpropagation algorithm. The loss calculation is based only on points with measurement values, and the loss is fed back to the measurement model. This process continuously optimizes the model, enabling it to more accurately reproduce the actual state of the wafer in subsequent predictions. Through this iterative optimization, the model gradually improves its predictive ability, providing important support for quality control and process optimization in semiconductor manufacturing.
[0136] In one example, the measurement model includes a text encoder, an image encoder, and a feature decoder;
[0137] The text encoder is used to extract features from text to obtain text vectors; the image encoder is used to extract features from image to obtain image vectors; and the feature decoder is used to generate wafer images based on the text vectors and image vectors.
[0138] Specifically, the measurement model includes two text encoders to extract text features, two image encoders to extract image features, and one feature decoder to generate a prediction wafer map. The text encoders, for the aforementioned text features T1 and T2, use a residual network as their basic architecture, employing one-dimensional convolution to extract features, performing non-linear mapping via ReLU activation, and then downsampling via average pooling before outputting the result. The specific structure includes a cascaded input layer, a residual block, and an average pooling layer to build the text encoder, resulting in vector T. v1 = [v1, v2, ..., v N ]、T v2 = [v1, v2, ..., v N ], where v N The text encoder represents the high-dimensional features extracted. The image encoder, for the aforementioned image features M1 and M2, uses a residual network as its basic architecture, employs two-dimensional convolution to extract features, performs non-linear mapping using the ReLU activation function, and outputs the features after downsampling via average pooling. The specific structure includes a cascaded input layer, residual block, and average pooling layer to build the image encoder, obtaining vector M. v1 =[v1 ’ v2 ’ ,…,v N ’ ]、M v2 =[v1 ’ v2 ’ ,…,v N ’ ], where v N ’ The high-dimensional features extracted by the image encoder represent the output vector T of the text encoder. v1 and T v2 By concatenating them together, we obtain the text concatenation vector T. v The output vector M of the image encoder v1 M v2 By concatenating the images together, we obtain the image concatenation vector M. v Then T v and M v By splicing them together, we obtain the fused feature T. Mv For the fusion feature T Mv Based on a residual network architecture, a feature decoder is constructed by recovering features layer by layer using two-dimensional convolutions and controlling the output range between 0 and 1 using the sigmoid function. Figure 4 This is a schematic diagram of the overall structure of the encoder. Figure 5 This is a schematic diagram of the overall structure of the residual block.
[0139] In the example above, the measurement model integrates a text encoder, an image encoder, and a feature decoder to achieve efficient processing and fusion of multimodal data, enabling prediction of measurement information at all locations.
[0140] Based on the aforementioned example, the device further includes:
[0141] The third acquisition module is used to acquire multiple verification samples, wherein each verification sample includes sensor data of adjacent verification wafers on the current machine, measurement data of adjacent verification wafers after passing the current machine, and measurement data of the next verification wafer after passing the previous machine.
[0142] The third processing module is used to perform a first normalization process on the sensor data of adjacent verification wafers in the current machine for each verification sample to obtain the text features of the adjacent verification wafers, which are used as the text features corresponding to the verification sample; perform a second normalization process on the measurement data of the next verification wafer after passing through the previous machine to obtain the first image features corresponding to the verification sample; perform a second normalization process on the measurement data of the previous verification wafer after passing through the current machine to obtain the second image features corresponding to the verification sample; and perform a second normalization process on the measurement data of the next verification wafer after passing through the current machine to obtain the third image features corresponding to the verification sample.
[0143] The second input module is used to input the text features, first image features, and second image features corresponding to the verification sample into the measurement model after each training process to obtain the second wafer image output by the measurement model.
[0144] The judgment module is used to calculate the second loss between the features of the second wafer image and the third image corresponding to the verification sample, and select the model with the smallest second loss as the measurement model.
[0145] In this example, after selecting all wafers within the analysis period, the wafers are sorted according to their processing time, and 5% of the wafers (i.e., verification wafers) are selected as the verification sample. Then, using the wafer ID as an index, sensor data of adjacent verification wafers (usually two adjacent verification wafers) are collected in the current machine, resulting in {'ID”:{'Si”:{t1 ’ :vi1 ’ ,t2 ’ :vi2 ’ ,t3 ’ :vi3 ’ ,…,tj ’ :vij ’}}}, where ID ’ Represents the current wafer number, Si ’ Representing sensor number i, tj’ Represents the time of data collection, vij ’ The representative sensor Si in tj ’ Data collected at specific time points. Measurement data is collected from the measurement stations that adjacent verification wafers pass through after the current machine and from the measurement stations that the next verification wafer passes through after the previous machine (the process machine preceding the current process machine), resulting in {'ID”:{'Pp”:{vp1 ’ ,vp2 ’ vp3 ’ VPN ’}}}, where ID ’ Pp represents the current wafer number. ’ Represents the p-th measurement item, VPN ’ This represents the measurement data of the nth measurement point of the p-th measurement item. The measurement data includes the measurement value and measurement coordinates, where the measurement coordinates are {'ID”:{'Pp”:{(xp1)}. ’ yp1 ’ ),(xp2 ’ yp2 ’ (xp3) ’ yp3 ’ ),…,(xpn ’ ypn ’ )}}}, where ID ’ Pp represents the current wafer number. ’ Represents the p-th measurement term, (xpn) ’ ypn ’ ) represents the measurement coordinates of the nth measurement point of the pth measurement item.
[0146] Optionally, after acquiring multiple verification samples, for the sensor data of adjacent verification wafers in each verification sample on the current machine, the aforementioned first normalization process is performed to obtain the text features of the adjacent verification wafers, which are used as the text features corresponding to the verification sample. For the measurement data of the next verification wafer after passing the previous machine in the adjacent verification wafers, the aforementioned second normalization process is performed to obtain the first image features corresponding to the verification sample. For the measurement data of the previous verification wafer after passing the current machine in the adjacent verification wafers, the aforementioned second normalization process is performed to obtain the second image features corresponding to the verification sample. For the measurement data of the next verification wafer after passing the current machine in the adjacent verification wafers, the aforementioned second normalization process is performed to obtain the third image features corresponding to the verification sample, so as to standardize the range of data and eliminate dimensional differences.
[0147] Accordingly, after every 100 iterations, the text features, first image features, and second image features corresponding to the verification sample are input into the measurement model to obtain the second wafer image output by the measurement model. The second loss between the second wafer image and the third image features corresponding to the verification sample is calculated, and the model with the smallest second loss is selected as the measurement model.
[0148] In the example above, the measurement model was validated by setting up validation samples, which improved the model's prediction accuracy and reliability.
[0149] Building upon the aforementioned example, the device further includes a removal module for removing samples from multiple training samples and multiple validation samples where the sensor data is constant or where the sensor data is missing beyond a preset threshold. Specifically, for data in the training and validation samples where the sensor data is constant, where the proportion of missing sensor data due to sensor failure is greater than 80%, and for anomalous data, all data (including FDC data and wafer measurement data) of the wafer is removed. Samples with constant sensor data typically lack useful information because constant values do not provide any meaningful variation information about process changes or equipment status, which may prevent the model from learning important features and patterns in the data. A sample with more than a preset threshold of missing sensor data may result in incomplete data, affecting the model's training and validation process. This missing data may introduce noise and bias, making it difficult for the model to accurately capture the true relationships within the data.
[0150] In the example above, by removing these low-quality samples, the quality of the dataset was significantly improved, enabling the model to focus on learning from more representative and informative samples, thus improving the model's robustness and stability.
[0151] In the measurement device provided in this embodiment, the first acquisition module first acquires sensor data of the wafer under test and the adjacent wafer before the wafer under test at the current machine, measurement data of the wafer under test after passing the previous machine, and measurement data of the adjacent wafer before the wafer under test after passing the current machine. Then, the first processing module performs a first normalization process on the sensor data to obtain text features corresponding to the wafer under test; performs a second normalization process on the measurement data to obtain image features corresponding to the wafer under test. Finally, the first input module inputs the text features and image features corresponding to the wafer under test into the measurement model corresponding to the current machine, and uses the output wafer image as the measurement result. The solution of this application processes the acquired sensor data as text data to obtain text features, processes the acquired measurement data as image data to obtain image features, and inputs the text features and image features into a multimodal measurement model to obtain the wafer image output by the model. By employing a multimodal image generation method to model the machine, full-point measurement information prediction is achieved, improving the accuracy of semiconductor virtual measurement.
[0152] Example 5
[0153] Figure 7 This is a schematic diagram of the structure of the electronic device provided in Embodiment 5 of this application, as shown below. Figure 7 As shown, the electronic device includes:
[0154] The electronic device includes a processor 291 and a memory 292; it may also include a communication interface 293 and a bus 294. The processor 291, memory 292, and communication interface 293 can communicate with each other via the bus 294. The communication interface 293 can be used for information transmission. The processor 291 can invoke logical instructions stored in the memory 292 to execute the methods described in the example above.
[0155] Furthermore, the logic instructions in the aforementioned memory 292 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.
[0156] The memory 292, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this application. The processor 291 executes functional applications and data processing by running the software programs, instructions, and modules stored in the memory 292, that is, it implements the methods in the above method examples.
[0157] The memory 292 may include a program storage area and a data storage area. The program storage area may store the operating system and application programs required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 292 may include high-speed random access memory and may also include non-volatile memory.
[0158] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method in any of the embodiments.
[0159] This application also provides a computer program product, which, when executed by a processor, implements the method in any of the embodiments.
[0160] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0161] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A measurement model training method, characterized in that, include: Multiple training samples are acquired, wherein each training sample includes sensor data of adjacent training wafers on the current machine, measurement data of the adjacent training wafers after passing the current machine, and measurement data of the next training wafer in the adjacent training wafers after passing the previous machine. For each training sample, the sensor data of the adjacent training wafers on the current machine are subjected to a first normalization process to obtain the text features corresponding to the training sample. The measurement data of the next training wafer in the adjacent training wafers after passing through the previous machine is subjected to a second normalization process to obtain the first image feature corresponding to the training sample; the measurement data of the previous training wafer in the adjacent training wafers after passing through the current machine is subjected to a second normalization process to obtain the second image feature corresponding to the training sample; and the measurement data of the next training wafer in the adjacent training wafers after passing through the current machine is subjected to a second normalization process to obtain the third image feature corresponding to the training sample. Based on the text features, first image features, second image features, and third image features corresponding to the multiple training samples, the measurement model is trained until the measurement model training is completed.
2. The method according to claim 1, characterized in that, The training process includes: The text features, first image features, and second image features corresponding to the training samples are input into the measurement model to obtain the first wafer image output by the measurement model. The first loss between the first wafer image and the third image features corresponding to the training samples is fed back to the measurement model until the first loss is less than a first threshold, at which point the measurement model is determined to have completed training.
3. The method according to claim 1, characterized in that, The method further includes: Multiple verification samples are acquired, wherein each verification sample includes sensor data of adjacent verification wafers on the current machine, measurement data of the adjacent verification wafers after passing the current machine, and measurement data of the next verification wafer in the adjacent verification wafers after passing the previous machine. For each adjacent verification wafer in a verification sample, the sensor data of the adjacent verification wafer on the current machine are subjected to the first normalization process to obtain the text features of the adjacent verification wafer, which are used as the text features corresponding to the verification sample. The measurement data of the next verification wafer after passing the previous machine are subjected to the second normalization process to obtain the first image features corresponding to the verification sample. The measurement data of the previous verification wafer after passing the current machine are subjected to the second normalization process to obtain the second image features corresponding to the verification sample. The measurement data of the next verification wafer after passing the current machine are subjected to the second normalization process to obtain the third image features corresponding to the verification sample. After each training process, the text features, first image features, and second image features corresponding to the verification sample are input into the measurement model to obtain the second wafer image output by the measurement model. Calculate the second loss between the second wafer image and the third image features corresponding to the verification sample, and select the model with the smallest second loss as the measurement model.
4. The method according to claim 3, characterized in that, The method further includes: Remove samples from the plurality of training samples and the plurality of validation samples where the sensor data is a constant value or where the sensor data is missing greater than a preset threshold.
5. The method according to claim 1, characterized in that, The first normalization process includes: Based on the number of processes of the current machine, calculate the average, maximum, minimum and sum of the sensor data of the training wafer under each process of the current machine; For any missing values in the average value of the sensor data of the training wafer under each process of the current machine, the missing values are filled using the average of the average values of the training wafer under all processes; for any missing values in the maximum value of the sensor data of the training wafer under each process of the current machine, the missing values are filled using the average of the maximum values under all processes; for any missing values in the minimum value of the sensor data of the training wafer under each process of the current machine, the missing values are filled using the average of the minimum values under all processes; for any missing values in the sum of the sensor data of the training wafer under each process of the current machine, the missing values are filled using the average of the sums under all processes. The average, maximum, minimum, and sum of the sensor data for each process of the current machine tool after filling are normalized to obtain the text features of the training wafer.
6. The method according to claim 1, characterized in that, The measurement data includes measurement values and measurement coordinates. The second normalization process includes: The measurement coordinates in the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine are mapped to a preset range to obtain processed measurement coordinates. The processed measurement coordinates are then mapped to the measurement values in the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine, respectively. The missing values in the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine are filled with their respective average values. Based on the number of measurement items, the maximum and minimum values of the measurement data of the next training wafer after passing the previous machine, the measurement data of the previous training wafer after passing the current machine, and the measurement data of the next training wafer after passing the current machine are calculated and normalized to obtain the first image feature, the second image feature, and the third image feature corresponding to the training sample.
7. The method according to any one of claims 1-6, characterized in that, The measurement model includes a text encoder, an image encoder, and a feature decoder; The text encoder is used to extract features from text to obtain a text vector; the image encoder is used to extract features from image to obtain an image vector; and the feature decoder is used to generate a wafer image based on the text vector and the image vector.
8. A measurement method, characterized in that, The measurement model obtained by the measurement model training method as described in any one of claims 1-7, wherein the measurement method includes: The sensor data of the wafer under test and the previous wafer adjacent to the wafer under test at the current machine station are acquired, the measurement data of the wafer under test after passing the previous machine station, and the measurement data of the previous wafer adjacent to the wafer under test after passing the current machine station are acquired. The sensor data of the wafer under test and the adjacent wafer before the wafer under test at the current machine are subjected to a first normalization process to obtain the text features corresponding to the wafer under test; the measurement data of the wafer under test after passing the previous machine and the measurement data of the adjacent wafer before the wafer under test after passing the current machine are subjected to a second normalization process to obtain the image features corresponding to the wafer under test. The text features and image features corresponding to the wafer to be tested are input into the measurement model corresponding to the current machine, and the output wafer image is used as the measurement result; wherein, the measurement model corresponding to the current machine is a pre-trained multimodal model, and the measurement model corresponding to the current machine is used to predict the wafer image after the wafer passes through the machine; wherein, the wafer image is wafer process data.
9. A measuring device, characterized in that, The measurement model obtained using the measurement model training method as described in any one of claims 1-7, wherein the measurement device comprises: The first acquisition module is used to acquire sensor data of the wafer under test and the previous wafer adjacent to the wafer under test on the current machine, measurement data of the wafer under test after passing the previous machine, and measurement data of the previous wafer adjacent to the wafer under test after passing the current machine. The first processing module is used to perform a first normalization process on the sensor data of the wafer under test and the adjacent wafer before the wafer under test at the current machine to obtain the text features corresponding to the wafer under test; and to perform a second normalization process on the measurement data of the wafer under test after passing the previous machine and the measurement data of the adjacent wafer before the wafer under test after passing the current machine to obtain the image features corresponding to the wafer under test. The first input module is used to input the text features and image features corresponding to the wafer to be tested into the measurement model corresponding to the current machine, and to output the wafer image as the measurement result; wherein, the measurement model corresponding to the current machine is a pre-trained multimodal model, and the measurement model corresponding to the current machine is used to predict the wafer image after the wafer passes through the machine; wherein, the wafer image is wafer process data.
10. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-8.
12. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-8.