Semiconductor process engineering data fitting method and device based on polynomial fitting, medium and product

By constructing a polynomial fitting model, the semiconductor process is automatically optimized, and the fluctuations caused by hardware aging and process changes are solved, and process efficiency and product quality are improved.

CN120408056APending Publication Date: 2025-08-01上海朋熙半导体股份有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510197993.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-21
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

There are fluctuations in the existing semiconductor manufacturing processes due to hardware aging, process changes and frequent tuning, resulting in low machine usage and waste of production capacity, and relying on manual tuning and inefficiency.

Method used

By constructing a polynomial fitting model, collecting semiconductor process data, extracting feature variables and target variables, determining optimization coefficients and functions, performing fitting calculations, automating the optimization process, and reducing manual intervention.

Benefits of technology

It improves the efficiency and accuracy of process data fitting, reduces manual intervention, and improves the efficiency and product quality of semiconductor manufacturing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408056A_ABST
    Figure CN120408056A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of semiconductor manufacturing, and discloses a polynomial fitting-based semiconductor process engineering data fitting method and device, a medium and a product. The method comprises the following steps: collecting semiconductor process engineering data, and constructing a data set; extracting a characteristic variable and a target variable in the data set, and constructing a polynomial; according to the characteristic variable and the target variable, determining an optimization coefficient in the fitting calculation process, and determining an optimization function in the fitting calculation process; according to the optimization coefficient, performing fitting calculation on a training set in the data set by using an optimization function to obtain an optimal polynomial; and verifying the optimal polynomial by using a verification set in the data set, and if the verification is passed, taking the optimal polynomial as a polynomial fitting result between the feature variable and the target variable. By adopting the scheme, the efficiency of fitting calculation can be improved, the accuracy of a fitting result is improved, manual intervention is avoided, and automatic fitting is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of semiconductor manufacturing technology, and in particular to a semiconductor process engineering data fitting method, equipment, medium and product based on polynomial fitting. Background Art

[0002] In recent years, with the rapid development of science and technology, the semiconductor industry has made rapid progress.

[0003] Existing semiconductor manufacturing processes face numerous challenges, including silicon wafer rejection due to the natural, unavoidable aging of hardware during the manufacturing process; fluctuations caused by systemic changes, such as process variations, equipment status changes, and variations in processed silicon wafers; low machine utilization due to frequent adjustments to formula parameters, leading to wasted production capacity; and experience-based tuning formulas that require continuous trial and error. Therefore, minimizing manual intervention and automating the semiconductor manufacturing process tuning process remains a pressing technical challenge in this field. Summary of the Invention

[0004] One purpose of this application is to provide a semiconductor process engineering data fitting method, apparatus, medium, and product based on polynomial fitting, at least to address the various problems caused by manual intervention in fitting calculations. This application constructs a polynomial, defines optimization coefficients, defines an optimization function, and uses a sample data set for fitting calculations to obtain a polynomial fitting result. This improves the efficiency of the fitting calculation and the accuracy of the fitting result, avoiding manual intervention and achieving automated fitting.

[0005] To achieve the above objectives, some embodiments of the present application provide the following aspects:

[0006] In a first aspect, some embodiments of the present application provide a semiconductor process engineering data fitting method based on polynomial fitting, the method comprising:

[0007] Collect semiconductor process engineering data and build data sets;

[0008] Extracting characteristic variables and target variables from the data set and constructing polynomials;

[0009] Determining an optimization coefficient in a fitting calculation process and an optimization function in a fitting calculation process according to the characteristic variables and the target variables;

[0010] According to the optimization coefficient, the optimization function is used to perform a fitting calculation on the training set in the data set to obtain an optimal polynomial;

[0011] Verify the optimal polynomial using the validation set in the dataset. If the verification passes, use the optimal polynomial as the polynomial fitting result between the feature variable and the target variable.

[0012] In a second aspect, some embodiments of the present application further provide an electronic device, which includes: one or more processors; and a memory storing computer program instructions, and when the computer program instructions are executed, the processors execute the steps of the method described above.

[0013] In a third aspect, some embodiments of the present application further provide a computer-readable medium, on which computer program instructions are stored, and the computer program instructions can be executed by a processor to implement the method described above.

[0014] In a fourth aspect, some embodiments of the present application further provide a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method described above are implemented.

[0015] Compared with the related art, in the solution provided by the embodiments of the present application, semiconductor process engineering data is collected to construct a dataset; feature variables and target variables in the dataset are extracted, and a polynomial is constructed; according to the feature variables and the target variables, optimization coefficients in the fitting calculation process are determined, and an optimization function in the fitting calculation process is determined; according to the optimization coefficients, the training set in the dataset is fitted using the optimization function to obtain an optimal polynomial; the optimal polynomial is verified using the validation set in the dataset. If the verification passes, the optimal polynomial is used as the polynomial fitting result between the feature variable and the target variable. In this technical solution, by constructing a polynomial, defining optimization coefficients, defining an optimization function, and performing fitting calculations using a sample dataset to obtain a polynomial fitting result, the efficiency of fitting calculations can be improved, the accuracy of the fitting result can be increased, manual intervention can be avoided, and automatic fitting can be achieved. Description of the Drawings

[0016] One or more embodiments are exemplarily illustrated by pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements, unless otherwise stated, and the drawings in the drawings do not constitute a proportional limitation.

[0017] Figure 1 It is an exemplary flowchart of a semiconductor process engineering data fitting method based on polynomial fitting provided by some embodiments of the present application;

[0018] Figure 2 It is a fitting flowchart provided by some embodiments of the present application;

[0019] Figure 3 A scatter plot of original data provided according to some embodiments of the present application;

[0020] Figure 4 A fitted curve and a scatter plot of data provided according to some embodiments of the present application;

[0021] Figure 5 An exemplary structural diagram of the electronic device is disclosed. Detailed implementation manners

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0023] The first embodiment

[0024] The first embodiment of the present application relates to a semiconductor process engineering data fitting method based on polynomial fitting. As Figure 1 shown, the method may include the following steps:

[0025] Step S101, collect semiconductor process engineering data and construct a data set;

[0026] Semiconductor process engineering may be a technical field involving the processing and manufacturing of semiconductor materials into semiconductor devices and integrated circuits, covering a series of complex processes from raw material preparation to chip packaging and testing. For example, specific process steps such as lithography, etching, and ion implantation all belong to the scope of semiconductor process engineering.

[0027] Semiconductor process engineering data may be various information generated in each link of semiconductor process engineering, such as quantitative indicators like temperature, pressure, time, current, voltage, material composition ratio, product yield, etc. These data reflect the process and product characteristics.

[0028] A data set may be a data collection formed by organizing and storing the collected various semiconductor process engineering data according to certain rules and formats, facilitating subsequent data analysis and processing. For example, it is stored in a table form, with each row representing a process experiment or production record, and each column representing a specific data indicator.

[0029] This solution can utilize various sensors, monitoring devices, and data recording systems to obtain relevant data information in real time from all aspects of the semiconductor process. For example, temperature data at specific positions during chip manufacturing can be collected through high-precision temperature sensors, or the performance parameters of products can be recorded using automated test equipment. The collected data is screened and cleaned (removing incorrect, duplicate, or incomplete data), and then organized in a specific structure and format to form a data set. For example, different types of data are classified and sorted into corresponding columns, and appropriate header identifiers are added to construct a structured data set that can be used for analysis.

[0030] Step S102: Extract the feature variables and target variables from the data set and construct a polynomial.

[0031] Feature variables can be factors in the data set that can affect or explain the changes in the target variable. In semiconductor processes, for example, the exposure time of lithography, the gas flow rate of etching, the initial temperature of the wafer, etc. Changes in these factors may affect the final performance of the product.

[0032] The target variable is the variable for which we hope to predict or understand its change pattern. For example, the final performance indicators of semiconductor products, such as the computing speed, power consumption, and yield rate of chips, which are the key objects of concern for process optimization.

[0033] A polynomial can be a mathematical expression composed of variables and coefficients, which describes the relationship between feature variables and target variables through combinations of variables of different degrees. In the semiconductor process scenario, a polynomial such as may be constructed, where y is the target variable, x1, x2, etc. are feature variables, and a i is the coefficient.

[0034] In this solution, the data columns of feature variables and target variables relevant to the research purpose can be selected from the data set. For example, according to process knowledge and research requirements, determine the lithography exposure time and etching gas flow rate (feature variables) that affect the chip yield rate (target variable), and then extract the corresponding data columns from the data set. Based on the possible relationship assumptions between feature variables and target variables, combined with mathematical principles, construct an appropriate polynomial expression. For example, by analyzing historical data and process principles, assume that there is a quadratic relationship between the target variable and certain feature variables, and thus construct a polynomial containing the quadratic terms of these feature variables.

[0035] Step S103: Determine the optimization coefficients in the fitting calculation process according to the feature variables and the target variables, and determine the optimization function in the fitting calculation process.

[0036] Fitting calculation can be to adjust the coefficients of a polynomial so that the constructed polynomial can describe the relationship between the feature variables and the target variable as accurately as possible. The process of calculating these adjusted coefficients is the fitting calculation.

[0037] Optimization coefficient can be a parameter used to control and adjust the update of polynomial coefficients during the fitting calculation. Different settings of the optimization coefficient will affect the speed and accuracy of fitting. For example, the learning rate is a common optimization coefficient, which determines the step size of the polynomial coefficient update in each iteration.

[0038] Optimization function can be a function used to measure the goodness of fit of a polynomial to the data. By minimizing (or maximizing) the value of this function, the optimal polynomial coefficients can be found. For example, the mean squared error function (MSE), which calculates the average of the squares of the errors between the predicted values (obtained from the polynomial) and the actual target variable values. The smaller the value, the better the fitting effect.

[0039] This solution can select appropriate optimization coefficient values and optimization function types based on the analysis of the characteristics of semiconductor process data, past experience, and mathematical principles. For example, by experimenting with different learning rate values, observing the fitting effect, and selecting the learning rate that can make the fitting process converge quickly and accurately as the optimization coefficient; according to the data distribution and research purpose, selecting the mean squared error function or other more suitable functions as the optimization function.

[0040] Step S104: According to the optimization coefficient, use the optimization function to perform fitting calculation on the training set in the dataset to obtain the optimal polynomial;

[0041] The dataset contains data sets for training and validation. The training set can be a part of the data divided from the dataset, specifically used to train the model (i.e., the polynomial here). By learning the training set data, the coefficients of the polynomial are adjusted. For example, 70% of the data in the dataset is used as the training set for fitting calculation.

[0042] The optimal polynomial can be the polynomial found after a series of fitting calculations that can achieve the best fitting effect on the training set data under the given optimization function, and its coefficient combination can make the value of the optimization function reach the minimum (or maximum, depending on the nature of the optimization function).

[0043] This solution can apply the selected optimization function to the fitting calculation process of the training set data to evaluate the fitting effect after each adjustment of the polynomial coefficients. For example, the mean squared error optimization function is used to calculate the error between the predicted value of each polynomial and the actual value of the target variable in the training set. The fitting calculation operation is carried out, and according to the rules determined by the optimization coefficients, the coefficients of the polynomial are continuously adjusted to make the value of the optimization function change in the direction of the minimum (or maximum). For example, the gradient descent algorithm is used to iteratively update the polynomial coefficients according to the optimization coefficients (such as the learning rate) for fitting calculation. After multiple iterations of fitting calculation, the polynomial and its coefficient combination with the best fitting effect on the training set under the current optimization conditions are determined.

[0044] Step S105: Use the validation set in the dataset to verify the optimal polynomial. If the verification passes, use the optimal polynomial as the polynomial fitting result between the feature variable and the target variable.

[0045] The validation set can be another part of the data divided from the dataset, different from the training set, and is used to test the generalization ability of the optimal polynomial obtained on the training set, that is, the adaptability to new data. For example, 30% of the data in the dataset is used as the validation set.

[0046] The polynomial fitting result is the polynomial that can accurately describe the relationship between the feature variable and the target variable determined after verification, and can be used for prediction and analysis.

[0047] This solution can input the validation set data into the optimal polynomial to calculate the predicted value of the polynomial. By comparing the predicted value of the optimal polynomial on the validation set with the actual value of the target variable in the validation set, an index similar or related to the optimization function (such as calculating the mean squared error on the validation set) is used to evaluate the performance of the optimal polynomial. If the evaluation index reaches a certain standard (such as the mean squared error is within an acceptable range), it is considered that the verification passes. After the verification passes, this optimal polynomial is determined as the mathematical model that finally describes the relationship between the feature variable and the target variable, and is used for subsequent practical applications such as process analysis and prediction.

[0048] The technical solution provided in this embodiment, by collecting semiconductor process engineering data, constructing a dataset, extracting features and target variables from it to construct a polynomial, determining optimization coefficients and functions for fitting calculation, and then obtaining the final polynomial fitting result through verification. This series of steps can deeply explore the potential relationship between semiconductor process parameters and product performance, establish an accurate mathematical model, provide a scientific and effective method and basis for semiconductor process optimization, product quality improvement, and performance prediction, and contribute to improving the efficiency and product quality of semiconductor manufacturing and promoting the technological progress of the semiconductor industry.

[0049] In one embodiment, after determining the optimization coefficient in the fitting calculation process according to the characteristic variable and the target variable, the method further includes:

[0050] Initial coefficients are defined to control coefficient distortion of the polynomial during the fitting calculation process.

[0051] Initial coefficients are the starting values set for the coefficients in a polynomial before the fitting calculation. These initial values are crucial to the entire fitting process, as the fitting algorithm typically starts with these initial values and iteratively adjusts the coefficients to better fit the polynomial to the data. For example, in a simple linear polynomial y = a0 + a1x1, the initial values of a0 and a1 are the initial coefficients. The selection of initial coefficients cannot be arbitrary; it must be based on a preliminary understanding of the data, past experience with similar problems, or some theoretical basis. If the initial coefficients are not chosen properly, the fitting process may fall into a local optimum or the coefficients may be distorted during the iteration process.

[0052] In this solution, specific initial values can be set for the coefficients of the polynomial based on relevant theoretical knowledge, preliminary feature analysis of the data, and past experience. For example, for polynomials constructed by the relationship between characteristic variables and target variables with clear physical meanings, the initial coefficients can be determined based on physical principles. If in a certain semiconductor process, it is known that a certain characteristic variable has an approximate linear relationship with the target variable, and the range of the proportional coefficient is roughly known from a physical point of view, a suitable value can be selected within this range as the initial coefficient. In the absence of a clear physical basis, some common initialization methods can also be used, such as random initialization, but the random values will be limited within a certain range to ensure that the initial coefficients are not too outrageous, thereby avoiding subsequent problems such as coefficient distortion.

[0053] This technical solution, by defining appropriate initial coefficients, can effectively guide the fitting calculation process in a reasonable direction, reduce the risk of coefficient distortion caused by improper coefficient initialization, and ensure the stability and accuracy of the fitting process. This helps improve the efficiency of polynomial fitting, more quickly find the optimal polynomial that accurately describes the relationship between characteristic variables and target variables, and thus improve the reliability of the entire model in depicting the relationship between semiconductor process and product performance.

[0054] In one embodiment, before determining the optimization function in the fitting calculation process, the method further includes:

[0055] A setting operation is received, and at least one bias term is added to the polynomial according to the setting operation.

[0056] Among them, the setting operation can be an interaction behavior between the user and the system. The user conveys their specific requirements for polynomial construction to the system through specific interfaces, instructions, configuration files, etc. For example, in a data analysis software, the user can complete the setting operation by checking relevant options, inputting specific parameters, or selecting a preset configuration scheme on the graphical interface; it can also be achieved by writing code, calling corresponding functions, and passing specific parameters. In the scenario of semiconductor process data fitting, the setting operation may be that the user decides the number, position, and form of the bias terms to be added to the polynomial based on their understanding of the process principle.

[0057] In a polynomial, a bias term is a constant term independent of the feature variables or a term in a specific form, and its role is to adjust the overall position of the polynomial or provide a basic offset. In a mathematical expression, the bias term usually appears as a constant or in a form independent of the feature variables. For example, in the simple linear polynomial y = a1x1 + b, b is the bias term. In the polynomial fitting of semiconductor processes, the bias term can be used to represent some basic effects that cannot be directly explained by the feature variables, such as inherent errors in the process, average effects of environmental factors, etc.

[0058] In this solution, the system can wait for and obtain the setting operation information input by the user. This may involve various technical means. For example, in a software system, by listening for events such as button clicks and text inputs on the user interface, or by reading relevant parameters in the configuration file; in code implementation, by receiving the setting information passed by the user through function parameters. The system will parse and verify the received setting operation to ensure that it meets the requirements and specifications of the system. According to the received setting operation, the bias term is introduced into the constructed polynomial. The specific implementation method will vary according to different programming languages and data processing frameworks. For example, in the numpy library of Python, the bias term can be added by modifying the coefficient array of the polynomial; in some machine learning libraries, there may be dedicated functions to complete this operation. When adding the bias term, it is necessary to consider the position and form of the bias term to ensure that it can accurately reflect the user's setting intention.

[0059] This technical solution can enhance the flexibility and expressive power of the polynomial model by receiving the setting operation and adding the bias term to the polynomial. The bias term can help the model better fit the data, especially when there is a certain basic offset in the data or the data cannot be fully explained by the feature variables. In semiconductor process data fitting, adding appropriate bias terms can improve the description accuracy of the relationship between the process and product performance of the model, enabling the model to more accurately capture the laws in the actual data, thereby providing a more reliable basis for process optimization and product quality control.

[0060] In one embodiment, semiconductor process engineering data is collected to construct a dataset, including:

[0061] Use the read_csv function of the pandas library to read semiconductor process engineering data from a file named dataset.csv and construct a dataset.

[0062] Among them, the pandas library is a powerful and widely used data analysis and processing library in Python. It provides rich data structures (such as Series and DataFrame) and data operation methods, and can efficiently process and analyze various types of data. When processing semiconductor process engineering data, pandas can conveniently perform operations such as data reading, cleaning, transformation, and statistical analysis. For example, it can easily handle missing values and duplicate data, and can also perform operations such as data filtering, sorting, and grouping, providing a good foundation for subsequent data mining and modeling.

[0063] The read_csv function is an important function in the pandas library, used to read data from a file stored in comma-separated values (CSV) format. A CSV file is a common text file format, where data is stored in a tabular form, each row represents a record, each column represents a field, and fields are separated by commas. The read_csv function can read the file content according to the file path and convert it into a pandas DataFrame object for convenient subsequent data processing and analysis. This function also supports various parameter settings, such as specifying file encoding, column names, delimiters, etc., to adapt to different formats of CSV files.

[0064] dataset.csv is a CSV file storing semiconductor process engineering data. This file may contain various data information in the semiconductor process, such as process parameters (temperature, pressure, time, etc.), equipment status data, product performance indicators (yield, electrical performance, etc.). Each row of the file represents a process experiment or production record, and each column corresponds to a specific data indicator. By reading this file, a dataset for subsequent analysis and modeling can be obtained.

[0065] The dataset can be a data collection read from the dataset.csv file and exists in the form of a pandas DataFrame object. A DataFrame is a two-dimensional tabular data structure, similar to a table in a database or an Excel spreadsheet, with row indexes and column indexes. In this dataset, each row can be regarded as a sample, and each column is a feature or attribute. The dataset is the basis for subsequent operations such as feature extraction and model training.

[0066] In this solution, a specific task is completed by calling the read_csv function in the pandas library, that is, reading data from the dataset.csv file. In Python code, it is usually achieved by importing the pandas library and using the functions in the library. For example: import pandas as pd, and then use pd.read_csv() to call this function. The read_csv function opens the file and parses the data content according to the specified file path (here it is dataset.csv). It can identify the file format, split the data into different rows and columns according to the rules of CSV files, and store them in the DataFrame object of pandas. During the reading process, the parameters of the function can also be set as needed, such as specifying the encoding format of the file, column names, etc. The data read from the dataset.csv file is converted into a complete dataset that can be used for subsequent analysis. The data read by the read_csv function is organized into a DataFrame object of pandas, and this object has rich attributes and methods, which can facilitate data operation and analysis. The process of constructing the dataset is actually to convert the original CSV file data into a structured data form, providing convenience for subsequent data processing and modeling.

[0067] In this technical solution, the read_csv function of the pandas library is used to read semiconductor process engineering data from the dataset.csv file and construct a dataset, which has the characteristics of high efficiency and convenience. The rich functions provided by the pandas library can quickly process a large amount of data and can flexibly process CSV files in different formats. The constructed dataset exists in the form of a DataFrame object, which is convenient for various data processing and analysis operations, provides a stable and reliable data basis for subsequent steps such as feature extraction and model training, helps to further explore the potential information in semiconductor process engineering data, and thus provides support for process optimization and product quality improvement.

[0068] In one embodiment, the optimization coefficient includes one or more of a learning rate, the number of iterations, a bias term range, and the power of gradient calculation.

[0069] The optimization coefficient can be a set of parameters used to control and adjust the update of model parameters during the fitting calculation process. These parameters will affect the performance and effect of the optimization algorithm and determine whether the model can quickly and accurately converge to the optimal solution. Different combinations of optimization coefficients will lead to different fitting results, and reasonable selection of optimization coefficients is crucial for constructing an accurate and effective model.

[0070] The learning rate can be a very important hyperparameter in optimization algorithms, which controls the step size of model parameter updates in each iteration. In optimization algorithms such as gradient descent, the learning rate determines the distance moved along the gradient direction. If the learning rate is set too large, the model parameters may skip the optimal solution, resulting in non-convergence or oscillation near the optimal solution; if the learning rate is set too small, the convergence speed of the model will be very slow, and more iteration times are required to achieve a better fitting effect. For example, in semiconductor process data fitting, an appropriate learning rate can adjust the coefficients of the polynomial to the optimal values faster and improve the fitting efficiency.

[0071] The number of iterations can refer to the number of loops executed by the optimization algorithm. In each iteration, the optimization algorithm calculates the gradient based on the current model parameters and updates the parameters according to the learning rate. The setting of the number of iterations needs to balance computational resources and fitting effects. If the number of iterations is too small, the model may not converge to the optimal solution, resulting in poor fitting effects; if the number of iterations is too large, although a better fitting effect may be obtained, it will increase the computational time and resource consumption. In the polynomial fitting of semiconductor process data, a reasonable number of iterations can ensure that the model fully learns the patterns in the data.

[0072] The bias term is a constant term or a term in a specific form in the polynomial that is independent of the feature variables. The bias term range specifies the interval within which the bias term can take values. Setting the bias term range can prevent the value of the bias term from being too large or too small, thus affecting the stability and generalization ability of the model. In semiconductor process data fitting, the setting of the bias term range can be combined with the actual situation of the process and prior knowledge, such as determining the reasonable value interval of the bias term according to the inherent error range in the process.

[0073] The power of gradient calculation can be used in the optimization function to adjust the way of gradient calculation. Usually, the gradient is obtained by calculating the derivative of the loss function with respect to the model parameters, and the power of gradient calculation can change the calculation form of the derivative. For example, in the common mean squared error loss function, the power is 2 when calculating the gradient; if the power is set to other values, different gradient calculation results will be obtained, thus affecting the way of updating model parameters. In semiconductor process data fitting, an appropriate power of gradient calculation can make the optimization algorithm more adaptable to the characteristics of the data and improve the fitting accuracy.

[0074] In this technical solution, by reasonably selecting optimization coefficients, including the learning rate, the number of iterations, the bias term range, and the power of gradient calculation, the fitting effect and optimization efficiency of the model can be significantly improved. A suitable learning rate enables the model to quickly converge to the optimal solution and avoid falling into local optima; an appropriate number of iterations can reduce the waste of computing resources while ensuring the fitting accuracy; a reasonable bias term range helps to improve the stability and generalization ability of the model; a suitable power of gradient calculation allows the optimization algorithm to better adapt to the distribution and characteristics of the data. In the fitting of semiconductor process data, by optimizing these coefficients, a relationship model between process parameters and product performance can be established more accurately, providing a more reliable basis for process optimization and product quality control.

[0075] In one embodiment, the optimization function employs the gradient descent algorithm of the least squares method.

[0076] Optimization function: A function used to measure the degree of difference between the model prediction result and the actual data. Its core objective is to find a set of optimal model parameters by minimizing (or maximizing) the value of this function, so that the model can fit the data as accurately as possible. In the scenario of semiconductor process data fitting, the optimization function can help us determine the coefficients of the polynomial, thereby establishing the best relationship between the feature variables and the target variables.

[0077] The least squares method is a mathematical optimization technique that finds the best function match for the data by minimizing the sum of the squares of the errors. The coefficients of the polynomial can be determined by the least squares method, making the sum of the squares of the errors between the polynomial prediction value and the actual target value the smallest.

[0078] The gradient descent algorithm is an iterative optimization algorithm used to find the minimum value of a function. Its basic idea is to continuously update the parameters along the negative gradient direction of the function, gradually approaching the minimum point of the function. In each iteration, the algorithm calculates the gradient of the function according to the current parameters, and then updates the parameters along the negative gradient direction with a certain step size (determined by the learning rate). As the iteration progresses, the parameters are continuously adjusted, and the function value gradually decreases until the stop condition is met (such as reaching the maximum number of iterations or the change in the function value is less than a certain threshold). When using the least squares method for polynomial fitting, the gradient descent algorithm can help us find the polynomial coefficients that minimize the sum of the squares of the errors.

[0079] This technical solution has various technical effects by using the gradient descent algorithm of the least squares method as the optimization function. First of all, the least squares method can mathematically ensure that the polynomial coefficients found minimize the sum of the squares of the errors between the predicted values and the actual values, thereby improving the fitting accuracy of the model. Secondly, the gradient descent algorithm is an iterative algorithm that can work effectively on large-scale datasets. By adjusting parameters such as the learning rate, the convergence speed and stability of the algorithm can be controlled. In the fitting of semiconductor process data, this optimization function can help us quickly and accurately find the polynomial that can describe the relationship between process parameters and product performance, providing a reliable basis for process optimization and product quality control. At the same time, the iterative nature of the gradient descent algorithm also enables us to gradually approach the optimal solution and improve the computational efficiency under limited computational resources.

[0080] In one embodiment, the gradient descent algorithm is used for the optimization function to find the minimum value of the residual.

[0081] The residual is the difference between the model predicted value and the actual observed value. The residual reflects the fitting degree of the model to the data. The smaller the residual, the closer the prediction of the model is to the actual value.

[0082] The optimization function is constructed based on the residual, and its purpose is to measure the overall fitting error of the model. Usually, a certain mathematical transformation of the residual is used as the value of the optimization function. For example, the common method is to sum the squares of the residuals to obtain the mean squared error (MSE). The smaller the value of the optimization function, the better the fitting effect of the model. The task of the gradient descent algorithm is to continuously adjust the model parameters to make the value of this optimization function reach the minimum.

[0083] In this solution, the gradient descent algorithm can be used to start from an initial parameter value and gradually adjust the parameters by continuous iteration, trying to find the parameter combination that can make the optimization function obtain the minimum value. In each iteration, the algorithm calculates the gradient of the optimization function under the current parameters, and then updates the parameters according to the direction and magnitude of the gradient, continuously moving in the direction of decreasing the value of the optimization function until the stop condition is met (such as reaching the maximum number of iterations, the change in the value of the optimization function is less than a certain threshold, etc.). At this time, it is considered that the minimum value point of the optimization function or a point close to the minimum value has been found. The purpose and use of the gradient descent algorithm are specifically for the optimization function constructed based on the residual, and it uses its own iterative mechanism to find the minimum value of this function, so that the predicted value of the model is as close as possible to the actual observed value and the fitting accuracy of the model is improved.

[0084] This technical solution has significant technical effects through the use of the gradient descent algorithm to find the optimization function for the minimum residual. In terms of computational efficiency, the gradient descent algorithm is an iterative algorithm that can quickly calculate the gradient and update parameters in each iteration. It is especially suitable for large-scale datasets and can find relatively optimal parameter solutions within a reasonable time. In terms of the model fitting effect, by minimizing the optimization function based on the residual, the constructed polynomial model can better fit the semiconductor process engineering data and accurately capture the relationship between the characteristic variables and the target variables. This helps to improve the understanding and prediction ability of the semiconductor process, provides a reliable basis for process optimization, product quality control, etc., and thus enhances the production efficiency and quality of semiconductor products.

[0085] Second Embodiment

[0086] The second embodiment of this application relates to a method for fitting semiconductor process engineering data based on polynomial fitting. The method may include the following steps:

[0087] The following uses a Python script for calculation description.

[0088] import numpy as np

[0089] import matplotlib.pyplot as plt

[0090] import pandas as pd

[0091] # Load the dataset / / Obtain the dataset for which the formula needs to be derived

[0092] df = pd.read_csv('dataset.csv')

[0093] # Extract the feature and target variables / / Determine the parameter ranges of the independent and dependent variables. For example: Y = a1X1 + a2X2 + b1 + b2, where b1 and b2 are offsets, Y is the dependent variable, and X1 and X2 are the independent variables X = df['X'].values

[0094] y = df['y'].values

[0095] # Define the optimization coefficients / / Set reasonable optimization coefficients according to the range of hyperparameters

[0096] rate = k1

[0097] inter = k2

[0098] max_offset = k3

[0099] min_offset = k4

[0100] gra = k5

[0101] # Define initial coefficients / / To prevent out-of-bounds, distortion, etc., initial coefficient values need to be defined

[0102] init1 = k6

[0103] init2 = k7

[0104] # Add bias terms / / Set whether to add 1 to multiple Offsets according to user needs. For example: Y = a1X1 + a2X2 + b1 + b2, where b1 and b2 are Offsets

[0105] X_b = np.c_[np.ones((max_offset, min_offset)), X]

[0106] # Define the optimization function / / Set up the deduction formula required for fitting. Here, the least squares method is implemented, and many other loss functions will also be provided

[0107] def gradient_descent(X, y, theta, learning_rate = rate, iterations = inter):

[0108] m = len(y)

[0109] history = {'cost': []}

[0110] for iteration in range(iterations):

[0111] gradients = gra / m * X.T.dot(X.dot(theta) - y)

[0112] theta = theta - learning_rate * gradients

[0113] cost = np.mean((X.dot(theta) - y) ** gra)

[0114] history['cost'].append(cost)

[0115] return theta, history

[0116] # Initialize parameters

[0117] theta_initial = np.random.randn(init1, init2)

[0118] # Run the optimization algorithm / / Output the values of a1, a2, b1, and b2 in the target formula Y = aX + b1 + b2

[0119] theta_best, history = gradient_descent(X_b, y, theta_initial, learning_rate = rate, iterations = inter);

[0120] Figure 2 The fitting flowchart provided according to some embodiments of the present application; combined Figure 2 With the corresponding code above, it can be summarized into the following steps:

[0121] 1. Load the dataset;

[0122] The purpose of this step is to obtain the dataset required for deriving the formula. The code uses the read_csv function of the pandas library to read data from a file named dataset.csv and stores it in a data frame named df. This dataset will serve as the basis for subsequent analysis and model training.

[0123] 2. Extract features and target variables;

[0124] In this step, it is necessary to determine the independent variable and the dependent variable from the loaded dataset. The code extracts the column named X from the data frame df as the feature variable (independent variable) and stores its values in the X array; at the same time, it extracts the column named y as the target variable (dependent variable) and stores its values in the y array. This is done to clarify the input and output variables when building the model later.

[0125] 3. Define optimization coefficients;

[0126] According to the reasonable range of hyperparameters, a series of optimization coefficients are set. These coefficients play an important role in the subsequent optimization algorithm, namely:

[0127] rate: Represents the learning rate, which controls the step size of parameter updates in the optimization algorithm and is assigned the value of k1 here.

[0128] inter: Represents the number of iterations, that is, the number of times the optimization algorithm is executed, and is assigned the value of k2.

[0129] max_offset and min_offset: May be used to determine the range of the bias term and are assigned the values of k3 and k4 respectively.

[0130] gra: May be used to control the power of gradient calculation or loss function, assigned with k5.

[0131] 4. Define initial coefficients;

[0132] To avoid problems such as out-of-bounds and distortion during the optimization process, initial values need to be set for the parameters of the model. In the code, k6 and k7 are used to assign values to init1 and init2 respectively, and these two initial coefficients will be used when initializing the model parameters later.

[0133] 5. Add bias terms;

[0134] According to the user's requirements, decide whether to add bias terms to the feature variables. In this example, the np.c_ function of the numpy library is used to combine an array of all 1s with the feature variable X column-wise to generate a new array X_b. The bias terms play a role similar to intercepts in the model. For example, in the formula Y = a1X1 + a2X2 + b1 + b2, b1 and b2 are the bias terms.

[0135] 6. Define the optimization function;

[0136] Here, an optimization function named gradient_descent is defined, which implements the gradient descent algorithm of the least squares method. The specific steps are as follows:

[0137] First, obtain the length m of the target variable y and initialize a dictionary history to record the loss value of each iteration.

[0138] Then, use a loop to perform a specified number (iterations) of iterations. In each iteration, calculate the gradient gradients at the current parameters. The formula for the gradient is gra / m * X.T.dot(X.dot(theta) - y), where X is the feature matrix and theta is the model parameter.

[0139] Next, update the model parameter theta according to the calculated gradient. The update formula is theta = theta - learning_rate * gradients, where learning_rate is the learning rate.

[0140] Finally, calculate the loss value cost at the current parameters. The loss function is np.mean((X.dot(theta) - y) ** gra), and add it to the history dictionary.

[0141] After the iteration ends, return the optimized parameter theta and the history dictionary recording the loss values.

[0142] 7. Initialize the parameters;

[0143] Randomly generate an array theta_initial with the shape of (init1, init2) using the np.random.randn function of the numpy library as the initial value of the model parameters.

[0144] 8. Run the optimization algorithm;

[0145] Call the previously defined gradient_descent function, passing in the feature matrix X_b, the target variable y, the initial parameter theta_initial, the learning rate rate, and the number of iterations inter for optimization calculation. Finally, obtain the optimized parameter theta_best and the history dictionary recording the loss values. These parameters correspond to the values of a1, a2, b1, and b2 in the target formula Y = aX + b1 + b2.

[0146] In the technical solution of this patent, an optimization algorithm for finding the minimum value of the residual by the gradient descent method is used. It iteratively moves in the opposite direction of the function gradient to gradually approach the minimum value. In each iteration, the gradient descent method updates the parameters according to the gradient information of the current point, so that the value of the objective function decreases.

[0147] Optimization formula: η is the search step size

[0148] Target formula:

[0149] Figure 3 To draw a scatter plot of the original data provided by some embodiments of the present application;

[0150] plt.figure(figsize=(10,6))

[0151] plt.scatter(X, y, c='b', label='Original data')

[0152] plt.xlabel('X')

[0153] plt.ylabel('y')

[0154] plt.title('Regression with Gradient Descent')

[0155] plt.legend()

[0156] plt.grid(True)

[0157] plt.show()

[0158] Figure 4 A fitted curve and a data scatter plot provided according to some embodiments of the present application;

[0159] plt.figure(figsize=(10,6))

[0160] plt.scatter(X,y,c='b',label='Original data')

[0161] plt.plot(X,X_b.dot(theta_best),c='r',label='Fitted line')

[0162] plt.xlabel('X')

[0163] plt.ylabel('y')

[0164] plt.title('Linear Regression with Fitted Line')

[0165] plt.legend()

[0166] plt.grid(True)

[0167] plt.show()

[0168] This embodiment provides a linear fitting method between a dependent variable and one or more independent variables. By finding the best-fitting straight line, the distance from the sample data points to the straight line is minimized.

[0169] In addition, some embodiments of the present application also provide an electronic device. The electronic device can be various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and so on. The electronic device can also be various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices.

[0170] The electronic device includes: one or more processors; and a memory storing computer program instructions, which when executed cause the processors to execute the steps of the method provided in any one or more of the above embodiments. Figure 5 An exemplary structural diagram of the electronic device is disclosed. As Figure 5As shown, the electronic device includes: one or more processors 501, a memory 502, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component is interconnected using different buses and can be mounted on a common motherboard or otherwise installed as needed. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory for displaying graphical information of a graphical user interface (GUI) on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories if needed. Similarly, multiple electronic devices can be connected, with each device providing part of the necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Among them, the components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit the implementation of the present application described and / or claimed herein.

[0171] The electronic device may further include: an input device 503 and an output device 504. The processor 501, the memory 502, the input device 503, and the output device 504 can be connected via a bus or other means. Figure 5 Taking connection via a bus as an example.

[0172] The input device 503 can receive input digital or character information and generate key signal inputs related to the user settings and function controls of the electronic device, such as input devices like a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 504 can include a display device, an auxiliary lighting device (such as a light-emitting diode, LED), and a haptic feedback device (such as a vibration motor), etc. The display device can include, but is not limited to, a liquid crystal display (LCD), a light-emitting diode (LED) display, and a plasma display. In some embodiments, the display device can be a touch screen.

[0173] To provide interaction with a user, the electronic device may be a computer. The computer has: a display device for displaying information to the user (e.g., a Cathode-Ray Tube (CRT) or a Liquid Crystal Display (LCD) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including voice input, speech input, or tactile input).

[0174] In the embodiments of the present application, a computer program / instructions is stored on a computer-readable medium. When the computer program / instructions are executed by a processor, the steps of the method provided in any one or more of the above embodiments are implemented. The computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist separately without being assembled into the device. The above computer-readable medium carries one or more computer-readable instructions.

[0175] The memory 502 can be used as a non-transitory computer-readable storage medium for storing non-transitory software programs, non-transitory computer-executable programs, and modules. The processor 501 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 502, so as to implement the program instructions / modules corresponding to the method provided in any one or more of the above embodiments of the present application.

[0176] The memory 502 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 502 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 502 may optionally include a memory remotely provided with respect to the processor 501, and these remote memories can be connected to the electronic device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0177] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, the computer-readable medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0178] The computer-readable medium includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of the computer's storage medium include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, magnetic disk storage, or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0179] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0180] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. For example, an Application-Specific Integrated Circuit (ASIC), a general-purpose computer, or any other similar hardware device can be used. In some embodiments, the software program of the present application can be executed by a processor to implement the above steps or functions. Similarly, the software program of the present application (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and the like. In addition, some steps or functions of the present application can be implemented by hardware, for example, as a circuit that cooperates with the processor to execute each step or function.

[0181] The computer program product provided by the embodiments of the present application includes one or more computer programs / instructions. When the computer program / instructions are executed by a processor, they wholly or partly generate the processes or functions described in the embodiments of the present application. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0182] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0183] The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference numerals in the claims should not be construed as limiting the claims involved. In addition, it is obvious that the term "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the apparatus claims may also be implemented by one unit or device through software or hardware. The terms "first", "second", etc. are only used for descriptive distinction and do not represent any specific order, nor can they be construed as indicating or implying relative importance.

[0184] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily mention changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A semiconductor process engineering data fitting method based on polynomial fitting, characterized in that, The method includes: Collecting semiconductor process engineering data and constructing a data set; Extracting feature variables and target variables from the data set and constructing a polynomial; Determining optimization coefficients in the fitting calculation process according to the feature variables and the target variables, and determining an optimization function in the fitting calculation process; Performing a fitting calculation on the training set in the data set using the optimization function according to the optimization coefficients to obtain an optimal polynomial; Verifying the optimal polynomial using the validation set in the data set. If the verification passes, using the optimal polynomial as the polynomial fitting result between the feature variables and the target variables.

2. The method according to claim 1, characterized in that, After determining the optimization coefficients in the fitting calculation process according to the feature variables and the target variables, the method further includes: Defining an initial coefficient for controlling the coefficient distortion of the polynomial in the fitting calculation process.

3. The method according to claim 1, characterized in that Before determining the optimization function in the fitting calculation process, the method further includes: Receiving a setting operation and adding at least one bias term to the polynomial according to the setting operation.

4. The method according to claim 1, wherein Collecting semiconductor process engineering data and constructing a data set, including: Reading semiconductor process engineering data from a file named dataset.csv using the read_csv function of the pandas library to construct a data set.

5. The method according to claim 1, characterized in that The optimization coefficients include one or more of a learning rate, the number of iterations, a bias term range, and the power of gradient calculation.

6. The method according to claim 1, characterized in that, The optimization function adopts the gradient descent algorithm of the least squares method.

7. The method according to claim 6, wherein An optimization function for finding the minimum value of the residual through the gradient descent algorithm.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; and A memory storing computer program instructions, which when executed cause the processor to execute the steps of the method according to any one of claims 1 to 7.

9. A computer-readable medium having computer programs / instructions stored thereon, characterized in that, The computer program / instructions, when executed by the processor, implement the steps of the method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program / instructions, characterized in that, The computer program / instructions, when executed by the processor, implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Virtual wafer three-dimensional model generation method and device, medium and program product

    CN120742626A