Wafer film thickness prediction method, device and system, storage medium and equipment

By deploying local film layer thickness prediction models locally in each wafer manufacturer and using central servers to train the global film layer thickness prediction model, the problem of poor prediction accuracy of wafer film layer thickness in the prior art is solved, and higher prediction accuracy and data privacy protection are achieved.

CN120218281APending Publication Date: 2025-06-27ZHEJIANG ICSPROUT SEMICONDUCTOR CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510361942.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-25
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, machine learning models are less accurate in predicting wafer film layer thickness, mainly because data sharing cannot be achieved between different wafer manufacturers.

Method used

Improve prediction accuracy by deploying local film thickness prediction models locally in each wafer manufacturer, and using a central server to train the global film thickness prediction model, and updating local model parameters.

Benefits of technology

It realizes that without sharing sensitive data, improves the accuracy of wafer film thickness prediction, reduces production costs, and improves production efficiency and product quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218281A_ABST
    Figure CN120218281A_ABST
Patent Text Reader

Abstract

The invention discloses a wafer film thickness prediction method, device and system, a storage medium and equipment. The method comprises the following steps: acquiring related data about thickness prediction of a first film layer; inputting the obtained relevant data about the first film layer thickness prediction into the mulching film layer thickness prediction model to obtain a first film layer thickness prediction result; wherein the mulching film layer thickness prediction model is obtained by updating model parameters of a global film layer thickness prediction model; the global film thickness prediction model is obtained by training production data provided by more than two wafer manufacturers. By adopting the scheme, the accuracy of model prediction can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semiconductor technology, and particularly to a method, device, system, storage medium, and equipment for predicting the thickness of a wafer film layer. Background Art

[0002] In the semiconductor field, it is often necessary to predict the thickness of the wafer film layer. Predicting the thickness of the wafer film layer can not only improve production efficiency and product quality, but also reduce production costs.

[0003] Currently, machine learning models are often used to predict the thickness of the wafer film layer, but the accuracy is poor. Summary of the Invention

[0004] The problem to be solved by the present invention is: how to improve the accuracy of model prediction.

[0005] To solve the above problems, an embodiment of the present invention provides a method for predicting the thickness of a wafer film layer, the method including:

[0006] Obtain relevant data for predicting the thickness of the first film layer;

[0007] Input the obtained relevant data for predicting the thickness of the first film layer into the local film layer thickness prediction model to obtain the first film layer thickness prediction result;

[0008] Wherein, the local film layer thickness prediction model is obtained after parameter update using the model parameters of the global film layer thickness prediction model; the global film layer thickness prediction model is trained using production data provided by two or more wafer manufacturing plants.

[0009] In a possible embodiment, the global film layer thickness prediction model is trained using the following method:

[0010] Obtain the production data provided by the two or more wafer manufacturing plants and the model parameters of the initial local film layer thickness prediction models of the two or more wafer manufacturing plants;

[0011] Generate an initial global film layer thickness prediction model based on the model parameters of the initial local film layer thickness prediction models of the two or more wafer manufacturing plants;

[0012] Perform aggregation processing on the production data provided by the two or more wafer manufacturing plants;

[0013] Train the initial global film layer thickness prediction model using the aggregated data to obtain the global film layer thickness prediction model.

[0014] In a possible embodiment, before inputting the obtained relevant data for predicting the thickness of the first film layer into the local film layer thickness prediction model, it further includes:

[0015] Obtain the model parameters of the global film thickness prediction model;

[0016] Update the initial local film thickness prediction model by using the model parameters of the global film thickness prediction model to obtain a local film thickness prediction model.

[0017] In a possible embodiment, the production data provided by the two or more wafer fabs obtained is encrypted data; the model parameters of the initial local film thickness prediction models of the two or more wafer fabs obtained are encrypted model parameters.

[0018] In a possible embodiment, the initial local film thickness prediction model and the initial global film thickness prediction model are the same machine learning model.

[0019] In a possible embodiment, both the initial local film thickness prediction model and the initial global film thickness prediction model are gradient boosting models.

[0020] In a possible embodiment, the initial local film thickness prediction model is obtained by training an initial machine learning model with the production data provided by the wafer fab to which it belongs.

[0021] In a possible embodiment, after aggregating the production data provided by the two or more wafer fabs, it further includes:

[0022] Screen the production data provided by the two or more wafer fabs.

[0023] In a possible embodiment, the relevant data for the first film thickness prediction includes: process parameter data, film property data, wafer position and environmental data.

[0024] An embodiment of the present invention further provides a wafer film thickness prediction device, and the method includes:

[0025] An acquisition unit, adapted to acquire relevant data for the first film thickness prediction;

[0026] A prediction unit, adapted to input the acquired relevant data for the first film thickness prediction into the local film thickness prediction model to obtain a first film thickness prediction result;

[0027] Wherein, the local film thickness prediction model is obtained after parameter update by using the model parameters of the global film thickness prediction model; the global film thickness prediction model is obtained by training with the production data provided by two or more wafer fabs.

[0028] An embodiment of the present invention further provides a wafer film layer thickness prediction system, the system comprising:

[0029] Two or more of the above-mentioned wafer film layer thickness prediction devices, each of the wafer film layer thickness prediction devices being distributed in different wafer manufacturing plants;

[0030] And a central server, connected to each of the wafer film layer thickness prediction devices, adapted to train a global film layer thickness prediction model by using production data provided by the two or more wafer manufacturing plants, and update the model parameters of the local film layer thickness prediction model through the model parameters of the global film layer thickness prediction model.

[0031] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the steps of any one of the above methods.

[0032] An embodiment of the present invention further provides an electronic device, comprising a memory and a processor, a computer program capable of running on the processor being stored on the memory, and the processor executing the steps of any one of the above methods when running the computer program.

[0033] Compared with the prior art, the technical solution of the embodiment of the present invention has the following advantages:

[0034] Applying the solution of the present invention, after obtaining relevant data on the first film layer thickness prediction, the obtained relevant data on the first film layer thickness prediction is input into the local film layer thickness prediction model to obtain the first film layer thickness prediction result. Since the local film layer thickness prediction model is obtained by updating the model parameters by using the model parameters of the global film layer thickness prediction model, and the global film layer thickness prediction model is trained by using production data provided by two or more wafer manufacturing plants, the accuracy of the local film layer thickness prediction model can be improved, and thus the accuracy of the film layer thickness prediction can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 is a flowchart of a method for predicting the thickness of a wafer film layer in an embodiment of the present invention;

[0036] Figure 2 is a flowchart of a method for training a global film layer thickness prediction model in an embodiment of the present invention;

[0037] Figure 3 is a schematic structural diagram of a wafer film layer thickness prediction system in an embodiment of the present invention. DETAILED DESCRIPTION

[0038] Currently, due to the fact that traditional film thickness measurement methods are usually time-consuming and costly, machine learning models are often used to predict the wafer film thickness. This can not only improve production efficiency and product quality but also reduce production costs.

[0039] However, when using machine learning models to predict the wafer film thickness, a large amount of production data is required. Due to data privacy and trade secret issues among different wafer manufacturers, data sharing cannot be achieved, resulting in poor accuracy in predicting the wafer film thickness.

[0040] To address this problem, the present invention provides a method for predicting the wafer film thickness. By using this method, each wafer manufacturer can use the local film thickness prediction model to predict the wafer film thickness. Since the model parameters of the local film thickness prediction model are obtained by updating the model parameters of the global film thickness prediction model, and the global film thickness prediction model is trained using production data provided by two or more wafer manufacturers, the model parameters of the local film thickness prediction model are more accurate, thereby improving the accuracy of predicting the wafer film thickness.

[0041] To make the above objects, features, and advantages of the present invention more apparent and understandable, the following detailed description of the specific embodiments of the present invention will be provided in conjunction with the accompanying drawings.

[0042] Figure 1 It is a flowchart of the method for predicting the wafer film thickness in an embodiment of the present invention. Referring to Figure 1 , the method may include:

[0043] Step 11, obtaining relevant data for predicting the thickness of the first film.

[0044] In a specific implementation, the first film can be any film on the wafer, such as a metal film, a semiconductor film, etc.

[0045] In a specific implementation, the relevant data for predicting the thickness of the first film includes relevant data that can be used to predict the thickness of the first film, including but not limited to process parameter data, film property data, wafer position and environment data. Among them, the process parameter data may include: deposition temperature, deposition time, gas flow rate, gas concentration, chamber pressure, etc.; the film property data may include: film reflectivity, film refractive index, actual film measurement value, film material, etc.; the wafer position and environment data may include: wafer position data in the furnace tube, environmental temperature, pressure, humidity data, etc.

[0046] Step 12, inputting the obtained relevant data for predicting the thickness of the first film into the local film thickness prediction model to obtain the prediction result of the thickness of the first film.

[0047] Among them, the local film thickness prediction model updates its model parameters by using the model parameters of the global film thickness prediction model; the global film thickness prediction model is trained using production data provided by more than two wafer fabs.

[0048] In a specific implementation, the model parameters of the local film thickness initial prediction model can be updated by using the model parameters of the global film thickness prediction model. Among them, the local film thickness initial prediction model is usually set in each wafer fab, while the global film thickness prediction model is set in the central server. The central server is communicatively connected to each wafer fab, so that it can obtain the production data of each wafer fab to train the global film thickness prediction model, and then use the global film thickness prediction model to update the model parameters of each local film thickness initial prediction model. In this way, each wafer fab can, without sharing sensitive data, have the central server train the global film thickness prediction model, enabling each wafer fab to cooperate to achieve virtual metrology, thereby obtaining the local film thickness prediction model, and thus improving the accuracy of the local film thickness prediction model.

[0049] In a specific implementation, the local film thickness initial prediction model can be obtained by training an initial machine learning model using the production data provided by the wafer fab to which it belongs, that is, using the data related to film thickness prediction in the local data that the wafer fab can obtain to train the initial machine learning model to obtain the local film thickness initial prediction model. Each wafer fab trains the local film thickness initial prediction model only using the data it can obtain. Since the data that each wafer fab can obtain is different, the model parameters of each local film thickness initial prediction model are not the same.

[0050] In a specific implementation, for each wafer fab, the same federated learning algorithm and initial machine learning model can be deployed. Each wafer fab can use local production data to obtain the local film thickness initial prediction model from the initial machine learning model. Among them, during the local training process, the model parameters of the local film thickness initial prediction model can be updated according to the local data so that the local film thickness initial prediction model can best adapt to the characteristics of the local data.

[0051] After obtaining the local film thickness initial prediction model of each wafer fab, the production data used to train the local film thickness initial prediction model and the model parameters of the local film thickness initial prediction model are sent to the central server, and the central server trains the global film thickness initial prediction model based on the received data to obtain the global film thickness prediction model.

[0052] Specifically, referring to Figure 2, embodiments of the present invention also provide a method for training a global film thickness prediction model, the method may include:

[0053] Step 21, obtain production data provided by the two or more wafer fabs and model parameters of the initial local film thickness prediction models of the two or more wafer fabs.

[0054] In a specific implementation, the two or more wafer fabs refer to some or all of the wafer fabs communicatively connected to the central server. Each wafer fab provides production data for training the initial local film thickness prediction model and model parameters of the initial local film thickness prediction model.

[0055] In an embodiment of the present invention, the production data provided by the two or more wafer fabs obtained by the central server is encrypted data; the model parameters of the initial local film thickness prediction models of the two or more wafer fabs obtained by the central server are encrypted model parameters.

[0056] That is to say, before each wafer fab sends the production data and model parameters to the central server, it can first encrypt the production data and model parameters to ensure that the data is not stolen or tampered with during the transmission process, thereby providing the security of data transmission and further protecting the commercial interests and data privacy of the manufacturers.

[0057] After receiving the data, the central server can obtain the real production data and model parameters through decryption.

[0058] Step 22, generate an initial global film thickness prediction model based on the model parameters of the initial local film thickness prediction models of the two or more wafer fabs.

[0059] In a specific implementation, the initial local film thickness prediction model and the initial global film thickness prediction model can be the same machine learning model. For example, both the initial local film thickness prediction model and the initial global film thickness prediction model can be gradient boosting models.

[0060] After obtaining the model parameters of the initial local film thickness prediction models of each wafer fab, the model parameters of the initial local film thickness prediction models of each wafer fab can be processed in various ways to obtain the model parameters of the initial global film thickness prediction model, and then the initial global film thickness prediction model can be obtained.

[0061] For example, the model parameters of the initial local film thickness prediction models of each wafer fab at the same position can be averaged to obtain the model parameters of the initial global film thickness prediction model at that position.

[0062] Taking the initial local film thickness prediction model and the initial global film thickness prediction model as y = ax + b as an example, the model parameters provided by wafer fab 1 are a1 and b1, the model parameters provided by wafer fab 2 are a2 and b2, ……, the model parameters provided by wafer fab n are an and bn. Then, in the initial global film thickness prediction model, a = (a1 + a2 + …… + an) / n, b = (b1 + b2 + …… + bn) / n. Based on the values of a and b, the initial global film thickness prediction model can be obtained.

[0063] Step 23: Aggregate the production data provided by the two or more wafer fabs.

[0064] In a specific implementation, after obtaining the production data provided by each wafer fab, the production data provided by each wafer fab can be simply summarized to form a production data set. Subsequently, the initial global film thickness prediction model can be trained using this production data set to obtain the global film thickness prediction model.

[0065] In some embodiments, after aggregating the production data provided by the two or more wafer fabs, the production data provided by the two or more wafer fabs can also be screened. For example, according to the characteristics of the current film, the production data with a relatively low correlation with the film thickness prediction can be removed to improve the training efficiency. Obvious abnormal data can also be removed from the production data set. The data after screening is used as the data for training the initial global film thickness prediction model.

[0066] Step 24: Train the initial global film thickness prediction model using the data after aggregation to obtain the global film thickness prediction model.

[0067] In a specific implementation, the initial global film thickness prediction model is trained using the production data set or the production data set after screening to obtain the global film thickness prediction model.

[0068] In a specific implementation, after obtaining the global film thickness prediction model, the central server can send the model parameters of the obtained global film thickness prediction model to each wafer fab, and each wafer fab updates the model parameters of the local film thickness prediction model according to the model parameters of the global film thickness prediction model.

[0069] In one embodiment, the model parameters of the local film thickness initial prediction model can be directly replaced with the model parameters of the global film thickness prediction model by each wafer fab, so as to obtain the local film thickness prediction model.

[0070] In a specific implementation, when the local data of a wafer fab changes, the wafer fab can send the changed data to the central server again to update the training samples of the central server, thereby optimizing the global film thickness prediction model, and finally optimizing the local film thickness prediction model. Thus, through multiple cycles, the overall performance of the global film thickness prediction model and the local film thickness prediction model can be gradually improved, and the prediction accuracy can be enhanced.

[0071] In an embodiment of the present invention, referring to Figure 1 , before inputting the obtained relevant data for the first film thickness prediction into the local film thickness prediction model, that is, before performing step 12, the method may further include the following steps:

[0072] Step 13, obtaining the model parameters of the global film thickness prediction model.

[0073] In a specific implementation, obtain the model parameters of the global film thickness prediction model sent by the central server. The model parameters of the global film thickness prediction model may be sent encrypted. Each wafer fab can decrypt the received data to obtain the model parameters of the global film thickness prediction model.

[0074] Step 14, updating the local initial film thickness prediction model with the model parameters of the global film thickness prediction model to obtain the local film thickness prediction model.

[0075] In a specific implementation, each wafer fab can directly use the model parameters of the global film thickness prediction model to replace the model parameters of the local initial film thickness prediction model, thereby obtaining the local film thickness prediction model, and subsequently using the local film thickness prediction model to predict the thickness of the first film.

[0076] Adopting the solution of the present invention, the present invention realizes the cooperative virtual metrology of wafer film thickness across wafer fabs through the federated learning algorithm. Each wafer fab does not need to share sensitive data, protecting data privacy and business secrets. At the same time, this method can improve the efficiency and accuracy of virtual metrology of wafer film thickness, shorten the product development cycle, reduce costs, and promote the technological innovation and development of the semiconductor industry.

[0077] To enable those skilled in the art to better understand and implement the present invention, the corresponding apparatuses, systems, electronic devices, and computer-readable storage media of the above method are described in detail below.

[0078] Figure 3 It is a schematic structural diagram of a wafer film thickness prediction system in an embodiment of the present invention. Referring to Figure 3, the wafer film thickness prediction system may include: more than two wafer film thickness prediction devices, and a central server 30.

[0079] Among them, each wafer film thickness prediction device is distributed in different wafer manufacturing factories. One wafer film thickness prediction device can be set in each wafer manufacturing factory. For example, for n wafer manufacturing factories, one wafer film thickness prediction device is respectively set in each wafer manufacturing factory, and one local film thickness prediction model is set in each wafer film thickness prediction device. This wafer film thickness prediction device can obtain relevant data for film thickness prediction in real time, and use the corresponding local film thickness prediction model to obtain the film thickness prediction result. Each local film thickness prediction model is obtained by training the local film thickness initial prediction model.

[0080] The central server 30 can be communicatively connected to n wafer manufacturing factories respectively. The n wafer manufacturing factories respectively send the production data for training the local film thickness initial prediction model and the model parameters of this local film thickness initial prediction model to the central server 30, and the central server 30 uses the received data to obtain the global film thickness prediction model.

[0081] In an embodiment of the present invention, each wafer film thickness prediction device may include: an acquisition unit and a prediction unit. Among them, the acquisition unit can acquire relevant data for the first film thickness prediction, and the prediction unit can input the acquired relevant data for the first film thickness prediction into the local film thickness prediction model to obtain the first film thickness prediction result.

[0082] In an embodiment of the present invention, the prediction unit may include: a model training subunit and a model inference subunit. Among them, the model training subunit can use the acquisition unit to acquire production data related to film thickness prediction, and use the acquired production data to train the initial machine learning model to obtain the local film thickness initial prediction model. The model inference subunit can use the global film thickness prediction model to update the model parameters of the local film thickness initial prediction model to obtain the local film thickness prediction model.

[0083] For example, referring to Figure 3 , the prediction unit of the wafer film thickness prediction device in wafer manufacturing factory 1 may include: a model training subunit 101 and a model inference subunit 102. The prediction unit of the wafer film thickness prediction device in wafer manufacturing factory 2 may include: a model training subunit 201 and a model inference subunit 202.... The prediction unit of the wafer film thickness prediction device in wafer manufacturing factory n may include: a model training subunit n01 and a model inference subunit n02.

[0084] In some embodiments, each wafer film layer thickness prediction device may further include an encryption unit. Wherein, the encryption unit may encrypt the model parameters of the local film layer thickness initial prediction model and the production data related to film layer thickness prediction obtained by the acquisition unit.

[0085] In a specific implementation, the encrypted data is sent to the central server 30. After training to obtain the global film layer thickness prediction model, the central server 30 will send the model parameters of the global film layer thickness prediction model to the corresponding wafer manufacturing factory in an encrypted manner. At this time, each wafer film layer thickness prediction device may further include a decryption unit, and the decryption unit may decrypt the model parameters sent by the central server 30, and then the model inference subunit may obtain the local film layer thickness prediction model according to the decrypted model parameters.

[0086] For example, referring to Figure 3 , the wafer film layer thickness prediction device in wafer manufacturing factory 1 may further include: an encryption unit 103 and a decryption unit 104. The wafer film layer thickness prediction device in wafer manufacturing factory 2 may further include: an encryption unit 203 and a decryption unit 204.... The wafer film layer thickness prediction device in wafer manufacturing factory n may further include: an encryption unit n03 and a decryption unit n04.

[0087] Regarding the above respective functional units, specific implementation may refer to the description of the corresponding steps above, which will not be elaborated here.

[0088] Adopting the solution of the present invention, on the one hand, data privacy protection in virtual measurement of wafer film layer thickness is achieved through the federated learning algorithm. Traditional virtual measurement methods for wafer film layer thickness often require sharing a large amount of production data, but this data may contain sensitive information such as manufacturing processes and product designs. Different from traditional methods, the present invention trains the local film layer thickness initial prediction model on the local devices of each wafer manufacturing factory without sharing the original data, thus protecting data privacy. On the other hand, each wafer manufacturing factory can jointly train to obtain the global film layer thickness prediction model without sharing sensitive data, thereby realizing cooperative virtual measurement and promoting industry cooperation and innovation.

[0089] In addition, before the model parameters of the local film layer thickness initial prediction model of each wafer manufacturing factory are sent to the central server for aggregation, the model parameters of each wafer manufacturing factory will undergo encryption processing to ensure that the data will not be stolen or tampered with during the transmission process. This encryption protection mechanism improves data security and further safeguards the commercial interest data privacy of the wafer manufacturing factory.

[0090] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and the computer program is executed by a processor to implement the steps of any one of the above methods.

[0091] In a specific implementation, the computer-readable storage medium may include: ROM, RAM, a magnetic disk, an optical disc, etc.

[0092] An embodiment of the present invention further provides an electronic device, which includes a memory and a processor. A computer program capable of running on the processor is stored on the memory. When the processor runs the computer program, it executes the steps of any of the above methods.

[0093] Regarding each module / unit included in the various devices and products described in the above embodiments, it can be a software module / unit, a hardware module / unit, or it can also be partly a software module / unit and partly a hardware module / unit. For example, for each device and product applied to or integrated into a chip, each module / unit included therein can all be implemented in a hardware manner such as a circuit. Or, at least some of the modules / units can be implemented in the form of a software program that runs on a processor integrated inside the chip, and the remaining (if any) part of the modules / units can be implemented in a hardware manner such as a circuit; for each device and product applied to or integrated into a chip module, each module / unit included therein can all be implemented in a hardware manner such as a circuit. Different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components of the chip module. Or, at least some of the modules / units can be implemented in the form of a software program that runs on a processor integrated inside the chip module, and the remaining (if any) part of the modules / units can be implemented in a hardware manner such as a circuit; for each device and product applied to or integrated into a terminal, each module / unit included therein can all be implemented in a hardware manner such as a circuit. Different modules / units can be located in the same component (such as a chip, a circuit module, etc.) or different components inside the terminal. Or, at least some of the modules / units can be implemented in the form of a software program that runs on a processor integrated inside the terminal, and the remaining (if any) part of the modules / units can be implemented in a hardware manner such as a circuit.

[0094] Although the present invention is disclosed as above, the present invention is not limited thereto. Any person skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention should be subject to the scope defined by the claims.

Claims

1. A wafer film thickness prediction method, characterized in that: include: Obtain relevant data on the prediction of the thickness of the first film layer; Inputting the acquired relevant data on the prediction of the thickness of the first film layer into the local film layer thickness prediction model to obtain the prediction result of the thickness of the first film layer; The local film thickness prediction model is obtained by updating the model parameters of the global film thickness prediction model; the global film thickness prediction model is trained using production data provided by more than two wafer manufacturing plants.

2. The wafer film thickness prediction method according to claim 1, characterized in that: The global film thickness prediction model is trained using the following method: Acquire production data provided by the two or more wafer manufacturing plants and model parameters of local film thickness prediction initial models of the two or more wafer manufacturing plants; Generate a global film thickness initial prediction model based on model parameters of local film thickness prediction initial models of more than two wafer manufacturing plants; Aggregate the production data provided by the two or more wafer manufacturing plants; The global film thickness initial prediction model is trained using the aggregated data to obtain the global film thickness prediction model.

3. The wafer film thickness prediction method according to claim 2, characterized in that: Before inputting the acquired relevant data on the first film thickness prediction into the local film thickness prediction model, the method further includes: Obtaining model parameters of the global film thickness prediction model; The local film thickness initial prediction model is updated by using the model parameters of the global film thickness prediction model to obtain a local film thickness prediction model.

4. The wafer film thickness prediction method according to claim 2, characterized in that: The production data obtained from the two or more wafer manufacturing plants are encrypted data; the model parameters of the local film layer thickness initial prediction model obtained from the two or more wafer manufacturing plants are encrypted model parameters.

5. The wafer film thickness prediction method according to claim 2, wherein: The local film thickness initial prediction model and the global film thickness initial prediction model are the same machine learning model.

6. The wafer film thickness prediction method according to claim 2, characterized in that: The local film thickness initial prediction model and the global film thickness initial prediction model are both gradient boosting models.

7. The wafer film thickness prediction method according to claim 2, characterized in that: The local film thickness initial prediction model is obtained by training an initial machine learning model using production data provided by the wafer manufacturing plant to which it belongs.

8. The wafer film thickness prediction method according to claim 2, characterized in that: After aggregating the production data provided by the two or more wafer manufacturing plants, the method further includes: The production data provided by the two or more wafer manufacturing plants are screened and processed.

9. The wafer film thickness prediction method according to claim 1, wherein: The data related to the prediction of the thickness of the first film layer include: process parameter data, film layer characteristic data, wafer position and environment data.

10. A wafer film thickness prediction device, characterized in that: include: An acquisition unit, adapted to acquire relevant data on the prediction of the thickness of the first film layer; A prediction unit, adapted to input the acquired relevant data on the prediction of the thickness of the first film layer into a local film layer thickness prediction model to obtain a prediction result of the thickness of the first film layer; The local film thickness prediction model is obtained by updating the model parameters of the global film thickness prediction model; the global film thickness prediction model is trained using production data provided by more than two wafer manufacturing plants.

11. A wafer film thickness prediction system, characterized in that: include: Two or more wafer film thickness prediction devices according to claim 10, each of the wafer film thickness prediction devices is distributed in different wafer manufacturing plants; And a central server is connected to each of the wafer film thickness prediction devices, suitable for using the production data provided by the two or more wafer manufacturing plants to train a global film thickness prediction model, and updating the model parameters of the local film thickness prediction model through the model parameters of the global film thickness prediction model.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that: The computer program is executed by a processor to implement the steps of the method according to any one of claims 1 to 9.

13. An electronic device comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, wherein: When the processor runs the computer program, the steps of the method according to any one of claims 1 to 9 are performed.