Integrated circuit core process modeling method and system based on personalized federal learning

By adopting a personalized federated learning framework for cross-level local training and model aggregation in integrated circuit manufacturing, the data distribution differences and data privacy problems between different manufacturing plants or equipment are solved, and efficient and accurate integrated circuit core process modeling is achieved, which improves the generalization ability and adaptability of the model.

CN120012693APending Publication Date: 2025-05-16SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510178750.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

It is difficult for the prior art to realize efficient modeling of the core manufacturing process of integrated circuits without sharing original data. Especially when there are significant differences in the process conditions and data distribution of different manufacturing plants or equipment, the models trained by centralized machine learning methods are insufficient in generalization capabilities and violate the manufacturer's requirements for data privacy.

Method used

Using a framework based on personalized federated learning, the global model is initialized through the server side and distributed to the client. The client performs cross-level local training to generate personalized and general local models, and generates the updated global model through model aggregation. This process is repeated until the preset convergence conditions are met.

Benefits of technology

It significantly improves the accuracy, stability, generalization ability and adaptability of the model, and realizes efficient and accurate modeling of the core manufacturing process of integrated circuits without sharing original data, while protecting data privacy and improving the adaptability and scalability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012693A_ABST
    Figure CN120012693A_ABST
Patent Text Reader

Abstract

The invention provides an integrated circuit core process modeling method and system based on personalized federal learning, and the method comprises the steps: a server initializes a global model, and transmits the global model to a plurality of clients; after the client side receives the global model, cross-level local training is executed, and local models are generated and comprise a personalized local model and a universal local model; the client uploads the local model to the server side, and the server side generates an updated global model through model aggregation; and repeating the steps until a preset convergence condition is met, and obtaining a global model corresponding to each core process flow of the integrated circuit. The method is used for realizing efficient and accurate core process modeling on the premise of protecting data privacy, a personalized model and a local model are generated through cross-level local training so as to improve the generalization ability of the model, and model convergence is accelerated and oscillation behaviors are reduced through an integrated optimization method, so that the modeling efficiency is improved. And high precision can still be kept under the condition that the data volume is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of integrated circuit manufacturing and computer science and technology, and in particular to a method and system for modeling an integrated circuit core process based on a personalized federated learning framework, as well as a corresponding computer terminal and a computer-readable storage medium. Background Art

[0002] Integrated Circuit (IC) manufacturing is the core of the modern electronics industry, and its technological progress has directly promoted the improvement of electronic product performance. In the IC manufacturing process, key processes such as lithography, etching, chemical mechanical polishing (CMP) and deposition directly affect the chip's dimensional accuracy, functional reliability and production efficiency. However, with the advancement of Moore's Law, the process size has gradually developed from micron to nanometer and even sub-nanometer, which has brought unprecedented challenges to the modeling and optimization of each manufacturing process.

[0003] The modeling of each manufacturing process link not only needs to consider the physical and chemical properties of the material, but also needs to combine the characteristics of different equipment, resulting in the modeling process involving multiple complex variables. For example, lithography is a key step in transferring the design pattern to the silicon wafer. Slight changes in process parameters may cause distortion of the mask pattern or change in line width, which directly affects chip performance. The etching process is responsible for forming precise structures on the silicon wafer, and its modeling needs to simulate the anisotropic etching behavior of the material. Chemical mechanical polishing is to eliminate the high non-uniformity of the silicon wafer surface and improve the flatness, involving complex mechanical and chemical action mechanisms. The deposition process is used to form a thin film of material on the silicon wafer, and its requirement for precise control of thickness and uniformity increases the complexity of the simulation. As the technology node continues to shrink (for example, from 28nm to 7nm or even 3nm), the impact of these variables on process results becomes more significant.

[0004] In order to build an accurate manufacturing process model, it is usually necessary to collect a large amount of data from different manufacturers. However, this data usually involves highly sensitive technical information, involving detailed contents of mask patterns, process conditions, and equipment parameters. Due to intellectual property protection and data security considerations, data sharing between manufacturers is very limited. In addition, data is also subject to potential malicious attack risks during transmission and storage. Traditional centralized machine learning methods require data to be centralized to a server for training. However, due to significant differences in process conditions and data distribution between different manufacturers or equipment, the models trained by centralized methods may perform poorly when generalized to new processes or equipment. For example, the etching gas composition or light source wavelength of lithography equipment used by different manufacturers may result in significantly different data distributions. In addition, centralized methods also require all data to be uploaded to a central server, which violates the manufacturer's requirements for data privacy. Therefore, how to achieve efficient modeling of the core manufacturing process of integrated circuits without sharing the original data and significantly improve the accuracy, stability, and adaptability of the model is a key issue to be solved.

[0005] At present, no description or report of similar technology to the present invention has been found, and similar information at home and abroad has not been collected yet. Summary of the invention

[0006] In view of the above-mentioned deficiencies in the prior art, the present invention provides an integrated circuit core process modeling method and system based on a personalized federated learning framework, and also provides a corresponding computer terminal and computer-readable storage medium for realizing efficient and accurate integrated circuit core manufacturing process modeling, which is suitable for various process steps such as lithography, etching, chemical mechanical polishing, and deposition.

[0007] According to one aspect of the present invention, a method for modeling an integrated circuit core process based on personalized federated learning is provided, comprising:

[0008] The server initializes the global model and sends the global model to multiple clients;

[0009] After receiving the global model, the client performs cross-level local training to generate local models, including: personalized local models and general local models;

[0010] The client uploads the local model to the server, and the server generates an updated global model through model aggregation;

[0011] Repeat the above steps until the preset convergence conditions are met, and obtain the global model corresponding to each core process flow of the integrated circuit.

[0012] Preferably, the above method further includes any one or more of the following:

[0013] -Standardize the client's process data and build a unified data structure; including:

[0014] Remove missing, incomplete and / or abnormal data through data cleaning;

[0015] Through data standardization, the data of different clients are normalized or standardized according to unified standards;

[0016] Expand the dataset through data augmentation methods;

[0017] -Through dynamic initialization, online synchronization and extended evaluation, new clients can be seamlessly connected, thereby achieving system adaptability and scalability; among which:

[0018] The dynamic initialization provides the new client with the initialization weights of the global model to adapt to the local process conditions;

[0019] The online synchronization supports the dynamic participation of new clients in model training without interrupting the existing federated learning process;

[0020] The extended evaluation dynamically adjusts the regularization weights by evaluating the data distribution and process conditions of a new client when the new client joins.

[0021] -Through additional model optimization, the global model of each core process flow of the corresponding integrated circuit is further optimized, including:

[0022] Through model pruning, redundant parameters are removed to improve model reasoning efficiency;

[0023] By quantizing the model, the model storage size is reduced and the running performance on resource-constrained devices is optimized;

[0024] Adopting model distillation, by using the global model as the teacher model, the performance of the client model is further improved;

[0025] -Protect client data privacy during data and model transmission, including:

[0026] Differential privacy is used to protect sensitive information of model parameters by adding noise.

[0027] Through secure multi-party computing, encrypted calculations are performed between the client and the server to avoid data leakage;

[0028] Through homomorphic encryption, the client's local model parameters are encrypted and uploaded to ensure data security;

[0029] -Analyze the process data characteristics of the client to provide optimization strategies for the training process; wherein the analysis includes:

[0030] Data distribution analysis, used to detect data characteristics such as skewness and kurtosis;

[0031] Process complexity assessment, used to dynamically adjust the model training depth based on the process parameters provided by the client;

[0032] Data quality assessment, used to provide data enhancement suggestions for clients with low-quality data;

[0033] - Real-time monitoring of model training status and performance, including:

[0034] Convergence monitoring, used to analyze the convergence speed and training stability of the model;

[0035] Performance evaluation, used to evaluate the prediction accuracy of global and local models;

[0036] Anomaly detection, used to automatically identify anomalies during training.

[0037] According to a second aspect of the present invention, there is provided an integrated circuit core process modeling system based on personalized federated learning, comprising: a model initialization module and a model aggregation module arranged on a server side and a cross-level local training module arranged on a client side; wherein:

[0038] The model initialization module is used to initialize the global model and send the global model to the local training models of multiple clients;

[0039] The cross-level local training model is used to perform local training after receiving the global model, generate a local model and send it to the model aggregation module, wherein the local model includes: a personalized local model and a general local model;

[0040] The model aggregation module is used to generate an updated global model through model aggregation to obtain a global model corresponding to each core process flow of the integrated circuit.

[0041] Preferably, the above system further includes any one or more of the following modules:

[0042] Data processing module, which is used to standardize the process data of the client and build a unified data structure;

[0043] Dynamic expansion module, which is used to support dynamic access of new clients to ensure system adaptability and scalability;

[0044] Model optimization module, which is used to prune, quantize, and distill the global model to improve model efficiency and performance;

[0045] Privacy protection module, which uses differential privacy and encryption technology to ensure privacy security during data transmission;

[0046] Process characteristic analysis module, which is used to analyze the characteristics of client process data and provide optimization strategies for model training.

[0047] According to a third aspect of the present invention, there is provided a computer terminal comprising a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, can be used to execute any of the above-described methods of the present invention, or to run any of the above-described systems of the present invention.

[0048] According to a fourth aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can be used to execute any of the methods described above in the present invention, or to run any of the systems described above in the present invention.

[0049] Due to the adoption of the above technical solution, the present invention has at least one of the following beneficial effects compared with the prior art:

[0050] Improve model generalization ability: Through the cross-level local training method, the model of the present invention can adapt to the process and equipment conditions of different manufacturers, significantly improving the generalization performance.

[0051] Improve training stability: The integrated optimization method of the present invention effectively avoids the oscillation of the loss curve during training and accelerates the convergence of the model.

[0052] Protecting data privacy: The present invention implements distributed training of the model through a federated learning framework, without sharing original data, thus avoiding the risk of data leakage.

[0053] Enhanced system scalability: The present invention supports a dynamic client access mechanism, which allows new devices or manufacturers to be added without restarting model training, significantly improving system adaptability.

[0054] Multi-process adaptability: The present invention is applicable to multiple core processes of integrated circuit manufacturing (such as lithography, etching, polishing and deposition), meeting diverse modeling requirements. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:

[0056] Figure 1 The present invention is a flowchart of an integrated circuit core process modeling method based on personalized federated learning in one embodiment of the present invention.

[0057] Figure 2The figure is a schematic diagram of the component modules of an integrated circuit core process modeling system based on personalized federated learning in one embodiment of the present invention.

[0058] Figure 3 This is a working architecture diagram of an integrated circuit core process modeling method based on personalized federated learning in a preferred embodiment of the present invention. DETAILED DESCRIPTION

[0059] The following is a detailed description of the embodiments of the present invention: This embodiment is implemented on the premise of the technical solution of the present invention, and a detailed implementation method and a specific operation process are given. It should be pointed out that for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, which all belong to the protection scope of the present invention.

[0060] In the prior art, traditional centralized machine learning methods require data to be centralized on a server for training. However, due to significant differences in process conditions and data distribution among different manufacturers or equipment, the models trained by centralized methods may perform poorly when generalized to new processes or equipment. In addition, centralized methods also require all data to be uploaded to a central server, which violates the manufacturer's requirements for data privacy. In response to the above problems, an embodiment of the present invention provides an integrated circuit core process modeling method based on personalized federated learning, which generates personalized models and general local models adapted to different clients through cross-level local training, and combines multiple model aggregation methods to effectively solve the problems of uneven data distribution and heterogeneous process conditions, significantly improve the accuracy, stability, generalization ability and adaptability of the global model, and achieve efficient and accurate modeling of the integrated circuit core manufacturing process without sharing the original data.

[0061] Specifically, Figure 1 As shown, the integrated circuit core process modeling method based on personalized federated learning provided by this embodiment may include:

[0062] S1, the server initializes the global model and sends the global model to multiple clients;

[0063] S2, after receiving the global model, the client performs cross-level local training to generate local models, including: personalized local models and general local models;

[0064] S3, the client uploads the local model to the server, and the server generates an updated global model through model aggregation;

[0065] S4, repeating steps S1 to S3 until a preset convergence condition is met, and obtaining a global model corresponding to each core process flow of the integrated circuit.

[0066] In some preferred embodiments, the global model refers to a specific model adapted to different manufacturing processes and designed according to task and precision requirements, including: chemical mechanical polishing modeling, lithography modeling, etching modeling and deposition modeling, etc. Its model types include: generative models and regression prediction models, and its model structures include: generative adversarial networks, variational autoencoders, Transformers, multi-layer perceptrons, long short-term memory networks, etc.

[0067] In some preferred implementations, the above S1, initializing the global model, may further include any one or more of the following operations:

[0068] S11, for scenarios with relatively uniform data distribution (i.e., scenarios where the mean deviation and standard deviation of data samples vary less than a set threshold), a random initialization method based on uniform distribution is used to initialize the global model; including:

[0069] Initialize the parameters of the global model to uniformly distributed random values ​​to avoid bias between the initial parameters, then:

[0070]

[0071] In the formula, represents the initialized global model parameters, and [a,b] are the upper and lower bounds of the uniform distribution.

[0072] S12: For scenarios with unknown data distribution or complex scenarios (i.e., the mean deviation and standard deviation of data samples vary by more than a set threshold), a random initialization method based on normal distribution is used to initialize the global model; including:

[0073] The parameters of the global model are initialized using normal distribution to ensure that the model weights are close to the zero center, then:

[0074]

[0075] Where μ is the mean of the normal distribution, usually 0, and σ 2 is the variance.

[0076] S13, for scenarios with similar process conditions, the global model is initialized using a migration initialization method based on an existing model, including:

[0077] Use a trained model (such as the model of the previous generation process) to initialize the global model parameters to accelerate convergence, then:

[0078]

[0079] In the formula, ω pretrained Represents the trained model parameters.

[0080] In some preferred implementations, the above S2, performing cross-level local training to generate a local model, may further include any one or more of the following operations:

[0081] S21, single-level local training: When the data distribution and process conditions of all clients are exactly the same, there is no need to generate a personalized model. At this time, a universal local model is directly generated through a single-level local training mechanism. In this case, all clients share the same training objective function, i.e., the loss function, which simplifies the model training process and significantly reduces the computing cost:

[0082]

[0083] In the formula, Denotes the loss function used to measure the general local model ω i Performance on local data.

[0084] S22, multi-level local training: When the data distribution of the client is inconsistent with the process conditions, it is necessary to generate personalized models for different client data. At this time, a personalized local model and a universal local model are generated through a multi-level local training mechanism; wherein, the multi-level local training mechanism includes: inner layer training and outer layer training; the inner layer training aims to generate a personalized local model for the unique process conditions and data distribution of the client; the outer layer training aims to generate a universal local model to integrate the data characteristics of multiple clients and enhance the robustness and adaptability of the model; the objective function in the multi-level local training process consists of a loss function term and a regularization term. The loss function term and the regularization term are integrated and optimized, and then:

[0085]

[0086] In the formula, Represents the loss function used to measure the personalized local model θ i Performance on local data, optimizing the adaptability of personalized local models to local process conditions and data characteristics; represents the regularization term, which is used to constrain the personalized local model θ according to the regularization coefficient λ i With the general local model ω i The degree of deviation between them is used to optimize the global consistency of the personalized local model.

[0087] S23, hybrid hierarchical local training: when the data similarity of some clients is high and the data difference of some clients is large, a hybrid hierarchical local training mechanism is used to generate personalized local models and general local models; among them, for clients with high data similarity, a single-level local training mechanism is used to directly generate a general local model; for clients with large data difference, a multi-level local training mechanism is used to generate a personalized local model through the inner layer training of the multi-level local training mechanism, and to generate a general local model through the outer layer training of the multi-level local training mechanism.

[0088] In some preferred implementations, the above S22, integrated optimization, may further include the following operations:

[0089] The loss function term and regularization term are optimized respectively through different optimization methods, so as to take into account the local adaptability and global consistency of the personalized local model.

[0090] In some preferred implementations, the above-mentioned optimization of the loss function term may further include:

[0091] Select a suitable optimization algorithm according to the task requirements to optimize the loss function terms, including but not limited to the following methods: Adaptive Moment Estimation (Adam), Stochastic Gradient Descent (SGD), Root Mean Square Propagation (RMSProp) and other optimization algorithms.

[0092] In some preferred implementations, the above-mentioned optimization of the regularization term may further include:

[0093] The regularization term is used to control the difference between the personalized local model and the global model to avoid over-personalization or deviation from the stability of the global model, including but not limited to the following methods: Gradient Descent (GD), adaptive regularization weight adjustment, dynamic adjustment of regularization weights according to the training stage, emphasizing personalized adaptation in the early stage, and balancing personalization and global consistency in the later stage.

[0094] In some preferred implementations, the above S3, the method of model aggregation, may further include: any one or more of the following operations: weighted average aggregation, dynamic personalized aggregation, hierarchical aggregation, contrastive learning aggregation, gradient fusion aggregation, entropy weight aggregation, fuzzy weight aggregation; wherein:

[0095] Weighted average aggregation dynamically adjusts weights based on client data volume or process complexity, including:

[0096] Dynamically adjust the aggregation weight according to the data volume, data quality or process complexity of each client, and combine the local model and the global model to generate a new global model:

[0097]

[0098] Where: is the global model after the t+1th iteration; is the global model after the tth iteration; β is the personalized aggregation parameter, which is used to adjust the contribution ratio of the global model and the local model; is the local model parameter of the i-th client, and its contribution is based on the amount of data n of N clients. i Dynamic adjustment.

[0099] Dynamic personalized aggregation dynamically adjusts aggregation parameters based on the similarity between the client and the global model, including:

[0100] In view of the large differences in client data distribution, dynamic personalized aggregation parameters are introduced:

[0101]

[0102] In the formula, is the global model after the t+1th iteration; is the global model after the tth iteration; α i To personalize the aggregation parameters, they are dynamically adjusted by the similarity between the client and the global model;

[0103] Aggregation based on contrastive learning uses the similarity between client models for priority aggregation, including:

[0104] During the aggregation process, the contrastive learning method is used to calculate the similarity between local models of different clients, and the models with higher similarity are aggregated first; where:

[0105] The formula for calculating the comparative similarity is:

[0106]

[0107] In the formula, ω i ,ω j They are represented as the local model parameters of clients i and j respectively;

[0108] Adjust the aggregation weight w according to similarity i for:

[0109]

[0110] In the formula, ω k is the local model parameter of client k, ∑ kΣ j sim(ω k,j ) is the normalized sum of similarities between all clients, which is used to calculate the aggregation weight;

[0111] The final aggregation formula is:

[0112]

[0113] Based on the layered aggregation, clients with similar process characteristics are first aggregated within the group and then globally aggregated, including:

[0114] The clients are grouped according to their process categories or data characteristics. The models in each group are first aggregated to generate the group model. Then, the group models are aggregated globally to generate the final global model. Among them:

[0115] The intra-group aggregation is:

[0116]

[0117] In the formula, is the model parameter within the group representing the kth group; group k represents the client set of the kth group;

[0118] The global aggregation is:

[0119]

[0120] In the formula, v k is the weight of each group, calculated according to the amount of data in the group.

[0121] Gradient-based aggregation updates the global model by fusing the client’s gradients;

[0122] Entropy weight aggregation: dynamically adjust the aggregation weight by calculating the entropy value output by the client model to enhance the robustness of the model;

[0123] Aggregation of fuzzy weights includes: using fuzzy logic reasoning to dynamically generate aggregate weights based on the fuzzy characteristics of client data (such as quality, process complexity), including but not limited to: if the client data quality is high and the process complexity is low, a higher weight is assigned; if the client data quality is low and the process complexity is high, a lower weight is assigned.

[0124] In some preferred embodiments, the above method may further include the following operations:

[0125] Standardize the client's process data and build a unified data structure; including:

[0126] Data cleaning: remove missing, incomplete or abnormal data;

[0127] Data standardization: normalize or standardize the data of different clients according to unified standards;

[0128] Data augmentation: Expanding the dataset through rotation, scaling, perturbation, etc.

[0129] In some preferred embodiments, the above method may further include the following operations:

[0130] Through dynamic initialization, online synchronization and extended evaluation, new clients can be seamlessly connected, thereby achieving system adaptability and scalability; among which:

[0131] Dynamic initialization provides the initialization weights of the global model for new clients to quickly adapt to local process conditions;

[0132] Online synchronization supports dynamic participation of new clients in model training without interrupting the existing federated learning process;

[0133] Extended evaluation, when new clients join, the regularization weights are dynamically adjusted by evaluating their data distribution and process conditions.

[0134] In some preferred embodiments, the above method may further include the following operations:

[0135] Through additional model optimization, the global model of each core process flow of the corresponding integrated circuit is further optimized, including:

[0136] Model pruning, which is used to remove redundant parameters to improve model reasoning efficiency;

[0137] Model quantization, which is used to reduce the model storage size and optimize the running performance on resource-constrained devices;

[0138] Model distillation, by using the global model as a teacher model, further improves the performance of the client model.

[0139] In some preferred embodiments, the above method may further include the following operations:

[0140] Protect client data privacy during data and model transmission, including:

[0141] Differential privacy, which is used to protect sensitive information of model parameters by adding noise;

[0142] Secure multi-party computing, which is used to perform encrypted calculations between clients and servers to avoid data leakage;

[0143] Homomorphic encryption is used to encrypt the client's local model parameters before uploading to ensure data security.

[0144] In some preferred embodiments, the above method may further include the following operations:

[0145] Analyze the process data characteristics of the client and provide optimization strategies for the training process; the analysis includes but is not limited to:

[0146] Data distribution analysis, used to detect data characteristics such as skewness and kurtosis;

[0147] Process complexity assessment, used to dynamically adjust the model training depth based on the process parameters provided by the client;

[0148] Data quality assessment is used to provide data enhancement suggestions to clients with low-quality data.

[0149] In some preferred embodiments, the above method may further include the following operations:

[0150] Real-time monitoring of model training status and performance, including but not limited to:

[0151] Convergence monitoring, used to analyze the convergence speed and training stability of the model;

[0152] Performance evaluation, used to evaluate the prediction accuracy of global and local models;

[0153] Anomaly detection, which is used to automatically identify anomalies during training (such as gradient explosion, training oscillations).

[0154] Based on the same inventive concept, an embodiment of the present invention also provides an integrated circuit core process modeling system based on personalized federated learning.

[0155] Specifically, Figure 2 As shown, the integrated circuit core process modeling system based on personalized federated learning provided by this embodiment may include: a model initialization module and a model aggregation module arranged on the server side and a cross-level local training module arranged on the client side; wherein:

[0156] The model initialization module is used to initialize the global model and send the global model to the local training models of multiple clients;

[0157] A cross-level local training model is used to perform local training after receiving the global model, generate a local model and send it to the model aggregation module, where the local model includes: a personalized local model and a general local model;

[0158] The model aggregation module is used to generate an updated global model through model aggregation to obtain a global model corresponding to each core process flow of the integrated circuit.

[0159] The integrated circuit core process modeling system based on personalized federated learning provided in the above embodiment of the present invention is further described in detail below in conjunction with the preferred implementation manner.

[0160] The integrated circuit core process modeling system based on personalized federated learning provided in this embodiment mainly includes the following modules:

[0161] The model initialization module is used to initialize the global model and send the global model to the local training models of multiple clients;

[0162] A cross-level local training model is used to perform cross-level local training after receiving the global model, generate a local model and send it to the model aggregation module, where the local model includes: a personalized local model and a general local model;

[0163] The model aggregation module is used to generate an updated global model through model aggregation to obtain a global model corresponding to each core process flow of the integrated circuit.

[0164] By continuously running the model initialization module, the cross-level local training model and the model aggregation module, the preset convergence conditions are met.

[0165] In some preferred implementations, the above-mentioned model initialization module is used to initialize the global model, including but not limited to:

[0166] For scenarios where data distribution is relatively uniform, a random initialization method based on uniform distribution is used to initialize the global model;

[0167] For scenarios with unknown or complex data distribution, a random initialization method based on normal distribution is used to initialize the global model;

[0168] For scenarios with similar process conditions, the global model is initialized using a migration initialization method based on existing models.

[0169] In some preferred embodiments, the cross-level local training module is used to perform cross-level local training to generate a personalized local model and a universal local model, including but not limited to the following functions:

[0170] Single-level training: Generates a common local model when all client data distributions and process conditions are consistent;

[0171] Multi-level training: Dynamically adjust the training level according to the process conditions and data distribution of different clients to generate personalized local models and general local models respectively;

[0172] Hybrid-level training: Combining single-level and multi-level training strategies, it can adapt to the situation where some clients have similar data and process conditions, while some clients have large differences;

[0173] In some preferred implementations, the model aggregation module is used to obtain a local model from a client and update a global model through model aggregation; the following aggregation methods are supported:

[0174] Weighted average aggregation: dynamically adjust weights based on client data volume or process complexity;

[0175] Dynamic personalized aggregation: Dynamically adjust aggregation parameters based on the similarity between the client and the global model;

[0176] Hierarchical aggregation: first perform intra-group aggregation on clients with similar process characteristics, and then perform global aggregation;

[0177] Contrastive learning aggregation: using the similarity between client models for priority aggregation;

[0178] Gradient fusion aggregation: Update the global model by fusing the client's gradients;

[0179] Entropy weight aggregation: Dynamically adjust the aggregation weight by calculating the entropy value output by the client model to enhance the robustness of the model.

[0180] In some preferred embodiments, the above system may further include:

[0181] The data processing module is used to standardize the process data of the client and build a unified data structure, including but not limited to the following processing steps:

[0182] Data cleaning: remove missing, incomplete or abnormal data;

[0183] Data standardization: normalize or standardize the data of different clients according to unified standards;

[0184] Data augmentation: Expand the data set through rotation, scaling, perturbation, etc. to improve the generalization ability of the model.

[0185] In some preferred embodiments, the above system may further include:

[0186] Dynamic expansion module is used to support seamless access of new clients and ensure the adaptability and scalability of the system. It includes but is not limited to the following functions:

[0187] Dynamic initialization: Provides initialization weights of the global model for new clients to quickly adapt to local process conditions;

[0188] Online synchronization: supports dynamic participation of new clients in model training without interrupting the existing federated learning process;

[0189] Extended evaluation: When new clients join, the regularization weights are dynamically adjusted by evaluating their data distribution and process conditions.

[0190] In some preferred embodiments, the above system may further include:

[0191] Model Optimization Module: used to perform additional optimization steps after the global model update, including but not limited to:

[0192] Model pruning: remove redundant parameters to improve model reasoning efficiency;

[0193] Model quantization: reduce the model storage size and optimize the running performance on resource-constrained devices;

[0194] Model distillation: Further improve the performance of client models by using the global model as a teacher model.

[0195] In some preferred embodiments, the above system may further include: a privacy protection module: used to protect the client data privacy during the data and model transmission process, including but not limited to:

[0196] Differential privacy: protect sensitive information of model parameters by adding noise;

[0197] Secure multi-party computing: Encrypted computing is performed between the client and the server to avoid data leakage;

[0198] Homomorphic encryption: The client's local model parameters are encrypted before uploading to ensure data security.

[0199] In some preferred embodiments, the above system may further include: a process characteristic analysis module: used to analyze the process data characteristics of the client and provide an optimization strategy for the training process.

[0200] The analysis includes but is not limited to:

[0201] Data distribution analysis: detect data characteristics such as skewness and kurtosis;

[0202] Process complexity assessment: Dynamically adjust the model training depth based on the process parameters provided by the client;

[0203] Data quality assessment: Provide data enhancement suggestions to clients with low-quality data.

[0204] Provide optimization strategies for the training process, including but not limited to: dynamically adjusting model training depth and providing data enhancement suggestions for clients with low-quality data

[0205] In some preferred embodiments, the above system may further include: a model monitoring module: used to monitor the model training status and performance in real time, including but not limited to:

[0206] Convergence monitoring: analyzing the convergence speed and training stability of the model;

[0207] Performance evaluation: Evaluate the prediction accuracy of the global model and the local model;

[0208] Anomaly detection: Automatically identify anomalies during training (such as gradient explosion, training oscillations).

[0209] The working contents of the system provided by the above embodiment of the present invention are further described in detail below.

[0210] Optionally, the above model initialization module includes but is not limited to the following work contents:

[0211] 1. Random initialization based on uniform distribution: Initializing the parameters of the global model to uniformly distributed random values ​​helps avoid bias between the initial parameters and is suitable for scenarios with relatively uniform data distribution.

[0212]

[0213] in, represents the initialized global model parameters, and [a,b] are the upper and lower bounds of the uniform distribution.

[0214] 2. Random initialization based on normal distribution: Use normal distribution to initialize the parameters of the global model. It is often used in scenarios with unknown or complex data distribution to ensure that the model weights are close to the zero center.

[0215]

[0216] Where μ is the mean of the normal distribution (usually 0), σ 2 is the variance.

[0217] 3. Migration initialization based on existing models: Using a trained model (such as the model of the previous generation process) to initialize the global model parameters can accelerate convergence and is suitable for scenarios with similar process conditions.

[0218]

[0219] Among them, ω pretrained Represents the parameters of the pre-trained model.

[0220] The above cross-level local training model includes but is not limited to the following work:

[0221] Single-level training: When the data distribution and process conditions of all clients are exactly the same, there is no need to generate a personalized model, and a general local model can be directly generated through single-level training. In this case, all clients share the same training objective function, namely the loss function, which simplifies the model training process and significantly reduces the computing cost.

[0222]

[0223] in Denotes the loss function, which measures the general local model ω i Performance on local data;

[0224] Multi-level training: When the client data distribution is inconsistent with the process conditions, it is necessary to generate personalized models for different client data. The inner layer training aims to generate personalized local models for the client's unique process conditions and data distribution. The outer layer training aims to generate a universal local model to integrate the data characteristics of multiple clients and enhance the robustness and adaptability of the model. The objective function in the cross-level training process consists of two parts: the loss function term and the regularization term:

[0225]

[0226] The loss function Measuring the personalized model θ i The performance of the model on local data, optimizing the adaptability of the personalized model to local process conditions and data characteristics; regularization term Constrain the personalized model θ according to the regularization coefficient λ i With the global model ω i The degree of deviation between them can enhance the global consistency of the model.

[0227] Hybrid hierarchical training: When some clients have high data similarity and some clients have large differences, a hybrid hierarchical training mechanism is adopted: clients with similar data use single-level training to directly generate a common local model; clients with large data differences use cross-level training to generate personalized local models through inner-layer training and generate a common model through outer-layer training.

[0228] In the above multi-level training, the loss function term and the regularization term are integrated and optimized. The integrated optimization method refers to dividing the cross-level optimization into two parts: the loss function term and the regularization term, and optimizing these two parts separately through different optimization methods, so as to take into account the local adaptability of the personalized model and the consistency of the global model.

[0229] In the above multi-level training, the optimization of the loss function terms is carried out by selecting a suitable optimization algorithm according to the task requirements, including but not limited to the following methods: Adaptive Moment Estimation optimization algorithm (Adaptive Moment Estimation, Adam), Stochastic Gradient Descent algorithm (Stochastic Gradient Descent, SGD), Root Mean Square Propagation optimization algorithm (Root Mean Square Propagation, RMSProp) and other different optimization algorithms.

[0230] In the above multi-level training, the optimization of the regularization term aims to control the difference between the personalized model and the global model, avoid over-personalization or deviation from the stability of the global model, including but not limited to the following methods: Gradient Descent (GD), adaptive regularization weight adjustment, dynamic adjustment of regularization weight according to the training stage, emphasizing personalized adaptation in the early stage, and balancing personalization and global consistency in the later stage.

[0231] The above model aggregation module can flexibly adapt to the needs of different scenarios, such as uniform data distribution, heterogeneous distribution, or significant differences in client characteristics, to further improve the generalization ability and robustness of the global model, including but not limited to the following work content:

[0232] 1. Weighted average aggregation: Dynamically adjust the aggregation weight according to the data volume, data quality or process complexity of each client, and generate a new global model by combining the local model and the global model;

[0233]

[0234] in: is the global model after the t+1th iteration; β is the personalized aggregation parameter, which is used to adjust the contribution ratio of the global model and the local model. ω i is the local model parameter of the i-th client, and its contribution is based on the amount of data n of N clients. i Dynamic adjustment;

[0235] 2. Dynamic personalized aggregation: In view of the large differences in client data distribution, dynamic personalized aggregation parameters are introduced:

[0236]

[0237] In the formula, is the global model after the t+1th iteration; is the global model after the tth iteration; α i To personalize the aggregation parameters, they are dynamically adjusted by the similarity between the client and the global model;

[0238] 3. Aggregation based on contrastive learning: During the aggregation process, the contrastive learning method is used to calculate the similarity between different client models, and the models with higher similarity are aggregated first:

[0239] Comparative similarity calculation formula:

[0240]

[0241] Adjust the aggregation weights based on similarity:

[0242]

[0243] Final aggregation formula:

[0244]

[0245] 4. Based on hierarchical aggregation: Group the clients by process category or data characteristics, aggregate them in each group to generate the group model, and then aggregate the group models globally to generate the final global model:

[0246] Aggregation within a group:

[0247]

[0248] Global aggregation:

[0249]

[0250] Among them, v k is the weight of each group, which can be calculated based on the amount of data in the group.

[0251] 5. Aggregation based on fuzzy weights: Use fuzzy logic reasoning to dynamically generate aggregation weights based on the fuzzy characteristics of client data (such as quality and process complexity), including but not limited to: if the client data quality is high and the process complexity is low, a higher weight is assigned; if the client data quality is low and the process complexity is high, a lower weight is assigned.

[0252] It should be noted that the steps in the method provided by the present invention can be implemented by using the corresponding components in the system, and those skilled in the art can refer to the technical solution of the system to implement the step flow of the method, and can also refer to the technical solution of the method to implement the composition of the system, that is, the embodiments in the system and the embodiments in the method can be understood as preferred examples of each other, which will not be elaborated here.

[0253] The technical solution provided by the above embodiment of the present invention is further described below in conjunction with a specific application example.

[0254] like Figure 3As shown in the figure, the integrated circuit lithography process modeling method based on personalized federated learning used in this specific application example mainly includes: initialization and distribution of the global model, cross-level local training, model uploading and model aggregation. Specifically, it includes the following steps:

[0255] Step S100, initializing the global model, randomly initializing the global model parameters and sending them to multiple clients, and the clients receive the global model and load it into the local model;

[0256] Step S200: The client performs cross-layer local training to generate a personalized model and a local model, wherein the inner-layer training generates a personalized model for local data and process conditions, and the outer-layer training generates a local model adapted to the global goal;

[0257] Step S300, the client uploads the trained local model to the server, and the server generates an updated global model through a model aggregation method;

[0258] Step S400, repeat the above steps until the preset convergence condition is met, and finally obtain the global model.

[0259] in:

[0260] Step S100, the methods of initializing the global model include the following:

[0261] S101, based on uniformly distributed random initialization, initializing the parameters of the global model to uniformly distributed random values;

[0262] S102, based on random initialization of normal distribution, initializing the parameters of the global model with normal distribution;

[0263] S103, based on the migration initialization of the existing model, the global model parameters are initialized using a trained model (such as the model of the previous generation process).

[0264] Step S200, cross-level local training includes the following steps:

[0265] S201, inner layer training generates a personalized local model by optimizing the objective function:

[0266]

[0267] in, represents the loss function of the model on local data, λ is the regularization coefficient, and constrains the personalized model θ i and local models The degree of deviation between

[0268] S202, outer layer training generates a universal local model, optimizes the global model adaptability, and updates local model parameters.

[0269] Step S300: The method of model aggregation includes but is not limited to:

[0270] S301, weighted average aggregation: dynamically adjust the aggregation weight according to the amount of data from each client. The calculation formula is as follows:

[0271]

[0272] in, is the global model after the t+1th iteration; β is the personalized aggregation parameter, which is used to adjust the contribution ratio of the global model and the local model. is the local model parameter of the i-th client, and its contribution is based on the amount of data n of N clients. i Dynamic adjustment;

[0273] S302, Dynamic Personalization Aggregation: Introducing Dynamic Personalization Parameter α i , dynamically adjust the aggregation weight according to the similarity between the client and the global model. The calculation formula is:

[0274]

[0275] where α i Dynamically adjusted by the similarity between the client and the global model; is the global model after the t+1th iteration, is the local model of the i-th client;

[0276] In step S400, it also includes supporting a dynamic expansion mechanism:

[0277] S401, dynamically accessing a new client: providing initialization parameters of a global model for the new client;

[0278] S402, the new client quickly adapts to local data and process conditions through inner layer training;

[0279] S403, balances personalization and global model consistency by dynamically adjusting regularization weights.

[0280] An embodiment of the present invention further provides a computer terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the processor can be used to execute any method of the above-mentioned embodiments of the present invention, or to run any system of the above-mentioned embodiments of the present invention.

[0281] Optionally, the memory is used to store programs; the memory may include volatile memory (English: volatile lememory), such as random-access memory (English: random-access memory, abbreviated: RAM), such as static random-access memory (English: static random-access memory, abbreviated: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic Random Access Memory, abbreviated: DDR SDRAM), etc.; the memory may also include non-volatile memory (English: non-volatile memory), such as flash memory (English: flash memory). The memory is used to store computer programs (such as applications, functional modules, etc. that implement the above method), computer instructions, etc., and the above computer programs, computer instructions, etc. can be partitioned and stored in one or more memories. And the above computer programs, computer instructions, data, etc. can be called by the processor.

[0282] The processor is used to execute the computer program stored in the memory to implement the various steps of the method or various modules of the system involved in the above embodiments. For details, please refer to the relevant descriptions in the above method and system embodiments.

[0283] The processor and the memory may be independent structures or integrated structures. When the processor and the memory are independent structures, the memory and the processor may be coupled and connected via a bus.

[0284] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it can be used to execute any method of the above embodiments of the present invention, or to run any system of the above embodiments of the present invention.

[0285] Among them, computer-readable media include computer storage media and communication media, wherein the communication media include any media that facilitates the transmission of computer programs from one place to another. The storage medium can be any available medium that can be accessed by a general or special-purpose computer. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in a user device. Of course, the processor and the storage medium can also be present in a communication device as discrete components.

[0286] The integrated circuit core process modeling method and system based on personalized federated learning provided by the above embodiment of the present invention are adapted to the modeling requirements of different manufacturing processes, and specific models are designed according to the tasks and accuracy requirements, such as chemical mechanical polishing modeling, lithography modeling, etching modeling and deposition modeling, wherein the model type can be a generative model and a regression prediction model, and the model structure can be a generative adversarial network, a variational autoencoder, a Transformer, a multi-layer perceptron and a long short-term memory network. The method of initializing the global model can adopt a random initialization method based on uniform distribution, a random initialization method based on normal distribution, and a migration initialization method based on an existing model. The method of cross-level local training can adopt single-level training, multi-level training and mixed-level training. In multi-level training, the optimization objective function includes a loss function term and a regularization term, wherein the loss function term is used to measure the performance of the personalized model on local data, and the regularization term is used to constrain the degree of deviation between the personalized model and the global model, and enhance the consistency of the model. The model aggregation method can adopt weighted average aggregation, dynamic personalized aggregation, contrastive learning aggregation, hierarchical aggregation, fuzzy weight aggregation, etc. At the same time, the method and system also support dynamic expansion, allowing new clients to join dynamically, and quickly adapt to new client data and process conditions through initialization of model weights, inner layer training and regularization weight adjustment.

[0287] The integrated circuit core process modeling method and system based on personalized federated learning provided by the above-mentioned embodiments of the present invention, through the cross-level local training method, the model can adapt to the process and equipment conditions of different manufacturers, and significantly improve the generalization performance; the integrated optimization method effectively avoids the oscillation phenomenon of the loss curve during training, and accelerates the convergence of the model; the distributed training of the model is realized through the federated learning framework, and there is no need to share the original data, avoiding the risk of data leakage; supports dynamic client access mechanism, and new equipment or manufacturers can be added without restarting model training, which significantly improves the system adaptability; it is suitable for multiple core processes of integrated circuit manufacturing (such as lithography, etching, polishing and deposition), and meets diverse modeling needs.

[0288] All matters not covered in the above embodiments of the present invention are well known in the art.

[0289] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art may make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.

Claims

1. A method for modeling integrated circuit core processes based on personalized federated learning, characterized in that: include: The server initializes the global model and sends the global model to multiple clients; After receiving the global model, the client performs cross-level local training to generate local models, including: personalized local models and general local models; The client uploads the local model to the server, and the server generates an updated global model through model aggregation; Repeat the above steps until the preset convergence conditions are met, and obtain the global model corresponding to each core process flow of the integrated circuit.

2. The integrated circuit core process modeling method based on personalized federated learning according to claim 1 is characterized in that: The initialization of the global model includes any one or more of the following: -For scenarios where the mean deviation and standard deviation of data samples are less than the set threshold, a random initialization method based on uniform distribution is used to initialize the global model; including: Initialize the parameters of the global model to uniformly distributed random values ​​to avoid bias between the initial parameters, then: In the formula, represents the initialized global model parameters, [a,b] are the upper and lower bounds of the uniform distribution; -For scenarios where data distribution is unknown or the mean deviation and standard deviation of data samples vary by more than the set threshold, a random initialization method based on normal distribution is used to initialize the global model; including: The parameters of the global model are initialized using normal distribution to ensure that the model weights are close to the zero center, then: In the formula, μ is the mean of the normal distribution, σ 2 is the variance; -For existing scenarios with the same or similar process conditions, the global model is initialized using the migration initialization method based on the existing model, including: Using a trained model to initialize the global model parameters to accelerate convergence, we have: In the formula, ω pretrained Represents the trained model parameters.

3. The integrated circuit core process modeling method based on personalized federated learning according to claim 1 is characterized in that: The performing of cross-level local training to generate a local model includes any one or more of the following: -When the data distribution and process conditions of all clients are exactly the same, a universal local model is directly generated through a single-level local training mechanism, and all clients share the same training objective function, i.e., loss function: In the formula, Denotes the loss function used to measure the general local model ω i Performance on local data; -When the data distribution of the client is inconsistent with the process conditions, a personalized local model and a universal local model are generated through a multi-level local training mechanism; wherein the multi-level local training mechanism includes: inner layer training and outer layer training; the inner layer training aims to generate a personalized local model for the unique process conditions and data distribution of the client; the outer layer training aims to generate a universal local model to integrate the data characteristics of multiple clients and enhance the robustness and adaptability of the model; the objective function in the multi-level local training process consists of a loss function term and a regularization term. The loss function term and the regularization term are integrated and optimized, and then: In the formula, Represents the loss function used to measure the personalized local model θ i Performance on local data, optimizing the adaptability of personalized local models to local process conditions and data characteristics; represents the regularization term, which is used to constrain the personalized local model θ according to the regularization coefficient λ i With the general local model ω i The degree of deviation between them is used to optimize the global consistency of the personalized local model; -When the mean deviation and standard deviation changes of data samples and process parameters of some clients are less than the set threshold, and when the mean deviation and standard deviation changes of data samples and process parameters of some clients are greater than the set threshold, a hybrid hierarchical local training mechanism is used to generate personalized local models and general local models; among them, for clients whose mean deviation and standard deviation are less than the set threshold, a single-level local training mechanism is used to directly generate a general local model; for clients whose mean deviation and standard deviation are greater than the set threshold, a multi-level local training mechanism is used to generate a personalized local model through the inner layer training of the multi-level local training mechanism, and a general local model is generated through the outer layer training of the multi-level local training mechanism.

4. The integrated circuit core process modeling method based on personalized federated learning according to claim 3 is characterized in that: Also includes any one or more of the following: - The integrated optimization includes: The loss function term and regularization term are optimized respectively through different optimization methods, so as to take into account the local adaptability and global consistency of the personalized local model; -Optimize the loss function terms, including: Select appropriate optimization algorithms to optimize loss function terms according to task requirements, including: adaptive moment estimation optimization algorithm, stochastic gradient descent algorithm and root mean square propagation optimization algorithm; -Optimize the regularization terms, including: The regularization term is used to control the difference between the personalized local model and the global model to avoid over-personalization or deviation from the stability of the global model, including: adaptive regularization weight adjustment through a gradient descent algorithm, dynamic adjustment of the regularization weight according to the training stage, emphasizing personalized adaptation in the early stage, and balancing personalization and global consistency in the later stage.

5. The integrated circuit core process modeling method based on personalized federated learning according to claim 1 is characterized in that: The model aggregation method includes: weighted average aggregation, dynamic personalized aggregation, hierarchical aggregation, contrastive learning aggregation, gradient fusion aggregation, entropy weight aggregation and / or fuzzy weight aggregation; wherein: The weighted average aggregation dynamically adjusts the weight according to the amount of client data or process complexity, including: Dynamically adjust the aggregation weight according to the data volume, data quality or process complexity of each client, and combine the local model and the global model to generate a new global model: Where: is the global model after the t+1th iteration; is the global model after the tth iteration; β is the personalized aggregation parameter, which is used to adjust the contribution ratio of the global model and the local model; ω i is the local model parameter of the i-th client, and its contribution is based on the amount of data n of N clients. i Dynamic adjustment; The dynamic personalized aggregation dynamically adjusts the aggregation parameters based on the similarity between the client and the global model, including: In view of the large differences in client data distribution, dynamic personalized aggregation parameters are introduced: In the formula, is the global model after the t+1th iteration; is the global model after the tth iteration; α i To personalize the aggregation parameters, they are dynamically adjusted by the similarity between the client and the global model; The aggregation based on contrastive learning uses the similarity between client models for priority aggregation, including: During the aggregation process, the contrastive learning method is used to calculate the similarity between local models of different clients, and the models with higher similarity are aggregated first; where: The formula for calculating the comparative similarity is: In the formula, ω i ,ω j They are represented as the local model parameters of clients i and j respectively; Adjust the aggregation weight w according to similarity i for: In the formula, ω k are the local model parameters of client k, Σ k ∑ j sim(ω k ,ω j ) is the normalized sum of similarities between all clients, which is used to calculate the aggregation weight; The final aggregation formula is: The hierarchical aggregation first performs intra-group aggregation on clients with similar process characteristics, and then performs global aggregation, including: The clients are grouped according to their process categories or data characteristics. The models in each group are first aggregated to generate the group model. Then, the group models are aggregated globally to generate the final global model. Among them: The intra-group aggregation is: In the formula, is the model parameter within the group representing the kth group; group k represents the client set of the kth group; The global aggregation is: In the formula, v k is the weight of each group, calculated according to the amount of data in the group; The gradient-based aggregation updates the global model by fusing the gradients of the clients; The entropy weight aggregation dynamically adjusts the aggregation weight by calculating the entropy value output by the client model to enhance the robustness of the model; The aggregation of fuzzy weights includes: using fuzzy logic reasoning to dynamically generate aggregation weights according to the fuzzy characteristics of client data; wherein, when the client data quality is high and the process complexity is low, a higher weight is assigned; when the client data quality is low and the process complexity is high, a lower weight is assigned.

6. The integrated circuit core process modeling method based on personalized federated learning according to any one of claims 1 to 5, characterized in that: Also includes any one or more of the following: -Standardize the client's process data and build a unified data structure; including: Remove missing, incomplete and / or abnormal data through data cleaning; Through data standardization, the data of different clients are normalized or standardized according to unified standards; Expand the dataset through data augmentation methods; -Through dynamic initialization, online synchronization and extended evaluation, new clients can be seamlessly connected, thereby achieving system adaptability and scalability; among which: The dynamic initialization provides the new client with the initialization weights of the global model to adapt to the local process conditions; The online synchronization supports the dynamic participation of new clients in model training without interrupting the existing federated learning process; The extended evaluation dynamically adjusts the regularization weights by evaluating the data distribution and process conditions of a new client when the new client joins. -Through additional model optimization, the global model of each core process flow of the corresponding integrated circuit is further optimized, including: Through model pruning, redundant parameters are removed to improve model reasoning efficiency; By quantizing the model, the model storage size is reduced and the running performance on resource-constrained devices is optimized; Adopting model distillation, by using the global model as the teacher model, the performance of the client model is further improved; -Protect client data privacy during data and model transmission, including: Differential privacy is used to protect sensitive information of model parameters by adding noise. Through secure multi-party computing, encrypted calculations are performed between the client and the server to avoid data leakage; Through homomorphic encryption, the client's local model parameters are encrypted and uploaded to ensure data security; -Analyze the process data characteristics of the client to provide optimization strategies for the training process; wherein the analysis includes: Data distribution analysis, used to detect data characteristics such as skewness and kurtosis; Process complexity assessment, used to dynamically adjust the model training depth based on the process parameters provided by the client; Data quality assessment, used to provide data enhancement suggestions for clients with low-quality data; - Real-time monitoring of model training status and performance, including: Convergence monitoring, used to analyze the convergence speed and training stability of the model; Performance evaluation, used to evaluate the prediction accuracy of global and local models; Anomaly detection, used to automatically identify anomalies during training.

7. An integrated circuit core process modeling system based on personalized federated learning, characterized in that: include: A model initialization module and a model aggregation module are arranged on the server side, and a cross-level local training module is arranged on the client side; wherein: The model initialization module is used to initialize the global model and send the global model to the local training models of multiple clients; The cross-level local training model is used to perform local training after receiving the global model, generate a local model and send it to the model aggregation module, wherein the local model includes: a personalized local model and a general local model; The model aggregation module is used to generate an updated global model through model aggregation to obtain a global model corresponding to each core process flow of the integrated circuit.

8. The integrated circuit core process modeling system based on personalized federated learning according to claim 7 is characterized in that: It also includes any one or more of the following modules: Data processing module, which is used to standardize the process data of the client and build a unified data structure; Dynamic expansion module, which is used to support dynamic access of new clients to ensure system adaptability and scalability; Model optimization module, which is used to prune, quantize, and distill the global model to improve model efficiency and performance; Privacy protection module, which uses differential privacy and encryption technology to ensure privacy security during data transmission; Process characteristic analysis module, which is used to analyze the characteristics of client process data and provide optimization strategies for model training.

9. A computer terminal comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it can be used to perform the method described in any one of claims 1 to 6, or to run the system described in any one of claims 7 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it can be used to perform the method described in any one of claims 1 to 6, or to run the system described in any one of claims 7 to 8.

Citation Information

Cited By

  • Hierarchical personalized federal learning method for oil field production prediction

    CN120163479A

  • A Hierarchical Personalized Federated Learning Method for Oilfield Production Prediction

    CN120163479B