Model correction method supporting large-scale parallel computing

Through large-scale parallel computing methods, dynamic adjustment of physical constraints and multi-scale model construction, the problems of insufficient accuracy and low efficiency in model correction are solved, and efficient and scalable model correction is achieved, which is suitable for fields such as semiconductor manufacturing and global climate simulation.

CN120633380AActive Publication Date: 2025-09-12上海芯无双仿真科技有限公司

Patent Information

Application Number
CN202510600510.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-09-12
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

The model correction methods in the existing technology suffer from insufficient accuracy, low efficiency and long calculation time due to fixed physical constraints and a single model. Especially when processing large-scale data sets, they cannot fully utilize hardware resources and it is difficult to take into account both global trends and local details.

Method used

Using large-scale parallel computing methods, through data sharding and dynamic physical constraint initialization, a multi-scale model is constructed to generate preliminary prediction results in parallel. Combined with multi-scale error feedback and cross-node collaborative evaluation, physical constraints are dynamically adjusted to achieve adaptive iterative correction. Finally, a global correction model is generated and the scalability is verified.

Benefits of technology

It significantly improves model accuracy and adaptability, shortens calculation time, improves resource utilization and computing speed, can process data scales from GB to TB, and adapts to industrial environments with different hardware configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633380A_ABST
    Figure CN120633380A_ABST
Patent Text Reader

Abstract

The invention provides a model correction method supporting large-scale parallel computing, which relates to the technical field of semiconductor manufacturing, and can accurately correct a model according to the distribution characteristics of input data and the variation trend of errors in an iteration process by dynamically adjusting physical constraints according to data distribution characteristics and error feedback in a correction process. The constraint condition of a physical model is optimized in real time, the precision and adaptability of the model are remarkably improved, a local optimal solution can be effectively avoided, and it is ensured that an output result better conforms to the physical law, so that higher prediction reliability is achieved in semiconductor manufacturing, meanwhile, the number of iterations caused by improper initial parameters is reduced, the calculation efficiency is optimized, and the reliability is improved. According to the method, collaborative optimization of global and local correction is achieved in parallel computing supporting tens of thousands of CPUs, the computing speed and the resource utilization rate are greatly increased, the robustness and expandability of the method are enhanced, the method can seamlessly process the data scale from the GB level to the TB level, and the method is suitable for industrial environments with different hardware configurations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of semiconductor manufacturing technology, and in particular to a model correction method supporting large-scale parallel computing. Background Art

[0002] The existing technology has the following technical problems when used:

[0003] Problem 1: In existing technologies, model correction generally uses fixed physical constraints, which cannot be dynamically adjusted based on the distribution characteristics of the input data or the error changes during the iteration process. When faced with complex, variable, and large-scale data sets, it is easy to fall into local optimal solutions due to improper initial constraint selection, causing the correction results to deviate from the actual physical laws, lack accuracy, and lead to overall low efficiency.

[0004] Problem two: existing technologies usually rely on a single model for correction, which makes it difficult to take into account both global trends and local details at the same time, resulting in long calculation times and low efficiency, especially when processing large-scale data. The error feedback mechanism of traditional methods lacks cross-node collaborative optimization capabilities, and cannot fully utilize hardware resources in large-scale parallel computing environments, resulting in insufficient scalability. Summary of the Invention

[0005] Technical problems solved

[0006] In response to the shortcomings of the existing technology, the present invention provides a model correction method that supports large-scale parallel computing and solves the following problems:

[0007] 1. Address the problems of insufficient correction accuracy and low efficiency caused by traditional fixed constraints;

[0008] 2. Address the problem of long calibration time and poor adaptability of a single model.

[0009] Technical Solution

[0010] To achieve the above objectives, the present invention is implemented through the following technical solutions: a model correction method supporting large-scale parallel computing, the method comprising the following steps:

[0011] Sp1: Data Sharding and Dynamic Physical Constraint Initialization: Preprocess and dynamically shard large-scale input data, distribute it to multiple parallel computing nodes, and adaptively generate initial constraints for the physical model based on the data distribution characteristics. Each node shares the constraints through a synchronization mechanism.

[0012] Sp2: Multi-scale parallel model construction and preliminary prediction: Based on dynamic physical constraints, a multi-scale model is constructed on each computing node, including a coarse-grained global model and a fine-grained local model. Preliminary prediction results are generated in parallel and integrated through distributed memory sharing.

[0013] Sp3: Multi-scale error feedback and cross-node collaborative evaluation: Compare the preliminary prediction results with the measured values, calculate the multi-scale error, and use a cross-node collaborative mechanism to dynamically adjust the error weight. Each node adaptively adjusts the correction granularity based on the local error, and the master node coordinates the optimization direction based on the global error.

[0014] Sp4: Adaptive Constraint Adjustment and Distributed Iterative Correction: Dynamically adjust physical constraints based on multi-scale error feedback, perform iterative correction in parallel at each node, and update parameters through an asynchronous parallel optimization algorithm until the error converges;

[0015] Sp5: Global model integration and scalability verification: Integrate the correction results of each node to generate a global correction model, and verify its scalability by simulating different data scales and conditions;

[0016] Sp6: Output optimization model and industrial application packaging: Output the final correction model, provide an adaptive parallel interface to support dynamic configuration of computing resources, and generate performance reports.

[0017] Preferably, the data sharding and dynamic physical constraint initialization include performing denoising and normalization processing on the input data to generate structured data, dynamically determining the sharding granularity and quantity according to the statistical characteristics of the data (such as variance and distribution shape), and generating initial constraint conditions (including variable ranges and boundary conditions of physical parameters) based on a physically meaningful adaptive algorithm, and achieving constraint synchronization between nodes through a message passing interface (MPI).

[0018] Preferably, in the multi-scale parallel model construction and preliminary prediction, a coarse-grained global model is constructed based on simplified physical equations to capture global trends, and a fine-grained local model is constructed based on high-resolution physical simulation to optimize local accuracy. Each node calculates the predicted values ​​of the coarse-grained and fine-grained models in parallel according to the allocated data slices, and integrates the multi-scale prediction results among the nodes through distributed memory sharing technology.

[0019] Preferably, in the multi-scale error feedback and cross-node collaborative evaluation, the preliminary prediction results are compared with the measured values ​​to calculate the global error (expressed by the root mean square error) and the local error (expressed by the weighted physical deviation), and the error weight is dynamically adjusted through the cross-node collaborative mechanism of distributed gradient aggregation (the weight changes adaptively according to the noise level of the local data). The master node performs real-time optimization coordination based on the global error and generates a correction direction that is broadcast to each node.

[0020] Preferably, in the adaptive constraint adjustment and distributed iterative correction, the physical constraint range is dynamically adjusted based on multi-scale error feedback and the error change trend is utilized (the adjustment strategy includes relaxing or tightening the parameter boundary), the optimal direction of the constraint adjustment is predicted by support vector regression or Bayesian optimization method, and the adjusted constraints are synchronously updated among the nodes to ensure the consistency of the iterative correction.

[0021] Preferably, the asynchronous parallel optimization algorithm in the adaptive constraint adjustment and distributed iterative correction includes the use of asynchronous stochastic gradient descent (AsynchronousSGD) or alternating direction multiplier method (ADMM) to implement parameter updates, each node independently performs local optimization and the master node periodically aggregates global parameters, and the convergence condition is set to the global error being less than a preset threshold or the number of iterations reaching an upper limit.

[0022] Preferably, the global model integration and scalability verification utilizes a distributed file system (such as HDFS) to store the correction results of each node, and synthesizes the global correction model through the master node. Scalability verification is performed by simulating multiple groups of test cases with data scales ranging from GB to TB. The verification indicators include model accuracy, computing time and parallel efficiency and must meet predefined performance standards.

[0023] Preferably, the output optimization model and the adaptive parallel interface in the industrial application package include providing configurable parallel node number parameters (users can dynamically adjust the CPU usage scale according to hardware resources), supporting mainstream parallel computing frameworks (such as MPI and Apache Spark), and encapsulating the interface into modular components to support seamless integration with industrial production systems.

[0024] Beneficial effects

[0025] The present invention provides a model correction method that supports large-scale parallel computing and has the following beneficial effects:

[0026] 1. The present invention dynamically adjusts physical constraints according to data distribution characteristics and error feedback during the correction process. It can optimize the constraints of the physical model (such as parameter range or boundary conditions) in real time according to the distribution characteristics of the input data (such as variance, noise level) and the changing trend of errors during the iteration process, significantly improving the accuracy and adaptability of the model. Compared with the traditional fixed constraint method, it can effectively avoid local optimal solutions and ensure that the output results are more in line with physical laws, thereby achieving higher prediction reliability in semiconductor manufacturing, while reducing the number of iterations caused by inappropriate initial parameters and optimizing computing efficiency.

[0027] 2. The present invention adopts the parallel construction of a coarse-grained global model and a fine-grained local model. The coarse-grained model quickly captures the overall rules, and the fine-grained model fine-tunes the local accuracy. The two are efficiently integrated through distributed memory sharing, which significantly shortens the time of traditional single model correction. As well as the multi-scale error feedback mechanism of cross-node collaboration, the collaborative optimization of global and local correction is achieved in the parallel computing that supports tens of thousands of CPUs. It not only greatly improves the computing speed and resource utilization, but also enhances the robustness and scalability of the method, enabling it to seamlessly process data scales from GB to TB, and adapt to industrial environments with different hardware configurations. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A diagram showing the steps of the method of the present invention;

[0029] Figure 2 It is the error convergence variation diagram of the dynamic constraint and the fixed constraint of the present invention;

[0030] Figure 3 A trajectory diagram for adjusting dynamic constraint parameters of the present invention;

[0031] Figure 4 This is a comparison chart of the multi-scale model error and calculation time of the present invention;

[0032] Figure 5 This is a diagram showing the multi-scale error distribution in collaborative optimization. DETAILED DESCRIPTION

[0033] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. Specific embodiment one:

[0035] like Figures 1 to 5 As shown, a model correction method supporting large-scale parallel computing includes the following steps:

[0036] Sp1: Data Sharding and Dynamic Physical Constraint Initialization: Preprocess and dynamically shard large-scale input data, distribute it to multiple parallel computing nodes, and adaptively generate initial constraints for the physical model based on the data distribution characteristics. Each node shares the constraints through a synchronization mechanism.

[0037] Sp2: Multi-scale parallel model construction and preliminary prediction: Based on dynamic physical constraints, a multi-scale model is constructed on each computing node, including a coarse-grained global model and a fine-grained local model. Preliminary prediction results are generated in parallel and integrated through distributed memory sharing.

[0038] Sp3: Multi-scale error feedback and cross-node collaborative evaluation: Compare the preliminary prediction results with the measured values, calculate the multi-scale error, and use a cross-node collaborative mechanism to dynamically adjust the error weight. Each node adaptively adjusts the correction granularity based on the local error, and the master node coordinates the optimization direction based on the global error.

[0039] Sp4: Adaptive Constraint Adjustment and Distributed Iterative Correction: Dynamically adjust physical constraints based on multi-scale error feedback, perform iterative correction in parallel at each node, and update parameters through an asynchronous parallel optimization algorithm until the error converges;

[0040] Sp5: Global model integration and scalability verification: Integrate the correction results of each node to generate a global correction model, and verify its scalability by simulating different data scales and conditions;

[0041] Sp6: Output optimization model and industrial application packaging: Output the final correction model, provide an adaptive parallel interface to support dynamic configuration of computing resources, and generate performance reports.

[0042] The above steps further include the following:

[0043] The data sharding and dynamic physical constraint initialization include performing denoising and normalization processing on the input data to generate structured data, dynamically determining the sharding granularity and number according to the statistical characteristics of the data (such as variance and distribution shape), and generating initial constraint conditions (including variable ranges and boundary conditions of physical parameters) based on a physically meaningful adaptive algorithm. Constraint synchronization is achieved between nodes through a message passing interface (MPI). By dynamically adjusting the physical constraints according to data distribution characteristics and error feedback during the correction process, the constraints of the physical model (such as parameter ranges or boundary conditions) can be optimized in real time according to the distribution characteristics of the input data (such as variance, noise level) and the changing trend of errors during the iteration process, significantly improving the accuracy and adaptability of the model. Compared with the traditional fixed constraint method, it can effectively avoid local optimal solutions and ensure that the output results are more in line with physical laws, thereby achieving higher prediction reliability in semiconductor manufacturing, while reducing the number of iterations caused by improper initial parameters and optimizing computing efficiency.

[0044] In the multi-scale parallel model construction and preliminary prediction, a coarse-grained global model is constructed based on simplified physical equations to capture global trends, and a fine-grained local model is constructed based on high-resolution physical simulation to optimize local accuracy. Each node calculates the prediction values ​​of the coarse-grained and fine-grained models in parallel according to the allocated data slices, and integrates the multi-scale prediction results between nodes through distributed memory sharing technology. In the multi-scale error feedback and cross-node collaborative evaluation, the preliminary prediction results are compared with the measured values ​​to calculate the global error (expressed as root mean square error) and local error (expressed as weighted physical deviation). The error weight is dynamically adjusted through the cross-node collaborative mechanism of distributed gradient aggregation (the weight changes adaptively according to the noise level of the local data). The master node performs real-time optimization coordination based on the global error and generates the correction direction and broadcasts it to each node. It adopts the parallel construction of the coarse-grained global model and the fine-grained local model. The coarse-grained model quickly captures the overall rules, and the fine-grained model fine-tunes the local accuracy. The two are efficiently integrated through distributed memory sharing, which significantly shortens the correction time of the traditional single model, and the multi-scale error feedback mechanism of cross-node collaboration. It realizes the collaborative optimization of global and local correction in the parallel computing that supports tens of thousands of CPUs, which not only greatly improves the computing speed and resource utilization, but also enhances the robustness and scalability of the method, enabling it to seamlessly process data scales from GB to TB, and adapt to industrial environments with different hardware configurations.

[0045] In the adaptive constraint adjustment and distributed iterative correction, the physical constraint range is dynamically adjusted based on multi-scale error feedback and the error change trend (the adjustment strategy includes relaxing or tightening the parameter boundary), the optimal direction of the constraint adjustment is predicted through support vector regression or Bayesian optimization method, and the adjusted constraints are synchronously updated between each node to ensure the consistency of the iterative correction. The asynchronous parallel optimization algorithm in the adaptive constraint adjustment and distributed iterative correction includes the use of asynchronous stochastic gradient descent (Asynchronous SGD) or alternating direction multiplier method (ADMM) to implement parameter update. Each node independently performs local optimization and the master node periodically aggregates global parameters. The convergence condition is set to that the global error is less than a preset threshold or the number of iterations reaches an upper limit.

[0046] In the global model integration and scalability verification, a distributed file system (such as HDFS) is used to store the correction results of each node, and the global correction model is synthesized through the master node. Scalability verification is performed by simulating multiple groups of test cases with data scales ranging from GB to TB. The verification indicators include model accuracy, computing time and parallel efficiency and must meet predefined performance standards. The output optimization model and the adaptive parallel interface in the industrial application package include providing a configurable parallel node number parameter (users can dynamically adjust the CPU usage scale according to hardware resources), supporting mainstream parallel computing frameworks (such as MPI and Apache Spark), and encapsulating the interface as modular components to support seamless integration with industrial production systems. Specific embodiment two:

[0048] like Figures 1 to 5 As shown, based on the content in the above specific embodiments, the following contents are further disclosed:

[0049] In actual use of the above method steps, the algorithm content in each step is as follows:

[0050] The application of support vector regression in constraint adjustment is as follows:

[0051] SVR is used to predict the optimal direction of constraint adjustment, and its optimization goal is:

[0052]

[0053] The constraints are:

[0054] y i -(w T x i +b)≤∈+ξ i

[0055]

[0056]

[0057] Where: w: weight vector, representing the regression direction of the model, dimension and input feature x i consistent;

[0058] b: bias term, which adjusts the offset of the regression line;

[0059] ξ i , Slack variables, which represent the error tolerance of the upper and lower bounds;

[0060] C: penalty parameter, the trade-off between control error and model complexity (usually a positive real number, such as 1.0);

[0061] xi :Input feature vector, such as error change trend [e1, e2, ..., e m ];

[0062] y i : target value, such as the desired constraint adjustment (e.g., parameter boundary change ΔC);

[0063] ∈: allowable error range (∈-insensitiveloss), usually a small positive number (such as 0.1);

[0064] ||w|| 2 : The L2 norm square of the weight vector, indicating the smoothness of the model.

[0065] By collecting error feedback data (such as multi-scale error e i ), construct the training set {(x i ,y i )}, initialize SVR parameters (C,∈), select kernel function (such as RBF kernel), solve the optimization problem, obtain w and b, and predict the constraint adjustment direction ΔC=w T x+b, SVR, by analyzing the error change trend (such as the increase or decrease of error with iteration), fits a nonlinear regression model, and predicts the optimal adjustment direction of the physical constraints (such as relaxing or tightening the boundaries). Its tolerance error range ∈ ensures that the prediction is insensitive to noise. SVR introduces data-driven intelligent prediction in constraint adjustment, which is more accurate than traditional manual adjustment (error is reduced by about 10%-20%), reduces the number of iterations (saving about 30% computing time), and improves the adaptability and robustness of the correction. It is especially suitable for complex parameter optimization in semiconductor manufacturing.

[0066] The application of Bayesian optimization in constraint adjustment is as follows:

[0067] The core of Bayesian optimization is the surrogate model (such as Gaussian process) and the acquisition function:

[0068] f(x)=GP(μ(x),k(x,x′))

[0069] x * =argmax x EI(x)

[0070] in:

[0071]

[0072] Where: f(x): objective function, which represents the performance of the model after constraint adjustment (such as error minimization);

[0073] GP: Gaussian process, defines the mean function μ(x) and the covariance kernel k(x, x′);

[0074] μ(x): the predicted mean of the current point;

[0075] k(x, x′): kernel function (such as square exponential kernel), which measures the similarity between points;

[0076] x: input variable, such as constraint parameter adjustment value ΔC;

[0077] x * : The optimal adjustment value of optimization;

[0078] EI(x): Expected improvement of acquisition function, balancing exploration and exploitation;

[0079] f(x + ): Current optimal target value.

[0080] Initialize a small number of sample points (such as constraint adjustment values ​​and corresponding errors), fit the Gaussian process model, predict the distribution of f(x), calculate EI(x), select the next x for evaluation, and iteratively update. Bayesian optimization predicts the effect of constraint adjustment by constructing a proxy model, uses the acquisition function to search for the optimal solution globally, and gradually converges to the best adjustment direction. Its uncertainty estimation ability ensures the exploration of unknown areas. Bayesian optimization reduces the trial and error cost in constraint adjustment (sample efficiency is increased by more than 50%), quickly finds the optimal constraint (convergence speed is increased by about 40%), and significantly improves the correction accuracy and efficiency. It is particularly suitable for large-scale parallel environments with limited resources.

[0081] The application of asynchronous stochastic gradient descent (AsynchronousSGD) in iterative correction is as follows:

[0082] Parameter update formula:

[0083]

[0084] in:

[0085]

[0086] Where: θ t : parameter vector of the tth iteration, such as the model parameters [θ1, θ2, ..., θ m ];

[0087] η: learning rate, controls the update step size (e.g. 0.01);

[0088] The gradient of the loss function, based on local data (x i ,y i )calculate;

[0089] J(θ): global loss function, such as mean square error

[0090] L(θ;x i ,y i ): loss of a single sample, such as (y i -f(x i ;θ)) 2 ;

[0091] x i ,y i : Input and target values ​​for the local data shard.

[0092] Each node independently loads data shards and calculates local gradients The global parameter θ is updated asynchronously without waiting for other nodes. The master node periodically aggregates the parameters and checks the convergence conditions (such as J < ∈). Each node asynchronously calculates the gradient and updates the parameters based on local data. The master node coordinates global consistency. The asynchrony reduces waiting time and is suitable for large-scale parallel computing. AsynchronousSGD accelerates iterations (about 2-3 times faster) when supporting tens of thousands of CPUs. Asynchronous updates reduce communication bottlenecks, maintain high accuracy (error converges to within 1%), and significantly improve computational efficiency and scalability.

[0093] Application of Alternating Direction Method of Multipliers (ADMM) in Iterative Correction

[0094] Objective function:

[0095]

[0096] The constraints are: Ax+Bz=c

[0097] Update formula:

[0098]

[0099]

[0100] u k+1 =u k +Ax k+1 +Bz k+1 -c

[0101] Where: x: local parameter vector, such as model parameters within the node;

[0102] z: global parameter vector, coordinating the results of each node;

[0103] f(x): local loss function, such as ∑L(x; x i ,y i );

[0104] g(z): regularization term, such as λ||z||1;

[0105] A, B: constraint matrices, defining the relationship between local and global parameters;

[0106] c: Constraint constant vector, usually zero vector;

[0107] u: Lagrange multiplier vector, adjusting the degree of constraint violation;

[0108] ρ: penalty parameter, controlling the convergence speed (e.g. 1.0).

[0109] By initializing x, z, u, each node calculates the local f(x) and iteratively updates x k+1 (local optimization), z k+1 (global coordination), u k+1 (Constraint adjustment), check convergence, such as ||Ax+Bz-c||<∈, ADMM decomposes the problem into local and global optimization, each node updates the local parameters in parallel, and the master node coordinates global consistency to ensure convergence to the global optimum. ADMM ensures consistency in a distributed environment (error reduced to within 0.5%), supports large-scale parallel computing (efficiency increased by about 2 times), improves the stability and accuracy of correction, and is particularly suitable for complex industrial scenarios. Specific embodiment three:

[0111] like Figures 1 to 5 As shown, based on the content of the above specific embodiment, the following content is further disclosed: When the above entire method is actually used, it is applied in actual scenarios. The specific application description is as follows:

[0112] Transistor model correction in semiconductor manufacturing:

[0113] Background: In semiconductor manufacturing, transistor models (such as the BSIM model) are used to predict the electrical characteristics of devices (such as leakage current I d , threshold voltage V th ) to optimize process parameters. However, as the process shrinks to 3nm, the fluctuation of process parameters (such as oxide layer thickness and doping concentration) increases significantly, resulting in a surge in the amount of measurement data (up to TB level). Traditional correction methods are difficult to meet the needs due to their low computational efficiency and insufficient accuracy. This case uses a "model correction method that supports large-scale parallel computing" to correct the transistor model on a high-performance computing cluster to ensure prediction accuracy and industrial production efficiency.

[0114] Application environment and machine data:

[0115] Input data: Transistor measurement data set, containing 10 billion records (about 1TB), each record includes process parameters (such as gate length L g , oxide layer thickness Tox ) and electrical measurements (such as leakage current I d , threshold voltage V th ), the data noise level is about 5%;

[0116] Computing cluster: A high-performance computing cluster equipped with 10,000 CPUs, each of which is an Intel Xeon Platinum 8160 (24 cores, 2.1 GHz), with a total computing power of approximately 500 TFLOPS, 512 GB of memory per node, and HDFS for distributed storage;

[0117] Parallel frameworks: MPI (for synchronization), Apache Spark (data partitioning and integration);

[0118] Objective: To calibrate BSIM model parameters (such as flatband voltage V fb , mobility μ0), making the prediction error less than 1%.

[0119] The specific application steps and data processing are as follows:

[0120] Sp1: Data sharding and dynamic physical constraints initialization: 1TB of data is denoised (5% outliers are removed) and normalized (normalized to [0, 1]), according to the variance σ 2 = 0.02 and the distribution shape is dynamically divided into 10,000 copies, each about 100MB, distributed to 10,000 CPU nodes, based on physical meaning (device physical equations, such as Generate initial constraints, such as V fb ∈[-0.8, -0.6]V, synchronization is performed through MPI broadcast, which takes about 0.5 seconds for all nodes, and each node obtains consistent initial constraints with a synchronization error of <0.01%;

[0121] Sp2: Multi-scale parallel model construction and preliminary prediction: The coarse-grained model is based on the simplified BSIM equation (ignoring secondary effects), the fine-grained model includes complete BSIM parameters (such as surface scattering), and each node calculates the predicted value in parallel (such as I d ), the distributed memory shared integration takes 2 minutes, the coarse-grained prediction error is about 10%, the fine-grained error is about 5%, and the total calculation time is 180 seconds (10,000 CPUs in parallel). The preliminary prediction is completed, and the global I d The mean deviation is 8%;

[0122] Sp3: Multi-scale error feedback and cross-node collaborative evaluation: Calculate the global error (RMSE = 0.08) and local error (weighted deviation 0.05), adjust the error weight through distributed gradient aggregation (reducing the weight of shards with high noise by 20%), and the master node coordinates the optimization direction. The broadcast takes 1 second. After the error weight adjustment, the local correction granularity is reduced from 100 parameters to 50 parameters. The calculation takes 90 seconds, the error distribution is optimized, and the average deviation is reduced to 4%;

[0123] Sp4: Adaptive constraint adjustment and distributed iterative correction: Use SVR to predict the constraint adjustment direction (such as V fb Adjust 0.05V), the error trend analysis takes 30 seconds, and Asynchronous SGD is used to update the parameters (learning rate η=0.01η=0.01). After 10 iterations, the error converges to 0.009. The total iteration time is 300 seconds. Each node is updated independently, and the master node aggregation takes 2 seconds per round. The constraint range is optimized to V fb ∈[-0.75, -0.65]V, the prediction error is reduced to 0.9%;

[0124] Sp5: Global model integration and scalability verification: HDFS stores the results of each node (total data volume 1.2TB). The master node takes 60 seconds to synthesize the global model. The verification data scale increases from 100GB to 1TB. The accuracy is stable at 0.9%-1%, the parallel efficiency is 90%, and the calculation time increases linearly with the data scale (about 50 seconds for 100GB and about 300 seconds for 1TB). The model scalability is verified to meet TB-level requirements.

[0125] Sp6: Output optimization model and industrial application packaging: Output the corrected BSIM model (parameter file is about 50MB), provide an adaptive interface (supports adjustment of the number of CPUs to 5000-15000), generate a performance report (accuracy 0.9%, total time 12 minutes), the interface test takes 15 minutes under 5000CPU and 10 minutes under 15000CPU. The model is successfully deployed to the production system and seamlessly integrated with process simulation.

[0126] After the above steps, the following indicators are obtained:

[0127] Data indicators:

[0128] Accuracy: The prediction error was reduced from the initial 8% to 0.9%, which is better than the traditional method (about 3%);

[0129] Efficiency: Total calibration time is 12 minutes, while traditional serial methods take several hours (about 5 hours on a single machine).

[0130] Resource utilization: 10,000 CPU parallel efficiency 90%, memory usage <70%.

[0131] Dynamic physical constraint adaptive adjustment: initial constraint V fb ∈[-0.8, -0.6]V is adjusted to V by SVR fb ∈[-0.75, -0.65]V, the number of iterations is reduced from the traditional 15 to 10, avoiding local optimality (the error fluctuation of the traditional method is 2%), improving the accuracy to 0.9%, and saving about 33% of the computing time. In the 3nm process, the parameters fluctuate violently, and dynamic adjustment ensures that the prediction conforms to the physical laws (such as I d Deviation from the measured value <1nA);

[0132] Multi-scale parallel model construction and cross-node collaborative optimization: coarse-grained models capture global I d Trend (error 10% → 4%), fine-grained model fine-tuning local V th (Error 5% → 0.9%), cross-node collaboration takes less than 2 seconds per round, computing time is shortened from the traditional 5 hours to 12 minutes, an increase of about 25 times, robustness is enhanced (error is still <1% in a noisy environment), TB-level data processing is seamlessly completed, supporting large-scale cluster applications and meeting the efficiency needs of industrial production.

[0133] This case verified the effectiveness of the method through semiconductor transistor model correction. Dynamic physical constraint adaptive adjustment ensured high accuracy (0.9%), multi-scale parallel construction and collaborative optimization achieved efficient computing (12 minutes), and supported clusters of tens of thousands of CPUs, demonstrating excellent scalability. Compared with traditional methods, this method has significantly improved in accuracy, efficiency, and adaptability, fully demonstrating the creativity and practical value of the core technology.

[0134] At the same time, this method is applied to the correction of global climate models, including the following:

[0135] Scenario Background: Global climate models (such as the Community Earth System Model (CESM)) are used to predict climate variables such as temperature and precipitation. However, in high-resolution simulations, initial assumptions about parameters (such as cloud microphysical parameters and convection coefficients) deviate significantly from observed data, necessitating correction to improve prediction accuracy. Modern climate simulations generate enormous amounts of data (petabytes), and traditional correction methods are computationally complex and inefficient, making them inadequate. This case study utilizes a "model correction method supporting large-scale parallel computing" to calibrate CESM model parameters on a supercomputer cluster to optimize global temperature forecasts for the next 10 years.

[0136] Application environment and machine data:

[0137] Input data: Global climate observation dataset, containing 10 years of daily data (about 5PB), including variables such as temperature T, rainfall P, and wind speed V, with a spatial resolution of 0.25°×0.25°, a temporal resolution of 6 hours, and a noise level of about 3%;

[0138] Computing cluster: A supercomputer equipped with 20,000 CPUs (AMD EPYC 7763, 64 cores, 2.45 GHz), with a total computing power of approximately 2 PFLOPS, 1 TB of memory per node, and distributed storage using the Lustre file system;

[0139] Parallel framework: MPI (synchronization and communication), Apache Spark (data sharding and integration);

[0140] Objective: To calibrate CESM model parameters (such as cloud condensation nucleus concentration NcNc and convective relaxation time ττ) to reduce the temperature prediction error to less than 0.5℃.

[0141] The specific application steps and data processing are as follows:

[0142] Sp1: Data sharding and dynamic physical constraint initialization: 5PB data is denoised (removing 3% outliers, such as temperature mutations) and normalized (normalized to [-1, 1]), according to the variance σ 2 =0.15 Dynamic sharding is divided into 20,000 copies, each of which is about 250GB and distributed to 20,000 CPU nodes. Based on the physical meaning (such as the energy conservation equation Q = C p ΔT+L v Δq) generates initial constraints, such as NC∈[50,150]cm -3 ,synchronization through MPI broadcast takes about 1 second, each node obtains consistent constraints, and the ,synchronization error is <0.005%;

[0143] Sp2: Multi-scale parallel model construction and preliminary prediction: The coarse-grained model is based on the simplified radiation transfer equation (ignoring the secondary cloud effect), the fine-grained model includes the complete CESM physical process (such as cloud microphysics), and each node calculates the predicted value (such as the global average temperature T) in parallel. avg ), the distributed memory shared integration took 5 minutes, the coarse-grained prediction error was about 2°C, the fine-grained error was about 1°C, and the total calculation time was 300 seconds (20,000 CPUs in parallel). The preliminary prediction was completed, and the TavgTavg deviation after integration was 1.5°C;

[0144] Sp3: Multi-scale error feedback and cross-node collaborative evaluation: Calculate the global error (RMSE = 1.5°C) and local error (weighted deviation 0.8°C), adjust the error weight through distributed gradient aggregation (weight in high-noise areas is reduced by 15%), and the master node coordinates the optimization direction. The broadcast takes 2 seconds. After the error weight is adjusted, the local correction granularity is reduced from 50 parameters to 30 parameters. The calculation takes 150 seconds, the error distribution is optimized, and the average deviation is reduced to 0.9°C.

[0145] Sp4: Adaptive constraint adjustment and distributed iterative correction: Use Bayesian optimization to predict the constraint adjustment direction (such as N c Adjust 10,\text{cm}^{-3}), the error trend analysis takes 60 seconds, and the ADMM is used to update the parameters (penalty parameter ρ = 1.0). After 15 iterations, the error converges to 0.45℃. The total iteration time is 600 seconds. Each node is updated independently, and the master node aggregation takes 3 seconds per round. The constraint range is optimized to Nc∈[60,130]cm -3 , the prediction error dropped to 0.45℃;

[0146] Sp5: Global model integration and scalability verification: Lustre stores the results of each node (total data volume 5.5PB). The master node takes 120 seconds to synthesize the global model. The verification data scale increases from 1PB to 5PB. The accuracy is stable at 0.45℃-0.5℃, the parallel efficiency is 92%, and the calculation time increases linearly with the data scale (about 200 seconds for 1PB and about 600 seconds for 5PB). The model scalability is verified to meet PB-level requirements.

[0147] Sp6: Output optimization model and industrial application packaging: Output the corrected CESM model (parameter file is about 100MB), provide an adaptive interface (supports adjustment of the number of CPUs to 10,000-30,000), generate a performance report (accuracy 0.45℃, total time 20 minutes), the interface test takes 25 minutes under 10,000 CPUs and 18 minutes under 30,000 CPUs. The model is successfully deployed to the climate prediction system and is highly consistent with the observation data.

[0148] The data indicators are as follows:

[0149] Accuracy: The prediction error was reduced from an initial 1.5°C to 0.45°C, which is better than the traditional method (about 1°C);

[0150] Efficiency: Total calibration time is 20 minutes, while traditional serial methods take several days (about 72 hours, single-machine calculation);

[0151] Resource utilization: 20,000 CPU parallel efficiency 92%, memory usage <80%;

[0152] Dynamic physical constraint adaptive adjustment: initial constraint Nc∈[50,150]cm -3 Adjusted to [60, 130] cm by Bayesian optimization -3 The number of iterations is reduced from the traditional 20 to 15, avoiding local optimality (the error fluctuation of the traditional method is up to 0.8℃), improving the accuracy to 0.45℃, and saving about 25% of computing time. In climate simulation, the parameter uncertainty is high, and dynamic adjustment ensures that the prediction complies with energy conservation (such as temperature deviation <0.5℃), thereby improving the reliability of long-term predictions.

[0153] Multi-scale parallel model construction and cross-node collaborative optimization: The coarse-grained model captures the global TavgTavg trend (error 2℃→0.9℃), and the fine-grained model fine-tunes the local rainfall P (error 1℃→0.45℃). The cross-node collaboration takes less than 3 seconds per round: the calculation time is shortened from the traditional 72 hours to 20 minutes, an increase of about 200 times, and the robustness is enhanced (the error is still <0.5℃ in a noisy environment). PB-level data can be efficiently processed, supporting ultra-large-scale cluster applications to meet the real-time needs of climate research.

[0154] Summary: This case study demonstrated the versatility of our method through climate model calibration. Dynamic physical constraint adaptive adjustment achieved high accuracy (0.45°C). Multi-scale parallel construction and collaborative optimization significantly improved efficiency (20 minutes). A cluster of 20,000 CPUs demonstrated excellent scalability. Compared to traditional methods, this method offers significant advantages in accuracy, speed, and adaptability.

[0155] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further restrictions, an element defined by the statement "comprising a reference structure" does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0156] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A model correction method supporting large-scale parallel computing, characterized by: The method comprises the following steps: Sp1: Data Sharding and Dynamic Physical Constraint Initialization: Preprocess and dynamically shard large-scale input data, distribute it to multiple parallel computing nodes, and adaptively generate initial constraints for the physical model based on the data distribution characteristics. Each node shares the constraints through a synchronization mechanism. Sp2: Multi-scale parallel model construction and preliminary prediction: Based on dynamic physical constraints, a multi-scale model is constructed on each computing node, including a coarse-grained global model and a fine-grained local model. Preliminary prediction results are generated in parallel and integrated through distributed memory sharing. Sp3: Multi-scale error feedback and cross-node collaborative evaluation: Compare the preliminary prediction results with the measured values, calculate the multi-scale error, and use a cross-node collaborative mechanism to dynamically adjust the error weight. Each node adaptively adjusts the correction granularity based on the local error, and the master node coordinates the optimization direction based on the global error. Sp4: Adaptive Constraint Adjustment and Distributed Iterative Correction: Dynamically adjust physical constraints based on multi-scale error feedback, perform iterative correction in parallel at each node, and update parameters through an asynchronous parallel optimization algorithm until the error converges; Sp5: Global model integration and scalability verification: Integrate the correction results of each node to generate a global correction model, and verify its scalability by simulating different data scales and conditions; Sp6: Output optimization model and industrial application packaging: Output the final correction model, provide an adaptive parallel interface to support dynamic configuration of computing resources, and generate performance reports.

2. The model correction method supporting large-scale parallel computing according to claim 1, characterized in that: The data sharding and dynamic physical constraint initialization includes performing denoising and normalization processing on the input data to generate structured data, dynamically determining the sharding granularity and quantity according to the statistical characteristics of the data, generating initial constraint conditions based on a physically meaningful adaptive algorithm, and realizing constraint synchronization between nodes through a message passing interface.

3. The model correction method supporting large-scale parallel computing according to claim 1, characterized in that: In the multi-scale parallel model construction and preliminary prediction, a coarse-grained global model is constructed based on simplified physical equations to capture global trends, and a fine-grained local model is constructed based on high-resolution physical simulation to optimize local accuracy. Each node calculates the prediction values ​​of the coarse-grained and fine-grained models in parallel according to the allocated data slices, and integrates the multi-scale prediction results between nodes through distributed memory sharing technology.

4. The model correction method supporting large-scale parallel computing according to claim 1, characterized in that: In the multi-scale error feedback and cross-node collaborative evaluation, the preliminary prediction results are compared with the measured values ​​to calculate the global error and local error. The error weight is dynamically adjusted through the cross-node collaborative mechanism of distributed gradient aggregation. The master node performs real-time optimization coordination based on the global error and generates a correction direction that is broadcast to each node.

5. The model correction method supporting large-scale parallel computing according to claim 1, characterized in that: In the adaptive constraint adjustment and distributed iterative correction, the physical constraint range is dynamically adjusted based on multi-scale error feedback and the error change trend is utilized. The optimal direction of the constraint adjustment is predicted through support vector regression or Bayesian optimization method, and the adjusted constraints are synchronously updated among the nodes to ensure the consistency of the iterative correction.

6. The model correction method supporting large-scale parallel computing according to claim 1, characterized in that: The asynchronous parallel optimization algorithm in the adaptive constraint adjustment and distributed iterative correction includes using asynchronous stochastic gradient descent or alternating direction multiplier method to implement parameter update. Each node independently performs local optimization and the master node periodically aggregates global parameters. The convergence condition is set as the global error is less than a preset threshold or the number of iterations reaches an upper limit.

7. The model correction method supporting large-scale parallel computing according to claim 1, characterized in that: In the global model integration and scalability verification, a distributed file system is used to store the correction results of each node, and a global correction model is synthesized through the master node. Scalability verification is performed by simulating multiple groups of test cases with data scales ranging from GB to TB. The verification indicators include model accuracy, computing time and parallel efficiency, and must meet predefined performance standards.

8. The model correction method supporting large-scale parallel computing according to claim 1, characterized in that: The output optimization model and the adaptive parallel interface in the industrial application package include providing configurable parallel node number parameters, supporting mainstream parallel computing frameworks, and packaging the interface into modular components to support seamless integration with industrial production systems.

Citation Information

Patent Citations

  • System and method for establishing physics-based model

    CN117546170A

  • Regression model hybrid constraint optimization method and system for semiconductor measurement

    CN118690341A

  • Methods for adaptive optimization of enhanced oil recovery performance under uncertainty

    US20160145977A1

  • Non-linear multitask support vector machines

    US20230252359A1

Cited By

  • Full-process automatic regression test and diagnosis method for Lustre file system

    CN122195859A