Heterogeneous EDA data-based multi-task learning QoR prediction method and device

Through a multi-task learning framework, sharing the underlying encoder and dynamically allocating weights in the expert network, combined with the cross-task feature interaction matrix and adaptive learning rate optimization, the problems of error and insufficient generalization ability of non-independent and identically distributed data in QoR prediction are solved, and the prediction performance and efficiency are improved.

CN120671629APending Publication Date: 2025-09-19GUANGDONG UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510129003.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing QoR prediction methods exhibit high errors and low generalization ability when facing non-independent and identically distributed EDA data, and fail to fully utilize the heterogeneity and correlation of data, resulting in limited prediction quality and efficiency.

Method used

A multi-task learning framework is adopted to extract common features by sharing the underlying encoder, and multiple expert networks and gating networks are combined to dynamically assign weights. A cross-task feature interaction matrix is ​​introduced, and a joint loss function and adaptive learning rate optimization algorithm are used to optimize the congestion prediction and DRC violation prediction tasks.

Benefits of technology

It significantly improves the generalization and robustness of the model, optimizes congestion prediction and DRC violation prediction in the EDA process, reduces computing costs and resource waste, and improves prediction performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671629A_ABST
    Figure CN120671629A_ABST
Patent Text Reader

Abstract

The invention aims to provide a heterogeneous EDA data-based multi-task learning QoR prediction method and device, and the method comprises the steps: carrying out the data preprocessing, and obtaining the feature data of non-IID data in different design stages in an EDA process; constructing a multi-task learning model, training the multi-task learning model to optimize loss functions of congestion prediction and DRC violation prediction tasks, and adjusting model parameters to obtain a trained multi-task learning model; generating a prediction result by using the trained multi-task learning model, and optimizing layout and wiring design in an EDA process; according to the method, information sharing among different tasks is integrated through a multi-task learning framework, the relevance among the tasks is fully utilized, the negative migration phenomenon caused by the non-independent identical distribution (Non-IID) characteristic of data is effectively relieved, and therefore the generalization ability and robustness of the model are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital integrated circuit electronic design automation (EDA), and in particular to a multi-task learning QoR prediction method and device based on heterogeneous EDA data. Background Art

[0002] Electronic design automation (EDA) technology is a key support tool for integrated circuit (IC) design. In recent years, it has made significant progress in optimizing design processes and improving design efficiency. With the rapid development of semiconductor technology, the scale and complexity of integrated circuits have increased rapidly, placing higher demands on EDA technology.

[0003] Quality of Results (QoR) prediction aims to predict key design metrics, such as congestion, power consumption, and timing, by modeling characteristic data from each stage of the design process. This helps designers optimize their designs in a timely manner at an early stage. Current research has led to the emergence of a range of Quality of Results (QoR) prediction methods. These methods, based on statistics, machine learning, and deep learning, analyze and learn from historical data to predict the quality of future circuit designs. However, most existing QoR prediction methods assume that the input data conforms to the independent and identically distributed (IID) property.

[0004] In actual EDA applications, IC design data often exhibits significant non-independent and identically distributed (Non-IID) characteristics. This non-IID characteristic manifests itself in data heterogeneity, temporality, and spatiality. For example, during the layout and routing design phase, feature data extracted at different stages exhibit strong temporal and spatial correlations, and their distribution often shifts due to changes in design parameters. Furthermore, statistical distribution differences in data from different design stages, coupled with high correlations between features, can lead to performance degradation in traditional prediction models. While deep learning methods have alleviated the limitations of traditional models to some extent, their ability to handle non-IID data remains insufficient, often exhibiting high errors and poor generalization capabilities.

[0005] In summary, existing technologies focus on improving prediction quality and efficiency, and considerable work has been done in these areas. However, they overlook the potential impact of non-IID data, which may limit further improvements in prediction quality and efficiency. To overcome this limitation, we must consider how to utilize non-IID data to mitigate the impact of heterogeneous data on prediction results.

[0006] In view of this, a multi-task learning QoR prediction method based on heterogeneous EDA data needs to be developed urgently. Summary of the Invention

[0007] The object of the present invention is to provide a multi-task learning QoR prediction method and device based on heterogeneous EDA data to solve at least one technical problem in the prior art.

[0008] The technical solution of the present invention is:

[0009] A multi-task learning QoR prediction method based on heterogeneous EDA data, including:

[0010] Data preprocessing to obtain characteristic data of non-IID data in the EDA process at different design stages;

[0011] Constructing a multi-task learning model, optimizing the loss functions of congestion prediction and DRC violation prediction tasks by training the multi-task learning model, adjusting model parameters, and obtaining a trained multi-task learning model;

[0012] The trained multi-task learning model is used to generate prediction results to optimize the placement and routing design in the EDA process.

[0013] The multi-task learning model is constructed, including:

[0014] The shared underlying encoder extracts common features for multiple tasks, while each task has its own decoder to generate task-specific prediction results.

[0015] Multiple expert networks are used to process task-specific data distribution, and expert output weights are dynamically assigned through task-specific gating networks, thereby achieving flexible information sharing between tasks;

[0016] By introducing a cross-task feature interaction matrix, the features between tasks are dynamically shared and integrated, achieving a more flexible multi-task learning mechanism.

[0017] The common features of multiple tasks are extracted by sharing the underlying encoder, while each task has its own decoder, which can be expressed as follows:

[0018] z shared =f encoder (x);

[0019]

[0020] Among them, f encoder is a shared encoder, is the decoder for task i; x is the input feature vector of the model, Z shared is the extracted shared feature representation.

[0021] The method uses multiple expert networks to process task-specific data distribution and dynamically allocates expert output weights through task-specific gating networks, thereby achieving flexible information sharing between tasks, which can be expressed as follows:

[0022] g i (x)=softmax(W i x+b i );

[0023]

[0024] Among them, g ij represents the weight distribution of task i to expert j; represents the output of expert j; Wi is the weight matrix of the gate network of task i; bi is the bias vector of the gate network of task i; M represents the number of expert networks in the model; g ij (x) Dynamic weight allocation of task i to expert j, indicating the degree of dependence of task i on the output of expert j.

[0025] By introducing a cross-task feature interaction matrix, the features between tasks are dynamically shared and integrated, achieving a more flexible multi-task learning mechanism, including:

[0026] The feature vector x for each task i i The feature vectors of other tasks are dynamically combined through the interaction matrix α, and the formula is as follows:

[0027] Among them, z i is the fusion feature of task i, α ij is the weight in the interaction matrix, indicating the proportion of features obtained by task i from task j;

[0028] The learning formula of the interaction matrix is:

[0029] α ij =softmax(Wx j +b);

[0030] Among them, W and b are learnable parameters.

[0031] The multi-task learning model is trained to optimize the loss functions of the congestion prediction and DRC violation prediction tasks and adjust the model parameters, including:

[0032] A joint loss function is used to optimize the congestion prediction and DRC violation prediction tasks respectively, combining the mean square error (MSE) and cross entropy loss functions. The joint loss function combines the mean square error and cross entropy losses.

[0033] An adaptive learning rate optimization algorithm is used to dynamically adjust the learning rate during training, combined with an early stopping strategy to avoid overfitting.

[0034] The joint loss function combines the mean squared error and cross entropy loss, including:

[0035] L total =λ1L MSE +λ2L CrossEntropy ;

[0036] Among them, λ1 and λ2 are weight coefficients, L MSE is the mean square error loss function; L Cross-Entropy Loss is the cross entropy loss function.

[0037] The adaptive learning rate optimization algorithm includes:

[0038] The update formula of the adaptive learning rate optimization algorithm is:

[0039] m t =β1m t-1 +(1-β1)g t ;

[0040]

[0041]

[0042] Among them, g t is the gradient, η is the learning rate, β1 and β2 are momentum coefficients; m t It is the first-order moment estimate of the gradient, that is, the moving average of the gradient, which is used to track the direction and magnitude of the gradient and reflects the trend of the gradient. t is the second-order moment estimate of the gradient, that is, the moving average of the square of the gradient; θ t Represents the current model parameters; ∈ is a very small positive number.

[0043] An electronic device, comprising:

[0044] storage media;

[0045] A processing unit is used to execute the computer program by the processing unit when optimizing the layout and routing design in the EDA process to perform the steps of the multi-task learning QoR prediction method based on heterogeneous EDA data as described above.

[0046] A computer-readable storage medium comprising:

[0047] The computer readable storage medium stores a computer program;

[0048] When the computer program is run, it executes the steps of the multi-task learning QoR prediction method based on heterogeneous EDA data as described above.

[0049] The beneficial effects of the present invention include at least:

[0050] The method described in the present invention integrates information sharing between different tasks through a multi-task learning framework, fully utilizes the correlation between tasks, effectively alleviates the negative transfer phenomenon caused by the non-independent and identically distributed (Non-IID) characteristics of data, and thus improves the generalization ability and robustness of the model; through the combination of hard parameter sharing (HPS) and soft parameter sharing methods, it significantly optimizes congestion prediction and design rule checking (DRC) violation prediction in the EDA process. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1(a) is a schematic diagram of the feature generation process of the integrated circuit dataset;

[0052] Figure 1(b) is a schematic diagram of the acquisition time of dataset labels in the integrated circuit dataset;

[0053] Figure 2(a) is a schematic diagram of soft and hard sharing in multi-task learning;

[0054] Figure 2(b) is a schematic diagram of another type of soft and hard sharing in multi-task learning;

[0055] Figure 3 This is the architecture diagram of the hard shared network model;

[0056] Figure 4 This is the architecture diagram of the soft-sharing MMoE network model;

[0057] Figure 5 This is the architecture diagram of the soft-sharing cross-stitch network model. DETAILED DESCRIPTION

[0058] The present application will be further described below with reference to the accompanying drawings.

[0059] Electronic design automation (EDA) technology is an important tool for integrated circuit (IC) design. The literature (Chan, T.-B., Kahng, AB, Woo, "Revisiting inherent noise floors for interconnect prediction." the ACM / IEEE International Workshop on System-Level Interconnect Problems and Pathfinding, pp. 1–7, 2020) explains in the context of macro-layout that design QoR prediction must be useful and actionable. Each major design stage (such as physical implementation) is actually a step-by-step sequential process, in which the results of the current stage largely depend on the results of the previous stage. Taking actions based on the results of previous predictions sometimes changes the content of the new predictions, so requirements must be made with caution.

[0060] A range of QoR (Quality of Results) prediction methods have emerged in current research. These methods, based on techniques such as statistics, machine learning, and deep learning, predict the quality outcomes of future circuit designs by analyzing and learning from historical data. Some employ regression models, leveraging the relationship between input features and target quality metrics for prediction; others utilize deep learning models such as neural networks to learn complex nonlinear mapping relationships from large amounts of data to improve prediction accuracy and robustness. These methods provide circuit designers with important decision support and guidance, helping them better grasp quality during the design process, thereby improving design efficiency and success rates. Reference (Jung, J., Kahng, AB, Kim, S., Varadarajan, R.: METRICS2.1 and flow tuning in the IEEE CEDA robust design flow and OpenROAD. In: Proceedings of the ACM / IEEE International Conference on Computer-Aided Design (ICCAD), pp. 1–9 (2021) [J].) METRICS2.1 is a standard ML platform for EDA and IC design modeling. It supports analyzing how process parameter settings affect QoR results and can build machine learning applications to predict tool and process results and capture and collect operational data. However, as can be seen from the above literature, most existing QoR prediction methods assume that the input data conforms to the independent and identically distributed (IID) property.

[0061] Non-IID (non-IID) data refers to instances in a dataset where the samples do not satisfy the independent and identically distributed (IID) property. In machine learning, data is considered non-IID when the assumption that samples are independent and drawn from the same distribution no longer holds. In other words, data points are correlated, and their distribution may vary across different subsets or categories. Furthermore, potential temporal or spatial correlations in the data, or unstable distributions between training and test datasets, can also cause the IID assumption to break. Currently, research on non-IID data primarily focuses on decentralized machine learning, such as federated learning.

[0062] This invention aims to address the characteristics of non-independent and identically distributed data in the EDA process by proposing a multi-task learning framework that comprehensively utilizes methods such as hard parameter sharing and multi-gated mixture of experts (MMoE) to achieve efficient modeling of congestion prediction and DRC violation prediction. By sharing potential information between tasks, the present invention reduces the negative transfer problem caused by data heterogeneity and significantly improves the model's adaptability to complex design scenarios. The main purpose of this invention is to develop a general, multi-task learning-driven prediction framework that reduces the computational cost and resource waste of EDA design while optimizing prediction performance.

[0063] The present invention provides the following embodiments to illustrate the technical solutions of the present invention in detail. Specific embodiment 1:

[0065] The present invention provides an embodiment:

[0066] As shown in Figure 1, a multi-task learning QoR prediction method based on heterogeneous EDA data focuses on addressing the impact of non-independent and identically distributed (Non-IID) data on quality of results (QoR) prediction in EDA processes. By combining core modules such as data preprocessing and multi-task learning model building and optimization, the performance of congestion prediction and design rule checking (DRC) violation prediction is significantly improved as follows:

[0067] S1: Data Preprocessing: The data preprocessing module is responsible for extracting features of different design stages based on the characteristics of non-IID data in the EDA process. For congestion prediction, features such as Macro Region, cell_density, and RUDY are extracted; Macro Region provides an overall view of the layout level for analysis and optimization; cell_density reflects the distribution density of logic cells, and its calculation formula is:

[0068]

[0069] RUDY evaluates routing resource utilization, which can be calculated by the distribution of signal paths and pins.

[0070] For DRC violation prediction, features such as congestion_GR_horizontal_overflow are extracted, which describe the overflow in the horizontal and vertical directions in the congested area.

[0071] To address non-IID characteristics, data is grouped by design type (e.g., class a and class b) to ensure that the model can learn the differences in data distribution. In addition, all features are normalized to map feature values ​​to the same scale range to avoid the impact of different feature magnitudes on model training. The normalization formula is as follows:

[0072]

[0073] S2: Model architecture design: The core of this embodiment is the multi-task learning model, which mainly includes two methods: hard parameter sharing (HPS) and multi-gated mixture of experts (MMoE), which adapt to complex data distribution and task requirements.

[0074] S201: Hard Parameter Sharing: HPS extracts common features across multiple tasks by sharing the underlying encoder, while each task has its own decoder to generate task-specific predictions. As shown in Figures 2(a) and 2(b), this approach reduces model redundancy, improves training efficiency, and improves model generalization. Its mathematical representation is as follows:

[0075] z shared =f encoder (x)(0.3)

[0076]

[0077] Among them, f encoder is a shared encoder, is the decoder for task i.

[0078] S202: MMoE method: multiple expert networks are used to process task-specific data distribution, and expert output weights are dynamically allocated through task-specific gating networks to achieve flexible information sharing between tasks, such as Figure 3 As shown. It can more adaptably model the correlation between tasks while retaining the specificity of the tasks, thus adapting to more complex data distributions. Its formula is expressed as:

[0079] g i (x)=softmax(W i x+b1) (0.5)

[0080]

[0081] Among them, g ijrepresents the weight distribution of task i to expert j, represents the output of expert j.

[0082] S203: Cross-Stitch Network: By introducing a cross-task feature interaction matrix, the features between tasks are dynamically shared and integrated, achieving a more flexible multi-task learning mechanism, such as Figure 4 In traditional multi-task learning, the degree of feature sharing between different tasks is static. However, the Cross-Stitch Network learns an interaction matrix, which enables each task to dynamically adjust the proportion of features obtained from other tasks according to its own needs, thereby enhancing the synergy between tasks while retaining task specificity.

[0083] In the cross-stitch network, the feature vector x for each task i is i The feature vectors of other tasks are dynamically combined through the interaction matrix α, and the formula is as follows:

[0084]

[0085] Among them, z i is the fusion feature of task i, α ij is the weight in the interaction matrix, which represents the proportion of features that task i obtains from task j. The weight matrix α is automatically optimized through learning, and the degree of sharing can be dynamically adjusted according to task requirements.

[0086] The learning formula of the interaction matrix is:

[0087] α ij =softmax(Wx j +b) (0.8)

[0088] Here, W and b are learnable parameters, and the softmax function ensures that the weights are normalized across different tasks.

[0089] S3: Model training and optimization: The model training process adopts a joint loss function, combining the mean square error (MSE) and cross entropy loss functions to optimize the congestion prediction and DRC violation prediction tasks respectively.

[0090] The joint loss function combines the mean squared error (MSE) and cross-entropy loss:

[0091] L total =λ1L MSE +λ2L CrossEntropy (0.9)

[0092] Among them, λ1 and λ2 are weight coefficients used to balance the influence of the two loss functions.

[0093] In order to improve the training efficiency of the model and accelerate the convergence of the model, the present invention adopts an adaptive learning rate optimization algorithm (Adam). By dynamically adjusting the learning rate during the training process, the model can reach the global optimal solution more quickly, and combined with the early stopping strategy to avoid overfitting.

[0094] The Adam update formula is:

[0095] m t =β1m t-1 +(1-β1)g t (0.10)

[0096]

[0097] Among them, g t is the gradient, η is the learning rate, and β1 and β2 are momentum coefficients.

[0098] The working process of the method described in this embodiment is as follows:

[0099] Step 1: Input data preprocessing

[0100] Users input data from the EDA design process into the system, and the data preprocessing module completes feature extraction, grouping and normalization.

[0101] Step 2: Model initialization and training

[0102] The preprocessed data is input into the multi-task learning model. The loss function of the congestion prediction and DRC violation prediction tasks is optimized through training, and the model parameters are adjusted. The specific architecture is shown in Figures 2 to Figure 4 ;

[0103] Step 3: Output prediction results

[0104] After training, the model generates predictions, including congestion area distribution and DRC violations. These results can be used directly as a basis for design decisions and to optimize placement and routing designs in the EDA process.

[0105] Verification process:

[0106] 1. Experimental Platform and Setup

[0107] The experiment used the CircuitNet-N28 dataset and an RTL design based on the open-source Pulpino platform. Logic synthesis was performed using Synopsys Design Compiler to generate a 54-gate-level netlist. Back-end design was then completed using Cadence Innovus v20.10 and Voltus v20.10, resulting in a dataset of 10,242 layouts. The experiment was run on Ubuntu 20.04 using the Python programming language and implemented with a multi-task learning model based on the TensorFlow deep learning framework.

[0108] 2. Feature selection and experimental design

[0109] In the post-placement stage, the features include Macro Region, cell_density, RUDY_long, RUDY_short, and RUDY_pin_long. Among them, Macro Region provides an overall view of the layout level; cell_density reflects the density of logic cell distribution; RUDY series features focus on the utilization of wiring resources and their impact on signal propagation delay. In the post-GR stage,

[0110] Features such as congestion_eGR_horizontal_overflow and congestion_GR_vertical_overflow focus on routing congestion in the horizontal and vertical directions.

[0111] Experiments were divided into single-task and multi-task categories. In single-task experiments, all data was proportionally divided into training and test sets, without distinguishing between design categories. In multi-task experiments, data was divided into two groups, Class A and Class B, for training and testing, depending on whether the design included additional macro modifications. Furthermore, the model employed a fully convolutional neural network (FCN) encoder-decoder architecture to achieve congestion and DRC prediction.

[0112] 3. Evaluation indicators

[0113] Congestion prediction evaluation indicators include:

[0114] PSNR (Peak Signal-to-Noise Ratio): Evaluates the similarity between the predicted image and the actual congestion distribution, the higher the value, the better.

[0115] NRMSE (Normalized Root Mean Square Error): Measures the error between the predicted value and the true value. The lower the value, the better.

[0116] PeakNRMSE: Analyzes the sensitivity to local congestion. The lower the value, the better.

[0117] SSIM (structural similarity): reflects the visual consistency of the prediction results.

[0118] DRC violation prediction uses ROC and F1 as indicators, where the ROC curve area represents the model's ability to distinguish between positive and negative samples, and F1 is the balance point between the model's recall rate and precision rate.

[0119] 4. Experimental Results and Analysis

[0120] In congestion prediction experiments, grouped experiments generally outperformed ungrouped experiments in evaluation metrics. PSNR and NRMSE improved by 12% and 15%, respectively, demonstrating that grouped data is more conducive to the model capturing data characteristics. Although the SSIM metric showed little difference, this is related to the complexity of congestion prediction itself. In DRC violation prediction, models that differentiated between data groups showed little difference in ROC and F1 metrics, but the MMoE method with soft parameter sharing demonstrated an advantage in capturing complex distributions.

[0121] Multi-task experiments show that the hard parameter sharing method improves the PSNR of congestion prediction and the F1 metric of DRC prediction by over 15%, indicating that information sharing between tasks effectively reduces model redundancy. The MMoE model significantly improves the NRMSE and PeakNRMSE metrics by 20% and 18%, respectively, demonstrating its advantage in adapting to non-IID data.

[0122] The experimental results are as follows:

[0123] Table 1 Results of single task of congestion prediction model

[0124]

[0125] Table 2 Results of the DRC violation prediction model for a single task

[0126]

[0127] Table 3. Multi-task experimental results of congestion prediction model

[0128]

[0129] Table 3.2DRC violation prediction model multi-task experimental results

[0130]

[0131] Overall Advantages: Experimental results demonstrate that the method described in this example demonstrates significant advantages when processing non-IID data in the EDA process. Hard parameter sharing improves model training efficiency by sharing underlying features, while soft parameter sharing improves adaptability to complex distributions through flexible information sharing.

[0132] Specifically, this paper proposes a multi-task learning framework for QoR prediction in the EDA process, targeting non-independent and identically distributed (Non-IID) data. By deeply analyzing the heterogeneity and distribution skew of EDA data across various design stages, this framework effectively addresses the generalization issues of traditional prediction models in non-IID scenarios. Furthermore, this paper combines hard parameter sharing (HPS) and multi-gated mixture of experts (MMoE) approaches to leverage the potential correlations between tasks while preserving task specificity. In the hard parameter sharing architecture, the model extracts common features through a shared encoder, improving training efficiency and reducing redundancy. In the MMoE architecture, task-specific gating networks dynamically assign expert weights, enhancing adaptability to complex data distributions. Furthermore, through joint optimization of congestion prediction and DRC violation prediction tasks, this paper achieves significant improvements in metrics such as PSNR and NRMSE, achieving performance improvements of over 15% compared to traditional single-task models. Experimental results demonstrate that the hard parameter sharing approach excels in capturing common features, while the MMoE model is more suitable for specific scenarios with complex distributions.

[0133] It is foreseeable that the framework designed by the present invention is not only suitable for congestion prediction and DRC violation prediction, but also has good scalability and can be applied to other optimization tasks in the EDA design process (such as power consumption and timing prediction). This versatility enables the present invention to be used as a core module in EDA tools, providing powerful prediction support for the semiconductor design industry. By improving the prediction accuracy and computational efficiency of EDA tools, the present invention significantly reduces design rework and optimization costs. Its efficiency and robustness enable EDA tools to better adapt to the complex design requirements of modern integrated circuits. Specific embodiment 2:

[0135] The present invention also provides an embodiment:

[0136] An electronic device comprising: a storage medium and a processing unit; wherein the storage medium is used to execute the computer program through the processing unit when optimizing the layout and routing design in the EDA process, and perform the steps of the multi-task learning QoR prediction method based on heterogeneous EDA data as described in specific embodiment 1.

[0137] A computer-readable storage medium having a computer program stored therein; when the computer program is run, the computer program executes the steps of the multi-task learning QoR prediction method based on heterogeneous EDA data as described in specific embodiment 1.

[0138] In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. Furthermore, in the present invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination thereof.

[0139] The above disclosures are only a few specific implementation scenarios of the present invention, but the present invention is not limited thereto. Any changes that can be conceived by those skilled in the art should fall within the scope of protection of the present invention. The above invention numbers are for descriptive purposes only and do not represent the advantages or disadvantages of the implementation scenarios.

Claims

1. A multi-task learning QoR prediction method based on heterogeneous EDA data, characterized by: include: Data preprocessing to obtain characteristic data of non-IID data in the EDA process at different design stages; Constructing a multi-task learning model, optimizing the loss functions of congestion prediction and DRC violation prediction tasks by training the multi-task learning model, adjusting model parameters, and obtaining a trained multi-task learning model; The trained multi-task learning model is used to generate prediction results to optimize the placement and routing design in the EDA process.

2. The multi-task learning QoR prediction method based on heterogeneous EDA data according to claim 1 is characterized in that: The multi-task learning model is constructed, including: The shared underlying encoder extracts common features for multiple tasks, while each task has its own decoder to generate task-specific prediction results. Multiple expert networks are used to process task-specific data distribution, and expert output weights are dynamically assigned through task-specific gating networks, thereby achieving flexible information sharing between tasks; By introducing a cross-task feature interaction matrix, the features between tasks are dynamically shared and integrated, achieving a more flexible multi-task learning mechanism.

3. The multi-task learning QoR prediction method based on heterogeneous EDA data according to claim 2 is characterized by: The common features of multiple tasks are extracted by sharing the underlying encoder, while each task has its own decoder, which can be expressed as follows: z shared =f encoder (x); Among them, f encoder is a shared encoder, is the decoder for task i; x is the input feature vector of the model, Z shared is the extracted shared feature representation.

4. The multi-task learning QoR prediction method based on heterogeneous EDA data according to claim 2 is characterized by: The method uses multiple expert networks to process task-specific data distribution and dynamically allocates expert output weights through task-specific gating networks, thereby achieving flexible information sharing between tasks, which can be expressed as follows: g i (x)=softmax(W i x+b i ); Among them, g ij represents the weight distribution of task i to expert j; represents the output of expert j; W i is the weight matrix of task i; b i is the bias vector of the gating network for task i; M represents the number of expert networks in the model; g ij (x) Dynamic weight allocation of task i to expert j, indicating the degree of dependence of task i on the output of expert j.

5. The multi-task learning QoR prediction method based on heterogeneous EDA data according to claim 2 is characterized in that: By introducing a cross-task feature interaction matrix, the features between tasks are dynamically shared and integrated, achieving a more flexible multi-task learning mechanism, including: The feature vector x for each task i i The feature vectors of other tasks are dynamically combined through the interaction matrix α, and the formula is as follows: Among them, z i is the fusion feature of task i, α ij is the weight in the interaction matrix, indicating the proportion of features obtained by task i from task j; The learning formula of the interaction matrix is: α ij =softmax(Wx j +b); Among them, W and b are learnable parameters.

6. The multi-task learning QoR prediction method based on heterogeneous EDA data according to claim 1 is characterized in that: The multi-task learning model is trained to optimize the loss functions of the congestion prediction and DRC violation prediction tasks and adjust the model parameters, including: A joint loss function is used to optimize the congestion prediction and DRC violation prediction tasks respectively, combining the mean square error (MSE) and cross entropy loss functions. The joint loss function combines the mean square error and cross entropy losses. An adaptive learning rate optimization algorithm is used to dynamically adjust the learning rate during training, combined with an early stopping strategy to avoid overfitting.

7. The multi-task learning QoR prediction method based on heterogeneous EDA data according to claim 6 is characterized in that: The joint loss function combines the mean squared error and cross entropy loss, including: L total =λ1L MSE +λ2L CrossEntropy ; Among them, λ1 and λ2 are weight coefficients, L MSE is the mean square error loss function; L Cr oss-Entropy Loss is the cross entropy loss function.

8. The multi-task learning QoR prediction method based on heterogeneous EDA data according to claim 6 is characterized in that: The adaptive learning rate optimization algorithm includes: The update formula of the adaptive learning rate optimization algorithm is: m t =β1m t-1 +(1-β1)g t ; Among them, g t is the gradient, η is the learning rate, β1 and β2 are momentum coefficients; m t is the first moment estimate of the gradient; v t is the second-order moment estimate of the gradient; θ t Represents the current model parameters; ∈ is a positive number.

9. An electronic device comprising: storage media; A processing unit is configured to execute the computer program by the processing unit when optimizing the layout and routing design in the EDA process, and perform the steps of the multi-task learning QoR prediction method based on heterogeneous EDA data according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that include: The computer readable storage medium stores a computer program; When the computer program is run, it executes the steps of the multi-task learning QoR prediction method based on heterogeneous EDA data according to any one of claims 1 to 8.

Citation Information

Cited By

  • Aviation oil consumption self-adaptive prediction method, system and equipment fusing physical model and multi-task learning

    CN121980227A