Wafer level chip system design space construction and rapid parameter search method
By combining Bayesian optimization and intelligent design space exploration methods of graph neural networks, the problems of inefficient computing efficiency and insufficient global optimization capabilities in wafer-level chip system design are solved, and efficient design space exploration and performance optimization are achieved.
Patent Information
- Application Number
- CN202510366146.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-03-26
AI Technical Summary
In wafer-level chip system design, traditional design space exploration methods have problems such as inefficient computing efficiency, insufficient global optimization capabilities, difficulty in task mapping and resource scheduling, and lack of multimodal data processing capabilities.
Using an intelligent design space exploration method combining Bayesian optimization algorithm and graph neural network, iterative optimization is performed through initialization, feature extraction, solution space generation, Bayesian optimization and model update, and the solution space is generated and Bayesian optimization is performed to quickly search for the optimal design scheme.
It significantly improves the performance of wafer-level chip system design, especially suitable for high-complex system design and multi-task scheduling optimization, improves computing efficiency and global optimization capabilities, and is interpretable.
Smart Images

Figure CN120163113A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of wafer-level chip system architecture design, and particularly relates to a method for constructing a design space and rapidly searching for parameters for a wafer-level chip system architecture of a wafer-level chip system. Background Art
[0002] With the rapid development of integrated circuit technology, wafer-scale systems (WSS) have become an important direction in the field of high-performance computing. Compared with traditional chip design, wafer-level chips integrate multiple chip functions on the same wafer, significantly improving computing power and energy efficiency. However, this highly integrated design approach also brings huge design complexity, including aspects such as hardware-software co-design, task mapping and scheduling, performance optimization, and power management.
[0003] In the design of wafer-level chip systems, design space exploration (DSE) is a key task, whose goal is to find the optimal design that meets the constraints of power, performance, and area (PPA) in a wide design parameter space. However, with the expansion of the design scale, the dimension and complexity of the parameter space have also increased significantly. Traditional design space exploration methods mostly use heuristic algorithms or empirical-based manual tuning. Although these methods are effective in small-scale designs, when faced with the ultra-high-dimensional design space of wafer-level systems, they often show the following deficiencies:
[0004] (1) Low computational efficiency: The number of parameter combinations grows exponentially, resulting in the need for a large amount of computing resources for traditional methods to cover the design space; (2) Insufficient global optimization ability: Existing methods are often limited to local optimal solutions and lack the ability to effectively explore the global design space; (3) Difficulty in task mapping and resource scheduling: The mapping relationship between system-level tasks and hardware resources is complex and dynamically changing, and existing methods cannot effectively handle multi-task concurrency and hardware resource contention problems; (4) Lack of multi-modal data processing ability: The design of wafer-level chip systems involves hardware-software interaction, multi-level performance indicators, and multi-modal data (such as structured task graphs and hardware parameters), and traditional methods are difficult to efficiently integrate and utilize this information.
[0005] In recent years, the development of machine learning and optimization algorithms has provided new solutions for design space exploration. For example, the Bayesian optimization algorithm has excellent global search capabilities and can efficiently find the optimal solution in the case of high evaluation costs; the Graph Neural Network (GNN) is good at processing structured data and can be used to represent task graphs or hardware topologies. However, the application of these two technologies in wafer-level chip system design still faces challenges: Bayesian optimization may have a slow convergence speed in high-dimensional parameter spaces, and the design and computational efficiency of the acquisition function need to be improved; the training process of the graph neural network is complex, and how to optimize it in combination with the specific requirements of wafer-level design remains an open question. The integration and fusion of the two require the design of an efficient framework that also supports the interpretation and verification of design results.
[0006] Therefore, there is an urgent need for an intelligent design space exploration method that combines Bayesian optimization and graph neural networks, which can effectively process complex multi-modal input data, provide global optimization capabilities, and have both computational efficiency and interpretability to meet the requirements of wafer-level chip system design. Summary of the Invention
[0007] The purpose of the present invention is to provide a design space construction and fast parameter search method for the architecture of wafer-level chip systems. This method combines the Bayesian optimization algorithm and a joint model, and performs iterative optimization through initialization, feature extraction, solution space generation, Bayesian optimization, and model update. While ensuring efficiency, it can significantly improve the performance of wafer-level chip system design, especially suitable for high-complexity system design and multi-task scheduling optimization.
[0008] The present invention is implemented by adopting the following technical solutions:
[0009] The design space construction and parameter search method of the present invention mainly includes the following five stages:
[0010] 1. Initialization stage
[0011] In this stage, the computing tasks are represented by a task graph. Each node in the task graph represents a computing task, and the edges between the nodes represent the dependencies between the tasks. At the same time, the parameter quantization of the prefabricated parts is carried out, including dynamic parameters and static parameters.
[0012] 2. Feature extraction stage
[0013] Use a graph convolutional network to extract features from the task graph. Each task node contains node features, which mainly reflect the computing resource requirements of the task. The node features include, but are not limited to, information such as the computing latency, memory requirements, data transmission requirements, and computing load of the task. The edges between tasks contain edge features, which mainly describe the data flow dependencies and priorities between tasks. These features play a crucial role in task scheduling, affecting the order and dependencies of task execution.
[0014] The hardware characteristics of the die can be divided into two categories: static parameters and dynamic parameters. Static parameters include manufacturing process, number of cores, cache hierarchy structure, etc. These hardware features are usually determined by the physical layer of the chip and do not change with tasks. Dynamic parameters reflect the impact of task load on hardware performance, including the number of instruction cycles, memory bandwidth occupancy rate, power consumption curve, etc., and are usually obtained through simulation tools. Use a Transformer network to extract features from the prefabricated part parameters.
[0015] 3. Solution space generation stage
[0016] The main task of this stage is to input the features of the task graph and the features of the prefabricated part parameters into the joint model to generate the solution space. The joint model uses a dynamic graph neural network (D-GNN) for feature fusion. The dynamic graph neural network can extract the features of nodes and edges in the graph structure, while capturing the dependencies between tasks. Through graph convolutional operations, D-GNN can effectively extract the topological structure information of the task graph, providing a basis for subsequent task partitioning and hardware selection.
[0017] The joint model also adopts a cross-modal attention mechanism to handle the heterogeneity between task features and hardware features. In this process, the attention mechanism can automatically identify which task features are strongly correlated with which hardware features, thereby aligning the feature spaces of the two and further enhancing the expression ability of the model.
[0018] The generated solution space contains two key elements: the task partitioning scheme and the die selection sequence. The task partitioning scheme determines how each task is allocated on different dies, while the die selection sequence determines which dies are specifically used to execute the tasks.
[0019] 4. Bayesian optimization stage
[0020] The goal of this stage is to quickly search for the optimal design scheme based on the generated solution space through the Bayesian optimization algorithm. Bayesian optimization first needs to construct a Gaussian process (GP) as a surrogate model for predicting the performance of different design schemes. In the surrogate model, the input is the task partitioning scheme and the die selection sequence, and the output is the corresponding performance metrics, such as power consumption, latency, and area.
[0021] The Gaussian process uses a kernel function to model the uncertainty in the design space, and the hyperparameters of the surrogate model are updated by methods such as maximum marginal likelihood estimation.
[0022] Bayesian optimization uses methods such as expected improvement (EI) as the acquisition function. EI can guide the search direction according to the uncertainty of the current model and select the evaluation point that is most likely to bring performance improvement. By calling the simulator multiple times, Bayesian optimization can gradually converge to a design solution with optimal performance.
[0023] 5. Model update phase
[0024] After each optimization, the loss function is calculated based on the simulation results, and the weights of the graph neural network in the joint model are updated by the backpropagation algorithm. The loss function consists of two parts: one is the cross-entropy loss between the task partitioning scheme and the grain selection sequence, which aims to ensure the rationality of task partitioning and hardware selection; the other is the mean square error (MSE) based on the simulation results, which is used to measure the gap between the predicted performance metrics and the actual performance.
[0025] The hyperparameters of the surrogate model are updated by methods such as marginal likelihood gradient descent (MLE) to improve the prediction accuracy of the model. In each iteration, the joint model and the surrogate model are updated according to the new simulation results, gradually optimizing the design solution until the preset convergence conditions are met.
[0026] Compared with the prior art, the present invention is a method for constructing a design space and rapidly searching for parameters for a wafer-level chip system architecture. It has the following effects: revealing the time and space statistical characteristics of the application running on the wafer-level chip system, establishing the theory and method for the co-design of domain-specific software and hardware for the wafer-level chip architecture, forming the basic theory and method for the software development environment of the wafer-level chip, and providing theoretical support and architecture guidance for the design and implementation of domain-specific wafer-level chips. Description of the drawings
[0027] Figure 1 It is a schematic diagram of the initialization, feature extraction, and solution space generation process in the present invention
[0028] Figure 2 It is a schematic diagram of the Bayesian optimization process in the present invention
[0029] Figure 3 It is a flowchart of the implementation of the design space construction and rapid parameter search method of the present invention Detailed implementation manners
[0030] The following further elaborates on the present invention with reference to the accompanying drawings.
[0031] As shown in the append Figure 1As shown in [the figure], first, the input computing task and the prefabricated components are initialized. The computing task is represented as a directed acyclic graph, and the quantization parameters of the prefabricated components are obtained. The graph convolutional network and Transformer are respectively used to extract the features of the task graph and the prefabricated component parameters. The features are input into the encoder network with an attention layer, and finally, the solution space is output through the decoder network.
[0032] As shown in the appendix Figure 2 As shown in [the figure], first, the Gaussian model is pre-trained by combining the emulator and a small number of random solutions. Subsequently, the solutions output by the joint model are provided to the Gaussian model, and the Gaussian model outputs the evaluation values. Then, according to the distribution law of the performance in the solution space, the next evaluation point is selected and the performance is evaluated using the emulator. If the performance meets the requirements, the optimal solution is output; otherwise, the parameters of the joint model and the Gaussian model are updated according to the evaluation results of the emulator.
[0033] As shown in the appendix Figure 3 As shown in [the figure], the implementation process of the design space construction and fast parameter search method of the present invention includes an initialization stage, a feature extraction stage, a solution space generation stage, a Bayesian optimization stage, and a model update stage. The specific steps of each stage are as follows.
[0034] 1. Initialization stage:
[0035] Task graph construction: First, the computing task is represented as a graph, where each node represents a computing task. The dependencies between tasks are represented by the edges between nodes. The structure of the task graph can clearly show the execution order and dependencies of the tasks.
[0036] Quantization of prefabricated component parameters:
[0037] Static parameters: Static parameters are fixed by the hardware design, such as the manufacturing process, the number of cores of the chip, the cache hierarchy, etc. These characteristics do not change during the computing process.
[0038] Dynamic parameters: Dynamic parameters are the relationship between hardware performance and load during task execution, such as the number of instruction cycles, the memory bandwidth occupancy rate, power consumption, etc. These parameters are usually obtained through simulation tools and reflect the impact of the task on the hardware.
[0039] 2. Feature extraction stage:
[0040] Feature extraction of the task graph: The task graph is processed using a graph convolutional network (GCN). Each task node contains features reflecting the computing requirements of the task, such as computing latency, memory requirements, data transmission requirements, etc. These features reflect the computing resource requirements of each task.
[0041] Edge Features: The edges in the task graph connect different tasks, and the features of the edges mainly describe the data flow dependencies and priorities between tasks. Edge features are particularly important in task scheduling because they affect the order and dependencies of task execution.
[0042] Precast Feature Extraction: Use a Transformer network to process static and dynamic precast parameters and extract hardware characteristics, such as the hardware architecture characteristics of the chip and the impact of the load on the hardware performance. The key in this stage is to obtain accurate information on how hardware resources affect task execution.
[0043] 3. Solution Space Generation Stage:
[0044] Joint Model Input: Input the task features in the task graph and the precast parameter features of the hardware into the joint model together. The model fuses the features of the tasks and the hardware to generate a solution space.
[0045] Use D-GNN for feature fusion, extract the topological structure information of the task graph, and capture the dependencies between tasks. The Graph Neural Network (GNN) has excellent graph data processing capabilities and can effectively mine the complex dependency structures between tasks.
[0046] Cross-modal Attention Mechanism: Handle the heterogeneity between task features and hardware features through the cross-modal attention mechanism. The attention mechanism can automatically identify which task features are strongly correlated with which hardware features and align these features, thereby enhancing the expressive power of the model.
[0047] Generate Solution Space: The solution space contains two key elements:
[0048] Task Partitioning Scheme: Determine how to partition tasks and allocate them to different hardware units (dies) for execution.
[0049] Die Selection Sequence: Determine which specific dies to use to execute each task.
[0050] 4. Bayesian Optimization Stage:
[0051] Surrogate Model Construction: The core of Bayesian optimization lies in constructing a surrogate model, usually a Gaussian Process (GP) model, to predict the performance (such as power consumption, latency, etc.) of different design schemes. The surrogate model is obtained through simulation or experimental data.
[0052] Modeling of Gaussian Process: The Gaussian process models the uncertainty of the design space by selecting an appropriate kernel function and can predict performance metrics between the task partitioning scheme and the die selection sequence.
[0053] Acquisition function (EI): Bayesian optimization automatically selects the next evaluation plan that is most likely to improve performance through the Expected Improvement (EI).
[0054] Performance evaluation: Bayesian optimization guides the search direction based on the uncertainty of the current model. By continuously evaluating the results of the emulator, it gradually converges to the optimal design plan.
[0055] 5. Model update stage:
[0056] Simulation results and loss function: After each optimization, the actual performance of the task partitioning and die selection plan is calculated through the emulator. The simulation results are compared with the predicted values of the model, and the loss function is calculated.
[0057] The loss function consists of two parts:
[0058] Cross-entropy loss: It is used to optimize the rationality of the task partitioning plan and die selection sequence to ensure that the task assignment and hardware selection meet the actual requirements.
[0059] Mean Squared Error (MSE): It is used to measure the gap between the performance metrics (such as power consumption, latency, etc.) predicted by the model and the simulation results.
[0060] Backpropagation update: According to the loss function, the weights of the graph neural network in the joint model are updated using the backpropagation algorithm to gradually optimize the model.
[0061] Surrogate model update: The Gaussian process model in Bayesian optimization also needs to be updated. By using the Marginal Likelihood Gradient Descent (MLE) method, the hyperparameters of the surrogate model are adjusted to improve the accuracy of performance prediction.
Claims
1. A method for constructing a wafer-level chip system design space and quickly searching for parameters, wherein the method includes a Bayesian optimization algorithm and a joint model, characterized in that: (1) Initialization phase: After the computing task is input, it needs to be converted into a task graph. The task graph is a directed graph consisting of a set of nodes and edges. Each node represents a computing task, and the edge represents the dependency relationship between tasks. At the same time, the optional prefabricated component parameters are quantified, including static parameters and dynamic parameters. (2) Feature extraction stage: The task graph is analyzed to extract node features and edge features. Node features include the computing resource requirements of the task, and edge features include the data flow dependencies and priorities between tasks. At the same time, the hardware parameters of the prefabricated die are quantitatively modeled. The quantitative indicators include task execution latency, chip area, and power consumption. (3) Solution space generation stage: The task graph features and grain features are input into the joint model, and the features are fused through the dynamic graph neural network and the cross-modal attention mechanism, and the solution space consisting of the task division scheme and the grain selection sequence is output; (4) Bayesian optimization stage: construct a Gaussian process proxy model based on the solution space, use the expected improvement function to select evaluation points, and call the simulator to obtain system performance indicators; (5) Model update phase: The loss function is calculated based on the simulation results, and the graph neural network weights of the joint model are updated through back propagation. At the same time, the hyperparameters of the proxy model are updated, and steps (2)-(4) are iteratively executed until the preset convergence conditions are reached.
2. The method according to claim 1, characterized in that: The quantitative modeling of the prefabricated die in step (2) includes: static parameters including process technology, core number, cache hierarchy structure; dynamic parameters including the number of instruction cycles based on task load simulation, memory bandwidth occupancy and power consumption curve.
3. The method according to claim 1, characterized in that In step (2), the task graph feature extraction can use a deep learning model or a traditional machine learning algorithm to extract node features and edge features.
4. The method according to claim 1, characterized in that: The joint model in step (3) adopts an encoder-decoder architecture.
5. The method according to claim 4, characterized in that Encoder side: Use graph convolutional networks to extract high-dimensional features, and use multi-head attention layers to achieve cross-modal alignment of task features and grain parameters.
6. The method according to claim 4, characterized in that Decoder side: A reinforcement learning strategy is used to generate a task division scheme, and a gated recurrent unit is combined to output a grain selection sequence.
7. The method according to claim 1, characterized in that The implementation of Bayesian optimization in step (4) includes: the proxy model adopts a Gaussian process model, in which the parameters of the hyperparameter covariance function and the variance of the noise are updated by the maximum marginal likelihood estimation method, and the acquisition function adopts the expected improvement function Expected Improvement, EI, to guide the Bayesian optimization.
8. The method according to claim 1, characterized in that The model update rule in step (5) is as follows: the loss function of the joint model includes the cross entropy loss of the task division scheme and the grain selection sequence, as well as the mean square error of the simulation index, and the proxy model hyperparameters are updated by the marginal likelihood gradient descent method.
9. The method according to claim 1, characterized in that: Design space exploration of wafer-level chip system architecture for selecting prefabricated matching for different task mappings in complex application scenarios.
Citation Information
Patent Citations
On-chip and inter-chip interconnected neural network chip hardware architecture design method and system
CN115115043A
Software and hardware cooperation-based wafer-level chip system architecture design method and device
CN117057305A
Heuristic-based field-specific wafer-level chip design optimization method and system, and storage medium
CN118821709A
Co-design of a model and chip for deep learning background
US20240289607A1
Deep neural network hyperparameter optimization method, electronic device and storage medium
WO2021007812A1
Cited By
Method and system for evaluating computer hardware performance based on analogue simulation model
CN120353684A
A method and system for evaluating computer hardware performance based on simulation model
CN120353684B
System and method for checking and optimizing port yard design scheme questions and answers
CN121052241A
A system and method for port yard design scheme question and answer checking and optimization
CN121052241B