Wafer-level system application-oriented core particle architecture intelligent design method
Through the intelligent design method of core-particle architecture of graph neural networks and Bayesian optimization algorithms, the problem of high demand for computing resources in deep neural network models is solved, and intelligent automatic search of wafer-level chip design space is realized, which reduces time and cost, and improves design efficiency and performance optimization effects.
Patent Information
- Application Number
- CN202510358564.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively solve the high demand for computing resources by deep neural network models, especially in the balance between energy efficiency and performance, and lacks a space intelligent automatic search method for hardware implementation of deep neural network models.
The intelligent design method of core-grain architecture using graph neural network and Bayesian optimization algorithm is used to realize intelligent automatic search of design space by segmenting task maps, building design spaces, automatic search of Bayesian optimization algorithms for multi-grain targets, rapid performance estimation of graph neural networks and high-fidelity performance evaluation of cycle precision simulators.
It greatly reduces the time and labor cost required for automatic search of the wafer-level chip system architecture design space, and improves the effect of design efficiency and performance optimization.
Smart Images

Figure CN120217989A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of integrated chip design, and particularly to an intelligent design method for a hardware architecture for deep neural network models. Background Art
[0002] With the wide application of deep neural networks in the field of artificial intelligence, the demand for computing resources is increasing continuously. The training and inference processes of deep neural network models involve a large amount of computation and data processing, which usually rely on high-performance hardware facilities such as GPUs, TPUs, etc. However, with the growth of the complexity of deep neural network models and the scale of data sets, a single traditional integrated circuit architecture has gradually become difficult to meet this demand, especially in terms of the balance between energy efficiency and performance.
[0003] As an emerging integrated circuit design method, wafer-level chips are changing this situation. Wafer-level chips decompose complex systems into multiple functional modules, and each module is called a die. These dies can be flexibly combined according to requirements to form a high-performance heterogeneous computing platform. This modular design not only improves the manufacturing efficiency of the system, but also greatly reduces the production cost and complexity.
[0004] In the application of deep neural networks, wafer-level chips can significantly improve the computing efficiency and energy efficiency ratio. By combining dedicated computing dies with functional dies such as storage, deep neural networks can operate efficiently in a highly integrated wafer-level chip. Such a design can optimize the data transmission path, reduce the computing latency, and achieve better task parallel processing among multiple dies, thereby improving the overall performance of deep neural networks.
[0005] In addition, wafer-level chips also support customized design and can be specifically optimized for the requirements of different deep neural network models. For example, in some computationally intensive deep neural network applications, more computing dies can be configured to meet the requirements of efficient training and inference; while in storage-intensive applications, the proportion of storage dies can be increased to optimize the data access speed.
[0006] In order to adapt to the application scenarios of deep neural networks, wafer-level chips integrate heterogeneous resources with different functions and computing powers. When processing various deep neural network tasks, these chips exhibit high dynamic characteristics of services, and the high flexibility and high efficiency of system services require that the architecture must be able to evolve dynamically. Therefore, the automatic search for the design space of wafer-level chip die architectures is the primary scientific problem to be solved, and there is currently no intelligent automatic search method for the architecture design space for the hardware implementation of deep neural network models. Summary of the Invention
[0007] In view of the above-mentioned problems in the prior art, the present invention provides an intelligent design method for die architectures for wafer-level system applications, which can significantly reduce the time and labor costs required for the automatic search of the system architecture design space of wafer-level chips through an intelligent method.
[0008] To achieve the above object, the present invention adopts the following technical solutions:
[0009] An intelligent design method for die architectures for wafer-level system applications, wherein the method includes a graph neural network and a Bayesian optimization algorithm, and the implementation of the method includes the following steps:
[0010] Step 1: Split to generate a task graph and construct a design space in combination with a die prefabrication library;
[0011] Step 2: Automatically search the design space by the Bayesian optimization algorithm for multi-die objectives;
[0012] Step 3: Perform rapid performance estimation by a graph neural network;
[0013] Step 4: Perform high-fidelity performance evaluation by a cycle accurate simulator.
[0014] In step 1, the task graph is split by calculating the structure and computational complexity of the tasks to generate the task graph. The task splitter obtains the topological space of the hardware by calculating the structure and attributes of the tasks. Subsequently, the mapper generates accurate performance and efficiency predictions by simulating these topologies and searches for the best scheme for scheduling operations and data on the specified architecture. The task graph is split by hierarchy and computational complexity, and the relationship between the task graph and the die prefabrication library is used to construct the design space.
[0015] The task graph is generated by splitting the computational tasks, and the node information of the graph represents the computational and storage loads, and the edge information of the graph represents the communication load.
[0016] The die prefabrication library contains a variety of heterogeneous computational and storage dies, and the dies that meet the constraints in the die prefabrication library are mapped to the nodes of the graph by using the load information of the task graph as a constraint condition to construct a wide design space.
[0017] In step 2, the Bayesian optimization algorithm for multi-die objectives automatically searches the design space. Initial sampling is performed on the design space in the multi-die objectives by a Gaussian process surrogate model, and a prior data set is constructed according to the inherent physical information of the dies to train the model. Then, the design space is automatically searched iteratively. In each iteration, first, the surrogate model is updated to fit the data set, and then the posterior distribution of the entire space is calculated to find the Pareto optimal design.
[0018] The Gaussian process surrogate model mentioned above is a non-parametric probability model. By using known training data to model the objective function of the constraint conditions, it can provide predicted values for unknown inputs and estimates of their uncertainties. The Gaussian process assumes that the function values of any finite number of points follow a multivariate normal distribution, thus generating a prediction distribution for each point in the design space, with the rule f(x) ∼ GP(m(x), k(x, x′)).
[0019] In Step 3, the graph neural network interacts and transfers the grain information according to the fan-in and fan-out situations of each layer of the grains in the task graph, quickly evaluates the performance of the wafer-level chip mapped on the task graph load, and compares whether it meets the constraint conditions. If it meets, it is stored as the current optimal design space; if it does not meet, the Bayesian optimization algorithm is repeated.
[0020] The above-mentioned rapid evaluation is achieved by a graph neural network trained with a prior data set. The graph neural network inference can quickly obtain a low-fidelity performance evaluation result and efficiently evaluate the current Pareto optimal design.
[0021] In Step 4, the current optimal design space is input into the cyclic exact emulator for performance evaluation, and it is compared whether it meets the constraint conditions. If it meets, it is stored as the optimal design space; if it does not meet, the Bayesian optimization algorithm and the graph neural network are repeated.
[0022] The above-mentioned cyclic exact emulator is the main tool for obtaining high-fidelity performance evaluation. In traditional automatic design space search methods, all possible system architecture performances are iteratively evaluated within the design space, and the final result is selected, which is very time-consuming for wafer-level chip simulation.
[0023] During the mapping process of different computing tasks in the chiplet prefabrication library, complex design spaces of different scales will be generated. The method in the present invention can efficiently complete the intelligent automatic search of complex design spaces of different scales.
[0024] Compared with the prior art, the present invention has the following beneficial effects: Compared with the traditional iterative automatic design space search method, an intelligent method process based on Bayesian optimization and graph neural network is introduced, reducing the number of times of iteratively calling the cyclic exact emulator within the design space, and greatly reducing the time and labor costs required for automatic design space search. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 It is a schematic diagram of the intelligent automatic search method for the system architecture design space of the wafer-level chip in the present invention;
[0026] Figure 2 It is a schematic diagram of the automatic searcher for the multi-grain target design space based on the Bayesian optimization algorithm in the present invention;
[0027] Figure 3 It is a schematic diagram of an evaluation architecture for a round of grain information interaction based on a graph neural network algorithm in the present invention.
[0028] Figure 4 It is a flowchart for implementing the intelligent design method of the chiplet architecture of the present invention Specific implementation manners
[0029] The present invention will be further described in detail below with reference to the accompanying drawings.
[0030] See Figure 1 , an intelligent design method for a chiplet architecture for wafer-level system applications. Step 1, generate a task graph and construct an architecture design space through mapping with a chiplet prefabrication library; Step 2, use the Bayesian optimization algorithm as a design space automatic searcher to quickly find the architecture design space that meets the Pareto optimality during the iterative process; Step 3, use the graph neural network as an efficient performance evaluator for the architecture design space, quickly conduct a preliminary evaluation and compare it with the constraint conditions to obtain the current optimal design space; Step 4, input the current optimal design space that meets the constraints into a loop accurate simulator for high-fidelity performance evaluation and compare it with the constraint conditions to finally obtain the optimal design space.
[0031] See Figure 2 , the design space automatic searcher based on the Bayesian optimization algorithm in the present invention. The prior data set, that is, the initial sampling data, needs to randomly sample several architecture feature data from the design space. After obtaining the corresponding results through the black-box function calculation, a Gaussian process surrogate model is trained. The sample format is matrix data. As the dimension of the search space increases, the exponential explosion phenomenon will occur. Therefore, when considering the graph node attributes, factors with less influence on the results need to be pruned. After constructing the prior knowledge of the Bayesian optimization, the mean and variance of the surrogate model in the search space can be obtained. The mean can be regarded as the predicted mean of the surrogate model for the input, and the variance can be regarded as the posterior probability of the surrogate model. The acquisition function is based on these two to calculate the probability distribution of the next sampling point in the search space, and the point with the highest probability is selected according to the acquisition strategy and the target result value is obtained through calculation in the black-box function. Add this point to the initial sampling set, update the surrogate model, and make the surrogate model in the optimal domain range of the Gaussian process fit the original function. Repeat the above operations until the iteration times are exhausted or the optimization target reaches the ideal situation, terminate the operation of the algorithm program, output the optimized optimal design space, and record the optimization time.
[0032] See Figure 3, in the architecture evaluator based on the graph neural network algorithm in the present invention, the graph neural network is iteratively trained according to the prior data set in the design space automatic searcher. The design space first completes graph embedding, aligns with the input interface of the graph neural network, and inputs it into the graph convolutional layer to complete the graph information clustering process. Then, the graph structure information is reconstructed into a vector, and performance prediction is completed in the fully connected layer. The prediction result and the data set label are used to calculate the loss function result to complete the training process of backpropagation and gradient optimization. After the specified number of iterative training times ends, the graph neural network outputs the prediction result.
[0033] Specifically in implementation, it can be divided into the following steps.
[0034] Step 1: Task graph generation and design space construction
[0035] Introduce a task splitter, input the computing tasks, including the description of the computing tasks (such as deep learning inference tasks, encryption algorithms, etc.) and the data flow relationship of the tasks. According to the dependency relationship of the computing tasks, use a hierarchical strategy to divide the tasks into multiple subtasks. The task splitter analyzes the computing tasks and extracts the computing complexity of each subtask. Finally, the task graph is represented as a directed acyclic graph (DAG), where the nodes record the computing and storage loads, and the edges record the communication loads.
[0036] Construct the design space, initialize the die prefabrication library, which contains a variety of heterogeneous die modules with different computing, storage, and interconnection characteristics. Filter out suitable die modules according to the node attributes (computing, storage, bandwidth requirements) of the task graph and map them into the design space. Optimize the mapping strategy through a heuristic algorithm to improve performance and reduce power consumption.
[0037] By combining different die mapping strategies, construct a wide range of die architecture design spaces to cover a variety of possible architecture combinations.
[0038] Step 2: Automatically search the design space using the Bayesian optimization algorithm for multi-die targets
[0039] The input of the surrogate model is the solution in the design space, and the output is the performance metrics (such as computing throughput, power consumption, communication delay, etc.). Construct a prior data set according to the inherent physical information of the die (such as frequency, cache size, power consumption) to train the model.
[0040] Select the initial sampling points to form a sample set, and use the sample data to train the Gaussian model. The surrogate model uses the sample data to predict the performance distribution of the entire design space, forming the posterior distribution of performance estimation.
[0041] In each iteration, the surrogate model is continuously updated based on the data of the newly sampled points to improve the prediction accuracy. The surrogate model recalculates the posterior distribution of the design space based on the updated data. The optimal design point is selected by maximizing the Expected Improvement (EI) strategy.
[0042] A Pareto optimal solution set is constructed among metrics such as performance, power consumption, and latency to ensure that the optimal design that takes into account multi-objective optimization is searched.
[0043] Step 3: The graph neural network (GNN) performs fast performance estimation
[0044] The structural information of the task graph is converted into node features (computing load, storage requirements) and edge features (communication bandwidth, data transmission path) of the graph. The graph neural network (GNN) is adopted, and the graph convolution mechanism is introduced to aggregate the information of neighboring nodes and extract the feature representation of each node.
[0045] The GNN outputs the predicted values of the performance metrics of each node and completes the overall performance evaluation in the fully connected layer. Continuous training is carried out until the model converges to ensure the accuracy of performance prediction.
[0046] Step 4: The loop exact emulator performs high-fidelity performance evaluation
[0047] The optimal solution selected through Bayesian optimization and the graph neural network is used as the input and fed into the loop exact emulator.
[0048] In the exact simulation process, the emulator performs high-precision performance simulation according to the physical parameters, data transmission characteristics, and task load of the chiplet. During the simulation process, the key metrics (such as performance, energy consumption, area, etc.) of each design scheme are recorded.
[0049] If the simulation results meet the target performance and constraint conditions, the design scheme is stored.
[0050] If the conditions are not met, return to the Bayesian optimization and GNN evaluation stages to continue optimizing the design.
Claims
1. A chiplet architecture intelligent design method for wafer-level system applications, the hardware architecture intelligent design method includes a Bayesian optimization algorithm and a graph neural network, and is characterized by: Step 1: Segment and generate the task graph and build the design space in combination with the core particle prefab library; Step 2: The Bayesian optimization algorithm of multi-grain targets automatically searches the design space; Step 3: Use graph neural network for fast performance estimation; Step 4: Cycle-accurate simulator for high-fidelity performance evaluation.
2. The method for intelligent design of chiplet architecture for wafer-level system application according to claim 1, characterized in that: In the step 1, the computing task is divided. The task divider divides the task into multiple subtasks using a hierarchical strategy based on the dependencies of the computing tasks, analyzes the computing task, and extracts the computational complexity of each subtask.
3. The method for intelligent design of chiplet architecture for wafer-level system application according to claim 2, characterized in that: The task graph is generated by dividing the computing tasks, the node information of the graph represents the computing and storage loads, and the edge information of the graph represents the communication load.
4. The method for intelligent design of chiplet architecture for wafer-level system application according to claim 1, characterized in that: The core particle prefab library contains a variety of heterogeneous computing and storage core particles. The core particles that meet the constraints in the core particle prefab library are mapped to the nodes of the task graph by using the load information of the task graph as constraints, thereby constructing a wide design space.
5. The method for intelligent design of chiplet architecture for wafer-level system application according to claim 1, characterized in that: The Bayesian optimization algorithm of the multi-grain target in step 2 automatically searches the design space, performs initial sampling of the design space in the multi-grain target through a Gaussian process proxy model, constructs a priori data set training model based on the inherent physical information of the grains, and then iteratively and automatically searches the design space. In each iteration, the proxy model is first updated to fit the data set, and the posterior distribution of the entire space is calculated to find the Pareto optimal design.
6. The method for intelligent design of chiplet architecture for wafer-level system application according to claim 5, characterized in that: The Gaussian process surrogate model is a non-parametric probability model that models the objective function of the constraints using known training data, and can provide predicted values for unknown inputs and their uncertainty estimates. The Gaussian process assumes that the function values of any finite number of points obey a multivariate normal distribution, thereby generating a predictive distribution for each point in the design space, with the rule being f(x)~GP(m(x),k(x,x′)).
7. The method for intelligent design of chiplet architecture for wafer-level system application according to claim 1, characterized in that: In the step three, the graph neural network interactively transmits grain information according to the fan-in and fan-out of each level of the grain in the task graph, quickly evaluates the performance of the wafer-level chip mapped on the task graph load, and compares whether the constraints are met. If so, it is stored as the current optimal design space; if not, the Bayesian optimization algorithm is repeated.
8. The method for intelligent design of chiplet architecture for wafer-level system application according to claim 7, characterized in that: The rapid evaluation is achieved through a graph neural network trained with a priori data sets. Graph neural network reasoning can quickly obtain low-fidelity performance evaluation results and make an efficient evaluation of the current Pareto optimal design.
9. The method for intelligent design of chiplet architecture for wafer-level system application according to claim 1, characterized in that: In the step 4, the current optimal design space is input into the cycle precise simulator for performance evaluation to compare whether the constraints are met. If so, it is stored as the optimal design space; if not, the Bayesian optimization algorithm and the graph neural network are repeated.