Wafer-level chip load balancing optimization method and system based on K-Means clustering and simulated annealing
By using the combination method of K-Means clustering and simulated annealing algorithm in wafer-level chip systems, the network topology map mapping of AI large models is optimized, which solves the problem of load imbalance and improves the parallel operation efficiency of the system.
Patent Information
- Application Number
- CN202510510079.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-04-23
AI Technical Summary
When deploying AI models, existing wafer-level chip systems are difficult to achieve load balancing, resulting in a long overall parallel running time.
Using the load balancing optimization method based on K-Means clustering and simulated annealing, the edge weight matrix of node run time and edge communication delay is constructed, combined with the critical path algorithm and the K-Means++ strategy, the cluster mapping scheme is optimized to minimize the overall parallel running time of the system.
The load balancing of AI models on wafer-level chip systems is realized, which significantly shortens the overall parallel running time of the system and improves the utilization efficiency of hardware resources.
Smart Images

Figure CN120030977A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of electronic design automation technology for chips, and in particular to a wafer-level chip load balancing optimization method and system based on K-Means clustering and simulated annealing. In a wafer-level chip multi-core system, a large AI model network topology graph is evenly mapped to a chip core group architecture and the overall parallel running time is optimized. Background Art
[0002] In recent years, with the rapid development of artificial intelligence, the scale of AI large models has continued to expand, and the computing nodes and data transmission relationships involved have become increasingly complex. Therefore, effectively deploying AI large models into wafer-level chip systems has become an urgent problem to be solved. In order to improve computing efficiency without increasing hardware costs, a solution is needed that can evenly map the network topology of the AI large model to each core group of the wafer-level chip.
[0003] Existing traditional algorithms usually only focus on local optimization or perform scheduling based on a single strategy. It is difficult to take into account the operator node running time, data transmission delay, node dependencies, and wafer-level chip system architecture at the same time. As a result, the overall parallel running time of the on-wafer system is relatively long, making it difficult to achieve accurate load balancing. Summary of the invention
[0004] The present invention aims to solve the problem of unbalanced load of large AI models in wafer-level chip systems, and proposes a wafer-level chip load balancing optimization method and system based on K-Means clustering and simulated annealing. By comprehensively considering the running time of operator nodes, data transmission delay, node dependencies and wafer-level chip system architecture, the network topology of the large AI model is evenly mapped to each core group to achieve more efficient load balancing.
[0005] To achieve the above purpose, the technical solution adopted is: The present invention provides a wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing, which evenly maps the computing tasks of the AI large model to the core group of the wafer-level chip, including the following steps: Step 1: Calculate the node running time and edge communication delay in the AI big model, and build the edge weight matrix of the entire topology; Step 2: Calculate the critical path weight between any two nodes using the critical path algorithm; Step 3: Generate a corresponding two-dimensional coordinate grid according to the physical layout of the wafer-level chip core group; Step 4: Use the K-Means++ strategy to select the initial clustering center of the nodes and divide the nodes into corresponding clusters; Step 5: Map the clusters to the core groups of the wafer-level chip, and preferentially map clusters with high communication costs to adjacent physical locations; Step 6: Optimize the cluster mapping scheme by using the simulated annealing algorithm. The optimization goal is to minimize the overall parallel running time of the system. Step 7: Update the cluster center and repeat the iteration until the cluster center is stable, and obtain the optimal solution for mapping the AI large model network topology to each core group of the wafer-level chip.
[0006] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, step 1 specifically includes: Calculate each node Running time on a single core group ; Calculate Edges Communication delay ,in, is the amount of data transferred, is the system bandwidth; Constructing edge weights .
[0007] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, step 4 specifically includes: Randomly select a node as the first cluster center; For any remaining node, calculate its minimum weighted distance to all selected cluster centers ; by Randomly select the next cluster center for the probability distribution; Repeat the above steps until all cluster centers are selected; Based on the two-way minimum weight strategy, all nodes are divided into the cluster corresponding to the cluster center with the closest weight distance.
[0008] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, step 5 specifically includes: Calculate each cluster The cross-cluster communication cost , Representation Cluster The cluster center of Sort the clusters from high to low according to the communication cost; Clusters with high communication costs are preferentially mapped to adjacent physical locations in a two-dimensional coordinate grid.
[0009] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, each cluster is calculated The cross-cluster communication cost The expression is:
[0010] in, Represents edge The first summation counts the amount of data transferred from cluster The total amount of data flowing out to other clusters; the second summation counts the data flowing into the cluster from other clusters The total amount of data.
[0011] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, clusters with high communication costs are preferentially mapped to adjacent physical locations in a two-dimensional coordinate grid, including: For the first ranked cluster , select any initial position in the grid; For each subsequent cluster, the location with the minimum Manhattan distance to the last assigned location is selected among the remaining available locations in the grid.
[0012] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, the optimization process of the simulated annealing algorithm in step 6 includes: In each iteration of simulated annealing, the positions of the two clusters are randomly swapped. After each swap, the parallel running time of the system is recalculated. ; If the new solution is better than the current solution, then accept the new solution. Otherwise, accept new solutions; If the current temperature parameter reaches the set threshold, the iteration is stopped and the optimal solution is obtained after the iteration is completed.
[0013] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, the system parallel operation time The calculation includes: Using Clusters The intra-cluster critical path weight is used as the intra-cluster running time ; cluster and The Manhattan distance and cluster of the mapping kernel group on a two-dimensional coordinate grid and The product of the communication delays between ; Calculate the start and end time of each cluster based on cluster dependencies; The maximum value of all cluster end times is taken as the system parallel running time.
[0014] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, updating the cluster center in step 7 specifically includes: Calculate the number of nodes in the cluster The overall weight of , including the sum of the minimum bidirectional weights with other nodes and the weight from the root node to the node; Select the node with the smallest overall weight as the new cluster center.
[0015] Furthermore, the present invention also provides a wafer-level chip load balancing optimization system based on K-Means clustering and simulated annealing, which is used to implement the above-mentioned wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing, comprising: The edge weight calculation module is used to calculate the node running time and edge communication delay in the AI large model and build the edge weight matrix of the entire topology; A critical path weight calculation module is used to calculate the critical path weight between any two nodes using a critical path algorithm; A grid generation module, used to generate a corresponding two-dimensional coordinate grid according to the physical layout of the wafer-level chip core group; The initial clustering module is used to select the initial clustering center of the nodes using the K-Means++ strategy and divide the nodes into corresponding clusters; An initial mapping module is used to map clusters to core groups of wafer-level chips, and preferentially map clusters with high communication costs to adjacent physical locations; The simulated annealing module is used to optimize the cluster mapping scheme through the simulated annealing algorithm. The optimization goal is to minimize the overall parallel running time of the system; The iterative output module is used to update the cluster center and repeat the iteration until the cluster center is stable, thereby obtaining the optimal solution for mapping the network topology of the large AI model to each core group of the wafer-level chip.
[0016] The beneficial effects achieved by adopting the above technical solution are: The present invention comprehensively considers multiple constraints such as operator node running time, data transmission delay, node dependency, and wafer-level system architecture. By combining K-Means clustering and simulated annealing algorithms, it realizes load balancing of AI large model computing tasks on wafer-level chip systems, minimizes the overall parallel running time of the system, optimizes the utilization efficiency of hardware resources, and provides reliable technical support for AI computing in wafer-level chip systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings of the embodiments of the present invention, wherein the drawings are only used to illustrate some embodiments of the present invention, but not to limit all embodiments of the present invention thereto.
[0018] Figure 1 It is a flow chart of a wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The following will be combined with the drawings of specific embodiments of the present invention to clearly and completely describe the exemplary scheme of the embodiment of the present invention. Unless otherwise defined, the technical terms or scientific terms used in the present invention should be the common meanings understood by people with ordinary skills in the field.
[0020] This embodiment discloses a wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing, which evenly maps the computing tasks of the AI large model to the core groups of the wafer-level chip. Figure 1 As shown, specifically including: Input data: (1) AI big model topology diagram: It is represented by a directed acyclic graph, which records the relationship between the various computing tasks of the AI big model.
[0021] (2) Node running time: the running time of each task node on a single core group of the on-chip system.
[0022] (3) Edge data transmission volume: The amount of data transmitted between computing nodes.
[0023] (4) Wafer-level system architecture: used to construct a two-dimensional core group coordinate grid.
[0024] Step S1: Calculate the node running time and edge communication delay in the AI big model, and build the edge weight matrix of the entire topology.
[0025] First, estimate the running time of each computing node in the AI large model on a single core group:
[0026] At the same time, for each edge connecting two nodes , according to the amount of data transmitted by the edge and the system bandwidth, the communication delay between nodes is calculated:
[0027] in, For any edge The amount of data transfer, is the system bandwidth.
[0028] Taking into account the node running time and edge communication delay, the edge weight information between all connected nodes in the topology is calculated. The edge weight is calculated as follows:
[0029] Step S2: According to the edge weight information of step S1, the critical path weight between any two nodes is calculated using the critical path algorithm. The sum of the edge weights on the critical path is recorded as ,Note to subtract the running time of the intermediate nodes that are repeatedly counted. If the critical path is misjudged due to repeated counting, the subsequent clustering and mapping optimization may target the wrong path, resulting in that the parallel running time of the system is not truly minimized.
[0030] Step S3: Generate a corresponding two-dimensional coordinate grid based on the actual physical layout of the wafer-level chip core group. Core groups can be constructed according to the core group layout Coordinates ,Each coordinate point corresponds to a core group mapping position, preparing for subsequent task mapping.
[0031] Step S4: According to the edge weight matrix, use K-Means++ strategy to select Initial cluster centers, where The size of is equal to the number of core groups of the wafer-level chip system, and the nodes are divided into corresponding clusters. This step includes sub-steps S401-S405.
[0032] Step S401: Randomly select a node as the first cluster center .
[0033] Step S402: For any remaining node, calculate its minimum weight distance to all selected cluster centers. :
[0034] in, and Is a node With cluster center The bidirectional weight between .
[0035] Step S403: Randomly select the next cluster center for the probability distribution.
[0036] Step S404: Repeat steps S402-S403 until the selected Cluster centers.
[0037] Step S405: After the selection of cluster centers is completed, all nodes are divided into clusters corresponding to the cluster centers with the closest weight distances based on the two-way minimum weight strategy. Each node is assigned to the cluster center with the closest weight distance:
[0038] Step S5: Initially map the initially formed clusters to the wafer-level chip core group, and preferentially map clusters with high communication costs to adjacent physical locations.
[0039] Generated based on wafer-level chip system core group architecture Two-dimensional grid positions are used as candidate positions of the physical core group. At the same time, in order to reduce the communication delay between clusters, this embodiment designs an improved initial mapping strategy, the core idea of which is: ① Calculate the cross-cluster communication cost for each cluster to reflect the total amount of data interaction between the cluster and other clusters.
[0040] ② Clusters with higher communication costs are preferentially allocated to locations close to each other in the two-dimensional grid, thereby shortening the physical distance between them and reducing communication delays.
[0041] The specific process is as follows: (1) For each cluster (The cluster center is recorded as ), the cost of cross-cluster communication is calculated as:
[0042] in, Represents edge The amount of data transmission reflects the amount of data interaction between tasks. The first summation counts the data from cluster The total amount of data flowing out to other clusters; the second summation counts the data flowing into the cluster from other clusters The total amount of data.
[0043] (2) According to the calculated communication cost, sort all clusters from high to low according to the communication cost. Suppose the sorted cluster set is: And meet
[0044] (3) Based on the physical structure of the wafer-level chip system, generate a The 2D grid coordinates of the core group , where each Represents a candidate location.
[0045] (4) In accordance with The clusters are assigned to grid locations in the order of , requiring clusters with higher communication cost to be mapped to mutually adjacent physical locations to shorten the cross-cluster communication distance.
[0046] For the first cluster , directly select any initial position in the grid.
[0047] For each subsequent cluster , select the last allocated position from the remaining available positions The position with the minimum Manhattan distance (i.e. the position of the last cluster in the mapped cluster). The Manhattan distance formula is defined as:
[0048] Thus, the new mapping satisfy:
[0049] in, Represents a collection of grid locations that have not yet been assigned.
[0050] The initial mapping result is obtained based on the initial mapping scheme, and the communication volume between clusters is calculated at the same time. The initial mapping scheme is formed based on the strategy of placing clusters with large communication volume close to each other. The parallel running time of the system is calculated based on the critical path within the cluster, the communication delay between clusters, and the dependency between clusters. The calculation strategy is as follows: (1) Intra-cluster running time: For each cluster , whose running time is equal to the critical path weight within the cluster.
[0051]
[0052] (2) Communication delay between clusters: For two clusters with data interaction and , the communication delay of the on-chip system is:
[0053] in, Representation Cluster and The Manhattan distance of the mapping kernel group on the two-dimensional coordinate grid, Representation Cluster and The communication delay between them is equal to the amount of data divided by the bandwidth.
[0054] (3) Clustering start and end time: According to the graph structure, if clustering Dependency on other clusters The output is:
[0055]
[0056] (4) System parallel running time: the maximum end time of all clusters .
[0057] Step S6: Based on the initial mapping, the simulated annealing algorithm is introduced to optimize the cluster mapping scheme. The optimization goal is to minimize the overall parallel running time of the system. This process gradually explores a better mapping scheme by randomly exchanging the positions of clusters and recalculating the parallel running time of the system. During the optimization process: (1) In each iteration of simulated annealing, two clusters are randomly selected for exchange. After each exchange, the parallel running time of the system is recalculated. .
[0058] (2) If the new solution satisfies , accept the solution, otherwise, according to the Metropolis criterion, the probability Accept new solutions to escape local optimality.
[0059] (3) The iteration stops when the temperature reaches the set threshold. The threshold is set to one thousandth of the initial temperature. The optimal solution is obtained after the iteration ends.
[0060] (4) Adaptive temperature attenuation strategy, weighted random exchange strategy and local optimal strategy optimization algorithm are added to obtain local optimality and accelerate convergence. Several strategies are as follows: Adaptive temperature decay strategy: During simulated annealing, the temperature is gradually reduced to control the disturbance amplitude of the system. The standard way to decay the temperature is to multiply the current temperature by a constant decay factor. This scheme adopts an adaptive cooling strategy. When there is no significant improvement during the optimization process, the decay factor will accelerate; conversely, if the improvement is large, the decay factor will slow down. This strategy allows for large-scale exploration in the early stage and detailed optimization in the later stage.
[0061] Weighted random exchange strategy: In order to increase the probability of clusters with higher communication costs being exchanged, clusters are sorted according to the amount of communication with other clusters during simulated annealing. By weighted selection of clusters for exchange, clusters with higher communication costs are more likely to be selected to exchange positions, thereby accelerating the convergence of the algorithm.
[0062] Locally optimal strategy: When the temperature is lower than 10%, the number of exchanges is increased to obtain the local optimal result.
[0063] Step S7: Update the cluster center.
[0064] After each round of simulated annealing, the overall weight is calculated for all nodes in each cluster. , its overall weight It consists of two parts: ① Cluster removal All other nodes except The minimum sum of the bidirectional weights between .
[0065] ②Root node to node The weight of .
[0066] The formula is as follows:
[0067] The node with the smallest overall weight is used as the new cluster center, and all nodes are re-clustered according to the original strategy.
[0068] Step S8: Repeat the iteration to obtain the optimal solution.
[0069] Repeat steps S5 to S7 until the cluster center is stable, record the optimal solution under each cluster for comparison, and finally obtain the mapping solution with the shortest overall parallel running time of the system.
[0070] Corresponding to the above method, this embodiment also discloses a wafer-level chip load balancing optimization system based on K-Means clustering and simulated annealing, comprising: The edge weight calculation module is used to calculate the node running time and edge communication delay in the AI large model and build the edge weight matrix of the entire topology; A critical path weight calculation module is used to calculate the critical path weight between any two nodes using a critical path algorithm; A grid generation module, used to generate a corresponding two-dimensional coordinate grid according to the physical layout of the wafer-level chip core group; The initial clustering module is used to select the initial clustering center of the nodes using the K-Means++ strategy and divide the nodes into corresponding clusters; An initial mapping module is used to map clusters to core groups of wafer-level chips, and preferentially map clusters with high communication costs to adjacent physical locations; The simulated annealing module is used to optimize the cluster mapping scheme through the simulated annealing algorithm. The optimization goal is to minimize the overall parallel running time of the system; The iterative output module is used to update the cluster center and repeat the iteration until the cluster center is stable, thereby obtaining the optimal solution for mapping the network topology of the large AI model to each core group of the wafer-level chip.
[0071] Unless otherwise specifically stated, the relative steps, numerical expressions and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.
[0072] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0073] The units and method steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. A person of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation is not considered to be beyond the scope of the present invention.
[0074] Those skilled in the art will appreciate that all or part of the steps in the above method can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a disk or an optical disk. Optionally, all or part of the steps in the above embodiment can also be implemented using one or more integrated circuits, and accordingly, each module / unit in the above embodiment can be implemented in the form of hardware or in the form of software function modules. The present invention is not limited to any specific form of combination of hardware and software.
[0075] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing, which evenly maps the computing tasks of the AI large model to the core group of the wafer-level chip, characterized in that: The following steps are involved: Step 1: Calculate the node running time and edge communication delay in the AI big model, and build the edge weight matrix of the entire topology; Step 2: Calculate the critical path weight between any two nodes using the critical path algorithm; Step 3: Generate a corresponding two-dimensional coordinate grid according to the physical layout of the wafer-level chip core group; Step 4: Use the K-Means++ strategy to select the initial clustering center of the nodes and divide the nodes into corresponding clusters; Step 5: Map the clusters to the core groups of the wafer-level chip, and preferentially map clusters with high communication costs to adjacent physical locations; Step 6: Optimize the cluster mapping scheme by using the simulated annealing algorithm. The optimization goal is to minimize the overall parallel running time of the system. Step 7: Update the cluster center and repeat the iteration until the cluster center is stable, and obtain the optimal solution for mapping the AI large model network topology to each core group of the wafer-level chip.
2. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 1, characterized in that: Step 1 specifically includes: Calculate each node Run time on single core group ; Calculate Edges Communication delay ,in, is the amount of data transferred, is the system bandwidth; Constructing edge weights .
3. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 1, characterized in that: Step 4 specifically includes: Randomly select a node as the first cluster center; For any remaining node, calculate its minimum weighted distance to all selected cluster centers ; by Randomly select the next cluster center for the probability distribution; Repeat the above steps until all cluster centers are selected; Based on the two-way minimum weight strategy, all nodes are divided into the cluster corresponding to the cluster center with the closest weight distance.
4. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 1, characterized in that: Step 5 specifically includes: Calculate each cluster The cross-cluster communication cost , Representation Cluster The cluster center of Sort the clusters from high to low according to the communication cost; Clusters with high communication costs are preferentially mapped to adjacent physical locations in a two-dimensional coordinate grid.
5. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 4 is characterized in that: Calculate each cluster The cross-cluster communication cost The expression is: ; in, Represents edge The first summation counts the amount of data transferred from cluster The total amount of data flowing out to other clusters; the second summation counts the data flowing into the cluster from other clusters The total amount of data.
6. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 4, characterized in that: Prioritizing mapping clusters with high communication costs to adjacent physical locations in a two-dimensional coordinate grid includes: For the first ranked cluster , select any initial position in the grid; For each subsequent cluster, the location with the minimum Manhattan distance to the last assigned location is selected among the remaining available locations in the grid.
7. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 1, characterized in that: Step 6 The optimization process of the simulated annealing algorithm includes: In each iteration of simulated annealing, the positions of the two clusters are randomly swapped. After each swap, the parallel running time of the system is recalculated. ; If the new solution is better than the current solution, then accept the new solution. Otherwise, accept new solutions; If the current temperature parameter reaches the set threshold, the iteration is stopped and the optimal solution is obtained after the iteration is completed.
8. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 7, characterized in that: System parallel running time The calculation includes: Using Clusters The intra-cluster critical path weight is used as the intra-cluster running time ; cluster and The Manhattan distance and cluster of the mapping kernel group on a two-dimensional coordinate grid and The product of the communication delays between ; Calculate the start and end time of each cluster based on cluster dependencies; The maximum value of all cluster end times is taken as the system parallel running time.
9. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 1, characterized in that: Updating the cluster center in step 7 specifically includes: Calculate the number of nodes in the cluster The overall weight of , including the sum of the minimum bidirectional weights with other nodes and the weight from the root node to the node; Select the node with the smallest overall weight as the new cluster center.
10. A wafer-level chip load balancing optimization system based on K-Means clustering and simulated annealing, characterized in that: The method for optimizing wafer-level chip load balancing based on K-Means clustering and simulated annealing according to any one of claims 1 to 9 comprises: The edge weight calculation module is used to calculate the node running time and edge communication delay in the AI large model and build the edge weight matrix of the entire topology; A critical path weight calculation module is used to calculate the critical path weight between any two nodes using a critical path algorithm; A grid generation module, used to generate a corresponding two-dimensional coordinate grid according to the physical layout of the wafer-level chip core group; The initial clustering module is used to select the initial clustering center of the nodes using the K-Means++ strategy and divide the nodes into corresponding clusters; An initial mapping module is used to map clusters to core groups of wafer-level chips, and preferentially map clusters with high communication costs to adjacent physical locations; The simulated annealing module is used to optimize the cluster mapping scheme through the simulated annealing algorithm. The optimization goal is to minimize the overall parallel running time of the system; The iterative output module is used to update the cluster center and repeat the iteration until the cluster center is stable, thereby obtaining the optimal solution for mapping the network topology of the large AI model to each core group of the wafer-level chip.
Citation Information
Patent Citations
On-chip network mapping method based on improved simulated annealing algorithm
CN108173760A
WSN clustering routing method based on improved fuzzy C-means clustering
CN115843085A
Heuristic-based field-specific wafer-level chip design optimization method and system, and storage medium
CN118821709A
Method and system for dynamically allocating computing resources of intelligent control chip
CN119759580A
Method and apparatus for detecting defect pattern on wafer based on unsupervised learning
US20200380655A1
Cited By
MoE expert deployment system and method based on wafer-level chip
CN121981181A