Wafer-level Chip Load Balancing Optimization Method and System Based on K-Means Clustering and Simulated Annealing

By adopting the combination method of K-Means clustering and simulated annealing algorithm in wafer-level chip systems, the load imbalance problem of AI large-model network topology map in wafer-level chip systems is solved, achieving more efficient load balancing and shorter parallel running time.

CN120030977BActive Publication Date: 2025-06-20HENAN SONGSHAN LAB IND RES INST CO LTD LUOYANG BRANCH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510510079.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-06-20
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The prior art is difficult to realize the balanced mapping of AI large-model network topology maps in wafer-level chip systems, resulting in unbalanced load and long overall parallel running time.

Method used

Using K-Means clustering and simulated annealing methods, we construct an edge weight matrix by calculating the node's running time and edge communication delay, using the critical path algorithm to calculate the key path weights between nodes, generate a two-dimensional coordinate grid, perform K-Means++ cluster division, and optimize the cluster mapping scheme through the simulated annealing algorithm to achieve load balancing.

Benefits of technology

The load balancing of AI large-model computing tasks on wafer-level chip systems is realized, which significantly reduces the overall parallel running time of the system and improves the utilization efficiency of hardware resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030977B_ABST
    Figure CN120030977B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of electronic design automation of chips, and particularly relates to a wafer-level chip load balancing optimization method and system based on K-Means clustering and simulated annealing. An edge weight matrix of the entire topology is constructed; the critical path weight between any two nodes is calculated using the critical path algorithm; according to the physical layout of the wafer-level chip core groups, a corresponding two-dimensional coordinate grid is generated; the K-Means++ strategy is used to select the initial clustering centers of the nodes and divide the nodes into the corresponding clusters; the clusters are mapped onto the core groups of the wafer-level chip, and the clusters with high communication costs are preferentially mapped to adjacent physical positions; the mapping scheme of the clusters is optimized through the simulated annealing algorithm, and the optimization goal is to minimize the overall parallel running time of the system; the clustering centers are updated and the iteration is repeated until the clustering centers are stable, and the optimal mapping scheme is output. The present invention evenly maps the network topology diagram of the AI large model onto each core group to achieve more efficient load balancing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of electronic design automation of chips, and particularly relates to a wafer-level chip load balancing optimization method and system based on K-Means clustering and simulated annealing. In a multi-core system of wafer-level chips, an AI large model network topology map is evenly mapped into a chip core group architecture to optimize the overall parallel running time. Background Art

[0002] In recent years, with the rapid development of artificial intelligence, the scale of AI large models has been continuously expanding, and the computational nodes and data transmission relationships involved have become more and more complex. Therefore, effectively deploying an AI large model into a wafer-level chip system has become an urgent problem to be solved. In order to improve the computing efficiency without increasing the hardware cost, a solution that can evenly map the network topology map of the AI large model to each core group of the wafer-level chip is needed.

[0003] In existing traditional algorithms, usually only local optimization or scheduling based on a single strategy is concerned, and it is difficult to take into account the operator node running time, data transmission delay, node dependency relationship, and wafer-level chip system architecture at the same time, resulting in a relatively long overall parallel running time of the on-chip system and it is difficult to achieve precise load balancing. Summary of the Invention

[0004] The present invention aims to solve the problem of unbalanced load of the AI large model in the wafer-level chip system, and proposes a wafer-level chip load balancing optimization method and system based on K-Means clustering and simulated annealing. By comprehensively considering the operator node running time, data transmission delay, node dependency relationship, and wafer-level chip system architecture, the network topology map of the AI large model is evenly mapped onto each core group to achieve more efficient load balancing.

[0005] To achieve the above object, the technical solutions adopted are as follows:

[0006] The present invention provides a wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing, which evenly maps the computing tasks of the AI large model to the core groups of the wafer-level chip, including the following steps:

[0007] Step 1: Calculate the running duration of nodes and edge communication delays in the AI large model, and construct an edge weight matrix for the entire topology;

[0008] Step 2: Use the critical path algorithm to calculate the critical path weights between any two nodes;

[0009] Step 3: Generate a corresponding two-dimensional coordinate grid according to the physical layout of the wafer-level chip core group;

[0010] Step 4: Use the K-Means++ strategy to select the initial clustering center of the nodes and divide the nodes into corresponding clusters;

[0011] Step 5: Map the clusters to the core groups of the wafer-level chip, and preferentially map clusters with high communication costs to adjacent physical locations;

[0012] Step 6: Optimize the cluster mapping scheme by using the simulated annealing algorithm. The optimization goal is to minimize the overall parallel running time of the system.

[0013] Step 7: Update the cluster center and repeat the iteration until the cluster center is stable, and obtain the optimal solution for mapping the AI ​​large model network topology to each core group of the wafer-level chip.

[0014] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, step 1 specifically includes:

[0015] Calculate each node Running time on a single core group ;

[0016] Calculate Edges Communication delay ,in, is the amount of data transferred, is the system bandwidth;

[0017] Constructing edge weights .

[0018] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, step 4 specifically includes:

[0019] Randomly select a node as the first cluster center;

[0020] For any remaining node, calculate its minimum weighted distance to all selected cluster centers ;

[0021] by Randomly select the next cluster center for the probability distribution;

[0022] Repeat the above steps until all cluster centers are selected;

[0023] Based on the two-way minimum weight strategy, all nodes are divided into the cluster corresponding to the cluster center with the closest weight distance.

[0024] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, step 5 specifically includes:

[0025] Calculate the cross-cluster communication cost for each cluster ; let be the clustering center of cluster ; Sort the clusters in descending order according to the communication cost

[0026] ;

[0027] Map the clusters with high communication cost to adjacent physical positions in the two-dimensional coordinate grid preferentially

[0028] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, the expression for calculating the cross-cluster communication cost for each cluster is as follows :

[0029] wherein, represents the data transmission volume of edge ; the first summation term counts the total data volume flowing out from cluster to other clusters; the second summation term counts the total data volume flowing into cluster from other clusters

[0030] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, mapping the clusters with high communication cost to adjacent physical positions in the two-dimensional coordinate grid preferentially includes

[0031] For the cluster ranked first , select any initial position in the grid

[0032] For each subsequent cluster, select the position with the minimum Manhattan distance from the previous assigned position among the remaining available positions in the grid

[0033] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, the optimization process of step 6 of the simulated annealing algorithm includes

[0034] In each iteration of the simulated annealing, randomly swap the positions of two clusters, and after each swap, recalculate the system parallel running time ;

[0035] If the new solution is better than the current solution, accept the new solution; otherwise, accept the new solution according to the probability ;

[0036] Stop the iteration when the current temperature parameter reaches the set threshold, and the optimal solution is obtained after the iteration ends

[0037] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, the calculation of the system parallel running time includes:

[0038] Using the weight of the critical path within the cluster as the running time within the cluster ;

[0039] The product of the Manhattan distance of the mapping kernel groups of clusters and on the two-dimensional coordinate grid and the communication delay between clusters and is used as the cross-cluster communication delay ;

[0040] According to the cluster dependency relationship, calculate the start and end times of each cluster;

[0041] Take the maximum value of the end times of all clusters as the system parallel running time.

[0042] According to the wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing of the present invention, further, the specific update of the clustering center in step 7 includes:

[0043] Calculate the overall weight of each node within the cluster, including the sum of the minimum values of the two-way weights with other nodes and the weight from the root node to this node;

[0044] Select the node with the minimum overall weight as the new clustering center.

[0045] Further, the present invention also provides a wafer-level chip load balancing optimization system based on K-Means clustering and simulated annealing, which is used to implement the above-mentioned wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing, and includes:

[0046] An edge weight calculation module, which is used to calculate the node running duration and edge communication delay in the AI large model, and construct an edge weight matrix of the entire topology;

[0047] A critical path weight calculation module, which is used to calculate the critical path weight between any two nodes by using the critical path algorithm;

[0048] A grid generation module, which is used to generate a corresponding two-dimensional coordinate grid according to the physical layout of the wafer-level chip kernel group;

[0049] An initial clustering division module, which is used to select the initial clustering center of the nodes by using the K-Means++ strategy and divide the nodes into the corresponding clusters;

[0050] An initial mapping module, which is used to map clusters to the core groups of a wafer-level chip, and preferentially maps the clusters with high communication costs to adjacent physical locations;

[0051] A simulated annealing module, which is used to optimize the mapping scheme of clusters through the simulated annealing algorithm, and the optimization goal is to minimize the overall parallel running time of the system;

[0052] An iterative output module, which is used to update the clustering center and repeat the iteration until the clustering center is stable, so as to obtain the optimal mapping scheme of the AI large model network topology map to each core group of the wafer-level chip.

[0053] Adopting the above technical solution, the beneficial effects obtained are as follows:

[0054] The present invention comprehensively considers multiple constraint factors such as the running time of operator nodes, data transmission delay, node dependency relationships, and wafer-level system architectures. By combining the K-Means clustering and simulated annealing algorithms, it realizes the load balancing of AI large model computing tasks on the wafer-level chip system, minimizes the overall parallel running time of the system, optimizes the utilization efficiency of hardware resources, and provides reliable technical support for AI computing of the wafer-level chip system. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings of the embodiments of the present invention will be briefly introduced below. Among them, the drawings are only used to show some embodiments of the present invention, rather than limiting all embodiments of the present invention thereto.

[0056] Figure 1 It is a schematic flowchart of a wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] In the following, the exemplary solutions of the embodiments of the present invention will be clearly and completely described in conjunction with the drawings of the specific embodiments of the present invention. Unless otherwise defined, the technical terms or scientific terms used in the present invention should have the ordinary meaning understood by those of ordinary skill in the art.

[0058] This embodiment discloses a wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing, which evenly maps the computing tasks of an AI large model to the core groups of a wafer-level chip, as Figure 1 shown, and specifically includes:

[0059] Input data:

[0060] (1) AI large model topology graph: represented by a directed acyclic graph, which records the relationships between various computing tasks of the AI large model.

[0061] (2) Node running time: The running time of each task node on the single-core group of the on-chip system.

[0062] (3) Edge data transfer volume: The amount of data transferred between computing nodes.

[0063] (4) Wafer-level system architecture: Used to construct a two-dimensional core group coordinate grid.

[0064] Step S1: Calculate the node running duration and edge communication delay in the AI large model, and construct the edge weight matrix of the entire topology.

[0065] First, estimate the running duration of each computing node in the AI large model on the single-core group:

[0066]

[0067] At the same time, for each edge connecting two nodes , calculate the communication delay between the nodes according to the data volume transmitted by this edge and the system bandwidth:

[0068]

[0069] Among them, is the data transfer volume of any edge , is the system bandwidth.

[0070] Considering the node running duration and edge communication delay comprehensively, calculate the edge weight information between all connected nodes in the topology. The calculation method of the edge weight is as follows:

[0071]

[0072] Step S2: According to the edge weight information in Step S1, use the critical path algorithm to calculate the critical path weight between any two nodes. Denote the sum of the edge weights on the critical path as , note to subtract the running time of the intermediate nodes that are repeatedly counted. If the critical path is misjudged due to repeated calculation, subsequent clustering and mapping optimization may target the wrong path, resulting in the system parallel running time not being truly minimized.

[0073] Step S3: Generate the corresponding two-dimensional coordinate grid according to the actual physical layout of the wafer-level chip core group. For example, if the system contains core groups, coordinates can be constructed according to the core group layout , and each coordinate point corresponds to a core group mapping position, preparing for subsequent task mapping.

[0074] Step S4: According to the edge weight matrix, select using the K-Means++ strategy an initial clustering center, where its size is equal to the number of cores in the wafer-level chip system, and the nodes are partitioned into the corresponding clusters. This step includes sub-steps S401 - S405.

[0075] Step S401: Randomly select a node as the first clustering center .

[0076] Step S402: For any remaining node, calculate its minimum weighted distance from all the selected clustering centers :

[0077] where and are the two-way weights between node and clustering center .

[0078] Step S403: Randomly select the next clustering center with as the probability distribution.

[0079] Step S404: Repeat steps S402 - S403 until clustering centers are selected.

[0080] Step S405: After the selection of the clustering centers is completed, based on the two-way minimum weight strategy, all nodes are partitioned into the clusters corresponding to the clustering center with the closest weighted distance. Assign each node to the clustering center with the closest weighted distance:

[0081] Step S5: Perform an initial mapping between the initially formed clusters and the wafer-level chip core groups, and preferentially map the clusters with high communication costs to adjacent physical locations.

[0082] Generate two-dimensional grid positions as candidate positions for the physical core groups according to the wafer-level chip system core group architecture. At the same time, to reduce the communication delay between clusters, this embodiment designs an improved initial mapping strategy, and its core idea is:

[0083] ① Calculate the inter-cluster communication cost for each cluster, which reflects the total amount of data interaction between this cluster and other clusters.

[0084] ② Preferentially assign the clusters with higher communication costs to positions close to each other in the two-dimensional grid, so as to shorten the physical distance between them and reduce the communication delay.

[0085] The specific process is as follows:

[0086] (1) For each cluster (whose clustering center is denoted as ), the cross-cluster communication cost is calculated as follows:

[0087]

[0088] Among them, represents the data transmission volume of edge , reflecting the data interaction volume between tasks. The first summation term counts the total data volume flowing out from cluster to other clusters; the second summation term counts the total data volume flowing into cluster from other clusters.

[0089] (2) According to the calculated communication cost, all clusters are sorted in descending order of communication cost. Let the sorted cluster set be:

[0090] and satisfy

[0091] (3) According to the physical structure of the wafer-level chip system, a two-dimensional grid coordinate set containing nuclear groups is generated, where each represents a candidate position.

[0092] (4) In the order of , the clusters are assigned to the grid positions, requiring that the clusters with larger communication costs be mapped to physically adjacent positions to shorten the cross-cluster communication distance.

[0093] For the first cluster , any initial position in the grid is directly selected.

[0094] For each subsequent cluster , among the remaining available positions, the position with the minimum Manhattan distance from the previous assigned position (i.e., the position of the last cluster in the already mapped clusters) is selected. Among them, the Manhattan distance formula is defined as:

[0095] In this way, the new mapping satisfies:

[0096] Among them, represents the set of grid positions that have not been assigned yet.

[0097] According to the initial mapping scheme, the initial mapping result is obtained, and at the same time, the communication volume between clusters is calculated. According to the strategy of placing close positions with large communication volume, an initial mapping scheme is formed. According to the intra-cluster critical path, inter-cluster communication delay, and inter-cluster dependency relationship, the system parallel running time is calculated, and its calculation strategy is as follows:

[0098] (1) Intra-cluster running time: For each cluster , its running time is equal to the critical path weight within the cluster.

[0099]

[0100] (2) Inter-cluster communication delay: For two clusters and with data interaction, the communication delay in the on-chip system is:

[0101] where represents the Manhattan distance of the mapping kernel groups of clusters and on the two-dimensional coordinate grid, and represents the communication delay between clusters and , with a magnitude of the data volume divided by the bandwidth.

[0102] (3) Clustering start and end times: According to the graph structure, if clustering depends on the output of other clustering , then there is:

[0103]

[0104]

[0105] (4) System parallel running time: The maximum value of the end times of all clusters .

[0106] Step S6. Based on the initial mapping, introduce the simulated annealing algorithm to optimize the mapping scheme of clusters, with the optimization goal of minimizing the overall parallel running time of the system. This process explores a better mapping scheme step by step by randomly swapping the positions of clusters and recalculating the system parallel running time. During the optimization process:

[0107] (1) In each iteration of simulated annealing, randomly select two clusters to swap. After each swap, recalculate the system parallel running time .

[0108] (2) If the new scheme satisfies , accept the scheme; otherwise, accept the new scheme with a probability according to the Metropolis criterion to jump out of the local optimum.

[0109] (3) Stop the iteration when the temperature reaches the set threshold, which is set to one-thousandth of the initial temperature. The optimal solution is obtained after the iteration ends.

[0110] (4) Add an adaptive temperature decay strategy, a weighted random exchange strategy, and a local optimal strategy to optimize the algorithm, obtaining the local optimum and accelerating convergence. The several strategies are as follows:

[0111] Adaptive temperature decay strategy:

[0112] During the simulated annealing process, the temperature gradually decreases to control the perturbation amplitude of the system. The standard way of temperature decay is to multiply the current temperature by a constant decay factor. This scheme adopts an adaptive cooling strategy. When there is no significant improvement during the optimization process, the decay factor accelerates; conversely, if the improvement is large, the decay factor slows down. This strategy can conduct a large-scale exploration in the initial stage and perform detailed optimization in the later stage.

[0113] Weighted random exchange strategy:

[0114] To increase the probability of exchanging clusters with a larger communication cost, during the simulated annealing process, sort the clusters according to their communication volume with other clusters. Select clusters for exchange through weighted selection. Clusters with a larger communication cost are more likely to be selected to exchange positions, thus accelerating the convergence of the algorithm.

[0115] Local optimal strategy:

[0116] When the temperature is below 10%, increase the number of exchanges to obtain the local optimal result.

[0117] Step S7: Update the cluster center.

[0118] After each round of simulated annealing, calculate the overall weight for all nodes within each cluster. For any node in the cluster , its overall weight consists of two parts:

[0119] ① The sum of the minimum values of the two-way weights between all other nodes in the cluster except and .

[0120] ② The weight from the root node to node .

[0121] The formula is as follows:

[0122]

[0123] And use the node with the minimum overall weight as the new cluster center, and re-cluster all nodes according to the original strategy.

[0124] Step S8: Repeat the iteration to obtain the optimal solution.

[0125] Repeat the process of steps S5 to S7 until the clustering centers are stable, record the optimal solutions for each type of clustering for comparison, and finally obtain the mapping solution with the shortest overall parallel running time of the system.

[0126] Correspondingly, this embodiment also discloses a wafer-level chip load balancing optimization system based on K-Means clustering and simulated annealing, including:

[0127] An edge weight calculation module, configured to calculate the node running duration and edge communication delay in the AI large model, and construct an edge weight matrix for the entire topology;

[0128] A critical path weight calculation module, configured to calculate the critical path weight between any two nodes by using the critical path algorithm;

[0129] A grid generation module, configured to generate a corresponding two-dimensional coordinate grid according to the physical layout of the wafer-level chip core group;

[0130] An initial clustering division module, configured to select the initial clustering centers of the nodes by using the K-Means++ strategy and divide the nodes into the corresponding clusters;

[0131] An initial mapping module, configured to map the clusters to the core groups of the wafer-level chip, and preferentially map the clusters with high communication costs to adjacent physical positions;

[0132] A simulated annealing module, configured to optimize the mapping solution of the clusters by using the simulated annealing algorithm, and the optimization goal is to minimize the overall parallel running time of the system;

[0133] An iterative output module, configured to update the clustering centers and repeat the iteration until the clustering centers are stable, and obtain the optimal solution for mapping the AI large model network topology to each core group of the wafer-level chip.

[0134] Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0135] Each embodiment in this specification is described in a progressive manner. The key points of each embodiment are the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description in the method part.

[0136] The units and method steps of the examples described in connection with the embodiments disclosed in this document can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art can use different methods to implement the described functions for each specific application, but such implementation is not considered to exceed the scope of the present invention.

[0137] Those of ordinary skill in the art can understand that all or part of the steps in the above methods can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, each module / unit in the above embodiments can be implemented in the form of hardware or in the form of a software functional module. The present invention is not limited to any specific form of the combination of hardware and software.

[0138] Finally, it should be noted that the above-described embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments or easily conceive of changes, or make equivalent replacements for some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be determined by the protection scope of the claims.

Claims

1. A wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing, which evenly maps the computing tasks of the AI ​​large model to the core group of the wafer-level chip, characterized in that: The following steps are involved: Step 1: Calculate the node running time and edge communication delay in the AI ​​big model, and build the edge weight matrix of the entire topology; Step 2: Calculate the critical path weight between any two nodes using the critical path algorithm; Step 3: Generate a corresponding two-dimensional coordinate grid according to the physical layout of the wafer-level chip core group; Step 4: Use the K-Means++ strategy to select the initial clustering center of the nodes and divide the nodes into corresponding clusters; Step 5: Map the clusters to the core groups of the wafer-level chip, and preferentially map clusters with high communication costs to adjacent physical locations; Step 6: Optimize the cluster mapping scheme by using the simulated annealing algorithm. The optimization goal is to minimize the overall parallel running time of the system. Step 7: Update the cluster center and repeat the iteration until the cluster center is stable, and obtain the optimal solution for mapping the AI ​​large model network topology to each core group of the wafer-level chip.

2. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 1, characterized in that: Step 1 specifically includes: Calculate each node Run time on single core group ; Calculate Edges Communication delay ,in, is the amount of data transferred, is the system bandwidth; Constructing edge weights .

3. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 1, characterized in that: Step 4 specifically includes: Randomly select a node as the first cluster center; For any remaining node, calculate its minimum weighted distance to all selected cluster centers ; by Randomly select the next cluster center for the probability distribution; Repeat the above steps until all cluster centers are selected; Based on the two-way minimum weight strategy, all nodes are divided into the cluster corresponding to the cluster center with the closest weight distance.

4. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 1, characterized in that: Step 5 specifically includes: Calculate each cluster The cross-cluster communication cost , Representation Cluster The cluster center of Sort the clusters from high to low according to the communication cost; Clusters with high communication costs are preferentially mapped to adjacent physical locations in a two-dimensional coordinate grid.

5. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 4 is characterized in that: Calculate each cluster The cross-cluster communication cost The expression is: ; in, Represents edge The first summation counts the amount of data transferred from cluster The total amount of data flowing out to other clusters; the second summation counts the data flowing into the cluster from other clusters The total amount of data.

6. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 4, characterized in that: Prioritizing mapping clusters with high communication costs to adjacent physical locations in a two-dimensional coordinate grid includes: For the first ranked cluster , select any initial position in the grid; For each subsequent cluster, the location with the minimum Manhattan distance to the last assigned location is selected among the remaining available locations in the grid.

7. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 1, characterized in that: Step 6 The optimization process of the simulated annealing algorithm includes: In each iteration of simulated annealing, the positions of the two clusters are randomly swapped. After each swap, the parallel running time of the system is recalculated. ; If the new solution is better than the current solution, then accept the new solution. Otherwise, accept new solutions; If the current temperature parameter reaches the set threshold, the iteration is stopped and the optimal solution is obtained after the iteration is completed.

8. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 7, characterized in that: System parallel running time The calculation includes: Using Clusters The intra-cluster critical path weight is used as the intra-cluster running time ; cluster and The Manhattan distance and cluster of the mapping kernel group on a two-dimensional coordinate grid and The product of the communication delays between ; Calculate the start and end time of each cluster based on cluster dependencies; The maximum value of all cluster end times is taken as the system parallel running time.

9. The wafer-level chip load balancing optimization method based on K-Means clustering and simulated annealing according to claim 1, characterized in that: Updating the cluster center in step 7 specifically includes: Calculate the number of nodes in the cluster The overall weight of , including the sum of the minimum bidirectional weights with other nodes and the weight from the root node to the node; Select the node with the smallest overall weight as the new cluster center.

10. A wafer-level chip load balancing optimization system based on K-Means clustering and simulated annealing, characterized in that: The method for optimizing wafer-level chip load balancing based on K-Means clustering and simulated annealing according to any one of claims 1 to 9 comprises: The edge weight calculation module is used to calculate the node running time and edge communication delay in the AI ​​large model and build the edge weight matrix of the entire topology; A critical path weight calculation module is used to calculate the critical path weight between any two nodes using a critical path algorithm; A grid generation module, used to generate a corresponding two-dimensional coordinate grid according to the physical layout of the wafer-level chip core group; The initial clustering module is used to select the initial clustering center of the nodes using the K-Means++ strategy and divide the nodes into corresponding clusters; An initial mapping module is used to map clusters to core groups of wafer-level chips, and preferentially map clusters with high communication costs to adjacent physical locations; The simulated annealing module is used to optimize the cluster mapping scheme through the simulated annealing algorithm. The optimization goal is to minimize the overall parallel running time of the system; The iterative output module is used to update the cluster center and repeat the iteration until the cluster center is stable, thereby obtaining the optimal solution for mapping the network topology of the large AI model to each core group of the wafer-level chip.

Citation Information

Patent Citations

  • Method and system for dynamically allocating computing resources of intelligent control chip

    CN119759580A

  • Method and apparatus for detecting defect pattern on wafer based on unsupervised learning

    US20200380655A1