Design space exploration method for inter-core connection layout in multi-core processor

By applying a deep learning model based on Transformer architecture in multi-core processors, a ring-shaped connection path layout solution with low latency and high throughput is quickly searched, which solves the problems of poor scalability, high latency and limited bandwidth of the connection method between cores of the multi-core processor when scale expansion, and achieves more efficient system performance and applicability of larger system size.

CN120067034APending Publication Date: 2025-05-30GUANGDONG UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510133201.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing multi-core inter-core connection methods show poor scalability, high latency and bandwidth limitation when scaling up, especially in 2.5D integrated chips, it is difficult to effectively utilize the wiring resources of the intermediary layer.

Method used

Using a deep learning model based on Transformer architecture, through the formation of the training set and the iteration of the greedy algorithm, a ring-shaped connection path layout solution with low latency and high throughput is quickly searched, which is suitable for larger-scale multi-core processors and 2.5D integrated chips.

Benefits of technology

It realizes faster design space exploration, improves the overall performance of multi-core processor systems, can effectively utilize the wiring resources of the intermediary layer, and is suitable for larger system sizes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067034A_ABST
    Figure CN120067034A_ABST
Patent Text Reader

Abstract

The invention aims to provide a design space exploration method for inter-core connection layout in a multi-core processor. The method comprises the following steps: acquiring a system size; the system size is input into a trained deep learning model based on a Transform framework; and the deep learning model based on the Transform framework outputs a layout scheme. According to the method, searching of the annular connection path layout scheme of the multi-core processor on a larger scale can be rapidly completed, the application of the annular connection path layout scheme can be expanded to a 2.5 D integrated chip, and wiring resources of an intermediate layer network (NoI) are fully utilized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and particularly to a method for exploring the design space of the connection layout between cores in a multi-core processor. Background Art

[0002] A multi-core processor refers to a single processor integrating two or more complete cores. As the computing speed of single-core chips gets faster and faster, the heat dissipation problem gradually becomes an insurmountable obstacle. Under the dual effects of continuous increase in power consumption and decreasing performance return of single-core processors, multi-core processors have gradually replaced single-core processors and become the cornerstone of building computer systems. By dividing multi-threaded applications into multiple tasks and executing them in parallel by different cores, multi-core processors can complete more tasks within a given clock cycle, achieving a greater energy consumption ratio. On this basis, changing the number of processor cores according to application requirements can provide different performances, endowing multi-core processors with good scalability. Especially in the context of the continuous popularization of data-intensive applications, high-performance computing systems can integrate dozens or even hundreds of computing cores on a single chip. However, increasing the number of cores brings great challenges to efficient communication between cores while bringing more powerful computing capabilities. These communication bottlenecks between cores may lead to limited data bandwidth of the entire system-on-chip. Therefore, as the number of on-chip computing cores continues to increase, constructing a scalable, low-latency, and high-bandwidth communication structure according to the communication requirements of the multi-core architecture to connect each core becomes a key factor in improving the overall performance of multi-core systems.

[0003] The work on inter-core connections can be divided into bus-based methods and network-based methods. The latter can be further divided into router-based connections and routerless connections. Among them, the bus-based connection method uses a centralized communication system to achieve data transmission and exchange by connecting each core to the bus. This method is easy to implement, but has poor scalability. As the number of system cores increases, the bus load will quickly saturate, resulting in a significant decline in the overall system performance. The router-based connection method uses a decentralized communication system to provide the switching capabilities of multi-path and parallel communication through a router with a complex structure. However, the router with a complex structure will bring huge area and power consumption overheads. On this basis, in 2018, Chen Lizhong et al. proposed a routerless connection method. Motivated by reducing the huge hardware overhead brought by routers, the inter-core interconnection and interoperability are realized through a large number of directed and overlapping ring connection paths, so that the data packet can only reach the next position along the ring connection path or pop up to the current core in each clock cycle, significantly reducing the data transmission delay of the system. In 2020, the team further proposed a design space exploration method based on deep reinforcement learning for the design space of the previous proposed connection method, which is convenient for finding the layout scheme of the ring connection path. However, this design space exploration method only supports the search of on-chip network (NoC) cores and has a slow search speed. Summary of the Invention

[0004] The object of the present invention is to provide a design space exploration method for the layout of inter-core connections in a multi-core processor, which can quickly complete the search for the layout scheme of the ring connection path of a larger-scale multi-core processor and can extend its application to 2.5D integrated chips, making full use of the wiring resources of the interposer network (NoI).

[0005] A design space exploration method for the layout of inter-core connections in a multi-core processor, comprising:

[0006] Obtaining the system size;

[0007] Inputting the system size into a trained deep learning model based on the Transformer framework;

[0008] The deep learning model based on the Transformer framework outputs a layout scheme.

[0009] Preferably, after obtaining the system size, it further includes training a deep learning model based on the Transformer framework, specifically:

[0010] Using the greedy algorithm to form a training set;

[0011] Selecting different system sizes and different search starting points to search the design space;

[0012] Iterate in a loop until the exit condition is met.

[0013] Preferably, the formation of the training set using the greedy algorithm further includes:

[0014] For each piece of data in the original training set, remove the first circular connection path, and use the remaining set of circular connection paths as the new state to input into the greedy algorithm to search for a new loop, and combine to form a new piece of data;

[0015] Repeat the iteration until the specified number of times is reached or the combined state is an invalid state.

[0016] Preferably, the input of the system size into the trained deep learning model based on the Transformer framework includes:

[0017] Maintain a vocabulary according to the required system size N, where the vocabulary contains all circular connection paths within the circular connection path design space of size N and their corresponding encodings;

[0018] The encoding method is implemented by traversing all circular connection paths and incrementing step by step starting from 1 in the traversal order. For ' <unk>The encoding of 0 represents the digital encoding of all ring connection paths that are not in the vocabulary.

[0019] Map the digital encoding to the content encoding of the specified dimension through the torch.nn.Embedding() function.

[0020] Learn the corresponding position encoding according to the order of combination of ring connection paths, and add the content encoding and the position encoding to form the total input encoding.

[0021] Preferably, the output layout scheme of the deep learning model based on the Transformer framework includes:

[0022] Call the trained deep learning model based on the Transformer framework, predict an optimal ring connection path in the current state and combine it with the initial ring connection path to form a new state.

[0023] The prediction basis is to make the average number of hops of the system decrease the most in the current state.

[0024] Repeat the call-prediction-combination steps until the state reaches the end criterion, and output the current state as the optimal state searched in the current design space, that is, the required ring connection path layout scheme with low latency and high throughput.

[0025] Preferably, the repeating the call-prediction-combination steps until the state reaches the end criterion and outputting the current state as the optimal state searched in the current design space includes:

[0026] Group the cores of the multi-core processor as needed. The grouped cores are the cores on the same on-chip network and the cores on different on-chip networks.

[0027] If all the cores involved in the current ring connection path are in the same on-chip network, then regard the current ring connection path as a NoC ring connection path and route it within the on-chip network.

[0028] If the cores involved in the current ring connection path include cores on different on-chip networks, then regard the current ring connection path as a NoI ring connection path and route it within the interposer network.

[0029] A design space exploration system for the connection layout between cores in a multi-core processor, including:

[0030] A data acquisition module for acquiring the system size.

[0031] A data transmission module for inputting the system size into the trained deep learning model based on the Transformer framework.

[0032] A data processing module for outputting a layout scheme by the deep learning model based on the Transformer framework.

[0033] The beneficial effects of the present invention are as follows: 1. The present invention defines a new deep learning prediction model based on the Transformer architecture, achieving faster design space exploration; 2. The present invention expands the application scope of design space exploration, enabling it to be applied to the design space exploration task of the inter-core connection layout in 2.5D integrated chips, making full use of the wiring resources of the interposer; 3. The present invention can be applied to larger system sizes. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present invention and used together with the specification to explain the principles of the present invention.

[0035] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0036] Figure 1 It is a flowchart of a method for exploring the design space of the inter-core connection layout in a multi-core processor according to the present invention;

[0037] Figure 2 It is a schematic diagram of the non-routing connection method according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0039] It should be noted that all directional indications (such as up, down, left, right, front, back...) in the embodiments of the present invention are only used to explain the relative positional relationship and movement conditions between components in a specific posture (as shown in the drawings). If the specific posture changes, the directional indications will also change accordingly.

[0040] In addition, the descriptions involving "first", "second", etc. in the present invention are for descriptive purposes only, and should not be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0041] The work on inter-core connection can be divided into bus-based methods and network-based methods. The latter can be further divided into router-based connections and routerless connections. Among them, the bus-based connection method adopts a centralized communication system, and realizes data transmission and exchange by connecting each core to the bus. This method is easy to implement, but has poor scalability. As the number of system cores increases, the bus load will quickly saturate, resulting in a significant decline in the overall performance of the system; the router-based connection method adopts a decentralized communication system, and provides the switching ability of multi-path and parallel communication through a router with a complex structure. However, the router with a complex structure will bring huge area and power consumption overhead. On this basis, in 2018, Chen Lizhong et al. proposed a routerless connection method. Motivated by reducing the huge hardware overhead brought by routers, the inter-core interconnection and interoperability are realized through a large number of directed and overlapping ring connection paths, so that data packets can only reach the next position along the ring connection path or pop up to the current core in each clock cycle, significantly reducing the data transmission delay of the system. In 2020, the team further proposed a design space exploration method based on deep reinforcement learning for the design space of the previous proposed connection method, which is convenient for finding the layout scheme of the ring connection path. However, this design space exploration method only supports the search of on-chip network (NoC) cores and has a slow search speed.

[0042] The present invention defines a new deep learning prediction model based on the Transformer architecture to achieve faster design space exploration; the present invention expands the application scope of design space exploration so that it can be applied to the design space exploration task of the inter-core connection layout of 2.5D integrated chips, making full use of the wiring resources of the interposer layer; the present invention can be applied to larger system sizes.

[0043] Embodiment 1

[0044] A method for exploring the design space of the inter-core connection layout in a multi-core processor, referring to Figure 1 , includes:

[0045] S100, obtaining the system size;

[0046] S200. Input the system size into the trained deep learning model based on the Transformer framework.

[0047] S300. The deep learning model based on the Transformer framework outputs a layout scheme.

[0048] In the present invention, a large number of directed and overlapping loop connection paths are used to connect each core pair, so that there are no isolated nodes in the entire system. The attached Figure 2 is a schematic diagram of this non-routing connection method. It is stipulated that the cores of the multi-core system are arranged in a square matrix, and the cores are aligned vertically, horizontally, and evenly distributed; it is stipulated that the loop connection path is a regular rectangle, that is, the cores in the upper left corner and the lower right corner of the loop connection path cannot be in the same row or the same column; the hop count between core pairs is defined as the Manhattan distance between two core pairs along the given loop connection path, and the average hop count of the system is the average value of such Manhattan distances between all core pairs. The larger this average value, the greater the system data transmission delay; the system size is defined as the number of cores in each row (the number of cores in each column is also the same); given the system size, the design space of the system is determined accordingly, which is the set of all non-repeating loop connection paths. The size of this design space increases sharply with the increase of the system size; the layout state of the loop connection path (hereinafter referred to as the state for short) is defined as the set of different loop connection paths existing in the current system. That is, the state is a subset of the design space.

[0049] The purpose of the present invention is to accelerate the search speed for the design space of the non-routing loop connection path layout between the cores of a multi-core processor, improve the scalability of the search method, make full use of the wiring resources, and propose a design space search method using a deep learning model based on the Transformer architecture to search for a loop connection path layout scheme with low latency and high throughput, so as to improve the overall performance of the multi-core processor system. The processing cores in the multi-core processor targeted by this invention patent are arranged according to the mesh rule. For example, 9 cores are arranged in a 3X3 pattern. The loop connection path (abbreviated as loop) is required to be a regular rectangle with the head and tail connected, which will connect the cores on the path. The requirement for the rectangle is that the core in the upper left corner and the core in the lower right corner are not in the same row and not in the same column.

[0050] Preferably, after S100 obtains the system size, it further includes S110, training the deep learning model based on the Transformer framework, specifically:

[0051] S111. Use the greedy algorithm to form a training set.

[0052] S112. Select different system sizes and different search starting points to search the design space.

[0053] Complete the design space search round by round according to different system sizes and different search starting points. In each round, a new circular connection path is added on the basis of the previous round until the jump-out condition is met. The present invention gradually increases step by step to complete the search of the design space for a given size and starting point. The main role of the deep learning model of the Transformer framework is to predict the next optimal circular connection path for the currently obtained search result. Therefore, the same model can be applied to different system sizes and search starting points, but a complete search is for a fixed size and starting point.

[0054] S113, iterate until the jump-out condition is met.

[0055] Preferably, S111, forming a training set using the greedy algorithm further includes:

[0056] For each piece of data in the original training set, remove the first circular connection path, and use the remaining set of circular connection paths as a new state to input the greedy algorithm to search for a new loop and combine it to form a new piece of data;

[0057] Repeat the iteration until the specified number of times is reached or the combined state is an invalid state.

[0058] The deep learning model based on the Transformer architecture involved in the present invention requires a large amount of training data so that it can learn an "intuition" for selecting the optimal circular connection path in the current state. The original data set is formed by the greedy algorithm. Select some different system sizes and different search starting points to complete the exploration of the design space, and this process consumes relatively more time. The implementation method of the greedy algorithm is to select the circular connection path that reduces the average number of hops the most each time and combine it with the previous state until the end condition is reached. The data volume of this original data set cannot support model training. Therefore, the following data set expansion method is used: for each piece of data in the original data set, remove the first circular connection path, and use the remaining set of circular connection paths as a new state to input the greedy algorithm to search for a new loop and combine it to form a new piece of data. Repeat this way until the specified number of times is reached or the combined state is an invalid state. For example, an original piece of data contains three circular connection paths A, B, and C. When expanding the data, remove path A, combine paths B and C as a new state, search for path D through the greedy algorithm, and combine the three circular connection paths B, C, and D as a new piece of data. This expansion method can fully capture the context information of the search process and is beneficial for the deep learning model to learn the "intuition" for selecting the optimal circular connection path.

[0059] Preferably, S200, inputting the system size into the trained deep learning model based on the Transformer framework includes:

[0060] Maintain a vocabulary according to the required system size N, where the vocabulary contains all the loop-connected paths and their corresponding encodings within the loop-connected path design space of size N;

[0061] N is the maximum system size that the model can support and needs to be given before training. For example, for a 3X3 mesh arrangement of cores, N is 3.

[0062] The encoding method is implemented by traversing all the loop-connected paths and incrementing step by step starting from 1 in the traversal order. For ' <unk>'Encoded as 0, it represents the digital encoding of all circular connection paths not in the vocabulary;

[0063] Map the digital encoding to the content encoding of the specified dimension through the torch.nn.Embedding() function;

[0064] Learn the corresponding position encoding according to the order of combination of circular connection paths, and add the content encoding and the position encoding to form the total input encoding.

[0065] The design space exploration method involved in the present invention is based on the Transformer architecture. In the present invention, the current state is used as the input of the Transformer. Since the data form of the current state cannot be directly used as the input of the Transformer, the following modifications are made. A vocabulary is maintained according to the required system size N, and the vocabulary contains all circular connection paths in the circular connection path design space of size N and their corresponding encodings. The encoding method is to traverse all circular connection paths and implement it by incrementing step by step starting from 1 in the traversal order. Additionally, for ' <unk>'Encoded as 0, it represents the digital encoding of all ring connection paths not in the vocabulary. Then, through the torch.nn.Embedding() function, the digital encoding is mapped to the content encoding of the specified dimension. After that, according to the order of combination of the ring connection paths, the position encoding is learned, and the content encoding and the position encoding are added together to form the total input encoding.

[0066] Preferably, in S300, the output layout scheme of the deep learning model based on the Transformer framework includes:

[0067] S310, call the trained deep learning model based on the Transformer framework, predict an optimal ring connection path in the current state and combine it with the initial ring connection path to form a new state;

[0068] S320, the prediction basis is to make the system average hop count decrease the most in the current state;

[0069] The definition of the hop count is the absolute difference between the shortest horizontal and vertical coordinates between two node pairs (the Manhattan distance between two nodes with a path length of 1 between adjacent nodes), and the average hop count is the average of the hop counts of all node pairs.

[0070] S330, repeat the call-prediction-combination steps until the state reaches the end criterion, and output the current state as the optimal state searched in the current design space, that is, the required ring connection path layout scheme with low latency and high throughput.

[0071] In the embodiment of the present invention, starting from a given system size, search begins from an initial ring connection path, which can be changed. For the convenience of description, it is set here as a clockwise ring connection path passing through all the cores in the outermost circle. Call the trained deep learning model based on the Transformer framework, predict an optimal ring connection path in the current state and combine it with the initial ring connection path to form a new state, and the prediction basis is to make the system average hop count decrease the most in the current state. Repeat the above call-prediction-combination steps until the state reaches the end criterion, and output the current state as the optimal state searched in the current design space, that is, the required ring connection path layout scheme with low latency and high throughput.

[0072] Preferably, in S330, repeating the call-prediction-combination steps until the state reaches the end criterion and outputting the current state as the optimal state searched in the current design space includes:

[0073] Group the cores of the multi-core processor as needed. The grouped cores are the cores on the same on-chip network and the cores on different on-chip networks;

[0074] If all the cores involved in the current ring connection path are in the same on-chip network, then the current ring connection path is regarded as the NoC ring connection path and routed within the on-chip network;

[0075] If the cores involved in the current ring connection path are included in cores of different on-chip networks, then the current ring connection path is regarded as the NoI ring connection path and routed within the interposer network.

[0076] In the embodiment of the present invention, the cores of the multi-core processor are grouped as needed. After grouping, there are two states for the core pairs, namely cores on the same on-chip network and cores on different on-chip networks. All the ring connection paths in the ring connection path layout searched by the design space exploration method claimed in the present invention are judged. If all the cores involved in the ring connection path are in the same on-chip network, then the ring connection path is regarded as the NoC ring connection path and routed within the on-chip network; if the cores involved in the ring connection path are included in cores of different on-chip networks, then the ring connection path is regarded as the NoI ring connection path and routed within the interposer network. In this way, the design space exploration method claimed in the present invention can extend the application scope to 2.5D integrated chips.

[0077] Embodiment 2

[0078] A design space exploration system for the connection layout between cores in a multi-core processor, comprising:

[0079] A data acquisition module, configured to acquire the system size;

[0080] A data transmission module, configured to input the system size into a trained deep learning model based on the Transformer framework;

[0081] A data processing module, configured to output a layout scheme based on the deep learning model of the Transformer framework.

[0082] The present invention defines a new deep learning prediction model based on the Transformer architecture to achieve faster design space exploration; the present invention expands the application scope of the design space exploration so that it can be applied to the design space exploration task of the connection layout between cores of 2.5D integrated chips, making full use of the wiring resources of the interposer; the present invention can be applied to larger system sizes.

[0083] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.< / unk> < / unk> < / unk>

Claims

1. A design space exploration method for inter-core connection layout in a multi-core processor, characterized in that: include: Get system size; Inputting the system size into the trained deep learning model based on the Transformer framework; The deep learning model output layout solution based on the Transformer framework.

2. The design space exploration method for inter-core connection layout in a multi-core processor according to claim 1, characterized in that: After obtaining the system size, the process also includes training a deep learning model based on the Transformer framework, specifically: Use greedy algorithm to form training set; Select different system sizes and different search starting points to search the design space; The loop iterates until the exit condition is met.

3. The design space exploration method for inter-core connection layout in a multi-core processor according to claim 1, characterized in that: The method of using the greedy algorithm to form a training set also includes: For each piece of data in the original training set, remove the first loop connection path, and use the remaining loop connection path set as the new state input into the greedy algorithm to search for a new loop, and combine to form a new piece of data; Repeat until the specified number of iterations is reached or the combined state is an invalid state.

4. The design space exploration method for inter-core connection layout in a multi-core processor according to claim 1, characterized in that: The inputting the system size into the trained deep learning model based on the Transformer framework includes: Maintaining a vocabulary table according to the required system size N, the vocabulary table containing all the annular connection paths and corresponding codes in the annular connection path design space of size N; The encoding method is to traverse all the ring connection paths, starting from 1 and gradually adding 1 in the traversal order. <unk> 'The code is 0, indicating the digital code of all the circular connection paths that are not in the vocabulary;< / unk> Use torch.nn.Embedding() function to map digital encoding to content encoding of specified dimension; The corresponding position codes are learned according to the sequence of the combination of the annular connection paths, and the content code and the position code are added together to form the total input code.

5. The design space exploration method for inter-core connection layout in a multi-core processor according to claim 1, characterized in that: The deep learning model output layout solution based on the Transformer framework includes: Call the trained deep learning model based on the Transformer framework to predict an optimal ring connection path in the current state and combine it with the initial ring connection path to form a new state; The prediction basis is to reduce the average number of system hops the most in the current state; Repeat the call-prediction-combination steps until the state reaches the end criterion, and output the state at this time as the optimal state searched in the current design space, that is, the required low-latency, high-throughput ring connection path layout solution.

6. The design space exploration method for inter-core connection layout in a multi-core processor according to claim 5, characterized in that: The steps of repeatedly calling, predicting and combining are repeated until the state reaches the end criterion, and the state at this time is output as the optimal state searched in the current design space, including: The cores of the multi-core processor are grouped as needed, and the grouped cores are cores on the same network on chip and cores on different networks on chip; If all cores involved in the current ring connection path are in the same on-chip network, the current ring connection path is used as a NoC ring connection path and routed in the on-chip network; If the cores involved in the current ring connection path are included in cores of different on-chip networks, the current ring connection path is used as a NoI ring connection path and is routed in the interposer network.

7. A design space exploration system for inter-core connection layout in a multi-core processor, characterized in that: include: A data acquisition module, used to obtain system dimensions; A data transmission module, used for inputting the system size into a trained deep learning model based on a Transformer framework; The data processing module is used for the deep learning model output layout solution based on the Transformer framework.