Micro-architecture design space exploration method based on deep learning

Through the deep learning-based microarchitecture design space exploration method, the graph attention network and federated learning algorithm are used to solve the cumbersome and time-consuming problem of the existing microarchitecture design process, and more efficient and accurate microarchitecture parameter configuration is achieved.

CN119990013AActive Publication Date: 2025-05-13GUANGDONG UNIV OF TECH

Patent Information

Application Number
CN202510079982.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-19
Publication Date
2025-05-13
Estimated Expiration
2045-01-19

AI Technical Summary

Technical Problem

The existing microarchitecture design process is cumbersome and time-consuming, making it difficult to achieve optimal global performance.

Method used

The microarchitecture design space exploration method based on deep learning is adopted. By obtaining the target processor microarchitecture parameters, building design space, clustering design points, creating directed ring-free subgraphs, using the graph attention network to build a prediction model, obtaining the optimal parameters that meet the design goals, and optimizing the global prediction model through federated learning algorithms.

Benefits of technology

The efficiency and flexibility of microarchitecture design is improved, and the accuracy and generalization capabilities of design space exploration are improved through graph attention network and federated learning algorithms, achieving better microarchitecture parameter configuration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990013A_ABST
    Figure CN119990013A_ABST
Patent Text Reader

Abstract

The invention discloses a micro-architecture design space exploration method based on deep learning. The method comprises the following steps: acquiring micro-architecture parameters of a target processor, forming a design space, acquiring performance indexes of different design points in the design space, and dividing key modules based on the design points; creating a directed acyclic sub-graph according to the key module, and constructing a prediction model by using a graph attention network to obtain a parameter optimal solution meeting a design target in the directed acyclic sub-graph; splicing the directed acyclic sub-graphs into a global graph to obtain interconnection parameters of key modules, introducing federal learning to train a global prediction model, and predicting performance data of an optimal solution of the evaluation parameters based on the global graph; and when a performance evaluation result meets a preset performance standard, outputting a target configuration parameter of the target processor under each configuration dimension. According to the method, the micro-architecture design space is explored in the graph embedding space, decision support is provided for chip design, and the design efficiency and flexibility are improved by adopting the thought that different design stages and design modes are coordinated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of processor microarchitecture design, and more specifically, to a microarchitecture design space exploration method based on deep learning. Background Art

[0002] In today's era of big data and the Internet of Things, important breakthroughs have been made in the research of algorithm models for data-intensive applications such as neural networks and graph computing. Data-intensive applications often rely on the storage and processing of massive amounts of data, but traditional computer hardware facilities have gradually been unable to meet the computing needs of the above application scenarios. The reason for this is that, on the one hand, Moore's Law has slowed down, the size of transistors is about to reach the physical limit, and the speed of improvement in computing power and energy efficiency is difficult to maintain; on the other hand, no matter how the chip integration is improved, it always takes a huge proportion of energy to move data to the processor, and the potential of hardware computing resources is difficult to fully explore.

[0003] The effect of exploring the microarchitecture design space determines the market competitiveness of the processor, but the exploration process is complex and time-consuming, which poses a great challenge to designers. First, the processor runs many loads during evaluation and it is time-consuming. Second, the microarchitecture design involves many complex circuit modules. The design and evaluation of the register transfer level (circuit) is time-consuming, and its simulation speed is several orders of magnitude slower than that of the physical machine. It takes months to complete the complete evaluation of one load. Faced with the increasingly complex microarchitecture design and the exponentially increasing design space, brute force exploration of the entire space is not possible. Therefore, how to adopt the idea of ​​​​coordinating different design stages and design patterns in microarchitecture design, consider the mutual influence of different stages, and achieve parameter optimization is the focus of attention. Summary of the invention

[0004] In order to solve the above technical problems, the present invention proposes a micro-architecture design space exploration method based on deep learning, which aims to solve the problems that the existing micro-architecture design process is cumbersome and time-consuming, and it is difficult to achieve global performance optimization.

[0005] The present invention provides a micro-architecture design space exploration method based on deep learning, comprising:

[0006] Obtaining microarchitecture parameters of a target processor, forming a design space according to the microarchitecture parameter combination, obtaining performance indicators of different design points in the design space, and clustering the design points to obtain key module partitioning results;

[0007] Create a directed acyclic subgraph based on the design points and corresponding performance indicators in the key modules, obtain the embedded representation of the design points through graph representation learning, and use the graph attention network to build a prediction model to obtain the optimal solution of parameters in the directed acyclic subgraph that meets the design goals;

[0008] Each directed acyclic subgraph is spliced ​​into a global graph to obtain the interconnection parameters between key modules, a federated learning algorithm is introduced to initialize a global prediction model, and the global prediction model is used to predict the optimal solution set of micro-architecture parameters based on the global graph;

[0009] When the performance evaluation result meets the preset performance standard, the target configuration parameters of the target processor in each configuration dimension are output.

[0010] In this solution, the target processor microarchitecture parameters are obtained, a design space is formed according to the microarchitecture parameter combination, and performance indicators of different design points in the design space are obtained, specifically:

[0011] Collecting design goals of a target processor, determining micro-architecture parameters involved based on the design goals, generating different feature vectors according to the micro-architecture parameter combinations, using the feature vectors as design points, and aggregating all design points to form a design space;

[0012] Obtain a large amount of hardware test instance data, and use the micro-architecture parameters corresponding to the design points to filter the corresponding performance results in the hardware test instance data, obtain the corresponding performance indicator labels according to the performance results, and obtain the performance indicator label set corresponding to each design point;

[0013] Constructing a search window, calculating the Pearson correlation coefficient between the design goal and the performance indicator label set corresponding to each design point in the search window, and obtaining the performance indicator label of each design point that meets the correlation standard after traversing the entire design space;

[0014] The performance indicators of different design points in the design space are determined according to the performance indicator labels obtained from each design point.

[0015] In this scheme, the design points are clustered to obtain the key module division results, specifically:

[0016] Obtaining design points with performance indicator labels, clustering them using a clustering algorithm, randomly selecting an initial cluster center, selecting the Mahalanobis distance as a metric function, and calculating a membership matrix from the design point to the initial cluster center according to the metric function;

[0017] Obtain the maximum membership of the design point. When the maximum membership is greater than the preset hard clustering standard, the design point is assigned to the cluster corresponding to the maximum membership. When the maximum membership of the sample does not meet the preset hard clustering standard, determine whether the maximum membership of the design point for each cluster meets the preset soft clustering standard. If multiple clusters meet the standard, the design point is assigned to multiple clusters.

[0018] Iterative clustering is performed by updating the cluster center. When the maximum number of iterations is met, the cluster division corresponding to the last clustering result is selected to generate the key module division result.

[0019] In this solution, a directed acyclic subgraph is created based on the design points and corresponding performance indicators in the key modules, and the embedded representation of the design points is obtained through graph representation learning, specifically:

[0020] Generate a node set according to design points with performance indicator labels in the key module, generate an edge structure set according to the operation associations between the design points, and create a directed acyclic subgraph corresponding to the key module based on the node set and the edge structure set;

[0021] In the directed acyclic subgraph, the initial weight of each node is calculated according to a preset Pareto rank algorithm, and the design point nodes in the directed acyclic subgraph are pre-trained and embedded using meta-path embedding to obtain the meta-path of each design point node;

[0022] In the meta-path, a graph attention mechanism is used to obtain the importance of each meta-path neighbor to the design point node, and the weight of the meta-path neighbor is obtained based on the importance. The meta-path neighbors are weighted averaged based on the initial weight and the weight, and the output of the design point node after the graph attention aggregates the neighbor features to generate an embedded representation of the design point.

[0023] In this solution, a prediction model is constructed using a graph attention network to obtain the optimal solution of parameters in a directed acyclic subgraph that meets the design objectives. Specifically:

[0024] Using the design target to read the corresponding load, running the load through a simulator to simulate the design points in each key module, outputting the performance index evaluation value of each design point under the load, and constructing a training set according to the performance index evaluation value;

[0025] Introducing the idea of ​​global regret into the graph attention network, constructing a prediction model, using the training set for training, and obtaining an embedded representation of the design point based on the training set and the directed acyclic subgraph of the key module;

[0026] The weights of the edges corresponding to the design points are read according to the embedded representation, and the weights are converted into global regrets. The initial optimal path is searched by greedy strategy, and then the initial optimal path is iteratively optimized in combination with the 2-Opt heuristic algorithm to obtain the optimal solution of the parameters that meet the design objectives of each design point.

[0027] In this scheme, each directed acyclic subgraph is spliced ​​into a global graph to obtain the interconnection parameters between key modules, specifically:

[0028] According to the operational associations between key modules, the corresponding directed acyclic subgraphs are spliced ​​to generate a global graph, and an encoder is constructed using a multi-layer graph convolutional neural network to learn and represent the global graph;

[0029] Constructing an adjacency matrix in the global graph to obtain the performance correlation of key modules in the micro-architecture, and using average pooling and maximum pooling to assign different weights to neighbor nodes;

[0030] By aggregating neighbor nodes, the representation of global graph nodes is updated, the updated global graph node representation is imported into the decoder, the RNN network is used to reorganize features, and high-dimensional features are output. The high-dimensional features are used to obtain the interconnection parameters between each key module subgraph, and the edge structure of the global graph is set according to the interconnection parameters.

[0031] In this solution, a federated learning algorithm is introduced to initialize the global prediction model, specifically:

[0032] The prediction model corresponding to each key module is used as a local model, and the error value between the performance indicator prediction value of the local model and the performance indicator evaluation value in the training set is obtained. Federated learning is introduced to retrain the prediction model of each key module. After training, local models are selectively selected and uploaded to the server for model aggregation, and the average error of each local model is calculated.

[0033] The error of the global prediction model after aggregation is obtained as a reference value, the local model is scored according to the average error to generate a confidence level, the local model whose average detection error is greater than the reference value is screened, and the score of the screened local model is set to zero;

[0034] The local models with non-zero scores are uploaded and compressed according to the confidence level, each local model is aggregated, and the parameters of the local model and the global prediction model are updated. The optimal model parameters are obtained through iteration and the global prediction model is output.

[0035] In this solution, the global prediction model is used to predict the optimal solution set of micro-architecture parameters based on the global graph, specifically:

[0036] The optimal solution of the parameters of the design points output by the prediction model corresponding to each key module is mapped to the global graph to guide the training of the global prediction model. By obtaining the global regret of the edges corresponding to the key module nodes, the greedy strategy is used to search for the initial optimal path. The initial optimal path is then iteratively optimized in combination with the 2-Opt heuristic algorithm to obtain the optimal solution of the parameters of each key module node.

[0037] The microarchitecture parameters of the design points are configured according to the optimal solution parameters of each key module node and integrated into an optimal solution set of microarchitecture parameters. The optimal solution set of microarchitecture parameters is placed in a virtual prototype platform to obtain simulation verification data. If the preset design goals are achieved, the target configuration parameters of the target processor in each configuration dimension are output.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] The present invention explores the micro-architecture design space in the graph embedding space, provides decision support for chip design, and adopts the idea of ​​​​coordinating different design stages and design modes to improve design efficiency and flexibility. The message passing mechanism of the graph attention network effectively learns the relevant information between nodes and captures the gap between each edge and the optimal strategy. The weight is obtained according to the attention mechanism to solve the global regret of each edge, and then a better initial solution is obtained through the greedy search strategy, which reduces the number of heuristic iterations and improves the prediction accuracy of the prediction model, thereby improving the overall design space exploration efficiency. And through federated learning, the integrated learning of local models is realized, so that the global prediction model obtains better accuracy and generalization ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments or exemplary embodiments of the present invention, the drawings required for use in the embodiments or exemplary descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained according to the drawings without paying creative work.

[0041] Figure 1 A flowchart of a deep learning-based microarchitecture design space exploration method is shown;

[0042] Figure 2 A flow chart is shown for obtaining the optimal solution of parameters in a directed acyclic subgraph that meets the design objectives;

[0043] Figure 3 A flowchart of introducing a federated learning algorithm to initialize a global prediction model is shown;

[0044] Figure 4 A schematic diagram of a deep learning based microarchitecture design space exploration system is shown. DETAILED DESCRIPTION

[0045] In order to more clearly understand the above-mentioned purpose, features and advantages of the present invention, the present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be noted that the embodiments of the present application and the features in the embodiments can be combined with each other without conflict.

[0046] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited to the specific embodiments disclosed below.

[0047] Figure 1 A flowchart of a deep learning-based microarchitecture design space exploration method is shown.

[0048] like Figure 1 As shown, this embodiment provides a micro-architecture design space exploration method based on deep learning, including:

[0049] S102, obtaining micro-architecture parameters of a target processor, forming a design space according to the micro-architecture parameter combination, obtaining performance indicators of different design points in the design space, and clustering the design points to obtain key module division results;

[0050] S104, creating a directed acyclic subgraph according to the design points and corresponding performance indicators in the key modules, obtaining the embedded representation of the design points through graph representation learning, and using the graph attention network to build a prediction model to obtain the optimal solution of the parameters in the directed acyclic subgraph that meets the design goals;

[0051] S106, splicing the directed acyclic subgraphs into a global graph to obtain interconnection parameters between key modules, introducing a federated learning algorithm to initialize a global prediction model, and using the global prediction model to predict an optimal solution set of micro-architecture parameters based on the global graph;

[0052] S108, when the performance evaluation result meets the preset performance standard, the target configuration parameters of the target processor in each configuration dimension are output.

[0053] It should be noted that the design goal of the target processor is collected, and the micro-architecture parameters involved are determined based on the design goal. For formal representation, different feature vectors are generated according to the combination of the micro-architecture parameters, and the feature vectors are used as design points, and all design points are aggregated to form a design space; a large amount of hardware test instance data is obtained, and the micro-architecture parameters corresponding to the design points are used to filter the corresponding performance results in the hardware test instance data, and the corresponding performance indicator labels are obtained according to the performance results, and the performance indicator label set corresponding to each design point is obtained, including execution time, number of cycles, number of instructions per cycle, energy consumption, power, area and energy consumption delay product, etc. A search window is constructed, and the Pearson correlation coefficient between the design goal and the performance indicator label set corresponding to each design point is calculated in the search window, and the performance indicator labels of each design point that meet the correlation standard are obtained after traversing the entire design space; the performance indicators of different design points in the design space are determined according to the performance indicator labels obtained from each design point.

[0054] It should be noted that the design points with performance index labels are obtained, and clustering is performed using the K-means clustering algorithm, and the initial cluster center is randomly selected. The K-means clustering algorithm is preferably optimized according to the genetic algorithm to set the best initial cluster center. The Mahalanobis distance is selected as the metric function, and the membership matrix from the design point to the initial cluster center is calculated according to the metric function; the maximum membership of the design point is obtained, and when the maximum membership is greater than the preset hard clustering standard, the design point is assigned to the cluster corresponding to the maximum membership. When the maximum membership of the sample does not meet the preset hard clustering standard, it is determined whether the maximum membership of the design point for each cluster meets the preset soft clustering standard. If multiple clusters meet the standard, the design point is assigned to multiple clusters; iterative clustering is performed by updating the cluster center, and when the maximum number of iterations is met, the cluster partition corresponding to the last clustering result is selected to generate the key module partition result. Since the processor prototype data usually contained in the same design point is repeated, the design point is soft clustered and divided into multiple clusters that meet the standard to supplement the hard clustering result and improve the partition accuracy.

[0055] It should be noted that, in the key module, a node set is generated according to the design points with performance indicator labels, an edge structure set is generated according to the operation instruction association between the design points, and a directed acyclic subgraph corresponding to the key module is created based on the node set and the edge structure set; the corresponding load is read using the design target, and the load is run through the simulator to simulate the design points in each key module, and the performance indicator evaluation value of each design point under the load is output.

[0056] In the Pareto algorithm, among all the design points, a group of design points that are not dominated by other design points is called the Pareto optimal solution set; for each design point, its initial non-dominated set is an empty set, and other design points are traversed: if the performance index evaluation value of other design points is higher than the performance index evaluation value of the design point, the other design points are added to the non-dominated set of the design point; if the performance index evaluation value of other design points is lower than the performance index evaluation value of the design point, the number of design points that can dominate the design point is increased by one; after the traversal is completed, if the number of design points that can dominate the design point is zero, the design point is added to the frontier set, and its corresponding Pareto level is set to the first level; each design point in the frontier set is traversed according to the non-dominated set until the number that can be dominated is zero, and the Pareto level of each design point is output; wherein, for each subsequent round of traversal, the Pareto level of the design point in the frontier set of each round is increased by one level based on the previous round. The above steps are used to calculate the initial weight of each node in the directed acyclic subgraph, and the meta-path embedding is used to pre-train and embed the design point nodes in the directed acyclic subgraph to obtain the meta-path of each design point node; preferably, Metapath2vec++ is used to pre-train a high-dimensional representation space to embed the design point nodes in the directed acyclic subgraph.

[0057] In the graph attention network, in each step of message passing, each node receives information from its adjacent nodes, and through some calculations, generates new messages to send to other nodes. Each layer operates on the graphical representation of the data and outputs a new graphical representation. The graph attention mechanism is used in the meta-path to obtain the importance of each meta-path neighbor to the design point node, and the weight of the meta-path neighbor is obtained based on the importance. The meta-path neighbors are weighted averaged based on the initial weight and the weight, and the output of the design point node after the graph attention aggregates the neighbor features to generate an embedded representation of the design point.

[0058] Figure 2 A flow chart for obtaining the optimal solution of parameters in a directed acyclic subgraph that meets the design objectives is shown.

[0059] According to an embodiment of the present invention, a prediction model is constructed using a graph attention network to obtain the optimal solution of parameters that meet the design objectives in a directed acyclic subgraph, specifically:

[0060] S202, reading the corresponding load using the design target, running the load through a simulator to simulate the design points in each key module, outputting the performance index evaluation value of each design point under the load, and constructing a training set according to the performance index evaluation value;

[0061] S204, introducing the idea of ​​global regret into the graph attention network, constructing a prediction model, using the training set for training, and obtaining an embedded representation of the design point according to the training set and the directed acyclic subgraph of the key module;

[0062] S206, reading the weights of the edges corresponding to the design points according to the embedded representation, converting the weights into global regrets, searching for the initial optimal path using a greedy strategy, and then iteratively optimizing the initial optimal path in combination with the 2-Opt heuristic algorithm to obtain the optimal solution of the parameters for each design point that meets the design objectives.

[0063] It should be noted that since the greedy strategy is prone to fall into the local optimum, in order to improve the performance of the heuristic algorithm of the greedy strategy. The common method is to measure the future cost of an action by setting regret, and the action will generate immediate rewards to alleviate the short-sighted characteristics of the greedy combinatorial optimization algorithm. In the directed acyclic subgraph, the regret is extended to the weight of each edge relative to the global optimal solution. The global regret represents the weight of each edge in the directed acyclic subgraph, including the weight obtained by the graph attention mechanism and the initial weight of the node. The gap between each edge and the optimal solution is expressed in the form of weight. The greedy strategy evaluates the pros and cons of each route through the global regret, selects the path with the lowest cost in turn, and finally constructs a better design point meta-path to obtain the optimal solution of the parameters that meet the design objectives of each design point.

[0064] The basic idea of ​​the 2-Opt algorithm is to improve the quality of the solution by adjusting the two edges on the path. First, two points i and k in the path are selected and the path between the two points is flipped. The length of the original path and the flipped path are compared. If the flipped path is better, the flip is accepted and the above process is repeated. Otherwise, the flip is abandoned and a new pair of points is searched. The 2-Opt algorithm has a high search efficiency and can find an approximate solution in a short time.

[0065] According to the operational associations between key modules, the corresponding directed acyclic subgraphs are spliced ​​to generate a global graph, and an encoder is constructed using a multi-layer graph convolutional neural network to learn and represent the global graph; an adjacency matrix is ​​constructed in the global graph to obtain the performance associations of key modules in the microarchitecture, and different weights are assigned to neighbor nodes using a combination of average pooling and maximum pooling; the representation of the global graph nodes is updated through the aggregation of neighbor nodes, and the updated global graph node representation is imported into the decoder, and the RNN network is used for feature reorganization to output high-dimensional features. The high-dimensional features are used to obtain the interconnection parameters between the subgraphs of each key module, and the edge structure of the global graph is set according to the interconnection parameters between the key modules, and the associated modules are used as nodes in the global graph.

[0066] Figure 3 A flowchart for introducing a federated learning algorithm to initialize a global prediction model is shown.

[0067] According to an embodiment of the present invention, a federated learning algorithm is introduced to initialize the global prediction model, specifically:

[0068] S302, taking the prediction model corresponding to each key module as a local model, obtaining the error value between the performance indicator prediction value of the local model and the performance indicator evaluation value in the training set, introducing federated learning to retrain the prediction model of each key module, selectively selecting local models after training and uploading them to the server for model aggregation, and calculating the average error of each local model;

[0069] S304, obtaining the error of the aggregated global prediction model as a reference value, scoring the local models according to the average error to generate confidence, screening the local models whose average detection error is greater than the reference value, and setting the scores of the screened local models to zero;

[0070] S306, upload and compress the local models with non-zero scores according to the confidence level, aggregate each local model, update the parameters of the local model and the global prediction model, obtain the optimal model parameters through iteration, and output the global prediction model.

[0071] It should be noted that the integrated learning of local models is realized through federated learning to obtain superior accuracy and generalization ability. In federated learning, local models are dynamically selected for uploading, and unsatisfactory local models with zero scores are suppressed, which optimizes the aggregation speed of the model. When the parameters of the global prediction model are updated, they are re-divided into unsatisfactory local models for iterative training. After completing multiple rounds of uploading and sending, the model is trained until convergence, and the construction of the global prediction model is completed. The optimal solution of the parameters of the design points output by the prediction model corresponding to each key module is mapped to the global graph to guide the training of the global prediction model. By obtaining the global regret of the corresponding edge of the key module node, the initial optimal path is searched by the greedy strategy, and then the initial optimal path is iteratively optimized in combination with the 2-Opt heuristic algorithm to obtain the optimal solution of the parameters of each key module node; the micro-architecture parameters of the design points are configured according to the optimal solution of the parameters of each key module node, integrated into the optimal solution set of the micro-architecture parameters, and the optimal solution set of the micro-architecture parameters is placed in the virtual prototype platform to obtain simulation verification data, such as using Proteus software on the BOOM processor for verification. If the preset design goal is achieved, the target configuration parameters of the target processor under each configuration dimension are output.

[0072] Figure 4 A schematic diagram of a deep learning based microarchitecture design space exploration system is shown.

[0073] This embodiment provides a micro-architecture design space exploration system 4 based on deep learning, including: a data acquisition module 401, a key module division module 402, a simulator module 403, a local prediction model module 404, a global prediction model module 405 and a simulation verification module 406;

[0074] The data acquisition module is responsible for acquiring the target processor microarchitecture parameters, forming a design space, and acquiring performance indicators of different design points in the design space;

[0075] The key module partitioning module is responsible for clustering the design points to obtain the key module partitioning results, and creating a directed acyclic subgraph according to the design points in the key modules and the corresponding performance indicators;

[0076] The simulator module is responsible for simulating each design point by running the load, outputting the performance index evaluation value of each design point under the load, and constructing a training set;

[0077] The local prediction model module is responsible for obtaining the embedded representation of the design point through graph representation learning, building and training the prediction model using the graph attention network and the training set, and obtaining the optimal solution of the parameters that meet the design objectives;

[0078] The global prediction model module is responsible for splicing each directed acyclic subgraph into a global graph to obtain the interconnection parameters between key modules, introducing a federated learning algorithm to initialize the global prediction model, and using the global prediction model to predict the optimal solution set of micro-architecture parameters based on the global graph;

[0079] The simulation verification module is responsible for verifying the optimal solution set of micro-architecture parameters on the BOOM processor using Proteus software. When the performance evaluation result meets the preset performance standard, the target configuration parameters of the target processor in each configuration dimension are output according to the optimal solution set of micro-architecture parameters.

[0080] This embodiment provides a computer-readable storage medium, which includes a micro-architecture design space exploration method program based on deep learning. When the micro-architecture design space exploration method program based on deep learning is executed by a processor, the steps of the micro-architecture design space exploration method based on deep learning are implemented.

[0081] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.

[0082] In addition, all functional units in the embodiments of the present invention may be integrated into one processing unit, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated units may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0083] Those skilled in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions, and the aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps of the above method embodiments; and the aforementioned storage medium includes: mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), disks or optical disks, and other media that can store program codes.

[0084] Alternatively, if the above-mentioned integrated unit of the present invention is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention can be essentially or partly reflected in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as mobile storage devices, ROM, RAM, magnetic disks or optical disks.

[0085] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.

Claims

1. A microarchitecture design space exploration method based on deep learning, characterized in that: The following steps are involved: Obtaining microarchitecture parameters of a target processor, forming a design space according to a combination of the microarchitecture parameters, obtaining performance indicators of different design points in the design space, and clustering the design points to obtain key module partitioning results; Create a directed acyclic subgraph based on the design points and corresponding performance indicators in the key modules, obtain the embedded representation of the design points through graph representation learning, and use the graph attention network to build a prediction model to obtain the optimal solution of parameters in the directed acyclic subgraph that meets the design goals; Each directed acyclic subgraph is spliced ​​into a global graph to obtain the interconnection parameters between key modules, a federated learning algorithm is introduced to initialize a global prediction model, and the global prediction model is used to predict the optimal solution set of micro-architecture parameters based on the global graph; When the performance evaluation result meets the preset performance standard, the target configuration parameters of the target processor in each configuration dimension are output.

2. The method for exploring microarchitecture design space based on deep learning according to claim 1, characterized in that: Obtain the target processor microarchitecture parameters, form a design space according to the microarchitecture parameter combination, and obtain performance indicators of different design points in the design space, specifically: Collecting design goals of a target processor, determining micro-architecture parameters involved based on the design goals, generating different feature vectors according to the micro-architecture parameter combinations, using the feature vectors as design points, and aggregating all design points to form a design space; Obtain a large amount of hardware test instance data, and use the micro-architecture parameters corresponding to the design points to filter the corresponding performance results in the hardware test instance data, obtain the corresponding performance indicator labels according to the performance results, and obtain the performance indicator label set corresponding to each design point; Constructing a search window, calculating the Pearson correlation coefficient between the design goal and the performance indicator label set corresponding to each design point in the search window, and obtaining the performance indicator label of each design point that meets the correlation standard after traversing the entire design space; The performance indicators of different design points in the design space are determined according to the performance indicator labels obtained from each design point.

3. The microarchitecture design space exploration method based on deep learning according to claim 1, characterized in that: Cluster the design points to obtain the key module division results, specifically: Obtaining design points with performance indicator labels, clustering them using a clustering algorithm, randomly selecting an initial cluster center, selecting the Mahalanobis distance as a metric function, and calculating a membership matrix from the design point to the initial cluster center according to the metric function; Obtain the maximum membership of the design point. When the maximum membership is greater than the preset hard clustering standard, the design point is assigned to the cluster corresponding to the maximum membership. When the maximum membership of the sample does not meet the preset hard clustering standard, determine whether the maximum membership of the design point for each cluster meets the preset soft clustering standard. If multiple clusters meet the standard, the design point is assigned to multiple clusters. Iterative clustering is performed by updating the cluster center. When the maximum number of iterations is met, the cluster division corresponding to the last clustering result is selected to generate the key module division result.

4. The method for exploring microarchitecture design space based on deep learning according to claim 1, characterized in that: Create a directed acyclic subgraph based on the design points and corresponding performance indicators in the key modules, and obtain the embedded representation of the design points through graph representation learning, specifically: Generate a node set according to design points with performance indicator labels in the key module, generate an edge structure set according to the operation associations between the design points, and create a directed acyclic subgraph corresponding to the key module based on the node set and the edge structure set; In the directed acyclic subgraph, the initial weight of each node is calculated according to a preset Pareto rank algorithm, and the design point nodes in the directed acyclic subgraph are pre-trained and embedded using meta-path embedding to obtain the meta-path of each design point node; In the meta-path, a graph attention mechanism is used to obtain the importance of each meta-path neighbor to the design point node, and the weight of the meta-path neighbor is obtained based on the importance. The meta-path neighbors are weighted averaged based on the initial weight and the weight, and the output of the design point node after the graph attention aggregates the neighbor features to generate an embedded representation of the design point.

5. The method for exploring microarchitecture design space based on deep learning according to claim 1, characterized in that: The prediction model is constructed using the graph attention network to obtain the optimal solution of parameters in the directed acyclic subgraph that meets the design goals. Specifically: Using the design target to read the corresponding load, running the load through a simulator to simulate the design points in each key module, outputting the performance index evaluation value of each design point under the load, and constructing a training set according to the performance index evaluation value; Introducing the idea of ​​global regret into the graph attention network, constructing a prediction model, using the training set for training, and obtaining an embedded representation of the design point based on the training set and the directed acyclic subgraph of the key module; The weights of the edges corresponding to the design points are read according to the embedded representation, and the weights are converted into global regrets. The initial optimal path is searched by greedy strategy, and then the initial optimal path is iteratively optimized in combination with the 2-Opt heuristic algorithm to obtain the optimal solution of the parameters that meet the design objectives of each design point.

6. The method for exploring microarchitecture design space based on deep learning according to claim 1, characterized in that: Each directed acyclic subgraph is spliced ​​into a global graph to obtain the interconnection parameters between key modules, specifically: According to the operational associations between key modules, the corresponding directed acyclic subgraphs are spliced ​​to generate a global graph, and an encoder is constructed using a multi-layer graph convolutional neural network to learn and represent the global graph; Constructing an adjacency matrix in the global graph to obtain the performance correlation of key modules in the micro-architecture, and using average pooling and maximum pooling to assign different weights to neighbor nodes; By aggregating neighbor nodes, the representation of global graph nodes is updated, the updated global graph node representation is imported into the decoder, the RNN network is used to reorganize features, and high-dimensional features are output. The high-dimensional features are used to obtain the interconnection parameters between each key module subgraph, and the edge structure of the global graph is set according to the interconnection parameters.

7. The method for exploring microarchitecture design space based on deep learning according to claim 1, characterized in that: The federated learning algorithm is introduced to initialize the global prediction model, specifically: The prediction model corresponding to each key module is used as a local model, and the error value between the performance indicator prediction value of the local model and the performance indicator evaluation value in the training set is obtained. Federated learning is introduced to retrain the prediction model of each key module. After training, local models are selectively selected and uploaded to the server for model aggregation, and the average error of each local model is calculated. The error of the global prediction model after aggregation is obtained as a reference value, the local model is scored according to the average error to generate a confidence level, the local model whose average detection error is greater than the reference value is screened, and the score of the screened local model is set to zero; The local models with non-zero scores are uploaded and compressed according to the confidence level, each local model is aggregated, and the parameters of the local model and the global prediction model are updated. The optimal model parameters are obtained through iteration and the global prediction model is output.

8. The method for exploring microarchitecture design space based on deep learning according to claim 1, characterized in that: The global prediction model is used to predict the optimal solution set of micro-architecture parameters based on the global graph, specifically: The optimal solution of the parameters of the design points output by the prediction model corresponding to each key module is mapped to the global graph to guide the training of the global prediction model. By obtaining the global regret of the edges corresponding to the key module nodes, the greedy strategy is used to search for the initial optimal path. The initial optimal path is then iteratively optimized in combination with the 2-Opt heuristic algorithm to obtain the optimal solution of the parameters of each key module node. The microarchitecture parameters of the design points are configured according to the optimal solution parameters of each key module node and integrated into an optimal solution set of microarchitecture parameters. The optimal solution set of microarchitecture parameters is placed in a virtual prototype platform to obtain simulation verification data. If the preset design goals are achieved, the target configuration parameters of the target processor in each configuration dimension are output.

Citation Information

Patent Citations

  • Lithium ion battery capacity estimation method based on graph neural network

    CN115856633A

  • Collaborative filtering recommendation method based on micrograph neural network architecture search

    CN117009661A

  • Memory Grounded Conversational Reasoning and Question Answering for Assistant Systems

    US20200410012A1

Cited By

  • Rotor type design method and system based on improved Bayesian optimization algorithm

    CN122087966A