Cooperative exploration method and device for mapping scheme and architecture of core particle accelerator

Through the coordinated exploration method of the mapping scheme of the core accelerator and the architecture, the optimal layer-pipeline space mapping scheme is generated, which solves the high packaging cost and high power consumption problems faced by the core accelerator in large-scale DNN applications, and achieves better performance and higher energy efficiency.

CN120216447APending Publication Date: 2025-06-27TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311811114.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In large-scale deep neural network (DNN) applications, core-grain accelerators face problems such as high packaging costs, high power consumption and low bandwidth, making it difficult to maximize the advantages of core-grain technology and minimize its disadvantages.

Method used

A method of collaborative exploration of the mapping scheme and architecture of the core particle accelerator is proposed. By obtaining the candidate values ​​of configurable architectural parameters, framework settings information and neural network model, a candidate collaboration scheme of the layer-pipeline space mapping scheme and architecture is generated, and its monetary cost, power consumption cost and delay are determined to achieve the optimal collaboration scheme.

Benefits of technology

Through collaborative exploration methods, it is possible to optimize monetary costs while considering power consumption costs and performance delays, and obtain better performance and higher energy efficiency, solving the high packaging costs and high power consumption problems faced by core accelerators in large-scale DNN applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216447A_ABST
    Figure CN120216447A_ABST
Patent Text Reader

Abstract

The invention relates to a core particle accelerator mapping scheme and architecture collaborative exploration method and device. The method comprises the following steps: acquiring candidate values of configurable architecture parameters of the core particle accelerator, framework setting information and a neural network model; according to the candidate values of the configurable architecture parameters, the framework setting information and the neural network model, generating a layer-pipeline space mapping scheme of the core particle accelerator and a candidate collaborative scheme of the architecture; the currency cost, the power consumption cost and the delay of the candidate cooperation scheme are determined; and determining an evaluation value of the candidate cooperation scheme according to the currency cost, the power consumption cost and the delay of the candidate cooperation scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method for collaborative exploration of a mapping scheme and an architecture of a chiplet accelerator, a device for collaborative exploration of a mapping scheme and an architecture of a chiplet accelerator, an electronic device, and a storage medium. Background Art

[0002] In the post-Moore era, chiplet technology can integrate more and more transistors on a single accelerator with a higher yield (i.e., qualification rate), so as to meet the huge computing demands brought about by the rapid development of artificial intelligence. However, chiplet technology also brings higher packaging costs and costly D2D (Die-to-Die) interfaces. Compared with on-chip interconnections, D2D interfaces require a larger area, consume higher power, and provide lower bandwidth. Maximizing the advantages of chiplet technology and minimizing its disadvantages are crucial for developing chiplet accelerators for large-scale deep neural networks (DNNs), which pose challenges to both mapping schemes and architectures. Summary of the Invention

[0003] The present disclosure provides a technical solution for collaborative exploration of a mapping scheme and an architecture of a chiplet accelerator.

[0004] According to one aspect of the present disclosure, there is provided a method for collaborative exploration of a mapping scheme and an architecture of a chiplet accelerator, including:

[0005] Obtaining candidate values of configurable architecture parameters of a chiplet accelerator, framework setting information, and a neural network model;

[0006] Generating a candidate collaborative scheme of a layer-pipeline space mapping scheme and an architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model;

[0007] Determining the monetary cost, power consumption cost, and latency of the candidate collaborative scheme;

[0008] Determining an evaluation value of the candidate collaborative scheme according to the monetary cost, power consumption cost, and latency of the candidate collaborative scheme.

[0009] In a possible implementation, the configurable architecture parameters include at least some of the following: the bandwidth of the on-chip network, the Die-to-Die communication bandwidth, the total bandwidth of the memory, the total number of computing cores in the X direction of the chiplet accelerator, the total number of computing cores in the Y direction of the chiplet accelerator, the number of chiplets divided in the X direction of the chiplet accelerator, the number of chiplets divided in the Y direction of the chiplet accelerator, the number of multiply-accumulate operations in the processing unit array of the computing core, and the size of the global buffer of the computing core.

[0010] In a possible implementation,

[0011] obtain a neural network model, including: obtaining a plurality of neural network models;

[0012] Generating a candidate cooperation plan for the layer-pipeline space mapping scheme and the architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model includes: generating a candidate cooperation plan for the layer-pipeline space mapping scheme and the architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and any one of the plurality of neural network models.

[0013] In a possible implementation, the layer-pipeline space mapping scheme in any candidate cooperation plan includes the mapping schemes of each network layer in the neural network model. The mapping scheme of any network layer includes a partitioning attribute, a core group attribute, and a data flow attribute. Among them, the partitioning attribute is used to partition the input feature map, weight, and output feature map of the network layer. The core group attribute represents information about the computing cores used to calculate the network layer. The data flow attribute represents the data source of the input feature map of the network layer, the weight, and the destination of the output feature map of the network layer.

[0014] In a possible implementation, the partitioning attribute in the mapping scheme of any network layer includes the number of partitions in the height dimension of the output feature map, the number of partitions in the width dimension of the output feature map, the number of partitions in the batch size dimension, and the number of partitions in the channel number dimension of the output feature map.

[0015] In a possible implementation, generating the candidate cooperation plan for the layer-pipeline space mapping scheme and the architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model includes:

[0016] Generating a candidate cooperation plan for the layer-pipeline space mapping scheme and the architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model, and using a preset simulated annealing operator;

[0017] Among them, the preset simulated annealing operator includes at least one of the following:

[0018] The first simulated annealing operator is used to randomly select a network layer and change the partitioning attribute in the mapping scheme of the network layer;

[0019] The second simulated annealing operator is used to randomly select a network layer and randomly exchange two computing cores in the core group attribute in the mapping scheme of the network layer;

[0020] The third simulated annealing operator is used to randomly select two network layers and randomly swap two computing cores in the kernel group attributes of the mapping schemes of the two network layers;

[0021] The fourth simulated annealing operator is used to randomly select two network layers, remove a computing core from the kernel group attributes of one of the two network layers in the mapping scheme, and add the removed computing core to the kernel group attributes of the other network layer in the mapping scheme of the two network layers;

[0022] The fifth simulated annealing operator is used to randomly select a network layer, randomly select a non - negative item in the data flow attributes of the mapping scheme of the network layer, and randomly determine the updated value of the non - negative item within the value range [0, D].

[0023] In a possible implementation manner, determining the power consumption cost of the candidate cooperation scheme includes:

[0024] Determining the power consumption cost of the candidate cooperation scheme according to the number of operations of each component in the architecture in the candidate cooperation scheme and the corresponding unit power consumption.

[0025] In a possible implementation manner, determining the latency of the candidate cooperation scheme includes:

[0026] Determining the computation time of the multiply - accumulate operations in the candidate cooperation scheme;

[0027] Determining the ratio of the maximum value of the data access volume of the memory in the candidate cooperation scheme to the access bandwidth;

[0028] Determining the latency of the candidate cooperation scheme according to the computation time and the ratio.

[0029] In a possible implementation manner, determining the monetary cost of the candidate cooperation scheme includes:

[0030] Determining the monetary cost of the candidate cooperation scheme according to the values of the configurable architecture parameters in the candidate cooperation scheme.

[0031] In a possible implementation manner, the monetary cost of the candidate cooperation scheme includes at least one of the following:

[0032] The die manufacturing cost of the candidate cooperation scheme, the memory cost of the candidate cooperation scheme, the packaging cost of the candidate cooperation scheme.

[0033] In a possible implementation manner, the determining the monetary cost of the candidate cooperation scheme according to the values of the configurable architecture parameters in the candidate cooperation scheme includes:

[0034] For any die in the architecture of the candidate collaboration solution, determine the silicon area of the die and the yield of the die, and obtain the manufacturing cost per unit area of the process technology node corresponding to the die;

[0035] Determine the die manufacturing cost of the die according to the silicon area of the die, the yield of the die, and the manufacturing cost per unit area of the process technology node corresponding to the die.

[0036] In a possible implementation, determining the yield of the die includes:

[0037] Determine the yield of the die according to the silicon area, defect density, and clustering parameter of the die.

[0038] In a possible implementation, the determining the monetary cost of the candidate collaboration solution according to the value of the configurable architecture parameter in the candidate collaboration solution includes:

[0039] Determine the memory cost of the candidate collaboration solution according to the bandwidth of the memory, the bandwidth of the memory unit, and the monetary cost of the memory unit in the architecture of the candidate collaboration solution.

[0040] In a possible implementation, the determining the monetary cost of the candidate collaboration solution according to the value of the configurable architecture parameter in the candidate collaboration solution includes:

[0041] Determine the packaging cost of the candidate collaboration solution according to the total silicon area of all dies in the architecture of the candidate collaboration solution, the scaling factor, the packaging yield, and the monetary cost per unit area of the substrate.

[0042] In a possible implementation, the determining the evaluation value of the candidate collaboration solution according to the monetary cost, power consumption cost, and latency of the candidate collaboration solution includes:

[0043] Obtain the first weight corresponding to the monetary cost, the second weight corresponding to the power consumption cost, and the third weight corresponding to the latency;

[0044] Determine the evaluation value of the candidate collaboration solution according to the monetary cost, power consumption cost, and latency of the candidate collaboration solution, and the first weight, the second weight, and the third weight.

[0045] In a possible implementation, after determining the evaluation value of the candidate collaboration solution, the method further includes:

[0046] Determine the optimal collaboration solution of the die accelerator according to the evaluation values of multiple candidate collaboration solutions.

[0047] In a possible implementation, the die accelerator is an inference accelerator for a deep neural network.

[0048] According to one aspect of the present disclosure, there is provided a co - exploration device for a mapping scheme and architecture of a die accelerator, including:

[0049] An acquisition module, configured to acquire candidate values of configurable architecture parameters of the die accelerator, framework setting information, and a neural network model;

[0050] A generation module, configured to generate a candidate co - scheme of the layer - pipeline space mapping scheme and architecture of the die accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model;

[0051] A first determination module, configured to determine the monetary cost, power consumption cost, and latency of the candidate co - scheme;

[0052] A second determination module, configured to determine an evaluation value of the candidate co - scheme according to the monetary cost, power consumption cost, and latency of the candidate co - scheme.

[0053] In a possible implementation, the configurable architecture parameters include at least some of the following: the bandwidth of the on - chip network, the die - to - die communication bandwidth, the total bandwidth of the memory, the total number of computing cores in the X direction of the die accelerator, the total number of computing cores in the Y direction of the die accelerator, the number of die divided in the X direction of the die accelerator, the number of die divided in the Y direction of the die accelerator, the number of multiply - accumulate operations in the processing unit array of the computing core, and the size of the global buffer of the computing core.

[0054] In a possible implementation,

[0055] The acquisition module is configured to: acquire a plurality of neural network models;

[0056] The generation module is configured to: generate a candidate co - scheme of the layer - pipeline space mapping scheme and architecture of the die accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and any one of the plurality of neural network models.

[0057] In a possible implementation, the layer - pipeline space mapping scheme in any candidate co - scheme includes mapping schemes of each network layer in the neural network model. The mapping scheme of any network layer includes a partitioning attribute, a kernel group attribute, and a data flow attribute. Among them, the partitioning attribute is used to partition the input feature map, weight, and output feature map of the network layer. The kernel group attribute represents information about the computing cores used to calculate the network layer. The data flow attribute represents the data source of the input feature map of the network layer, the weight, and the destination of the output feature map of the network layer.

[0058] In a possible implementation, the partitioning attributes in the mapping scheme of any network layer include the number of partitions in the height dimension of the output feature map, the number of partitions in the width dimension of the output feature map, the number of partitions in the batch size dimension, and the number of partitions in the channel number dimension of the output feature map.

[0059] In a possible implementation, the generating module is used for:

[0060] Generating a candidate cooperation scheme between the layer-pipeline space mapping scheme of the die accelerator and the architecture according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model, and using a preset simulated annealing operator;

[0061] Wherein, the preset simulated annealing operator includes at least one of the following:

[0062] The first simulated annealing operator is used to randomly select a network layer and change the partitioning attributes in the mapping scheme of the network layer;

[0063] The second simulated annealing operator is used to randomly select a network layer and randomly swap two computing cores in the kernel group attributes in the mapping scheme of the network layer;

[0064] The third simulated annealing operator is used to randomly select two network layers and randomly swap two computing cores in the kernel group attributes in the mapping schemes of the two network layers;

[0065] The fourth simulated annealing operator is used to randomly select two network layers, remove a computing core from the kernel group attributes in the mapping scheme of one of the two network layers, and add the removed computing core to the kernel group attributes in the mapping scheme of the other network layer of the two network layers;

[0066] The fifth simulated annealing operator is used to randomly select a network layer, randomly select a non-negative term in the data flow attributes in the mapping scheme of the network layer, and randomly determine the updated value of the non-negative term within the value range [0, D].

[0067] In a possible implementation, the first determining module is used for:

[0068] Determining the power consumption cost of the candidate cooperation scheme according to the number of operations and the corresponding unit power consumption of each component in the architecture in the candidate cooperation scheme.

[0069] In a possible implementation, the first determining module is used for:

[0070] Determining the computing time of the multiply-accumulate operation in the candidate cooperation scheme;

[0071] Determine the ratio of the maximum data access volume of the memory in the candidate collaboration plan to the access bandwidth;

[0072] Determine the latency of the candidate collaboration plan according to the calculation time and the ratio.

[0073] In a possible implementation manner, the first determination module is configured to:

[0074] Determine the monetary cost of the candidate collaboration plan according to the value of the configurable architecture parameter in the candidate collaboration plan.

[0075] In a possible implementation manner, the monetary cost of the candidate collaboration plan includes at least one of the following:

[0076] The die manufacturing cost of the candidate collaboration plan, the memory cost of the candidate collaboration plan, and the packaging cost of the candidate collaboration plan.

[0077] In a possible implementation manner, the first determination module is configured to:

[0078] For any die in the architecture of the candidate collaboration plan, determine the silicon area and the yield of the die, and obtain the unit area manufacturing cost of the process technology node corresponding to the die;

[0079] Determine the die manufacturing cost of the die according to the silicon area of the die, the yield of the die, and the unit area manufacturing cost of the process technology node corresponding to the die.

[0080] In a possible implementation manner, the first determination module is configured to:

[0081] Determine the yield of the die according to the silicon area, the defect density, and the clustering parameter of the die.

[0082] In a possible implementation manner, the first determination module is configured to:

[0083] Determine the memory cost of the candidate collaboration plan according to the bandwidth of the memory, the bandwidth of the memory unit, and the monetary cost of the memory unit in the architecture of the candidate collaboration plan.

[0084] In a possible implementation manner, the first determination module is configured to:

[0085] Determine the packaging cost of the candidate collaboration plan according to the total silicon area of all dies in the architecture of the candidate collaboration plan, the scaling factor, the packaging yield, and the unit area monetary cost of the substrate.

[0086] In a possible implementation manner, the second determination module is configured to:

[0087] Obtain a first weight corresponding to the currency cost, a second weight corresponding to the power consumption cost, and a third weight corresponding to the latency.

[0088] Determine an evaluation value of the candidate cooperation scheme according to the currency cost, the power consumption cost, and the latency of the candidate cooperation scheme, and the first weight, the second weight, and the third weight.

[0089] In a possible implementation, the device further includes:

[0090] A third determination module, configured to determine an optimal cooperation scheme of the chiplet accelerator according to the evaluation values of multiple candidate cooperation schemes.

[0091] In a possible implementation, the chiplet accelerator is an inference accelerator for a deep neural network.

[0092] According to one aspect of the present disclosure, there is provided an electronic device, including: one or more processors; a memory for storing executable instructions; wherein, the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.

[0093] According to one aspect of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented.

[0094] According to one aspect of the present disclosure, there is provided a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, and when the computer-readable code runs in an electronic device, the processor in the electronic device executes the above method.

[0095] In the embodiments of the present disclosure, by obtaining candidate values of configurable architecture parameters, framework setting information, and a neural network model of a chiplet accelerator, generating a candidate cooperation scheme of a layer-pipeline space mapping scheme and an architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model, determining the currency cost, the power consumption cost, and the latency of the candidate cooperation scheme, and determining the evaluation value of the candidate cooperation scheme according to the currency cost, the power consumption cost, and the latency of the candidate cooperation scheme, co-exploration of the layer-pipeline space mapping scheme and the architecture of the chiplet accelerator is realized, which not only considers the power consumption cost and performance (latency), but also considers the currency cost. The cooperation scheme obtained by exploration using the embodiments of the present disclosure can achieve better performance and higher energy efficiency.

[0096] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure.

[0097] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] The accompanying drawings herein are incorporated into and constitute a part of this specification, and these drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.

[0099] Figure 1a A schematic diagram showing a single-chip is shown.

[0100] Figure 1b A schematic diagram showing a chip including two dielets is shown.

[0101] Figure 2 A schematic diagram showing a hardware template of a dielet accelerator provided by an embodiment of the present disclosure is shown.

[0102] Figure 3 A schematic diagram showing a layer-pipeline space mapping scheme for parsing encoded in an optimization space in an embodiment of the present disclosure is shown.

[0103] Figure 4 A flowchart showing a co-exploration method for a mapping scheme and an architecture of a dielet accelerator provided by an embodiment of the present disclosure is shown.

[0104] Figure 5a A schematic diagram showing a co-exploration framework for a mapping scheme and an architecture of a dielet accelerator provided by an embodiment of the present disclosure is shown.

[0105] Figure 5b A schematic diagram showing a mapping engine in a co-exploration framework for a mapping scheme and an architecture of a dielet accelerator provided by an embodiment of the present disclosure is shown.

[0106] Figure 6 A block diagram showing a co-exploration device for a mapping scheme and an architecture of a dielet accelerator provided by an embodiment of the present disclosure is shown.

[0107] Figure 7 A block diagram showing an electronic device 1900 provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0108] The various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Like reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0109] As used herein, the term "exemplary" means "serving as an example, instance, or illustration". Any embodiment illustrated as "exemplary" herein is not necessarily to be construed as superior or better than other embodiments.

[0110] As used herein, the term "and / or" is merely a description of an associated relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the term "at least one" as used herein means any one of a plurality or any combination of at least two of a plurality. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set composed of A, B, and C.

[0111] In addition, to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0112] As deep neural networks (DNNs) solve increasingly complex problems, their scale and complexity have grown rapidly, leading to increased computing and storage requirements. Although related technologies have obtained large-scale single-chip accelerators with tens of billions of transistors by applying more advanced technologies and increasing the single-chip size, the end of Moore's Law and limited photomask sizes pose significant challenges to further integration of transistors.

[0113] Chiplet technology uses advanced packaging technologies to combine small functional chips, providing a solution to overcome the above limitations and achieve continuous transistor integration. Chiplet-based DNN inference accelerators, such as DNN inference accelerators with 36 chiplets, have emerged.

[0114] Figure 1a A schematic diagram of a single-chip is shown. Figure 1b A schematic diagram of a chip including 2 chiplets is shown.

[0115] There are the following four advantages in introducing chiplet technology. First, it improves the overall yield of the chip. By dividing a large-scale single-chip into smaller chiplets, the overall yield of the chip can be significantly improved. For example, at the 7nm technology node, the yields of 800mm 2 chips are approximately 18% respectively, and 200mm 2The yield of the chip is approximately 75%. Secondly, the chiplet technology can expand the area of the chip because, compared with a monolithic chip, an advanced substrate (such as an organic substrate or a silicon interposer) can achieve a larger area (for example, the size of the photomask can be 858mm 2 ). Thirdly, the chiplet technology enables heterogeneous integration. Different from logic circuits, analog circuit IP (such as designs or technologies related to analog circuits) has not significantly benefited from the performance and density improvements brought by technological progress. Therefore, manufacturing logic circuits using advanced technologies while producing analog IP with various IO (Input / Output) functions using older processes can save the expensive manufacturing, design, and intellectual property-related costs associated with advanced technologies. Fourthly, the chiplet technology can reuse individual chiplets to develop multiple computing chips, and the scale or target application of each computing chip can be different, which is a significant advantage of the chiplet technology. This method can greatly reduce the huge Non-Recurring Engineering (NRE) costs and the time traditionally required to develop different chips for each scale and scenario.

[0116] Introducing the chiplet technology also includes four disadvantages, all of which are caused by the D2D interface. First, the D2D interface increases energy consumption. Compared with the data transmission cost of less than 0.1 pJ / bit for on-chip lines, the energy consumed by the D2D link is several to dozens of times higher. Secondly, the relatively small inter-chip communication bandwidth may reduce performance. Compared with abundant on-chip interconnection resources, the inter-chip bandwidth is smaller because of the limited number of available IO pins around each chiplet. Thirdly, the D2D interface requires more area. Different from on-chip lines that hardly occupy silicon area, the D2D interface requires a specific analog physical layer and controller. Fourthly, the large number of interconnection requirements between chiplets increases the packaging cost of the substrate. Compared with a monolithic chip that only requires a basic fan-out substrate, advanced packaging solutions require an organic substrate or a silicon interposer with dozens of layers.

[0117] The chiplet technology brings new challenges in architecture design and DNN space mapping.

[0118] For architecture design, the main challenge is to determine the optimal chiplet granularity. Although the chiplet technology can achieve a larger accelerator scale and a higher yield, it also brings higher packaging costs and Die-to-Die (D2D) interconnection costs. Compared with on-chip interconnections, the D2D interface requires a larger area, consumes higher power, and provides lower bandwidth. Using more smaller chiplets can improve the yield but increase the cost of the chiplet accelerator; using fewer larger chiplets can reduce the cost of the chiplet accelerator but lower the yield. How to make a trade-off remains an unsolved problem.

[0119] For DNN mapping, the main challenge comes from the fact that chiplet technology brings greater scale and higher D2D interconnect costs.

[0120] As the scale of chiplet accelerators continues to increase, it becomes increasingly difficult to maintain high utilization and energy efficiency. Layer-sequential (LS) mapping exhibits limited scalability, while Layer-Pipeline mapping shows higher potential. To address the challenges brought about by the increasing scale of chiplet accelerators, the academic and industrial communities widely adopt Layer-Pipeline mapping in large-scale accelerators, that is, mapping multiple layers onto the accelerator in space. The core of Layer-Pipeline mapping lies in Spatial Mapping (SPM), which determines which part of which layer is assigned to which computing core, and this has a significant impact on the performance and energy efficiency of large-scale accelerators. For example, using the same core to execute different parts of different layers may have different utilization and energy efficiency. In addition, allocating the same workload to different computing cores may affect the performance and energy efficiency of network communication. Although Layer-Pipeline spatial mapping is important, most current strategies are still heuristic. In related technologies, the problems and optimization space of Layer-Pipeline spatial mapping have not been clearly defined, fully explored, or completely understood, which leads to the inability to fully optimize Layer-Pipeline spatial mapping. And as the scale of accelerators and the complexity of DNN architectures increase, this becomes increasingly important.

[0121] In addition, compared with on-chip wires, D2D links tend to consume more energy and provide lower bandwidth. Therefore, designing a spatial mapping strategy that can automatically reduce D2D communication costs is crucial for enhancing the performance and efficiency of chiplet accelerators and fully leveraging the advantages provided by chiplet technology.

[0122] The above analysis shows that the use of chiplet technology brings complex trade-off problems and challenges. Therefore, maximizing the advantages of chiplet technology and minimizing its disadvantages are crucial for developing chiplet accelerators such as DNNs. This goal not only poses challenges to architecture design but also requires more efficient mapping schemes to effectively utilize larger-scale chiplet accelerators while reducing the expensive communication costs between chiplets.

[0123] To solve the technical problems similar to those described above, the embodiments of the present disclosure provide a configurable hardware template for a chiplet accelerator, a Layer-Pipeline spatial mapping encoding method, a method for co-exploring a mapping scheme and an architecture of a chiplet accelerator, a co-exploration framework for a mapping scheme and an architecture of a chiplet accelerator, a co-exploration device for a mapping scheme and an architecture of a chiplet accelerator, an electronic device, and a storage medium.

[0124] Figure 2 FIG. shows a schematic diagram of a hardware template of a chiplet accelerator provided by an embodiment of the present disclosure. The hardware template is general and highly configurable. This highly configurable hardware template is conducive to the subsequent co-exploration of the mapping scheme and architecture of the chiplet accelerator.

[0125] As Figure 2 shown, the hardware template of the chiplet accelerator proposed by the embodiment of the present disclosure may include two different types of chiplets: IO chiplets and computing chiplets. In Figure 2 the example shown, a mesh NoC (Network-on-Chip) interconnects the computing cores in all computing chiplets and the controllers in the IO chiplets, thus allowing communication between any computing core and computing core, computing core and DRAM, and DRAM and computing core. For inter-chiplet communication, the D2D transmitter (TX) within the chiplet independently encodes the data and forwards the encoded data to the corresponding D2D receiver (RX) in another chiplet. The D2D receiver decodes the data and continues to transmit it using the NoC. This inter-chiplet communication is fully automatic and transparent to both the source and destination. This mesh NoC can improve the scalability of the hardware template, so that it can be applied to large-scale chiplet accelerators.

[0126] In addition, this heterogeneous architecture allows any number of computing chiplets, thus further improving the scalability of the entire hardware template. Otherwise, if we separately equip each computing chiplet with an IO-related physical layer (PHY) and a controller, it will occupy the edge area of the chip and the IO pins, which may affect the routing of the D2D link and thus affect the scalability of the system.

[0127] It should be noted that in addition to adopting a mesh topology, the hardware template of the chiplet accelerator in the embodiment of the present disclosure may also adopt other topological structures, which are not limited herein.

[0128] Regarding the architecture of the computing chiplet:

[0129] In the embodiment of the present disclosure, each computing chiplet may include any number of computing cores interconnected by a mesh NoC. To enhance the scalability of the hardware template of the chiplet accelerator, the D2D interfaces can be placed around the chiplet, and the number of D2D interfaces can be equal to the number of computing cores on each side. This arrangement enables the computing chiplet to form a larger-scale network with other chiplets. In Figure 2 the example shown, there are 4 cores on each side, so we can place four D2D interfaces on each side of the computing chiplet.

[0130] The computing core is the key component responsible for performing calculations in the entire chiplet accelerator, and its architecture can be as Figure 2As shown in (b). The communication unit of the computing core may include a DMA (Direct Memory Access) module and a NoC router to communicate with other computing cores and DRAM. The control unit of the computing core is mainly responsible for managing the computing tasks and task progress information based on static compilation instructions, as well as managing the reception and transmission of data or messages between other computing cores or DRAM. The global buffer (GLB) of each computing core is globally visible throughout the die accelerator. Each computing core can read data from the global buffer of other computing cores or write data to the global buffer of other computing cores on the premise that the data is valid or the address is writable. The Processing Element array (PE array) can be responsible for computing general matrix multiplication and / or convolution, and the vector unit can be responsible for vector and / or scalar operations. In each calculation, the Processing Element array reads the weights and input feature maps (Ifmap) of the workload tile from the global buffer. The output feature map (Ofmap) and / or partial sum values can be directly written back to the global buffer or post-processed (such as batch normalization and ReLU operations) within the vector unit. At the same time, the vector unit can be independently invoked for vector and / or scalar operations.

[0131] Regarding the architecture of the I / O die:

[0132] The I / O die is equipped with a series of I / O functions to enable interaction with DRAM, the host system, or other input sources (such as cameras). All input data from the host system or other input sources can be first loaded into DRAM and then loaded and processed by the computing core. The DRAM controller is also connected to multiple routers within the entire mesh NoC to match the bandwidth of DRAM and the network, ensuring the full utilization of DRAM bandwidth.

[0133] Regarding the configurable parameters:

[0134] The hardware template of the die accelerator provided by the embodiments of the present disclosure has high configurability and provides a series of configurable architecture parameters. These configurable architecture parameters may include the bandwidth of the NoC, the D2D communication bandwidth, the total DRAM bandwidth, the total number of cores in the X direction (for example, Figure 2 the total number of cores in the X direction is 8), the total number of cores in the Y direction (for example, Figure 2 the total number of cores in the Y direction is 8), the number of dies divided in the X direction (X Cut )(for example, Figure 2 the number of dies divided in the X direction is 2), the number of dies divided in the Y direction (Y Cut )(for example, Figure 2The number of dielets divided in the Y direction is 2), the number of multiply-accumulate (MAC) operations in the processing unit array within a single computing core, and the size of the global buffer for each computing core. It is worth mentioning that the microarchitecture of the processing unit array and its corresponding data flow have been widely studied in related technologies. In the embodiments of the present disclosure, the processing unit array may adopt a classic NVDLA architecture and the corresponding data flow. Of course, it may also be replaced with other microarchitectures with different data flows, which are not limited herein.

[0135] The layer-pipelined space mapping encoding method provided by the embodiments of the present disclosure will be introduced below. Figure 3 A schematic diagram showing the layer-pipelined space mapping scheme for encoding in the parsing optimization space in the embodiments of the present disclosure.

[0136] A. Encoding Format and Parsing Method

[0137] In the embodiments of the present disclosure, a layer-centric encoding method is proposed to describe the layer-pipelined space mapping scheme. The encoded layer-pipelined space mapping scheme includes two important pieces of information: (1) the division of each layer and the allocation of partitioned workloads to specific computing cores; (2) the data sources and destinations of the workloads on each computing core. This encoding method has strong versatility and can seamlessly adapt to different NoC topologies and the microarchitectures of computing cores.

[0138] A DNN can be regarded as a Directed Acyclic Graph (DAG), and each layer in the DNN can be regarded as a node in the directed acyclic graph. In layer-pipelined mapping, the graph (or subgraph) can be mapped onto the computing core array simultaneously, where different computing core groups can be used to compute different layers in the DNN. The on-chip interconnection can be responsible for transmitting feature maps between layers with dependencies. The layer-pipelined space mapping scheme can be used to determine which computing core specifically computes which parts of which layers. Figure 3 The upper left corner shows a DAG of a DNN containing two layers. In the layer sequence (LS) mapping, all 6 cores are used for layer-by-layer calculation, while in the layer-pipelined mapping, some of the 6 cores are used to compute the first layer, and the remaining cores are used to compute the second layer. The feature maps between the two layers can be transmitted through the NoC without accessing the DRAM.

[0139] Assume that an N-layer DNN DAG needs to be spatially mapped onto a dielet accelerator in a layer-pipelined manner. The dielet accelerator includes a computing core group CG, and the computing core group CG includes M computing cores and D DRAMs. The layers in the directed acyclic graph can form a layer group LG. In Figure 3In the example shown, the directed acyclic graph is composed of two convolutional layers (Layer1 and Layer2, denoted as L1 and L2 in Figure 3 ), and these two convolutional layers can form a layer group LG. Additionally, in Figure 3 the example shown, the compute kernel group CG includes 6 compute kernels and 2 DRAMs.

[0140] In the encoding format proposed in the embodiments of the present disclosure, the layer-pipeline space mapping scheme LMS of any layer in the layer group LG can be composed of the mapping schemes MS of each layer in the layer group LG. The mapping scheme MS of the i-th layer Layer i can include 3 attributes: the partitioning attribute Part i =(H i , W i , B i , K i ); the compute kernel group attribute where nc i represents the number of compute kernels in the compute kernel group CG i ; the data flow attribute FD i =(IF i , WGT i , OF i ), and in one example, -1 ≤ IF i , WGT i , OF i ≤ D. Figure 3 The left side of shows that the layer-pipeline space mapping scheme LMS of a certain layer in the layer group LG includes the mapping schemes MS1 and MS2 of two layers in the layer group LG.

[0141] Among them, the partitioning attribute Part i can be used to divide the i-th layer Layer i into approximately equal nc i parts along the 4 dimensions of the output feature map of the i-th layer Layer i . As shown in Figure 3 , the 4 dimensions include: the height dimension of the output feature map, the width dimension of the output feature map, the batch size dimension, and the number of channels dimension of the output feature map. Correspondingly, the partitioning attribute Part i can include the number of divisions of the height dimension of the output feature map (H i ), the number of divisions of the width dimension of the output feature map (W i ), the number of divisions of the batch size dimension (B i ), and the number of divisions of the number of channels dimension of the output feature map (K i ). Among them, the number of divisions of the batch size dimension (B i) can represent the number of samples processed in a pipeline stage. The number of divisions (K) of the channel dimension of the output feature map i ) can be equal to the number of convolutional weight kernels. Based on this division scheme of the output feature map, the division schemes of the input feature map and the weights can be uniquely determined according to the characteristics of different types of layers. In Figure 3 the example shown, Layer1 is a convolutional layer. The number of divisions (B1) of the batch size dimension corresponding to Layer1 is equal to 2, the number of divisions (K1) of the channel dimension of the output feature map is equal to 2, the number of divisions (H1) of the height dimension of the output feature map is equal to 1, and the number of divisions (W1) of the width dimension of the output feature map is equal to 1. That is, in this example, the height dimension and the width dimension of the output feature map of Layer1 are not divided. This example shows how to divide the output feature map of the convolutional layer Layer1 based on the division attribute Part1, and how the division scheme of the output feature map of the convolutional layer Layer1 derives the corresponding input feature map and weight division schemes.

[0142] Kernel group attribute CG i includes information for calculating the computation kernels of the i-th layer Layer i . The kernel group attribute CG i is represented in the order of the computation kernels. For example, (C1, C2) ≠ (C2, C1). Each computation kernel in the kernel group attribute CG i can be any computation kernel in the computation kernel group CG. In Figure 3 the example shown, CG1 = (2, 1, 5, 4).

[0143] Then, we can establish a correspondence rule to map each partitioned workload (PW) of the i-th layer Layer i to the corresponding computation kernel in the kernel group attribute CG i . First, we can assign a unique 4D ID, such as (h, w, b, k), to each partitioned workload PW according to its position in the output feature map cube, where h ∈ [0, H i ), w ∈ [0, W i ), b ∈ [0, B i ), k ∈ [0, K i ). Then, the 4D ID of the partitioned workload PW can be converted into a numerical ID (NID). For example, NID = h × W i × B i × K i + w × B i × K i+b×K i +k. The numerical ID of each partition workload PW can correspond to the core group attribute CG i to which it is assigned. For example, in Figure 3 , the first partition workload among the 4 partition workloads of Layer1 (denoted as PW 1-0 in Figure 3 ) has a 4D ID of (0, 0, 0, 0) and a numerical ID of 0. This partition workload can be mapped to the first computing core C2 in CG1.

[0144] Data flow attribute FD i can represent the data source IF i of the input feature map of the i-th layer Layer i , the weight WGT i , and the destination OF i of the output feature map of the i-th layer Layer i . The data flow attribute can be divided into two categories: the first category is those that need to be explicitly managed (can be non-negative values in the data flow attribute FD i ), and the second category is those that do not need to be explicitly managed or directly do not exist (can be -1 in the data flow attribute FD i ).

[0145] Scenarios that require explicit management are as follows: (1) For the output feature map, when the subsequent layer and the current layer are not in the same layer group or the output of the current layer is the output of the entire DNN, it is necessary to explicitly manage the temporary storage of the output of this layer to a specific DRAM. (2) For the input feature map, explicit management is only required when the input of the current layer is the input of the entire DNN; otherwise, the data can be obtained from the DRAM storing the output feature map of the previous layer. (3) For the weight, as long as a layer has a weight, explicit management is required. Therefore, the primary problem in explicitly managing the data flow is to determine which DRAM to store data in and which DRAM to obtain data from. In an example, when the numerical value in the data flow attribute FD i is greater than 0, this numerical value can represent the ID of the DRAM. At the same time, 0 can represent a special case - interleaving, in which case we can evenly distribute the data across all DRAMs to fully and evenly utilize the available bandwidth of each DRAM. For example, as Figure 3 shows, since Layer1 has the input and weight of the entire DNN, IF1 and WGT1 can be non-negative numbers. As shown in the example in Figure 3 , FD1 of Layer1 is (1, 1, -1), indicating that the input feature map and weight of Layer1 originate from DRAM 1.

[0146] If two layers with a dependency relationship are in the same layer group, there is no need for explicit operations on the output feature map of the previous layer and the input feature map of the next layer. This is because the destination of each partition of the previous layer and the data source of each partition of the next layer can be directly derived based on the partitioning attribute Part i and the core group attribute CG i of these two layers. In addition, for layers without weights, their WGT i value can be -1. For example, in the Figure 3 shown example, based on Part1, CG1, Part2, and CG2, the data communication dependency relationship between the computing cores corresponding to Layer1 and Layer2 can be directly derived. Therefore, OF1 and IF2 can be -1.

[0147] As Figure 3 shown, based on the partitioning attribute Part i , the core group attribute CG i , and the data flow attribute FD i in the mapping scheme of each layer in the layer group LG, the actual mapping scheme of the workload of each partition of each layer in the layer group LG to the computing core group CG can be parsed.

[0148] B. Spatial Computation

[0149] As Figure 3 shown, each layer-pipeline spatial mapping scheme can be regarded as a point in the optimization space respectively. Mapping N layers to a die accelerator including M computing cores and D DRAMs results in a rather large optimization space, which is extremely complex to calculate. Therefore, we can conservatively approximate the lower bound size of the optimization space as layer-pipeline spatial mapping schemes, where is the binomial coefficient,

[0150] C. Uncovering Hidden Optimization Opportunities

[0151] First, the different partitioning attributes Part i of each layer affect two aspects: (1) NoC traffic. Different partitioning schemes lead to different data requirements for each computing core, resulting in differences in NoC traffic, even if the NoC has multicast capabilities. For example, in Figure 3 , in the case of Part1 = (1, 1, 2, 2), each computing core in CG1 requires half of the input feature map and weights. However, if Part1 is changed to (1, 1, 1, 4), each computing core in CG1 requires the entire input feature map and only 1 / 4 of the weights. (2) Optimization space within the computing core: Different partitioning schemes result in different partition workloads, which in turn affect the optimal data flow scheme within the computing core.

[0152] Second, the number and positions of compute cores in each core group property CG i may vary. The number of compute cores affects the computation time of each layer, and thus the computation time of the entire pipeline, because the slowest stage may cause pipeline stalls. The positions of compute cores can significantly affect the data traffic volume and congestion level of the NoC.

[0153] Third, different data flow properties FD i affect the bandwidth utilization and access patterns of different DRAM and NoC communications. As Figure 3 shown, the input feature map and weights of Layer1 are read from DRAM1, while the weights of Layer2 are read from DRAM1 and the output feature map of Layer2 is written to DRAM2. In this case, the bandwidth requirements and access patterns of each DRAM are not balanced, but the total number of hops of the NoC is relatively small. If all positive values in FD1 and FD2 become 0 (interleaved), the DRAM bandwidth usage will become more balanced. However, since some data will interact with more remote DRAMs, the total number of hops of the NoC will increase.

[0154] Considering that there are sufficient optimization opportunities to optimize the NoC traffic and congestion, exploring the optimization space defined by us is of greater significance for the chiplet scenario. As mentioned above, for the chiplet scenario, it may be difficult to maintain the same D2D bandwidth as the on-chip link bandwidth, resulting in a lower bandwidth for some links in the chiplet accelerator than others. In addition, the D2D energy consumption is significantly higher than that of on-chip communication. Therefore, the chiplet scenario poses great requirements for better utilization of on-chip communication resources and minimization of D2D traffic, which can be well optimized by exploring the optimization space defined by us.

[0155] In the embodiments of the present disclosure, a layer-centric encoding method is proposed to represent the optimization space of the layer-pipeline space mapping scheme in the chiplet accelerator and calculate its huge size. In this encoding method, the optimization space of the layer-pipeline space mapping scheme in the chiplet accelerator is clearly and systematically defined. This method is significantly superior to existing heuristic strategies and analyzes the potential optimization opportunities hidden within this optimization space. These opportunities become particularly important in the post-Moore's law chiplet era.

[0156] Figure 4A flowchart showing a co - exploration method for the mapping scheme and architecture of a chiplet accelerator provided by an embodiment of the present disclosure. In a possible implementation, the execution subject of the co - exploration method for the mapping scheme and architecture of the chiplet accelerator can be a co - exploration device for the mapping scheme and architecture of the chiplet accelerator. For example, the co - exploration method for the mapping scheme and architecture of the chiplet accelerator can be executed by a terminal device, a server, or other electronic devices. Among them, the terminal device can be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle - mounted device, or a wearable device, etc. In some possible implementations, the co - exploration method for the mapping scheme and architecture of the chiplet accelerator can be implemented by a processor calling computer - readable instructions stored in a memory.

[0157] In an embodiment of the present disclosure, by obtaining candidate values of configurable architecture parameters of a chiplet accelerator, framework setting information, and a neural network model, generating a candidate co - solution for the layer - pipeline space mapping scheme and architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model, determining the monetary cost, power consumption cost, and latency of the candidate co - solution, and determining the evaluation value of the candidate co - solution according to the monetary cost, power consumption cost, and latency of the candidate co - solution, the co - exploration of the layer - pipeline space mapping scheme and architecture of the chiplet accelerator is realized. It not only considers the power consumption cost and performance (latency), but also considers the monetary cost. The co - solution obtained by exploring using the embodiment of the present disclosure can achieve better performance and higher energy efficiency.

[0158] As Figure 4 shown, the co - exploration method for the mapping scheme and architecture of the chiplet accelerator includes steps S11 to S14.

[0159] In step S11, obtain candidate values of configurable architecture parameters of a chiplet accelerator, framework setting information, and a neural network model.

[0160] In step S12, generate a candidate co - solution for the layer - pipeline space mapping scheme and architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model.

[0161] In step S13, determine the monetary cost, power consumption cost, and latency of the candidate co - solution.

[0162] In step S14, determine the evaluation value of the candidate co - solution according to the monetary cost, power consumption cost, and latency of the candidate co - solution.

[0163] In a possible implementation, the die accelerator is an inference accelerator for a deep neural network (DNN). That is, the mapping scheme and architecture co-exploration method of the die accelerator can be used to co-explore the layer-pipelining space mapping scheme and architecture of the inference die accelerator for a deep neural network.

[0164] Of course, the die accelerator can also be used as an accelerator for other types of neural networks, which is not limited here. The embodiments of the present disclosure are described by taking a deep neural network as an example.

[0165] In the embodiments of the present disclosure, the configurable architecture parameters can represent the configurable parameters in the architecture of the die accelerator. As the configurable architecture parameters change, the architecture of the die accelerator will change.

[0166] In a possible implementation, the configurable architecture parameters include at least some of the following: the bandwidth of the network-on-chip (NoC), the die-to-die (D2D) communication bandwidth, the total bandwidth of the memory, the total number of computing cores in the X direction of the die accelerator, the total number of computing cores in the Y direction of the die accelerator, the number of dies divided in the X direction of the die accelerator, the number of dies divided in the Y direction of the die accelerator, the number of multiply-accumulate operations in the processing unit array of the computing core, and the size of the global buffer of the computing core.

[0167] In one example, when the memory uses DRAM, the total bandwidth of the memory is the DRAM total bandwidth.

[0168] In this implementation, the total number of computing cores in the X direction of the die accelerator can represent the number of computing cores in the X direction of the die accelerator. For example, in Figure 2 the example shown, the total number of computing cores in the X direction of the die accelerator is 8.

[0169] In this implementation, the total number of computing cores in the Y direction of the die accelerator can represent the number of computing cores in the Y direction of the die accelerator. For example, in Figure 2 the example shown, the total number of computing cores in the Y direction of the die accelerator is 8.

[0170] In this implementation, the number of dies divided in the X direction of the die accelerator can represent the number of dies in the X direction of the die accelerator. For example, in Figure 2 the example shown, the number of dies divided in the X direction of the die accelerator is 2.

[0171] In this implementation, the number of dies divided in the Y direction of the die accelerator can represent the number of dies in the Y direction of the die accelerator. For example, in Figure 2In the example shown, the number of dielets divided in the Y direction in the dielet accelerator is 2.

[0172] In this implementation, the number of multiply-accumulate operations in the processing unit array of the computing core can refer to the number of multiply-accumulate operations in the processing unit array of a single computing core.

[0173] In a possible implementation, a list of candidate values for each configurable architecture parameter can be obtained. For example, a list of candidate values for the bandwidth of the NoC, a list of candidate values for the D2D communication bandwidth, a list of candidate values for the total DRAM bandwidth, a list of candidate values for the total number of computing cores in the X direction in the dielet accelerator, a list of candidate values for the total number of computing cores in the Y direction in the dielet accelerator, a list of candidate values for the number of dielets divided in the X direction in the dielet accelerator, a list of candidate values for the number of dielets divided in the Y direction in the dielet accelerator, a list of candidate values for the number of multiply-accumulate operations in the processing unit array of the computing core, and a list of candidate values for the size of the global buffer of the computing core can be obtained respectively. Among them, the list of candidate values for each configurable architecture parameter can be flexibly set according to requirements and is not limited herein.

[0174] In the embodiments of the present disclosure, the framework setting information can represent the setting information of the collaborative exploration framework of the mapping scheme and architecture of the dielet accelerator. In a possible implementation, the framework setting information can include optimization target information, constraint conditions, hyperparameters, and other relevant setting information.

[0175] Figure 5a A schematic diagram showing the collaborative exploration framework of the mapping scheme and architecture of the dielet accelerator provided by the embodiments of the present disclosure. Figure 5b A schematic diagram showing the mapping engine in the collaborative exploration framework of the mapping scheme and architecture of the dielet accelerator provided by the embodiments of the present disclosure. In Figure 5a In the example shown, the input of the collaborative exploration framework can include candidate values of configurable architecture parameters, framework setting information, and a DNN model.

[0176] In a possible implementation, obtaining a neural network model includes: obtaining a plurality of neural network models; generating a candidate collaborative scheme of the layer-pipeline space mapping scheme and architecture of the dielet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model, including: generating a candidate collaborative scheme of the layer-pipeline space mapping scheme and architecture of the dielet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and any one of the plurality of neural network models.

[0177] In one example, multiple DNN models can be obtained, and a layer-pipelining space mapping scheme and a candidate cooperation scheme for the architecture of the chiplet accelerator can be generated based on the candidate values of configurable architecture parameters, framework setting information, and any one of the multiple DNN models. In this implementation manner, any candidate cooperation scheme can correspond to any one of the multiple DNN models.

[0178] Considering that the same chiplet accelerator can be used to accelerate different deep neural networks in different scenarios, in this implementation manner, multiple DNN models are input into the cooperation exploration framework, thereby enabling the support for design space exploration (DSE) of multiple DNNs.

[0179] In the embodiments of the present disclosure, any candidate cooperation scheme may include a layer-pipelining space mapping candidate scheme and a candidate architecture.

[0180] In a possible implementation manner, the layer-pipelining space mapping scheme in any candidate cooperation scheme includes the mapping schemes of each network layer in the neural network model, and the mapping scheme of any network layer includes a partitioning attribute, a core group attribute, and a data flow attribute, where the partitioning attribute is used to partition the input feature map, weight, and output feature map of the network layer, the core group attribute represents the information of the computing cores used to calculate the network layer, and the data flow attribute represents the data source of the input feature map of the network layer, the weight, and the destination of the output feature map of the network layer.

[0181] As Figure 3 shown, each layer-pipelining space mapping scheme can be used as a point in the optimization space.

[0182] In Figure 3 the example shown, the neural network model is a DNN, and the DNN includes 2 network layers. Then, the layer-pipelining space mapping scheme in any candidate cooperation scheme may include the mapping schemes of the 2 network layers in the DNN. For example, the layer-pipelining space mapping scheme LMS in any candidate cooperation scheme may include the mapping scheme MS1 of network layer L1 and the mapping scheme MS2 of network layer L2 in the DNN.

[0183] The mapping scheme of the i-th network layer Layer i in the DNN may include a partitioning attribute Part i and a core group attribute CG i and a data flow attribute FD i . The partitioning attribute Part i can be used to partition the network layer Layer iPartition the input feature map, weights, and output feature map; kernel group attribute CG i can represent the information of the computing kernel used to calculate network layer Layer i ; data flow attribute FD i can represent the data source of the input feature map of network layer Layer i , the weights, and the destination of the output feature map of network layer Layer i .

[0184] In a possible implementation, the partitioning attributes in the mapping scheme of any network layer include the number of partitions in the height dimension of the output feature map, the number of partitions in the width dimension of the output feature map, the number of partitions in the batch size dimension, and the number of partitions in the channel number dimension of the output feature map.

[0185] For example, the partitioning attribute Part i of network layer Layer i =(H i , W i , B i , K i ), where H i can represent the number of partitions along the height dimension of the output feature map of network layer Layer i , W i can represent the number of partitions along the width dimension of the output feature map of network layer Layer i , B i can represent the number of partitions along the batch size dimension, and K i can represent the number of partitions along the channel number dimension of the output feature map of network layer Layer i .

[0186] In a possible implementation, generating the candidate co - scheme of the layer - pipeline space mapping scheme and the architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model includes: generating the candidate co - scheme of the layer - pipeline space mapping scheme and the architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model, and using a preset simulated annealing operator; wherein, the preset simulated annealing operator includes at least one of the following: a first simulated annealing operator for randomly selecting a network layer and changing the partitioning attribute in the mapping scheme of the network layer; a second simulated annealing operator for randomly selecting a network layer and randomly swapping two computing cores in the core group attribute in the mapping scheme of the network layer; a third simulated annealing operator for randomly selecting two network layers and randomly swapping two computing cores in the core group attribute in the mapping schemes of the two network layers; a fourth simulated annealing operator for randomly selecting two network layers, removing a computing core from the core group attribute in the mapping scheme of one of the two network layers, and adding the removed computing core to the core group attribute in the mapping scheme of the other of the two network layers; a fifth simulated annealing operator for randomly selecting a network layer, randomly selecting a non - negative term in the data flow attribute in the mapping scheme of the network layer, and randomly determining the updated value of the non - negative term within the value range [0, D].

[0187] In an example, the first simulated annealing operator can be represented by OP1. The first simulated annealing operator can be used to randomly select a network layer from each network layer of the DNN and change the partitioning attribute Part i in the mapping scheme of this network layer, while still satisfying the constraint conditions corresponding to the partitioning attribute Part i in the framework setting information.

[0188] In an example, the second simulated annealing operator can be represented by OP2. The second simulated annealing operator can be used to randomly select a network layer from each network layer of the DNN and randomly swap the positions of two computing cores in the core group attribute CG i in the mapping scheme of this network layer (equivalent to swapping the partition workloads of these two computing cores).

[0189] In an example, the third simulated annealing operator can be represented by OP3. The third simulated annealing operator can be used to randomly select two network layers from each network layer of the DNN and randomly swap two computing cores in the core group attribute in the mapping schemes of these two network layers (equivalent to swapping the partition workloads of these two computing cores).

[0190] In one example, the fourth simulated annealing operator can be represented as OP4. The fourth simulated annealing operator can be used to randomly select two network layers from the various network layers of the DNN, remove a computing core from the kernel group attributes in the mapping scheme of one of the two network layers, and add the removed computing core to the kernel group attributes in the mapping scheme of the other network layer. After the operation of the fourth simulated annealing operator is completed, the partitioning attributes of the two network layers can be randomly updated to match their new kernel group attributes.

[0191] In one example, the fifth simulated annealing operator can be represented as OP5. The fifth simulated annealing operator can be used to randomly select one network layer from the various network layers of the DNN and randomly select a non - negative term in the data flow attribute FD i in it, and randomly determine the updated value of the non - negative term within the value range [0, D], where D can represent the number of DRAMs in the architecture in the candidate cooperation scheme.

[0192] By utilizing these 5 simulated annealing operators, various attributes in the mapping schemes of the respective network layers can be transformed into any state that satisfies the constraint conditions. For example, by using the fourth simulated annealing operator, the number of computing cores in CG1 in Figure 3 can be modified to any value between 1 and 5. Therefore, by using 5 simulated annealing operators, the exploration process in simulated annealing can be promoted, enabling the full exploration of the optimization space and obtaining a solution close to the optimal one.

[0193] In addition, by utilizing these 5 simulated annealing operators, the collaborative exploration framework can not only explore the optimization space to solve the trade - off problem mentioned above, but also automatically optimize D2D link communication. Since D2D links often have a small bandwidth and high energy consumption, during the iteration process, if the simulated annealing operation increases the use of more D2D links, it is more likely to significantly reduce performance and energy efficiency, making it less likely for the modified mapping scheme to be accepted. On the other hand, the simulated annealing operation that reduces the use of D2D links is more likely to be accepted. Therefore, the entire exploration process essentially optimizes D2D communication. Moreover, this mapping technique can help in architecture design by equipping the chiplet accelerator with a smaller D2D bandwidth, reducing its area overhead and obtaining the benefits of improved yield caused by the chiplet, while only incurring a minimal loss in performance and energy efficiency.

[0194] In a possible implementation, determining the evaluation value of the candidate collaboration solution according to the monetary cost, power consumption cost, and latency of the candidate collaboration solution includes: obtaining a first weight corresponding to the monetary cost, a second weight corresponding to the power consumption cost, and a third weight corresponding to the latency; determining the evaluation value of the candidate collaboration solution according to the monetary cost, power consumption cost, and latency of the candidate collaboration solution, and the first weight, the second weight, and the third weight.

[0195] In one example, the evaluation value of the candidate collaboration solution can be MC α ×E β ×D γ , where MC represents the monetary cost, α represents the first weight, E represents the power consumption cost, β represents the second weight, D represents the latency, and γ represents the third weight.

[0196] In another example, the evaluation value of the candidate collaboration solution can be αMC + βE + γD.

[0197] Of course, the calculation method of the evaluation value of the candidate collaboration solution can be determined flexibly and is not limited here.

[0198] Among them, the monetary cost can be determined only by the architecture of the chiplet accelerator; the power consumption cost and latency are affected not only by the architecture of the chiplet accelerator, but also by the specific DNN workload and the corresponding mapping strategy.

[0199] In a possible implementation, the mapping engine can adopt a graph partitioning algorithm based on dynamic programming and a layer-pipelined space mapping scheme exploration algorithm based on simulated annealing to optimize the mapping of the DNN to the architecture in the candidate collaboration solution. For the i-th DNN input to the collaborative exploration architecture, this optimization process can adopt the optimization objective and can determine E based on the values of the configurable architecture parameters in the candidate collaboration solution i and D i . In one example, the overall power consumption cost of the candidate architecture of the chiplet accelerator can be The overall latency can be where n can represent the number of DNNs in the input collaborative exploration architecture.

[0200] In a possible implementation, determining the power consumption cost of the candidate collaboration solution includes: determining the power consumption cost of the candidate collaboration solution according to the number of operations of each component in the architecture of the candidate collaboration solution and the corresponding unit power consumption.

[0201] For example, the number of operations of each component in each computing core can be calculated, such as the number of accesses to different-level buffers, the number of multiply-accumulate operations with different precisions, etc.

[0202] In this implementation, the number of operations of each component in the architecture of the candidate cooperation solution can be determined, as well as the unit power consumption of each operation of each component. For any component in the architecture of the candidate cooperation solution, the number of operations of the component can be multiplied by the corresponding unit power consumption to obtain the power consumption cost of the component. The power consumption costs of all components can be added up to obtain the power consumption cost of the candidate cooperation solution.

[0203] In a possible implementation, determining the latency of the candidate cooperation solution includes: determining the calculation time of the multiply-accumulate operations in the candidate cooperation solution; determining the ratio of the maximum data access volume of the memory in the candidate cooperation solution to the access bandwidth; and determining the latency of the candidate cooperation solution according to the calculation time and the ratio.

[0204] In this implementation, the calculation time of each multiply-accumulate operation in the candidate cooperation solution can be determined, and the ratio of the maximum data access volume of each memory (such as DRAM) in the candidate cooperation solution to the access bandwidth can be determined. The maximum value among the calculation time and the ratio can be determined as the latency of the candidate cooperation solution.

[0205] The latency related to a small batch (a smaller batch size) can be used to illustrate the performance of the chiplet accelerator in latency-sensitive scenarios; the latency related to a large batch (a larger batch size) can be used to illustrate the performance of the chiplet accelerator in throughput-sensitive scenarios.

[0206] In a possible implementation, determining the monetary cost of the candidate cooperation solution includes: determining the monetary cost of the candidate cooperation solution according to the values of the configurable architecture parameters in the candidate cooperation solution.

[0207] In this implementation, when determining the monetary cost of any candidate cooperation solution, only the values of the configurable architecture parameters in the candidate cooperation solution need to be considered.

[0208] In a possible implementation, the monetary cost of the candidate cooperation solution includes at least one of the following: the chiplet manufacturing cost of the candidate cooperation solution, the memory cost of the candidate cooperation solution, and the packaging cost of the candidate cooperation solution.

[0209] In this implementation, the chiplet manufacturing cost of the candidate cooperation solution can be the sum of the chiplet manufacturing costs of all chiplets in the architecture of the candidate cooperation solution. For example, the chiplet manufacturing cost of the candidate cooperation solution can be the sum of the chiplet manufacturing costs of all computing chiplets and all IO chiplets in the architecture of the candidate cooperation solution.

[0210] In this implementation, the memory cost of the candidate collaboration solution can be the sum of the memory costs of each memory in the architecture of the candidate collaboration solution. For example, the memory can include DRAM (Dynamic Random Access Memory), etc., which is not limited here. When the memory uses DRAM, the memory cost can also be referred to as the DRAM cost.

[0211] In this implementation, the packaging cost of the candidate collaboration solution can be related to the substrate in the architecture of the candidate collaboration solution.

[0212] In a possible implementation, determining the currency cost of the candidate collaboration solution according to the value of the configurable architecture parameter in the candidate collaboration solution includes: for any die in the architecture of the candidate collaboration solution, determining the silicon area and the yield of the die, and obtaining the unit area manufacturing cost of the process technology node corresponding to the die; determining the die manufacturing cost of the die according to the silicon area of the die, the yield of the die, and the unit area manufacturing cost of the process technology node corresponding to the die.

[0213] In this implementation, for any die in the architecture of the candidate collaboration solution, the silicon area of the die can be determined according to the silicon areas of each module in the die.

[0214] Among them, for any analog module in the die, the area of the analog module can be obtained from the data sheet corresponding to the analog module. Among them, the analog module can include PCIe PHY (Peripheral Component Interface Express Physical Layer), DDR PHY (Double Data Rate Physical Layer), D2D PHY (Die-to-Die Physical Layer), etc.

[0215] For any logic module in the die, the area of the logic module can be estimated based on the Verilog code and the evaluation process adopted during the die development process.

[0216] In an example, the manufacturing cost of the i-th die in the architecture of the candidate collaboration solution is Among them, can represent the silicon area of the i-th die, can represent the yield of the i-th die, C siliconcan represent the manufacturing cost per unit area corresponding to the i-th die. For example, in the case of adopting a 12nm process, C silicon = 0.085 $ / mm 2 .

[0217] In a possible implementation, determining the yield of the die includes: determining the yield of the die according to the silicon area, defect density, and clustering parameter of the die.

[0218] Among them, the defect density and the clustering parameter are two hyperparameters. The defect density can refer to the number of defects per unit area during the manufacturing process of a semiconductor wafer. These defects may be caused by various factors during the manufacturing process, such as material impurity, equipment error, environmental pollution, etc. The defect density is one of the important indicators to measure the quality of semiconductor products. The clustering parameter can be used to describe the distribution pattern of defects on the wafer. The clustering parameter can help identify whether the defects tend to cluster in a certain area or are evenly distributed. In semiconductor manufacturing, understanding and controlling the clustering of defects is very important for improving the reliability and yield of products.

[0219] In an example, the yield of the i-th die in the architecture of the candidate cooperation scheme Among them, can represent the silicon area of the i-th die, D can represent the defect density, and c can represent the clustering parameter. For example, D = 0.1, c = 10.

[0220] In a possible implementation, for any die, the yield of the die can be determined according to the silicon area of the die, the unit area of the die, and the yield of the unit area of the die.

[0221] In an example, the yield of the i-th die in the architecture of the candidate cooperation scheme Among them, can represent the silicon area of the i-th die, can represent the unit area of the i-th die, and Yield unit can represent the yield of the unit area of the i-th die. For example, in the case of adopting a process below 12nm,

[0222] In a possible implementation, the determining the monetary cost of the candidate cooperation scheme according to the value of the configurable architecture parameter in the candidate cooperation scheme includes: determining the memory cost of the candidate cooperation scheme according to the bandwidth of the memory, the bandwidth of the memory unit, and the monetary cost of the memory unit in the architecture of the candidate cooperation scheme.

[0223] In one example, the memory cost can be where DRAM bw can represent the bandwidth of the memory, and Unit bw can represent the bandwidth of the memory cell, and can represent the monetary cost of the memory cell. In one example, the memory uses GDDR6,

[0224] In one possible implementation, determining the monetary cost of the candidate cooperation scheme according to the value of the configurable architecture parameter in the candidate cooperation scheme includes: determining the packaging cost of the candidate cooperation scheme according to the total silicon area, scaling factor, packaging yield and the monetary cost per unit area of the substrate of all the die in the architecture in the candidate cooperation scheme.

[0225] In one example, the total silicon area of all the die in the architecture in the candidate cooperation scheme where can represent the silicon area of the i-th die in the architecture in the candidate cooperation scheme.

[0226] Since the substrate needs to accommodate the fanout of the input / output module and the interconnection wiring function, the substrate requires a larger area compared with the total silicon area of all the die. In this implementation, the area of the substrate can be estimated according to the total silicon area of all the die and the scaling factor. The scaling factor can be determined according to empirical values.

[0227] The monetary cost per unit area of different substrates may be different, and the monetary cost per unit area of different regions in the same substrate may be different. In addition, a larger substrate area requires a more complex manufacturing process, which will result in a higher monetary cost per unit area.

[0228] In one example, the packaging cost of the candidate cooperation scheme can be (Area tot ·f scale ) / Yield package ·C package , where Area tot can represent the total silicon area of all the die in the architecture in the candidate cooperation scheme, f scale can represent the scaling factor, Yield package can represent the packaging yield, and C package can represent the monetary cost per unit area of the substrate.

[0229] In one possible implementation, after determining the evaluation value of the candidate collaboration solution, the method further includes: determining the optimal collaboration solution of the chiplet accelerator according to the evaluation values of multiple candidate collaboration solutions.

[0230] For example, the candidate collaboration solution with the smallest evaluation value among all candidate collaboration solutions can be determined as the optimal collaboration solution of the chiplet accelerator.

[0231] Reference Figure 5a and Figure 5b , the collaborative exploration architecture of the mapping scheme and architecture of the chiplet accelerator provided by the embodiments of the present disclosure will be introduced below. The collaborative exploration architecture of the mapping scheme and architecture of the chiplet accelerator can be used to implement the collaborative exploration method of the mapping scheme and architecture of the chiplet accelerator.

[0232] As Figure 5a shown, the input of the collaborative exploration framework may include candidate values of configurable architecture parameters, framework setting information, and multiple DNN models.

[0233] The collaborative exploration framework may include an iterator, a currency cost evaluator, a mapping engine, and an instruction generator.

[0234] Among them, the iterator can be used to generate different candidate architectures of the chiplet accelerator according to the candidate values of the configurable architecture parameters. The currency cost evaluator can be used to evaluate the currency cost of different candidate architectures.

[0235] The mapping engine can be used to evaluate the power consumption cost and latency of the candidate collaboration solution according to the candidate architecture, DNN model, and actual workload in the candidate collaboration solution.

[0236] As Figure 5a shown, the collaborative exploration framework can save the exploration records of the candidate collaboration solutions and can obtain the analysis results of each candidate collaboration solution. Based on the analysis results (such as evaluation values) of each candidate collaboration solution, the optimal collaboration solution can be determined. The instruction generator can be used to convert the high-level information in the optimal collaboration solution into low-level hardware instructions, so as to deploy the layer-pipelined space mapping scheme (i.e., the optimal layer-pipelined space mapping scheme) in the optimal collaboration solution on the chiplet accelerator with the architecture (i.e., the optimal architecture) in the optimal collaboration solution.

[0237] The output of the collaborative exploration framework may include the currency cost of the optimal collaboration solution, the architecture (i.e., the optimal architecture) in the optimal collaboration solution, the power consumption cost and latency report of the optimal collaboration solution, and the low-level hardware instructions corresponding to the optimal collaboration solution.

[0238] As Figure 5bAs shown, the mapping engine may include a graph-level partitioning engine, a layer-pipeline space mapping exploration engine, and an evaluator. Among them, the graph-level partitioning engine may include a model parser and a graph partitioning engine. The layer-pipeline space mapping exploration engine may include a simulated annealing controller, a layer-pipeline space mapping scheme analyzer, and an in-core exploration engine.

[0239] The model parser can be used to process the description file of the DNN, generate the structure topology graph corresponding to the DNN, and extract the feature maps of each network layer of the DNN.

[0240] The graph partitioning engine can adopt a graph partitioning algorithm based on Density Peaks (DP) to partition the structure topology graph of the DNN, and partition the DNN into at least one layer group. Among them, each layer group includes at least one layer of the DNN. Among them, the graph partitioning algorithm based on density peaks can not only effectively partition the layer group, but also determine the number of samples (batch size) processed in each pipeline stage. The graph partitioning engine can generate multiple layer group partitioning candidate schemes, and can evaluate the cost of each layer group partitioning candidate scheme through the evaluator to determine the optimal layer group partitioning scheme.

[0241] The graph partitioning engine can send the optimal layer group partitioning scheme to the simulated annealing controller. The simulated annealing controller may include a layer-pipeline space mapping scheme generator. The layer-pipeline space mapping scheme generator can, based on the optimal layer group partitioning scheme explored by the graph partitioning engine, adopt an algorithm based on simulated annealing to explore the optimization space of the layer-pipeline space mapping scheme for each layer group. For any layer group, an initial layer-pipeline space mapping scheme of the layer group can be obtained based on a heuristic striping strategy. Then, simulated annealing iteration can be performed based on 5 simulated annealing operators.

[0242] The layer-pipeline space mapping scheme analyzer can analyze the layer-pipeline space mapping scheme of each layer group. The mapping scheme of any layer in any layer group may include partitioning attributes, kernel group attributes, and data flow attributes.

[0243] After obtaining the partitioning attributes of each layer in the layer group, the in-core exploration engine can schedule the partition workload. The in-core exploration engine can adopt tiling technology and loop reorder technology to perform an exhaustive search on the scheduling scheme of the partition workload. After finding the optimal scheme of the in-core data flow of each partition workload, the optimal scheme of the in-core data flow can be sent to the evaluator together with other analysis information for overall evaluation. If the overall cost is low, the change is accepted; otherwise, it is not accepted. The probability that the scheme is accepted decreases as the number of iterations increases.

[0244] The evaluator can evaluate the candidate collaborative solutions in two aspects: in-core evaluation and global evaluation. For example, the evaluation objects of the evaluator can include NoC and D2D traffic, DRAM access times, DRAM access patterns, in-core analysis results, etc.

[0245] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not elaborate further. Those skilled in the art can understand that in the above methods of the specific implementation manner, the specific execution order of each step should be determined according to its function and possible internal logic.

[0246] In addition, the present disclosure also provides a co-exploration device, an electronic device, a computer-readable storage medium, and a computer program product for the mapping scheme and architecture of the chiplet accelerator. The above can all be used to implement any co-exploration method for the mapping scheme and architecture of the chiplet accelerator provided by the present disclosure. The corresponding technical solutions and technical effects can be seen in the corresponding records in the method part and will not be elaborated further.

[0247] Figure 6 The block diagram of the co-exploration device for the mapping scheme and architecture of the chiplet accelerator provided by the embodiments of the present disclosure is shown. As Figure 6 shown, the co-exploration device for the mapping scheme and architecture of the chiplet accelerator includes:

[0248] An acquisition module 61, configured to acquire candidate values of configurable architecture parameters of the chiplet accelerator, framework setting information, and a neural network model;

[0249] A generation module 62, configured to generate a candidate collaborative solution for the layer-pipeline space mapping scheme and architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model;

[0250] A first determination module 63, configured to determine the monetary cost, power consumption cost, and latency of the candidate collaborative solution;

[0251] A second determination module 64, configured to determine the evaluation value of the candidate collaborative solution according to the monetary cost, power consumption cost, and latency of the candidate collaborative solution.

[0252] In a possible implementation manner, the configurable architecture parameters include at least some of the following: the bandwidth of the on-chip network, the die-to-die communication bandwidth, the total bandwidth of the memory, the total number of computing cores in the X direction of the chiplet accelerator, the total number of computing cores in the Y direction of the chiplet accelerator, the number of chiplets divided in the X direction of the chiplet accelerator, the number of chiplets divided in the Y direction of the chiplet accelerator, the number of multiply-accumulate operations in the processing unit array of the computing core, and the size of the global buffer of the computing core.

[0253] In a possible implementation,

[0254] the obtaining module is configured to: obtain a plurality of neural network models;

[0255] the generating module is configured to: generate a candidate cooperation scheme of the layer-pipeline space mapping scheme and the architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and any one of the plurality of neural network models.

[0256] In a possible implementation, the layer-pipeline space mapping scheme in any candidate cooperation scheme includes the mapping schemes of each network layer in the neural network model. The mapping scheme of any network layer includes a partitioning attribute, a core group attribute, and a data flow attribute. Among them, the partitioning attribute is used to partition the input feature map, weight, and output feature map of the network layer. The core group attribute represents information about the computing cores used to compute the network layer. The data flow attribute represents the data source of the input feature map of the network layer, the weight, and the destination of the output feature map of the network layer.

[0257] In a possible implementation, the partitioning attribute in the mapping scheme of any network layer includes the number of partitions in the height dimension of the output feature map, the number of partitions in the width dimension of the output feature map, the number of partitions in the batch size dimension, and the number of partitions in the channel number dimension of the output feature map.

[0258] In a possible implementation, the generating module is configured to:

[0259] generate a candidate cooperation scheme of the layer-pipeline space mapping scheme and the architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model, and adopt a preset simulated annealing operator;

[0260] wherein, the preset simulated annealing operator includes at least one of the following:

[0261] The first simulated annealing operator is used to randomly select a network layer and change the partitioning attribute in the mapping scheme of the network layer;

[0262] The second simulated annealing operator is used to randomly select a network layer and randomly exchange two computing cores in the core group attribute in the mapping scheme of the network layer;

[0263] The third simulated annealing operator is used to randomly select two network layers and randomly exchange two computing cores in the core group attributes in the mapping schemes of the two network layers;

[0264] The fourth simulated annealing operator is used to randomly select two network layers, remove a computing core from the core group attributes in the mapping scheme of one of the two network layers, and add the removed computing core to the core group attributes in the mapping scheme of the other network layer among the two network layers;

[0265] The fifth simulated annealing operator is used to randomly select a network layer, randomly select a non - negative term in the data flow attributes in the mapping scheme of the network layer, and randomly determine the updated value of the non - negative term within the value range [0, D].

[0266] In a possible implementation manner, the first determination module is configured to:

[0267] Determine the power consumption cost of the candidate cooperation scheme according to the number of operations and the corresponding unit power consumption of each component in the architecture of the candidate cooperation scheme.

[0268] In a possible implementation manner, the first determination module is configured to:

[0269] Determine the computing time of the multiply - accumulate operations in the candidate cooperation scheme;

[0270] Determine the ratio of the maximum value of the data access volume of the memory in the candidate cooperation scheme to the access bandwidth;

[0271] Determine the latency of the candidate cooperation scheme according to the computing time and the ratio.

[0272] In a possible implementation manner, the first determination module is configured to:

[0273] Determine the currency cost of the candidate cooperation scheme according to the values of the configurable architecture parameters in the candidate cooperation scheme.

[0274] In a possible implementation manner, the currency cost of the candidate cooperation scheme includes at least one of the following:

[0275] The die manufacturing cost of the candidate cooperation scheme, the memory cost of the candidate cooperation scheme, the packaging cost of the candidate cooperation scheme.

[0276] In a possible implementation manner, the first determination module is configured to:

[0277] For any die in the architecture of the candidate cooperation scheme, determine the silicon area and the yield of the die, and obtain the unit area manufacturing cost of the process technology node corresponding to the die;

[0278] Determine the die manufacturing cost according to the silicon area of the die, the yield rate of the die, and the unit area manufacturing cost of the process technology node corresponding to the die.

[0279] In a possible implementation manner, the first determination module is configured to:

[0280] Determine the yield rate of the die according to the silicon area, defect density, and clustering parameter of the die.

[0281] In a possible implementation manner, the first determination module is configured to:

[0282] Determine the memory cost of the candidate collaboration solution according to the bandwidth of the memory, the bandwidth of the memory unit, and the currency cost of the memory unit in the architecture of the candidate collaboration solution.

[0283] In a possible implementation manner, the first determination module is configured to:

[0284] Determine the packaging cost of the candidate collaboration solution according to the total silicon area of all dies, the scaling factor, the packaging yield rate, and the unit area currency cost of the substrate in the architecture of the candidate collaboration solution.

[0285] In a possible implementation manner, the second determination module is configured to:

[0286] Obtain a first weight corresponding to the currency cost, a second weight corresponding to the power consumption cost, and a third weight corresponding to the latency;

[0287] Determine the evaluation value of the candidate collaboration solution according to the currency cost, power consumption cost, and latency of the candidate collaboration solution, and the first weight, the second weight, and the third weight.

[0288] In a possible implementation manner, the device further includes:

[0289] A third determination module, configured to determine the optimal collaboration solution of the die accelerator according to the evaluation values of multiple candidate collaboration solutions.

[0290] In a possible implementation manner, the die accelerator is an inference accelerator for a deep neural network.

[0291] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation and technical effects can refer to the descriptions of the above method embodiments. For the sake of brevity, they are not described here again.

[0292] Embodiments of the present disclosure also provide a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above method is implemented. Among them, the computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium.

[0293] Embodiments of the present disclosure also propose a computer program, including computer-readable code, and when the computer-readable code runs in an electronic device, the processor in the electronic device executes the above method.

[0294] Embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code, and when the computer-readable code runs in an electronic device, the processor in the electronic device executes the above method.

[0295] Embodiments of the present disclosure also provide an electronic device, including: one or more processors; a memory for storing executable instructions; wherein, the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.

[0296] The electronic device may be provided as a terminal, a server or other forms of devices.

[0297] Figure 7 The block diagram of the electronic device 1900 provided by the embodiments of the present disclosure is shown. For example, the electronic device 1900 may be provided as a terminal device or a server. Referring to Figure 7 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to execute the above method.

[0298] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as the Microsoft server operating system (Windows Server TM ), the graphical user interface-based operating system launched by Apple Inc. (MacOS X TM ), the multi-user and multi-process computer operating system (UnixTM ),a free and open-source Unix-like operating system (Linux TM ),an open-source Unix-like operating system (FreeBSD TM ) or the like.

[0299] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, and the computer program instructions can be executed by a processing component 1922 of the electronic device 1900 to complete the above method.

[0300] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0301] The computer-readable storage medium may be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as a transitory signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0302] The computer-readable program instructions described herein may be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0303] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0304] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.

[0305] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions comprises a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0306] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0307] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.

[0308] The computer program product may be implemented specifically in the form of hardware, software, or a combination thereof. In an alternative embodiment, the computer program product is embodied as a computer storage medium. In another alternative embodiment, the computer program product is embodied as a software product, such as a Software Development Kit (SDK), etc.

[0309] The above descriptions of the various embodiments tend to emphasize the differences between the embodiments. The similarities or likenesses among them can be referred to each other. For the sake of brevity, they will not be elaborated herein.

[0310] If the technical solution of the embodiments of the present disclosure involves personal information, the product using the technical solution of the embodiments of the present disclosure has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of the embodiments of the present disclosure involves sensitive personal information, the product using the technical solution of the embodiments of the present disclosure has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to collect his or her personal information; or on the device for processing personal information, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0311] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A collaborative exploration method for the mapping scheme and architecture of a die accelerator, characterized in that Including: Obtaining candidate values of configurable architecture parameters of the chiplet accelerator, framework setting information, and a neural network model; Generating a candidate cooperation scheme of the layer-pipeline space mapping scheme and the architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model; Determining the monetary cost, power consumption cost, and latency of the candidate cooperation scheme; Determining an evaluation value of the candidate cooperation scheme according to the monetary cost, power consumption cost, and latency of the candidate cooperation scheme.

2. The method according to claim 1, wherein The configurable architecture parameters include at least some of the following: the bandwidth of the network-on-chip, the die-to-die communication bandwidth, the total bandwidth of the memory, the total number of computing cores in the X direction in the chiplet accelerator, the total number of computing cores in the Y direction in the chiplet accelerator, the number of chiplets divided in the X direction in the chiplet accelerator, the number of chiplets divided in the Y direction in the chiplet accelerator, the number of multiply-accumulate operations in the processing unit array of the computing core, and the size of the global buffer of the computing core.

3. The method according to claim 1, wherein Obtaining a neural network model includes: obtaining a plurality of neural network models; The generating a candidate cooperation scheme of the layer-pipeline space mapping scheme and the architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model includes: generating a candidate cooperation scheme of the layer-pipeline space mapping scheme and the architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and any one of the plurality of neural network models.

4. The method according to claim 1, characterized in that The layer-pipeline space mapping scheme in any candidate cooperation scheme includes the mapping schemes of each network layer in the neural network model. The mapping scheme of any network layer includes a partitioning attribute, a core group attribute, and a data flow attribute. Among them, the partitioning attribute is used to partition the input feature map, weight, and output feature map of the network layer. The core group attribute represents the information of the computing cores used to calculate the network layer. The data flow attribute represents the data source of the input feature map of the network layer, the weight, and the destination of the output feature map of the network layer.

5. The method according to claim 4, characterized in that, The partitioning attribute in the mapping scheme of any network layer includes the number of partitions in the height dimension of the output feature map, the number of partitions in the width dimension of the output feature map, the number of partitions in the batch size dimension, and the number of partitions in the channel number dimension of the output feature map.

6. The method according to claim 4 or 5, characterized in that, The generating a candidate cooperation scheme of the layer-pipeline space mapping scheme and the architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model includes: Generating a candidate cooperation scheme of the layer-pipeline space mapping scheme and the architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model, and using a preset simulated annealing operator; Wherein, the preset simulated annealing operator includes at least one of the following: The first simulated annealing operator is used to randomly select a network layer and change the partitioning attribute in the mapping scheme of the network layer; The second simulated annealing operator is used to randomly select a network layer and randomly swap two computing cores in the kernel group attributes of the mapping scheme of the network layer; The third simulated annealing operator is used to randomly select two network layers and randomly swap two computing cores in the kernel group attributes of the mapping schemes of the two network layers; The fourth simulated annealing operator is used to randomly select two network layers, remove a computing core from the kernel group attributes of the mapping scheme of one of the two network layers, and add the removed computing core to the kernel group attributes of the mapping scheme of the other network layer of the two network layers; The fifth simulated annealing operator is used to randomly select a network layer, randomly select a non - negative term in the data flow attributes of the mapping scheme of the network layer, and randomly determine the updated value of the non - negative term within the value range [0, D].

7. The method according to any one of claims 1 to 5, characterized in that, Determining the power consumption cost of the candidate cooperation scheme includes: Determining the power consumption cost of the candidate cooperation scheme according to the number of operations of each component in the architecture of the candidate cooperation scheme and the corresponding unit power consumption.

8. The method according to any one of claims 1 to 5, characterized in that Determining the latency of the candidate cooperation scheme includes: Determining the computing time of the multiply - accumulate operations in the candidate cooperation scheme; Determining the ratio of the maximum value of the data access volume of the memory in the candidate cooperation scheme to the access bandwidth; Determining the latency of the candidate cooperation scheme according to the computing time and the ratio.

9. The method according to claim 1 or 2, characterized in that, The determining the monetary cost of the candidate cooperation scheme includes: Determining the monetary cost of the candidate cooperation scheme according to the values of the configurable architecture parameters in the candidate cooperation scheme.

10. The method according to claim 9, wherein The monetary cost of the candidate cooperation scheme includes at least one of the following: The die manufacturing cost of the candidate cooperation scheme, the memory cost of the candidate cooperation scheme, the packaging cost of the candidate cooperation scheme.

11. The method according to claim 10, characterized in that, The determining the monetary cost of the candidate cooperation scheme according to the values of the configurable architecture parameters in the candidate cooperation scheme includes: For any die in the architecture of the candidate cooperation scheme, determining the silicon area of the die and the yield of the die, and obtaining the unit area manufacturing cost of the process technology node corresponding to the die; Determining the die manufacturing cost of the die according to the silicon area of the die, the yield of the die, and the unit area manufacturing cost of the process technology node corresponding to the die.

12. The method according to claim 11, wherein, Determining the yield of the die includes: Determining the yield of the die according to the silicon area of the die, the defect density, and the clustering parameter.

13. The method according to claim 10, wherein The determining the monetary cost of the candidate cooperation scheme according to the values of the configurable architecture parameters in the candidate cooperation scheme includes: Determining the memory cost of the candidate cooperation scheme according to the bandwidth of the memory, the bandwidth of the memory unit, and the monetary cost of the memory unit in the architecture of the candidate cooperation scheme.

14. The method according to claim 10, wherein The determining the monetary cost of the candidate cooperation scheme according to the values of the configurable architecture parameters in the candidate cooperation scheme includes: Determining the packaging cost of the candidate cooperation scheme according to the total silicon area of all dies in the architecture of the candidate cooperation scheme, the scaling factor, the packaging yield, and the unit area monetary cost of the substrate.

15. The method according to claim 1, wherein Determining an evaluation value of the candidate cooperation solution according to the currency cost, power consumption cost, and latency of the candidate cooperation solution includes: Obtaining a first weight corresponding to the currency cost, a second weight corresponding to the power consumption cost, and a third weight corresponding to the latency; Determining the evaluation value of the candidate cooperation solution according to the currency cost, power consumption cost, and latency of the candidate cooperation solution, and the first weight, the second weight, and the third weight.

16. The method according to claim 1, characterized in that, After determining the evaluation value of the candidate cooperation solution, the method further includes: Determining an optimal cooperation solution of the chiplet accelerator according to the evaluation values of multiple candidate cooperation solutions.

17. The method according to claim 1, wherein The chiplet accelerator is an inference accelerator for a deep neural network.

18. A co - exploration device for the mapping scheme and architecture of a die accelerator, characterized in that, Including: An acquisition module, configured to acquire candidate values of configurable architecture parameters of the chiplet accelerator, framework setting information, and a neural network model; A generation module, configured to generate a candidate cooperation solution of the layer-pipeline space mapping scheme and architecture of the chiplet accelerator according to the candidate values of the configurable architecture parameters, the framework setting information, and the neural network model; A first determination module, configured to determine the currency cost, power consumption cost, and latency of the candidate cooperation solution; A second determination module, configured to determine the evaluation value of the candidate cooperation solution according to the currency cost, power consumption cost, and latency of the candidate cooperation solution.

19. An electronic device, characterized in that, Including: One or more processors; A memory for storing executable instructions; Wherein, the one or more processors are configured to call the executable instructions stored in the memory to execute the method according to any one of claims 1 to 17.

20. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 17 is implemented.