Compiling method and system for near storage computing architecture, electronic equipment and storage medium
By generating a tree-like architecture abstract model and building a search space for compilation strategies, the compilation method for near-storage computing architecture solves the problem that the toolchain cannot adapt to different architectures, and realizes the universality and efficiency of deploying neural network models on any architecture.
Patent Information
- Application Number
- CN202510223888.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-27
AI Technical Summary
The toolchain that compiles and deploys neural network models on the existing near-storage computing architecture cannot adapt to different types of near-storage computing architectures, and cannot provide suitable optimal compilation strategies.
A compilation method and system for near-storage computing architecture is proposed. By obtaining the operator-level intermediate representation of the neural network model and the hardware configuration information of the near-storage computing architecture, a tree-like architecture abstract model is generated, a search space for compilation strategies is constructed, performance prediction is performed, candidate compilation strategies are determined, and target compilation strategies are selected through simulation simulation.
It realizes the universality of deploying any neural network model on any near-storage computing architecture, provides better target compilation strategies, and improves the efficiency and adaptability of compilation deployment.
Smart Images

Figure CN120144134A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of near-memory computing, and in particular, to a compilation method and system, an electronic device, and a storage medium for a near-memory computing architecture. Background Art
[0002] The near-memory computing architecture is an architecture that integrates a memory and a computing unit to reduce data transfer latency and improve computing efficiency. By placing the computing unit near the memory, this computing architecture maximizes the use of the high access bandwidth of the internal memory using high-bandwidth circuit integration technology, thereby achieving efficient data processing.
[0003] With the development of deep learning technology, how to compile and deploy neural network models on various hardware architectures is an important research direction. Currently, the toolchain for compiling and deploying neural network models on a near-memory computing architecture generally includes a compiler and a simulator. The compiler is responsible for converting the algorithm input into instructions that can be executed on the near-memory computing architecture, and the simulator is responsible for comparing the advantages and disadvantages of different compilation strategies that may be involved during the compilation process to select a suitable compilation strategy.
[0004] However, the current toolchains for compiling and deploying neural network models on a near-memory computing architecture are designed specifically for different near-memory computing architectures, which limits the compilation and deployment tools to a single type or a single level of near-memory computing architecture. They cannot adapt to different types of near-memory computing architectures and cannot provide an optimal compilation strategy adapted to different near-memory computing architectures. Summary of the Invention
[0005] In view of this, the present disclosure provides a compilation method and system, an electronic device, and a storage medium for a near-memory computing architecture, which can provide an adapted target compilation strategy for the deployment of any neural network model on any near-memory computing architecture, thereby realizing the deployment of any neural network model on any near-memory computing architecture and having high versatility.
[0006] According to one aspect of the present disclosure, a compilation method is provided, including: obtaining an operator-level intermediate representation of a neural network model to be deployed to a near-memory computing architecture and hardware configuration information of the near-memory computing architecture, where the hardware configuration information is used to indicate the hardware configuration of the near-memory computing architecture; generating a tree-like architecture abstraction model of the near-memory computing architecture based on the hardware configuration information of the near-memory computing architecture, where the tree-like architecture abstraction model is used to describe the near-memory computing architecture in a tree structure; constructing a search space of compilation strategies based on the operator-level intermediate representation of the neural network model and the tree-like architecture abstraction model of the near-memory computing architecture, where the compilation strategies are used to compile the operator-level intermediate representation; determining a plurality of candidate compilation strategies and compilation results corresponding to each candidate compilation strategy by performing performance prediction on the compilation strategies in the search space, where the compilation results include a complete instruction sequence obtained by compiling the operator-level intermediate representation according to the candidate compilation strategy; performing a simulation on the execution process of the compilation results corresponding to each candidate compilation strategy on the near-memory computing architecture based on the tree-like architecture abstraction model to obtain simulation results corresponding to each candidate compilation strategy, where the simulation results represent the total time required to complete the execution of the compilation results on the near-memory computing architecture; determining a target compilation strategy from the plurality of candidate compilation strategies according to the simulation results corresponding to each candidate compilation strategy, and determining the compilation result of the target compilation strategy as the target compilation result of the operator-level intermediate representation of the neural network model, so as to deploy the neural network model on the near-memory computing architecture based on the target compilation result.
[0007] In a possible implementation, the near-memory computing architecture includes multiple levels of hardware structures. One level of hardware structure in the near-memory computing architecture corresponds to one layer of nodes in the tree-like architecture abstraction model. Each layer in the tree-like architecture abstraction model includes at least one central node, and each central node is connected to at least one of the following types of child nodes: central nodes of the next level, storage nodes, computing nodes, and cache nodes in the level where the central node belongs; the central node represents the data transmission relationship between two adjacent levels; where the storage node represents the storage unit of the memory in the near-memory computing architecture and includes attribute information of the storage unit, the computing node represents the computing unit in the near-memory computing architecture and includes attribute information of the computing unit, and the cache node represents the cache unit in the near-memory computing architecture and includes attribute information of the cache unit.
[0008] In a possible implementation, a single compilation strategy includes a partitioning strategy and a corresponding mapping strategy. The partitioning strategy represents the partitioning result of partitioning the operator-level intermediate representation along multiple dimensions at multiple levels. The mapping strategy represents the way of mapping the partitioning result of the operator-level intermediate representation to the memory in the near-memory computing architecture. The partitioning result represents multiple sub-intermediate representations obtained by partitioning the operator-level intermediate representation; wherein, constructing a search space for the compilation strategy based on the operator-level intermediate representation of the neural network model and the tree-like architecture abstraction model of the near-memory computing architecture includes: based on the operator-level intermediate representation and the tree-like architecture abstraction model, in the order from the highest level to the lowest level of the memory in the near-memory computing architecture, partitioning the operator-level intermediate representation along multiple dimensions at multiple levels under the satisfaction of the first constraint condition to obtain multiple partitioning strategies; wherein, the first constraint condition includes that the product of the number of partitions along each dimension among multiple dimensions at each level is less than or equal to the number of hardware structures corresponding to each level in the memory, and the partitioning methods of different partitioning strategies are different under the first constraint condition; for each sub-intermediate representation partitioned by each partitioning strategy, partitioning each sub-intermediate representation into multiple data slices under the satisfaction of the second constraint condition, and mapping each data slice to the memory in the near-memory computing architecture under the satisfaction of the third constraint condition to obtain multiple mapping strategies corresponding to each partitioning strategy; wherein, the second constraint condition includes: the product of the size of each data slice and the data bit-width of each data slice is less than or equal to the product of the burst length corresponding to the memory and the column bit-width of the memory; the third constraint condition includes: the number of data slices mapped to the same row in the memory is less than or equal to the ratio of the total number of columns in that row in the memory to the burst length.
[0009] In a possible implementation, determining multiple candidate compilation strategies and the compilation results corresponding to each candidate compilation strategy by performing performance prediction on the compilation strategies in the search space includes: determining the performance upper limit corresponding to each compilation strategy in the search space based on the hardware data corresponding to the near-memory computing architecture, where the performance upper limit represents the upper limit of the duration required for the compilation result corresponding to the compilation strategy to be executed and completed on the near-memory computing architecture, and the hardware data includes the bandwidth of the memory, the row miss penalty, the number and performance of the computing units in the near-memory computing architecture; determining multiple initial candidate compilation strategies whose performance upper limits are higher than a preset performance threshold based on the performance upper limits corresponding to each compilation strategy in the search space; using a compiler to perform partial compilation on the operator-level intermediate representation based on each initial candidate compilation strategy and a specific instruction set to obtain partial instruction sequences compiled by each initial candidate compilation strategy, where the specific instruction set covers instructions executed on various near-memory computing architectures; using a performance predictor to perform performance prediction on the partial instruction sequences compiled by each initial candidate compilation strategy to obtain the predicted performance of the partial instruction sequences compiled by each initial candidate compilation strategy, where the predicted performance represents the duration required for the partial instruction sequences to be executed and completed on the near-memory computing architecture; determining the target predicted performance corresponding to each initial candidate compilation strategy based on the predicted performance of the partial instruction sequences compiled by each initial candidate compilation strategy and the number of loop executions corresponding to the partial instruction sequences, where the target predicted performance represents the duration required for the complete instruction sequence compiled based on the initial candidate compilation strategy to be executed and completed on the near-memory computing architecture; selecting multiple candidate compilation strategies from the multiple initial candidate compilation strategies according to the target predicted performance corresponding to each initial candidate compilation strategy; using the compiler to perform complete compilation on the operator-level intermediate representation based on the multiple candidate compilation strategies and the specific instruction set respectively to obtain the compilation results corresponding to each candidate compilation strategy.
[0010] In a possible implementation, determining the performance upper limit corresponding to each compilation strategy in the search space based on the hardware data corresponding to the near-memory computing architecture includes: analyzing each compilation strategy in the search space to obtain the analysis result corresponding to each compilation strategy, where the analysis result includes the total access data volume and the number of row misses to the memory under each compilation strategy, and the load calculation amount of the computing unit; determining the performance upper limit corresponding to each compilation strategy in the search space according to the analysis result corresponding to each compilation strategy and the hardware data.
[0011] In a possible implementation, using the performance predictor to perform performance prediction on partial instruction sequences compiled by each initial candidate compilation strategy to obtain the predicted performance of the partial instruction sequences compiled by each initial candidate compilation strategy includes: determining, based on the partial instruction sequences compiled by each initial candidate compilation strategy, the total number of accesses to the memory in the near-memory computing architecture, the number of row misses in continuously accessing the memory, and the amount of data accessed by the computing units in the near-memory computing architecture to the cache unit or register in the partial instruction sequences compiled by each initial candidate compilation strategy; and determining the predicted performance of the partial instruction sequences compiled by each initial candidate compilation strategy according to the total number of accesses to the memory in the near-memory computing architecture and the column-to-column latency, the number of row misses in continuously accessing the memory and the row miss penalty, and the amount of data accessed by the computing units in the near-memory computing architecture to the cache unit or register and the bus bandwidth in the partial instruction sequences compiled by each initial candidate compilation strategy.
[0012] In a possible implementation, the method uses a simulator to perform simulation on the execution process of the compilation results corresponding to each candidate compilation strategy in the near-memory computing architecture based on the tree-like architecture abstraction model to obtain the simulation results corresponding to each candidate compilation strategy; wherein, the simulator includes: a hardware state space, an instruction queue, and a virtual memory controller; wherein, the hardware state space is used to simulate and track the hardware states of each hardware component in the near-memory computing architecture based on the tree-like architecture abstraction model, and the hardware components include a storage unit, a computing unit, a cache unit, and a transmission bus; the instruction queue is used to store the complete instruction sequences compiled based on the candidate compilation strategies and maintain the data dependencies in the complete instruction sequences, the instruction queue includes multiple instruction groups, instructions independent of each other executed by different computing units are located in different instruction groups, each instruction group includes multiple sequentially executed instructions, and instructions in instruction groups without data dependencies can be issued in parallel; the virtual content controller is used to determine the issue time of the instructions in the instruction queue according to the hardware state space and control the update time of the hardware states of each hardware component in the hardware state space.
[0013] In a possible implementation, tracking the hardware states of the respective hardware components in the near-memory computing architecture includes: tracking the activation states of the respective storage units in the memory of the near-memory computing architecture to issue an activation command before an unactivated storage unit is accessed; tracking the currently activated row address in the storage unit to indicate whether a subsequent access causes a row miss and scheduling a precharge command and an activation command to activate the row to be accessed when a row miss occurs; wherein, multiple clock countdowns are used in the hardware state space to respectively track the hardware states of the respective hardware components, wherein the hardware state of the storage unit includes the time when the next command can be issued for the storage unit, and the hardware states of the other hardware components except the storage unit include an idle state or an occupied state; the clock countdowns corresponding to the respective hardware components are updated according to the hardware occupation duration of the hardware components involved in issuing the instructions.
[0014] In a possible implementation, using the emulator to simulate the execution process of the compilation results corresponding to the respective candidate compilation strategies in the near-memory computing architecture based on the tree-like architecture abstraction model to obtain the simulation results corresponding to the respective candidate compilation strategies includes: for the compilation result corresponding to any candidate compilation strategy, cyclically performing the following simulation operations: the instruction queue determines the set of instructions that can be simultaneously issued currently based on the data dependencies involved in the respective instruction groups, and sends the set of instructions to the virtual memory controller, and the set of instructions includes the instructions that are currently ranked first in the respective instruction groups without data dependencies; the virtual content controller determines the issue time of each instruction in the set of instructions according to the hardware state space, and selects the instruction with the earliest issue time as the target instruction to be issued currently; the virtual content controller controls the hardware states of the respective hardware components in the hardware state space to be updated to the hardware states corresponding to the issue time of the target instruction based on the issue time of the target instruction, and deletes the target instruction from the instruction queue; the virtual content controller updates the hardware states of the hardware components occupied by the target instruction in the hardware state space according to the hardware occupation duration of the hardware components occupied by the target instruction; in the case where all the instructions in the instruction queue have been issued and the hardware states of the respective hardware components in the hardware state space become idle states, the simulation result corresponding to the candidate compilation strategy is obtained.
[0015] In a possible implementation, determining the issue time of each instruction in the instruction set according to the hardware state space includes: for the i-th instruction in the instruction set, when the operand of the i-th instruction comes from a storage unit, based on the hardware state space, determining the activation state of the row where the operand of the i-th instruction is located in the storage unit, and according to the activation state of the row where the operand of the i-th instruction is located, the clock countdown of the storage unit, and the read latency for reading the operand from the storage unit, determining the arrival time of the operand of the i-th instruction at the computing unit; when the operand of the i-th instruction comes from a cache unit, determining the arrival time of the operand of the i-th instruction at the computing unit according to the release time of the cache unit, the read latency for reading the operand from the cache unit, and the idle time of the transmission bus between the cache unit and the computing unit; according to the arrival time of the operand of the i-th instruction at the computing unit and the idle time of the computing unit, determining the start processing time when the computing unit starts to process the operand of the i-th instruction; and based on the start processing time when the computing unit starts to process the operand of the i-th instruction, determining the issue time of the i-th instruction.
[0016] In a possible implementation, selecting the instruction with the earliest issue time as the target instruction to be issued currently includes: selecting the instruction with the earliest start processing time as the target instruction to be issued currently.
[0017] According to another aspect of the present disclosure, there is provided a compilation system, including: an acquisition module, configured to acquire an operator-level intermediate representation of a neural network model to be deployed to a near-memory computing architecture and hardware configuration information of the near-memory computing architecture, where the hardware configuration information is used to indicate the hardware configuration of the near-memory computing architecture; an abstraction module, configured to generate a tree-structured architecture abstraction model of the near-memory computing architecture based on the hardware configuration information of the near-memory computing architecture, where the tree-structured architecture abstraction model is used to describe the near-memory computing architecture in a tree structure; a construction module, configured to construct a search space of a compilation strategy based on the operator-level intermediate representation of the neural network model and the tree-structured architecture abstraction model of the near-memory computing architecture, where the compilation strategy is used to compile the operator-level intermediate representation; a prediction module, configured to determine multiple candidate compilation strategies and compilation results corresponding to each candidate compilation strategy by performing performance prediction on the compilation strategies in the search space, where the compilation results include a complete instruction sequence obtained by compiling the operator-level intermediate representation according to the candidate compilation strategy; a simulation module, configured to perform a simulation on the execution process of the compilation results corresponding to each candidate compilation strategy on the near-memory computing architecture based on the tree-structured architecture abstraction model, to obtain simulation results corresponding to each candidate compilation strategy, where the simulation results represent the total time required for the compilation results to be executed and completed on the near-memory computing architecture; a determination module, configured to determine a target compilation strategy from the multiple candidate compilation strategies according to the simulation results corresponding to each candidate compilation strategy, and determine the compilation result of the target compilation strategy as the target compilation result of the operator-level intermediate representation of the neural network model, so as to deploy the neural network model on the near-memory computing architecture based on the target compilation result.
[0018] According to another aspect of the present disclosure, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the above method.
[0019] According to another aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which a computer program is stored, where the computer program, when executed by a processor, implements the steps of the above method.
[0020] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, or a non-volatile computer-readable storage medium carrying the computer program, where the computer program, when executed by a processor, implements the steps of the above method.
[0021] According to aspects of the present disclosure, by using a tree - structured abstract model to describe a near - memory computing architecture, any near - memory computing architecture can be uniformly described in terms of time series using the tree - structured model. Furthermore, by combining the operator - level intermediate representation of a neural network model, a compilation strategy search space is constructed. Then, by predicting the performance of the compilation strategies in the search space, a plurality of candidate compilation strategies are determined. Subsequently, based on the abstract model, the emulator simulates the execution process of the compilation results on the near - memory computing architecture to obtain accurate simulation results corresponding to each candidate compilation strategy, and then determines the final target compilation strategy and target compilation result. This can provide an optimal target compilation strategy for deploying any neural network model on any near - memory computing architecture, thereby realizing the deployment of any neural network model on any near - memory computing architecture, and having high generality.
[0022] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and are used to explain the principles of the present disclosure together with the specification.
[0024] Figure 1 The flowchart showing a compilation method according to an embodiment of the present disclosure is shown.
[0025] Figure 2a 、 Figure 2b and Figure 2c Schematic diagrams showing three NDP architectures according to an embodiment of the present disclosure are shown respectively.
[0026] Figure 3a 、 Figure 3b and Figure 3c Schematic diagrams showing three tree - structured architecture abstract models according to an embodiment of the present disclosure are shown respectively.
[0027] Figure 4 The schematic flowchart showing operator partitioning and operator mapping according to an embodiment of the present disclosure is shown.
[0028] Figure 5 The schematic flowchart showing search space pruning and performance prediction according to an embodiment of the present disclosure is shown.
[0029] Figure 6 The schematic diagram showing an emulator architecture according to an embodiment of the present disclosure is shown.
[0030] Figure 7a and Figure 7b Schematic diagrams showing two DRAM access processes according to an embodiment of the present disclosure are shown respectively.
[0031] Figure 8 Shows a schematic framework diagram of a compilation method according to an embodiment of the present disclosure.
[0032] Figure 9 Shows a block diagram of a compilation system according to an embodiment of the present disclosure.
[0033] Figure 10 Shows a block diagram of an electronic device 1900 according to an embodiment of the present disclosure. Detailed implementation manners
[0034] The following will detail various exemplary embodiments, features, and aspects of the present disclosure with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.
[0035] As used herein, the terms "comprising", "including", "having", or variations thereof are open-ended and include one or more stated features, wholes, elements, steps, components, or functions, but do not exclude the presence or addition of one or more other features, wholes, elements, steps, components, functions, or groups thereof.
[0036] When an element is referred to as being "connected", "coupled", "responsive" or variations thereof to another element, it can be directly connected, coupled, or responsive to the other element, or intervening elements may be present.
[0037] Although the terms first, second, third, etc. may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Thus, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.
[0038] The term "exemplary" as used herein means "serving as an example, embodiment, or illustration". Any embodiment illustrated herein as "exemplary" should not necessarily be construed as being superior or better than other embodiments.
[0039] In addition, for a better illustration of the present disclosure, numerous specific details are given in the following detailed implementation manners. Those skilled in the art should understand that the present disclosure can be implemented without some specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.
[0040] Figure 1 Shows a flowchart of a compilation method according to an embodiment of the present disclosure. As Figure 1 shown, the method includes: steps S11 to step S16.
[0041] In step S11, obtain the operator-level intermediate representation of the neural network model to be deployed to the near-memory computing architecture and the hardware configuration information of the near-memory computing architecture, where the hardware configuration information is used to indicate the hardware configuration of the near-memory computing architecture.
[0042] In practical applications, for example, the ONNX (Open Neural Network Exchange, an open-source deep learning model exchange format) framework can be used to convert some or all of the operators in the neural network model into a series of operator-level intermediate representations (operator-level IR, Intermediate Representation); of course, those skilled in the art can also use other known conversion methods in the art to convert them into the corresponding operator-level intermediate representations, and the embodiments of the present disclosure do not limit this.
[0043] Among them, the hardware configuration information may include the architecture configuration of the overall near-memory computing architecture and the timing parameters of the memory in the near-memory computing architecture. Among them, the timing parameters are a series of parameters describing the performance of the memory. These parameters are in clock cycles and specify the timing requirements for read and write operations to the storage unit. Exemplarily, taking the memory in the near-memory computing architecture as DRAM (Dynamic Random Access Memory), the timing parameters of DRAM may include tRCD (RAS to CAS Delay), the delay from the row address to the column address, that is, the time from activating the row to issuing the column address; tCL (CAS Latency), the column address strobe delay, that is, the time from issuing the column address to data output; tRP (RAS Precharge Time), the row activation time, that is, the shortest time from row activation to precharge; tRAS (RAS Active Time), the row activation time, that is, the shortest time from row activation to precharge; tRC (Row Cycle Time), the row cycle time, that is, the shortest time between two row activations; tWR (Write Recovery Time), the write recovery time, that is, the time from the completion of the write operation to precharge; tRRD (RAS to RAS Delay), the row-to-row delay, that is, the shortest time between two row activations; tWTR (Write to Read Delay), the write-to-read delay, that is, the time from the completion of the write operation to the start of the read operation; tCCD (CAS to CAS Delay), the column-to-column delay, that is, the shortest time between two column operations; tRTP (Read to Precharge Delay), the read-to-precharge delay, that is, the time from the completion of the read operation to precharge; tRTP (Read to Precharge Delay), the read-to-precharge delay, that is, the time from the completion of the read operation to precharge; tPD (Precharge Delay), the precharge delay, that is, the time from issuing the precharge command to execution; tDAL (Data Access Latency), the data access delay, that is, the time from issuing the read command to data output; tDPL (Data Precharge Latency), the data precharge delay, that is, the time from issuing the precharge command to the completion of data precharge, etc. It should be understood that if other types of memories are used in the near-memory computing architecture, the corresponding timing parameters of the memory can be obtained, and the embodiments of the present disclosure do not limit this.
[0044] It should be understood that the architectural configurations of different near-memory computing architectures are different. Generally speaking, the architectural configuration may at least include: hierarchical structure information of the memory (such as the number of channels in DRAM, the number of column ranks in each channel, the number of chips (devices) in each rank, the number of memory banks in each device), the number of rows and columns and bit width of the memory cells in the memory, the level where the near-memory cells (i.e., including memory cells and computing units) are located, the number of computing units, computing parallelism and computing precision, design information of the cache unit (such as cache size, read / write latency, etc.), etc.
[0045] In the embodiments of the present disclosure, a memory cell may be the smallest memory cell corresponding to a computing unit in the memory. For example, in a bank-level NDP (Near DRAM Processing) architecture, the computing unit corresponds to a memory bank in DRAM, then a memory cell may represent a bank. If the computing unit in a device-level NDP architecture corresponds to a chip (device) in DRAM, then a memory cell may represent a device. This depends on the structural layout of the computing unit and the memory cell in the NDP architecture, and the embodiments of the present disclosure do not limit this; among them, the NDP structure is a near-memory computing architecture.
[0046] Exemplarily, Figure 2a A schematic diagram of a bank-level NDP architecture is shown. As Figure 2a shown, it may be a certain device under a certain rank of a certain channel. This device includes multiple bank-level near-memory cells, and each bank-level near-memory cell includes a Bank PU (i.e., a bank-level computing unit), a local cache unit (i.e., an input cache and an output cache), and a corresponding bank-level memory cell; Figure 2b A schematic diagram of a device-level NDP architecture is shown. As Figure 2b shown, each device may include multiple Device PUs (i.e., device-level computing units), multiple Device PUs share the same global cache, and a bank-level memory cell corresponding to each Device PU; Figure 2c A schematic diagram of another device-level NDP architecture is shown. As Figure 2c shown, a rank includes multiple device-level near-memory cells, and each device-level near-memory cell includes a Device PU, a local cache unit (i.e., an input cache and an output cache), and a corresponding device-level memory cell.
[0047] It should be understood that the hierarchical structures and layouts of different near-memory computing architectures can be different. For example, in some near-memory computing architectures, each Bank further includes multiple groups, so the hierarchical structure information of the above-mentioned memory can further include the number of groups in each Bank, and the computing unit can also correspond to the group-level storage unit in the Bank; in some near-memory computing architectures, the memory does not include the chip Device level. Those skilled in the art can obtain the corresponding hardware configuration information according to the actual type, structure, layout, etc. of the near-memory computing architecture, and the present disclosure does not limit the manner of obtaining the hardware configuration information.
[0048] In step S12, based on the hardware configuration information of the near-memory computing architecture, a tree-like architecture abstraction model of the near-memory computing architecture is generated, and the tree-like architecture abstraction model is used to describe the near-memory computing architecture in a tree structure.
[0049] In order to make the compilation and deployment of neural network models generally applicable to various near-memory computing architectures, the embodiments of the present disclosure design a unified architecture abstraction method, that is, using an abstract model with a tree structure to describe different near-memory computing architectures, which is equivalent to using an extensible tree structure to abstract near-memory computing architectures. This description method can be used to describe existing near-memory computing architectures and can also be easily extended to support new near-memory computing architectures. As described above, the near-memory computing architecture includes multiple levels of hardware structures, so a hardware structure at one level in the near-memory computing architecture can correspond to a layer of nodes in the tree-like architecture abstraction model, or rather, each layer of the tree-like architecture abstraction model can represent a hierarchical structure of the near-memory computing architecture (that is, a hierarchical structure of the memory in the near-memory computing architecture). For example, a certain near-memory computing architecture includes four hierarchical structures: channel, rank, device, and bank, then the tree-like architecture abstraction model can include four layers of nodes (specifically, four central nodes). If a new near-memory computing architecture includes a new hierarchical structure (such as group), new node levels can be introduced in the abstract model for extension, and if a certain hierarchical structure does not exist in the near-memory computing architecture, the corresponding layer of nodes can also be removed. For example, HBM (High Bandwidth Memory) does not include the chip device level, so the corresponding tree-like architecture abstraction model does not include nodes at this level.
[0050] Among them, each layer in the tree - like architecture abstraction model contains at least one central node, and each central node is connected to at least one of the following types of child nodes: central nodes of the next level (non - bottom - layer nodes), storage nodes (bottom - layer nodes), computing nodes (bottom - layer nodes), and cache nodes (bottom - layer nodes) in the level where the central node belongs; among them, the central node represents the data transmission relationship between two adjacent levels. Specifically, the central node can represent a data transmission interface from the lower level to the upper level (such as a data bus), and the data transmission interface can describe the data transmission relationship. Among them, the non - bottom - layer node means that this node is not the last - layer node, and the bottom - layer node means that this node is the last - layer node.
[0051] Among them, the storage node represents the storage unit of the memory in the near - memory computing architecture and contains the attribute information of the storage unit, that is, the storage node is the smallest storage unit corresponding to the processing unit. For example, in the bank - level NDP architecture, each storage node can represent a bank, and the attribute information contained in the storage node can at least include: read / write bit width, capacity, and timing parameters of the storage unit, etc.
[0052] Among them, the computing node represents the computing unit in the near - memory computing architecture and contains the attribute information of the computing unit, that is, the computing node can represent the computing unit of the corresponding level of the node. The attribute information contained in the computing node can at least include the computing parallelism, computing precision, working frequency, etc. of the computing unit. In addition, if a central node contains multiple storage nodes and computing nodes, then each computing node under the same central node can read / write data from multiple storage nodes under the same central node, and the storage nodes accessible to the computing nodes under the same central node can be configured by the user according to actual needs.
[0053] Among them, the cache node represents the cache unit in the near - memory computing architecture and contains the attribute information of the cache unit. This cache unit can be a global cache or a local cache, and the local cache is a non - globally - shared unit; the attribute information contained in the cache node can at least include: cache size (width and depth) and read / write latency of the cache unit, etc. It should be understood that the cache node can be configured as the local cache of the computing nodes under the same central node, or can be configured as the global cache shared by multiple processing units under different central nodes. Thus, the computing nodes under the same central node can share other cache nodes (i.e., local cache) under the same central node, while the computing units from different central nodes need to use the cache nodes of the upper level (i.e., global cache) for data sharing.
[0054] It should be understood that the computing nodes and cache nodes can be optional in the tree - structured abstract model, that is, some layers of the tree - structured abstract model may not contain computing nodes and / or cache nodes. By adjusting the configuration of nodes, the number of central nodes, the number of storage nodes, the number of computing nodes, and the number of cache nodes in each layer of the tree - structured abstract model, the generated tree - structured abstract model can cover various types of near - memory computing architectures.
[0055] Exemplarily, Figure 3a 、 Figure 3b and Figure 3c respectively show schematic diagrams of three tree - structured abstract models. Among them, Figure 3a is obtained by abstracting the bank - level NDP architecture shown in Figure 2a . Figure 3b is obtained by abstracting the device - level NDP architecture shown in Figure 2b . Figure 3c is obtained by abstracting the device - level NDP architecture shown in Figure 2c . Among them, the circle represents the central node, the triangle represents the storage node, the square represents the cache node, and the heptagon represents the computing node; from Figure 3a 、 Figure 3b and Figure 3c , it can be seen that the tree - structure can be used to uniformly describe any near - memory computing architecture, which is beneficial to making the subsequent compilation and deployment processes adapt to various near - memory computing architectures.
[0056] In step S13, based on the operator - level intermediate representation of the neural network model and the tree - structured abstract model of the near - memory computing architecture, a search space for the compilation strategy is constructed, and the compilation strategy is used to compile the operator - level intermediate representation.
[0057] Among them, a single compilation strategy can include a partitioning strategy, a corresponding mapping strategy, and a scheduling strategy. The partitioning strategy represents the partitioning result of partitioning the operator - level intermediate representation along multiple dimensions at multiple levels. The mapping strategy represents the way of mapping the partitioning result of the operator - level intermediate representation to the storage units in the near - memory computing architecture. The partitioning result represents multiple sub - intermediate representations obtained by partitioning the operator - level intermediate representation. The scheduling strategy represents the order of traversing and scheduling the multiple sub - intermediate representations obtained by partitioning the operator - level intermediate representation during the computing process. Since currently, usually a fixed number of several scheduling strategies are adopted, such as horizontal traversal, vertical traversal, Z - shaped traversal, etc., therefore, this embodiment of the present disclosure mainly details the generation of the partitioning strategy and the mapping strategy. The partitioning strategy and the mapping strategy in the same compilation strategy can be described as how to partition and allocate the operator - level intermediate representation to the processing units and storage units at multiple levels of the memory, or how to split the operator - level intermediate representation and map it to different columns and rows of the storage units, that is, the data layout.
[0058] To achieve high parallelism, existing NDP architectures typically partition data along matrix dimensions that do not require reduction, and the number of partitions is equal to the number of available processing units. Taking the matrix multiplication operator O = W × A commonly used in neural network models as an example, where the dimensions of the weight matrix W, the input activation matrix A, and the output matrix O are (M, K), (K, N), and (M, N) respectively. Existing partitioning strategies usually partition the operator-level intermediate representation of the above matrix multiplication operator along the M dimension. However, when the K dimension and the N dimension increase, the number of updates to the corresponding input cache also increases, resulting in more input data accessing the memory, thus hindering the realization of high computational parallelism.
[0059] Therefore, the embodiments of the present disclosure propose a new partitioning method that enables data to be partitioned along multiple dimensions and processed across multiple levels. The proposed partitioning encoding method sequentially describes how the operator-level intermediate representation (such as matrix multiplication operators and input / outputs) is partitioned at different memory (such as DRAM) levels, and the partitioning along different dimensions at each memory (such as DRAM) level, such as Figure 4 As shown, the operator-level IR can be sliced from the Channel level to the Bank level layer by layer through data slicing, and then each data slice (i.e., sub-operator) of the sub-intermediate representation of each block can be mapped to the storage array (i.e., storage unit) through data mapping (i.e., address assignment). Specifically, in a possible implementation manner, step S13, based on the operator-level intermediate representation of the neural network model and the tree-like architecture abstraction model of the near-memory computing architecture, constructing the search space of the compilation strategy may include:
[0060] Step S131, based on the operator-level intermediate representation and the tree-like architecture abstraction model, in the order from the highest level to the lowest level of the memory in the near-memory computing architecture, under the condition of satisfying the first constraint, partitioning the operator-level intermediate representation along multiple dimensions of the operator-level intermediate representation layer by layer to obtain multiple partitioning strategies; where the first constraint includes that the product of the number of partitions along each dimension among multiple dimensions at each level is less than or equal to the number of hardware structures corresponding to each level in the memory, and the partitioning methods of different partitioning strategies are different under the first constraint;
[0061] Step S132: For each sub-intermediate representation divided by each partitioning strategy, under the condition of satisfying the second constraint, each sub-intermediate representation is divided into multiple data slices, and under the condition of satisfying the third constraint, each data slice is mapped to a storage unit in the near-memory computing architecture, obtaining multiple mapping strategies corresponding to each partitioning strategy; wherein, the second constraint includes: the product of the size of each data slice and the data bit-width of each data slice is less than or equal to the product of the burst length corresponding to the memory and the column bit-width of the memory; the third constraint includes: the number of data slices mapped to the same row in the memory is less than or equal to the ratio of the total number of columns in that row of the memory to the burst length; the burst length is the number of columns in the memory continuously accessed by one access command.
[0062] Exemplarily, in the embodiments of the present disclosure, taking the compilation and deployment of a matrix multiplication operator on a bank-level NDP architecture with a parallelism of {#CH, #RA, #DE, #BA} as an example, the generation process of the above-mentioned partitioning strategy and mapping strategy is introduced, where #CH, #RA, #DE, and #BA respectively represent the number of channels, the number of ranks per channel, the number of devices per rank, and the number of banks per device; then the partitioning strategy can be expressed as:
[0063] Partitioning strategy = {Part CH (M CH , K CH , N CH ), Part RA (M RA , K RA , N RA ),
[0064] Part DE (M DE , K DE , N DE ), Part BK (M BK , K BK , N BK )};
[0065] Among them, the above-mentioned partitioning strategy represents a four-level partitioning of four DRAM levels. The partitioning starts from the highest channel level, and each level is partitioned along three dimensions M, K, and N. Specifically, Part CH (M CH , K CH , N CH ) means that the operator-level intermediate representation of the matrix multiplication operator is evenly divided into M CH , K CH and N CHcopies to obtain multiple intermediate representations at the channel level. Each intermediate representation at the channel level divided in the channel hierarchy will be assigned to a different channel, and it should satisfy M CH ×K CH ×N CH ≤ the number of channels to avoid allocating more channels than available. Part RA (M RA ,K RA ,N RA ) represents that each intermediate representation at the channel level is evenly divided into M RA 、K RA and N RA copies along the M, K, and N dimensions to obtain multiple intermediate representations at the rank level. Each intermediate representation at the rank level will be assigned to a different rank, and it should satisfy M RA ×K RA ×N RA ≤ the number of ranks in each channel to avoid allocating more ranks than available; Part DE (M DE ,K DE ,N DE ) represents that each intermediate representation at the rank level is evenly divided into M DE 、K DE and N DE copies along the M, K, and N dimensions to obtain multiple intermediate representations at the device level. Each intermediate representation at the device level will be assigned to a different device, and it should satisfy M DE ×K DE ×N DE ≤ the number of devices in each rank to avoid allocating more devices than available; Part BK (M BK ,K BK ,N BK ) represents that each intermediate representation at the device level is evenly divided into M BK 、K BK and N BK copies along the M, K, and N dimensions to obtain multiple intermediate representations at the bank level. Each intermediate representation at the bank level will be assigned to a different bank, and it should satisfy M BK ×K BK ×N BK ≤ the number of banks in each device to avoid allocating more banks than available.
[0066] It can be seen that from the channel level to the bank level, the operator-level intermediate representation of the matrix multiplication operator can be partitioned along multiple dimensions and distributed to multiple hierarchical structures in the DRAM, and a 12-tuple can be used to represent the partitioning result under the partitioning strategy: {{M, N, K} {CH,RA,DE,BK}}. Among them, the first constraint condition may include the above-mentioned M CH ×K CH ×N CH ≤ the number of channels, M RA ×K RA ×N RA ≤ the number of ranks in each channel, M DE ×K DE ×N DE ≤ the number of devices in each rank, M BK ×K BK ×N BK ≤ the number of banks in each device. It should be understood that under the premise of satisfying the above first constraint condition, the number of partitions of the operator-level intermediate representation of the matrix multiplication operator in each dimension of each level can be different (that is, the partitioning method can be different). Based on the different number of partitions in each dimension of each level, it means that multiple partitioning results can be obtained for the same operator-level intermediate representation, that is, multiple partitioning strategies can be obtained; among them, the sub-intermediate representation partitioned by each partitioning strategy can be the intermediate representation partitioned by the smallest level, for example, the above-mentioned bank-level intermediate representation.
[0067] After partitioning the operator-level intermediate representation of the matrix multiplication operator, the next step is to encode the way of mapping the partitioned sub-intermediate representation (equivalent to the weights, inputs, and outputs of the sub-matrix multiplication operator) to the rows and columns of the DRAM to obtain the mapping strategy. In the current DRAM, multiple columns in the DRAM can be continuously accessed using a single read / write command, and the number of such column accesses is called the burst length (BL), that is, the burst length is the number of columns continuously accessed by a single access command. Embodiments of the present disclosure regard multiple data columns corresponding to the burst length as the smallest data block read from the DRAM, and use this as the basis for data mapping. Then the mapping strategy can be expressed as:
[0068]
[0069] Among them, Map {W,A,O} represents the mapping strategy of the weight matrix, input activation matrix, and output matrix. Taking Map WFor example, the sub-weight matrix W divided out is first split into multiple data slices (i.e., the bank-level intermediate representation is first split into multiple data slices), and each data slice will be mapped to a minimum data block in the DRAM. The size of each data slice is Then the second constraint condition can include: That is to say, the size of each data switch should be less than or equal to the size of the minimum data block read in the DRAM so that each data slice can be stored in the space corresponding to the minimum data block; where represents the size of the data slice (i.e., the product of the size of the data slice in the M dimension and the size in the K dimension), the matrix data bit width represents the data bit width of the data switch, and the DRAM column bit width represents the column bit width of the storage unit in the DRAM; represents the number of data slices of the sub-weight matrix in the M dimension represents the number of data slices of the sub-weight matrix in the K dimension, and it is stipulated that a total of data slices will be mapped to the same row in the DRAM. Thus, the third constraint condition can include: That is, the number of data slices mapped to the same row in the storage unit is less than or equal to the ratio of the total number of columns per row of the storage unit to the burst length. Similarly, the size of each data slice in the multiple data switches into which each sub-input activation matrix A is split can be expressed as Then the second constraint condition can include: represents the number of data slices of the sub-input activation matrix in the M dimension represents the number of data slices of the sub-input activation matrix in the K dimension, and if the number of data slices will be mapped to the same row in the DRAM, then the third constraint condition can include: The size of each data slice in the multiple data switches into which each sub-output matrix is split is expressed as Then the second constraint condition can include: represents the number of data slices of the sub-output matrix in the M dimension represents the number of data slices of the sub-output matrix in the K dimension, and if the number of data slices will be mapped to the same row in the DRAM, then the third constraint condition can include:
[0070]
[0071] It should be understood that for each sub - intermediate representation in the partitioning result of the matrix multiplication operator under any partitioning strategy, a corresponding mapping method can be obtained, and the mapping result of the operator - level intermediate representation of the matrix multiplication operator can be encoded using a 12 - tuple. Under the condition of satisfying the above - mentioned second constraint condition, the sizes of the data slices partitioned by each sub - intermediate representation of each partitioning strategy can be different, and under the above - mentioned third constraint condition, the number of data slices mapped to the same row can also be different. Thus, there can be multiple mapping methods for mapping the partitioning result under each partitioning strategy to the rows and columns in the memory, and each mapping method can be encoded to obtain multiple mapping strategies corresponding to each partitioning strategy.
[0072] Optionally, in order to reduce the search space of the compilation strategy to a certain extent, some additional constraint conditions can be added to limit the encoding range of the above - mentioned mapping strategy. For example, for the above - mentioned matrix multiplication operator, the second constraint condition can further include: the size of the data slice in the dimension (i.e., the K - dimension) where reduction calculation needs to be performed (i.e., the size of the sub - weight matrix in the K - dimension, K B W and the size of the sub - input activation matrix in the K - dimension, K B A ) should be as close as possible to the computing parallelism of the computing unit, in order to make full use of the computing parallelism of the reduction module (e.g., adder tree) in the computing unit (PU) and improve the computing efficiency. Also, the number of data slices of the sub - weight matrix in the K - dimension, K R W and the number of data slices of the sub - input activation matrix in the K - dimension, K R A are set to the same value to match the granularity of reading weights and inputs from DRAM, thereby reducing the number of redundant reads. Of course, those skilled in the art can also set other required constraint conditions according to actual needs to limit the encoding range of the mapping strategy, and the embodiments of the present disclosure do not limit this.
[0073] It should be understood that multiple partitioning strategies can be generated based on the first constraint condition, and multiple mapping strategies under each partitioning strategy can be generated based on the second and third constraint conditions. Thus, the search space of the compilation strategy can be constructed; as described above, the search space can also include multiple selectable scheduling strategies, so that the subsequent determined target compilation strategy can also include a scheduling strategy.
[0074] In step S14, by predicting the performance of the compilation strategies in the search space, multiple candidate compilation strategies and the compilation results corresponding to each candidate compilation strategy are determined. The compilation results include the complete instruction sequence obtained by compiling the operator - level intermediate representation according to the candidate compilation strategy.
[0075] In practical applications, a performance predictor can be used to predict the performance of each compilation strategy in the search space, obtaining the predicted performance corresponding to each compilation strategy. Here, the predicted performance can characterize the time required for the complete instruction sequence obtained by compiling the operator-level intermediate representation using the compilation strategy to be executed and completed on the near-memory computing architecture. Then, the compilation strategies can be sorted according to the predicted performance, and the top K compilation strategies with the highest predicted performance (i.e., the top K with the shortest time) can be selected as candidate compilation strategies, and at the same time, the compilation results of each candidate compilation strategy can be obtained. Among them, the performance predictor can adopt performance prediction algorithms known in the art. Of course, performance prediction algorithms can also be designed independently, or machine learning models or deep learning models can be trained as performance predictors. The embodiments of the present disclosure do not limit this.
[0076] Considering the diversity in the selection of partitioning strategies, mapping strategies, and scheduling strategies, the data scale of the search space of compilation strategies is huge. The huge search space will bring a large amount of computation to performance prediction. Therefore, as Figure 5 shown, the performance upper limit of the near-memory computing architecture required under each compilation strategy can be analyzed based on the hardware data first, and the performance upper limit can be used as an indicator to dynamically reduce the search space (i.e., perform search space pruning). Furthermore, the performance predictor can be used to predict the performance of the instruction fragments (i.e., partial instruction sequences) compiled by the compilation strategies in the pruned search space. Thus, in a possible implementation manner, in step S14, by predicting the performance of the compilation strategies in the search space, multiple candidate compilation strategies and the compilation results corresponding to each candidate compilation strategy are determined, including:
[0077] Step S141: Based on the hardware data corresponding to the near-memory computing architecture, determine the performance upper limit corresponding to each compilation strategy in the search space. The performance upper limit characterizes the upper limit of the time required for the compilation result corresponding to the compilation strategy to be executed and completed on the near-memory computing architecture. The hardware data includes the bandwidth of the memory in the near-memory computing architecture, the row miss penalty (i.e., the latency required for accessing a new row when the row address changes), the number and performance of the computing units (such as the utilization rate of the computing units, etc.);
[0078] Step S142: Based on the performance upper limit corresponding to each compilation strategy in the search space, determine multiple initial candidate compilation strategies with a performance upper limit higher than the preset performance threshold;
[0079] Step S143: Use the compiler to partially compile the operator-level intermediate representation based on each initial candidate compilation strategy and a specific instruction set to obtain the partial instruction sequences compiled by each initial candidate compilation strategy. The specific instruction set covers the instructions executed on various near-memory computing architectures;
[0080] Step S144: Use a performance predictor to predict the performance of partial instruction sequences compiled by each initial candidate compilation strategy, and obtain the predicted performance of the partial instruction sequences compiled by each initial candidate compilation strategy. The predicted performance characterizes the time required for the partial instruction sequences to be executed and completed on the near-memory computing architecture.
[0081] Step S145: Based on the predicted performance of the partial instruction sequences compiled by each initial candidate compilation strategy and the number of loop executions corresponding to the partial instruction sequences, determine the target predicted performance corresponding to each initial candidate compilation strategy. The target predicted performance characterizes the time required for the complete instruction sequence compiled based on the initial candidate compilation strategy to be executed and completed on the near-memory computing architecture.
[0082] Step S146: Select multiple candidate compilation strategies from multiple initial candidate compilation strategies according to the target predicted performance corresponding to each initial candidate compilation strategy.
[0083] Step S147: Use a compiler to perform complete compilation on the operator-level intermediate representation based on multiple candidate compilation strategies and a specific instruction set, respectively, to obtain the compilation results corresponding to each candidate compilation strategy.
[0084] In step S141, based on the hardware data corresponding to the near-memory computing architecture, determine the performance upper limit corresponding to each compilation strategy in the search space, including: by analyzing each compilation strategy in the search space, obtain the analysis result corresponding to each compilation strategy. The analysis result includes the total access data volume and the number of row misses of the memory, and the load calculation amount of the computing unit; according to the analysis result corresponding to each compilation strategy and the hardware data, determine the performance upper limit corresponding to each compilation strategy in the search space.
[0085] It should be understood that different compilation strategies produce different compilation results. When different compilation results are executed on the near-memory computing architecture, the access data volume and the number of row misses generated for the memory are different, and the computing load for the computing unit to execute the compilation results is different. Therefore, a mapping model between the compilation strategy and the total access data volume, the number of row misses, and the load calculation amount can be established in advance. The mapping model can characterize the mapping relationship between the compilation strategy and the total access data volume, the number of row misses, and the load calculation amount. Furthermore, the partitioning strategy, mapping strategy, and scheduling strategy in the compilation strategy can be analyzed using the mapping model, so as to obtain the total access data volume, the number of row misses, and the load calculation amount generated under each compilation strategy. The embodiments of the present disclosure do not limit this.
[0086] Exemplarily, for the NDP architecture, the following performance upper limit formula can be used to determine the performance upper limit Perf corresponding to each compilation strategy in the search space according to the analysis result corresponding to each compilation strategy and the hardware datamax :
[0087]
[0088] Among them, the first item is equivalent to calculating the minimum time for the computing unit PU to access the DRAM; the second item calculates the total precharge time through the number of row misses and the miss penalty; the third item is equivalent to estimating the ideal latency of PU computing. Among them, since there is actually some overlap between the minimum time for the computing unit PU to access the DRAM calculated by the first item and the ideal latency of PU computing calculated by the third item, therefore, optionally, the performance upper bound Perf max can also be equal to This embodiment of the present disclosure does not limit this.
[0089] It can be understood that since the performance upper bound calculated using the above performance upper bound formula does not consider the latency caused by additional data movement, such as the latency caused by data movement between the PU, cache, and other PUs, therefore, if the performance upper bound corresponding to a certain compilation strategy is lower than the preset performance threshold, it means that the actual performance of the compilation result under this compilation strategy when executed on the near-memory computing architecture is also lower than this preset performance threshold, then it is considered that this compilation strategy does not meet the minimum performance requirements, so the further performance prediction for this compilation strategy can be skipped to narrow the search space and reduce the computational effort required for the subsequent performance predictor to perform performance prediction on the compilation strategies in the search space, thereby improving the computational efficiency. Thus, in step S142, multiple initial candidate compilation strategies with performance upper bounds higher than the preset performance threshold can be first determined from the search space, and then multiple candidate compilation strategies can be selected from the multiple initial candidate compilation strategies. It should be understood that those skilled in the art can customize the specific value of the above preset performance threshold, and this embodiment of the present disclosure does not limit this.
[0090] In step S143, those skilled in the art can design a compiler according to actual needs. Of course, existing compilers in the art can also be used as long as they can compile the operator-level intermediate representation according to the compilation strategy. It can be known that the compiler can compile the operator-level intermediate representation into an instruction sequence that can be executed on the near-memory computing architecture. Therefore, a corresponding instruction set needs to be provided to the compiler. The instruction sets in the prior art are mainly designed separately for different near-memory computing architectures and cannot be applied to different memory computing architectures. Thus, this embodiment of the present disclosure proposes the unified NDP instruction set (i.e., the specific instruction set) shown in Table 1 based on the tree-structured NDP abstraction model to cover all instruction operations that may be used in various NDP architectures and meet the generalization requirements for deploying any neural network model on any NDP architecture.
[0091] Table 1 Specific Instruction Set
[0092]
[0093] The specific instruction set shown in Table 1 mainly includes three types of instructions: calculation instructions, data copy instructions, and host access instructions. Among them, LB (Local Buffer) represents the local cache, and GB (Global Buffer) represents the global cache. It should be noted that the instructions listed in Table 1 will be further split into a series of DRAM commands (such as activation ACT commands, read READ commands, write WRITE commands, etc.) by the memory controller at runtime.
[0094] Among them, the calculation instructions allow the calculation unit (PU) to perform calculations using the input operand A from the memory (such as a bank of DRAM), and the operand B from the memory (such as another bank), the global cache, or the local input cache of the calculation unit. Among them, the MAC (Multiply-Accumulate) calculation instruction is a vectorized instruction that accesses multiple data in consecutive columns of a row in the memory using a single instruction, thereby reducing the overhead of the command and address buses. In addition, each MAC calculation instruction requires specific hardware support. For example, bank-level MAC-DRAM requires a dedicated data bus between the calculation unit and two different banks; MAC-LB requires the calculation unit to be equipped with a local cache. Therefore, if some hardware components (such as no global cache) are missing in the near-memory computing architecture described by the tree-like architecture abstraction model, the compiler will automatically exclude the instructions related to the missing hardware components during the compilation process.
[0095] Among them, the data copy instructions are used to exchange data within the NDP architecture and mainly include three types of instructions: the first type copies data between the registers and the local cache of the calculation unit, usually used to store and update temporary calculation results; the second type copies data between the local cache and the memory (such as DRAM), taking advantage of data locality to reduce memory access; the third type transfers data between the global cache and the memory (such as DRAM), enabling the calculation unit to obtain data from invisible addresses.
[0096] Among them, the host access instruction is used to control the data communication between the central processing unit (CPU) and the near-memory computing architecture. In addition to the conventional memory read and write instructions, four types of instructions are added to the instruction set to achieve high-parallel communication between the host and all computing units in the near-memory computing architecture. Among them, it can be assumed that the host can write to the registers, local caches, and global caches of the computing units in parallel, thereby reducing the internal data transfer in the near-memory computing architecture and improving the efficiency of each computing unit in the near-memory computing architecture to obtain input data. For data reading, since the near-memory computing architecture only needs to return the final result to the host, the instruction for transferring data from the result register to the host can be retained.
[0097] Through the above specific instruction set, different data streams of various near-memory computing architectures can be covered. Here, taking UPMEM memory (a memory embedded in a processor core) and AiM memory (Accelerator in Memory, a technology that directly embeds computing units into memory) as examples, the changes in instructions in different architectures are described. In the near-memory computing architecture of UPMEM, the WRITE-LB instruction can be used to prepare the input matrix, while in the near-memory computing architecture of AiM, since the computing unit lacks a local cache, this process is replaced by the WRITE-GB or WRITE-DRAM instruction. When the input data is ready, the MAC-LB, MAC-GB, or MAC-DRAM instruction is used to perform the multiply-accumulate calculation. It should be noted that in the near-memory computing architecture of AiM, the MAC-DRAM instruction reduces the number of active computing units to half of the number of memory banks in the DRAM (that is, one computing unit is activated by two DRAM memory banks), so compared with the MAC-GB instruction, it loses half of the parallelism of the computing units. When some intermediate calculation results are obtained, the REG2LB instruction can be used to store the intermediate calculation results from the register of the computing unit to the output cache. After obtaining the complete calculation result, the LB2DRAM instruction can be used to write part of the calculation result back to the DRAM for subsequent processing. In the AiM without an output cache, the LB2DRAM instruction can be replaced by the READ-REG and WRITE-REG instructions to return the complete calculation result to the host (such as the CPU).
[0098] As described above, in addition to pruning the search space of compilation strategies, a performance predictor is introduced to predict the performance of compilation strategies. Although the search space of compilation strategies has been pruned, the overhead of using cycle-accurate simulation to accurately predict the performance of each compilation strategy is still relatively large. And although the analysis based on the ideal performance upper limit is fast enough, due to the lack of evaluation of additional data movement latency, its accuracy is not high enough. Therefore, a fast and accurate performance predictor is needed to predict more accurate performance. However, in the case where the instruction sequence has not been compiled, it is difficult to obtain the additional data movement latency through theoretical analysis because it is affected by various factors, such as cache size, data flow, and possible instruction overlap in out-of-order execution. Therefore, a performance predictor can be embedded during the compiler's compilation of the operator-level intermediate representation. After the compiler compiles a partial instruction sequence, the performance predictor can be used to predict the performance of this partial instruction sequence, and then the performance of the complete instruction sequence can be deduced. In this way, the efficiency of performance prediction can be improved.
[0099] Thus, in step S144, the performance predictor is used to predict the performance of the partial instruction sequences compiled by each initial candidate compilation strategy, and the predicted performance of the partial instruction sequences compiled by each initial candidate compilation strategy is obtained, including:
[0100] Step S1441: Based on the partial instruction sequences compiled by each initial candidate compilation strategy, determine the total number of accesses to the memory in the near-memory computing architecture, the number of row misses in consecutive memory accesses, and the amount of data accessed by the computing units in the near-memory computing architecture to the cache unit or register during the execution of the partial instruction sequences compiled by each initial candidate compilation strategy;
[0101] Step S1442: According to the total number of accesses to the memory in the near-memory computing architecture and the column-to-column latency, the number of row misses in consecutive memory accesses and the row miss penalty, and the amount of data accessed by the computing units in the near-memory computing architecture to the cache unit or register and the bus bandwidth during the execution of the partial instruction sequences compiled by each initial candidate compilation strategy, determine the predicted performance of the partial instruction sequences compiled by each initial candidate compilation strategy.
[0102] Among them, the total number of read / write instructions for the memory in a partial instruction sequence can be analyzed (e.g., the total number of DRAM read / write commands triggered) to obtain the total number of memory read / writes (i.e., the total access times) that may come from computational instructions, data copy instructions, and / or host access instructions; the number of instructions that trigger a wrap-around access to the memory in the partial instruction sequence can also be analyzed to determine the number of row misses during consecutive read / write processes to the memory; and the number of instructions that access the host cache or registers in the partial instruction sequence can be analyzed to determine the amount of data accessed by the host to the cache or registers. Furthermore, in combination with tCDD (column-to-column delay) in the timing parameters of the memory, the known row miss penalties (such as precharge delay, activation delay, and row-to-row delay, etc.), and the bandwidth of the data transfer bus between the host and the near-memory computing architecture (which is known), the predicted performance of the partial instruction sequences compiled by each initial candidate compilation strategy can be predicted.
[0103] Exemplarily, the performance predictor can use the following performance prediction formula to determine the predicted performance Lat of the partial instruction sequences compiled by each initial candidate compilation strategy partial :
[0104] Lat partial = t CCD × (#READ + #WRITE) + #Row miss (READ) × t RowChange (READ)
[0105] + #Row miss (WRITE) × t RowChange (WRITE) + Lat Buffer / Reg (Host)
[0106] Among them, t CCD represents the column-to-column delay (i.e., tCDD), #READ and #WRITE respectively represent the total number of DRAM reads and writes from computational instructions, data copy instructions, and host access instructions; #Row miss (READ) represents the number of row misses during consecutive read processes, t RowChange (READ) is the row miss penalty during consecutive read processes (i.e., the additional delay when the row is hit), #Row miss (WRITE) represents the number of row misses during consecutive write processes, t RowChange (WRITE) is the row miss penalty during consecutive write processes; Lat Buffer / Reg (Host) takes into account the impact of host access instructions on the data bus, Lat Buffer / Reg(Host) = The amount of data accessed by the host from the cache or register / bus bandwidth. By performing performance prediction using the above performance prediction formula, the predicted performance can include an evaluation of the data movement latency, which can improve the accuracy of performance prediction for some instruction sequences.
[0107] In step S145, the predicted performance of the complete instruction sequence (i.e., the total latency of the complete instruction sequence when executed on the near-memory computing architecture) can be obtained based on the number of loop executions of the partial instruction sequences compiled by each initial candidate compilation strategy. For example, the total latency = the number of loop executions × Lat partial , that is, the predicted performance of the partial instruction sequence can be multiplied by the number of loop instructions of the partial instruction sequence to obtain the target predicted performance of the predicted complete instruction sequence. In practical applications, those skilled in the art can use any known analysis method in the art to estimate the number of loop executions of the partial instruction sequences compiled by each of them by analyzing the partitioning strategy, mapping strategy, and scheduling strategy in each initial candidate compilation strategy. The embodiments of the present disclosure do not limit this.
[0108] In step S146, after obtaining the target predicted performance corresponding to each initial candidate compilation strategy, for example, the initial candidate compilation strategies can be sorted in descending order according to the target predicted performance (equivalent to sorting the initial candidate compilation strategies in ascending order according to the execution duration of the instruction sequence, and the shorter the execution duration, the higher the target predicted performance), and the K initial candidate compilation strategies with the top K target predicted performance are selected as the K candidate compilation strategies; then, in step S147, the compiler can use the above specific instruction set to perform complete compilation on the operator-level intermediate representation according to each candidate compilation strategy to obtain the complete instruction sequences compiled by each candidate compilation strategy.
[0109] In step S15, based on the tree-like architecture abstraction model, the execution process of the compilation results corresponding to each candidate compilation strategy on the near-memory computing architecture is simulated to obtain the simulation results corresponding to each candidate compilation strategy. The simulation results represent the total duration required for the compilation results to be executed and completed on the near-memory computing architecture.
[0110] It should be understood that although the above performance predictor can be used to predict the performance of the complete instruction sequence (i.e., the execution duration of the complete instruction sequence on the near-memory computing architecture), the performance predicted by the performance predictor is still not accurate enough. Therefore, the total duration required for the compilation results corresponding to each candidate compilation strategy to be executed and completed on the near-memory computing architecture can be obtained by simulating the execution process of the compilation results corresponding to each candidate compilation strategy on the near-memory computing architecture, which is beneficial to accurately determining the target compilation strategy that is most suitable for the near-memory computing architecture.
[0111] In a possible implementation, an emulator can be used to simulate the execution process of the compilation results corresponding to each candidate compilation strategy in a near-memory computing architecture based on a tree architecture abstraction model, and obtain the simulation results corresponding to each candidate compilation strategy. Among them, the emulator includes a hardware state space, an instruction queue, and a Virtual Memory Controller (VMC). Among them, the hardware state space is used to simulate and track the hardware states of each hardware component in the near-memory computing architecture based on the tree architecture abstraction model. The hardware components include a storage unit, a computing unit, a cache unit, and a transmission bus (such as the bus for transmitting data between the computing unit and the cache unit). The instruction queue is used to store the complete instruction sequence compiled based on the candidate compilation strategy and maintain the data dependency relationship in the complete instruction sequence (the data dependency relationship can describe the order required for the execution between instructions); the virtual content controller is used to determine the issue time of the instructions in the instruction queue according to the hardware state space and update the hardware states of each hardware component in the hardware state space.
[0112] Among them, the update of the hardware state space can be triggered by a memory command and tracked and updated according to the timing constraint model of the memory itself. It should be understood that, given the tree architecture abstraction model of the near-memory computing architecture, each hardware component in the near-memory computing architecture can be simulated, and then each hardware state can be tracked based on the timing constraint model of the memory itself. Specifically, tracking the hardware states of each hardware component in the near-memory computing architecture includes: tracking the activation states of each storage unit of the memory in the near-memory computing architecture to issue an activation command before an unactivated storage unit is accessed; tracking the currently activated row address in the storage unit to indicate whether a subsequent access triggers a row miss and scheduling a precharge command and an activation command to activate the row to be accessed when a row miss is triggered.
[0113] Among them, multiple clock countdowns can be used in the hardware state space to track the hardware states of respective hardware components; the hardware state of the storage unit includes the time when the next command can be issued for the storage unit (i.e., the time when the next command (such as an ACT command, a PRE command, a READ command, and a WRITE command) can be issued for the storage unit can be indicated by a clock countdown), and the hardware states of other hardware components (i.e., computing units, cache units, transmission buses, etc.) other than the storage unit include an idle state or an occupied state (i.e., the time when the space state or occupied state of the hardware component can be indicated by a clock countdown), where the idle state is also an unused state, and the occupied state is also a state of being in use; the clock countdowns corresponding to respective hardware components are updated according to the hardware occupation duration of the hardware component involved in the issued instruction. By establishing a hardware state space to track the hardware states of different hardware components, accurate simulation can be provided for the execution process of an instruction sequence on a near-memory computing architecture, and the timing parameters in a memory (such as DRAM) can be accurately modeled during the simulation process.
[0114] Exemplarily, for the NDP architecture, the activation state of each bank of DRAM in the NDP architecture can be tracked. This is because it is not allowed to execute DRAM read / write commands when a bank is not activated. Therefore, before accessing an unactivated row, an ACT command should be issued first. Secondly, the address of the currently activated DRAM row can also be tracked to indicate whether the next memory access will trigger a row miss (i.e., the row address changes). In the case of a row miss, a precharge command (PRE command) and a subsequent activation command (ACT command) can be scheduled first to activate another row of the DRAM. Furthermore, considering the timing constraints between different DRAM commands, four DRAM timing countdowns can be maintained to indicate when a certain bank can issue the next ACT command, PRE command, READ command, and WRITE command. At the same time, clock countdowns are also used to track the idle states and occupied states of respective processing units (PUs), transmission buses, and cache units.
[0115] Among them, all the clock countdowns in the hardware state space can be updated based on a global clock. When the virtual memory controller issues a new instruction, the clock countdowns of the hardware components involved in this instruction will be reset and updated according to the DRAM commands converted from this instruction or the hardware occupation duration required to execute this instruction, so as to maintain the hardware states of each hardware component in the hardware state space. Among them, the hardware occupation durations (i.e., the durations for the hardware to execute instructions) of the hardware components involved in different instruction executions are different. Therefore, the mapping relationship between different instructions and the hardware occupation durations can be pre-constructed, so as to obtain the hardware occupation duration corresponding to each instruction based on this mapping relationship during the simulation process, and then update the clock countdowns of the corresponding hardware components based on the hardware occupation duration; or, since the DRAM commands converted from different instructions are also different, the clock countdowns of the hardware components involved in the DRAM commands can be updated according to the hardware occupation durations of the hardware components involved in the DRAM commands converted from the currently issued instruction. The embodiments of the present disclosure do not limit this.
[0116] It can be known that the data dependency relationship between instructions is the most notable factor in instruction-driven simulation. Therefore, the emulator should follow two basic rules when issuing and evaluating an instruction sequence: (1) The operands of the instruction should be prepared before the instruction is issued; (2) The memory space of the data cannot be reused until the instruction is completed. Existing memory-oriented DRAM emulators usually use memory barrier instructions to separate DRAM read / write commands to follow the data dependency relationship between instructions. However, the barrier instruction design is less efficient for the NDP architecture. On the one hand, in traditional memory, different DRAM devices in the same rank or different banks in the same device work together, but in the NDP architecture, different computing units work independently, which leads to a significant increase in the number of instructions. Therefore, using memory barrier instructions in this case will impose a great burden on the size of the code file; on the other hand, in the NDP architecture, the data dependency relationship may only occur within the local scope of certain instructions. For example, in bank-level NDP, the data dependency between writing input data and calculating using the input data only exists between the instructions of the same bank. However, using memory barrier instructions to force all banks to synchronize will result in pipeline bubbles and inefficient utilization of processing units.
[0117] In view of this, the embodiments of the present disclosure design an instruction queue to schedule the order of instructions, so as to ensure that the issued instructions and the corresponding simulation behaviors do not violate the data dependencies between the instructions. Among them, the instruction queue includes multiple instruction groups. The mutually independent instructions executed by different computing units are located in different instruction groups. Each instruction group includes multiple instructions to be executed sequentially. The instructions in the instruction groups without data dependencies can be issued in parallel. Or rather, the instruction queue is composed of different instruction groups, and each instruction group contains multiple instructions that need to be executed sequentially, such as "input - calculation - output" instructions. Among them, the mutually independent instructions between different computing units will be placed in different instruction groups. For example, in the bank - level NDP architecture, the instructions for different banks can be placed in different instruction groups. Therefore, different instruction groups can issue instructions in parallel. Of course, if data communication and synchronization are required between two computing units or between a computing unit and other hardware components, the data dependencies between different instruction groups can also be configured. That is, data dependencies can also be configured between instruction groups, and then the instructions in the instruction groups without data dependencies can be issued in parallel. In each iteration of the simulation, the instruction queue first selects the instruction groups without predecessors in the data dependency graph (this data dependency graph can describe the data dependencies between each instruction group in the instruction queue), that is, the instruction groups that do not depend on a certain instruction group to be executed first (equivalent to the instruction groups without data dependencies), and sends the first instruction (that is, the instruction arranged at the top) in each instruction group without predecessors to the virtual memory controller (VMC) to control the issuance of instructions based on the virtual memory controller.
[0118] Among them, the virtual memory controller can support two ways of issuing instructions, namely instruction reordering and out - of - order execution. The virtual memory controller can also be used to schedule DRAM commands and coordinate other hardware components, such as processing units, cache units, and buses. The virtual memory controller mainly completes two functions, that is, finding the earliest time when a given instruction can be issued, and updating the clock countdown of each hardware component in the hardware state space to reflect the impact of issuing this instruction on the hardware state of the relevant hardware components.
[0119] Exemplarily, the embodiments of the present disclosure provide Figure 6 a schematic diagram of an emulator architecture as shown in Figure 6As shown, the instruction queue contains N instruction groups (inst.Group1, inst.Group2, …, inst.GroupN) for storing instruction sequences. By checking the data dependencies of the instruction groups, the instruction queue sends the set of currently issuable instructions {inst1.1, inst2.1, …, instN.1} to the virtual memory controller VMC. An instruction issue queue can also be set in the VMC to store the instructions in the instruction set. The VMC can check the possible issue time (i.e., calculate the start time of the instruction) of each instruction in the instruction set according to the hardware state space, and select the instruction with the earliest issue time as the next instruction to be issued. For example, if the issue times of three instructions in the instruction issue queue are Ts1, Ts2, and Ts3 respectively, and Ts2 < Ts1 and Ts2 < Ts3, then the instruction corresponding to Ts2 with min Ts is currently sent, and since no other instructions can be issued before the earliest issue time, the hardware state space can be directly updated to the hardware state at the earliest issue time; then the VMC deletes (i.e., issues) the instruction corresponding to Ts2 from the instruction queue and updates the hardware state of the hardware components occupied by the instruction corresponding to Ts2; among them, the hardware state space can use clock countdown to indicate the time when the next command can be issued for the DRAM Bank under the current instruction, and use clock countdown to indicate the current hardware state of the bus Bus, cache unit Buf, and processing unit PU.
[0120] Based on the above emulator architecture, the above uses the emulator to simulate the execution process of the compilation results corresponding to each candidate compilation strategy in the near-memory computing architecture based on the tree-like architecture abstraction model, and the obtained simulation results corresponding to each candidate compilation strategy may include:
[0121] Step S151, for the compilation result corresponding to any candidate compilation strategy, loop through the following simulation operations in steps S1511 to S1514:
[0122] Step S1511, based on the data dependencies involved in each instruction group, the instruction queue determines the set of instructions that can be issued simultaneously and sends the instruction set to the virtual memory controller. The instruction set includes the instructions that are currently ranked first in each instruction group without data dependencies.
[0123] Step S1512, the virtual content controller determines the issue time of each instruction in the instruction set according to the hardware state space and selects the instruction with the earliest issue time as the target instruction to be issued currently.
[0124] Step S1513, the virtual content controller controls the hardware state of each hardware component in the hardware state space to be updated to the hardware state corresponding to the issuance time of the target instruction based on the issuance time of the target instruction, and deletes the target instruction from the instruction queue;
[0125] Step S1514, the virtual content controller updates the hardware state of the hardware component occupied by the target instruction in the hardware state space according to the hardware occupation time of the hardware component occupied by the target instruction;
[0126] Step S152, when all instructions in the instruction queue are issued and the hardware state of each hardware component in the hardware state space becomes an idle state, obtain the simulation result corresponding to the candidate compilation strategy.
[0127] In step S1511, if there is no data dependency between the instruction groups in the instruction queue, the instruction set consisting of the instructions currently ranked first in each instruction group in the instruction queue can be sent to the virtual memory controller; if there is a data dependency between some instruction groups in the instruction queue, the instruction set consisting of the instructions currently ranked first in each instruction group without a predecessor in the instruction queue can be sent to the virtual memory controller.
[0128] In step S1512, the arrival time T of operands from different sources to the computing unit can be checked. A and the idle time T of the computing unit itself F (PU), determines the earliest time T at which the computation unit can start computing S , to obtain the issuance time T of each instruction in the instruction set i Specifically, T S The latest calculation start time is regarded as the latest calculation start time, and the DRAM command is scheduled in an ASAP (As Soon As Possible) manner, and the issuance time of the earliest issued DRAM command (such as PRE, ACT, READ, etc.) is used as the issuance time of the corresponding instruction.
[0129] Specifically, in the above step S1512, the virtual content controller determines the issuing time of each instruction in the instruction set according to the hardware state space, which may include:
[0130] Step S15121, for the ith instruction in the instruction set, when the operand of the ith instruction comes from the storage unit, based on the hardware state space, determine the activation state of the row where the operand of the ith instruction is located in the storage unit, and determine the arrival time of the operand of the ith instruction at the computing unit according to the activation state of the row where the operand of the ith instruction is located, the clock countdown of the storage unit, and the read delay of reading the operand from the storage unit, where i is an integer greater than or equal to 1;
[0131] Step S15122: When the operand of the i-th instruction comes from a cache unit, determine the arrival time of the operand of the i-th instruction at the computing unit according to the release time of the cache unit, the read latency of reading the operand from the cache unit, and the idle time of the transmission bus between the cache unit and the computing unit.
[0132] Step S15123: Determine the start processing time when the computing unit starts to process the operand of the i-th instruction according to the arrival time of the operand of the i-th instruction at the computing unit and the idle time of the computing unit.
[0133] Step S15124: Based on the start processing time when the computing unit starts to process the operand of the i-th instruction, determine the issue time of the i-th instruction.
[0134] As described above, the compute instruction allows the computing unit to perform a computation (such as matrix multiplication) on the input operand A from the DRAM (e.g., one memory bank) and the operand B from the DRAM (e.g., another memory bank), the global cache, or the local input cache of the computing unit. Here, the operand can be understood as the data that the operator wants to compute. For example, the operand of the matrix multiplication operator can be the input activation matrix. Thus, for the operand from the DRAM bank (i.e., the operand of the instruction comes from the storage unit), the activation status of the row where the operand is located can be read to check whether an ACT command and / or a PRE command need to be inserted before the access. And when the activation status indicates that the row where the operand is located is not activated, the arrival time T of the operand at the computing unit can be obtained by superimposing the time indicated by the clock countdown of the DRAM for issuing the next command (such as an ACT command, a PRE command, etc.), the read latency of reading the operand from the storage unit, and the latency generated by inserting the PRE command and the ACT command. A (BK i );If the activation status indicates that the row where the operand is located is activated, the arrival time T of the operand at the computing unit can be obtained by superimposing the time indicated by the clock countdown of the DRAM for issuing the next command (such as a READ command, etc.) and the read latency of reading the operand from the storage unit. A (BK i )。
[0135] Exemplarily, Figure 7a A schematic diagram of a DRAM access process shown, such as Figure 7aAs shown, Src1:Bank indicates that operand 1 comes from a bank of DRAM; Src2:Bank indicates that operand 2 comes from another bank of DRAM; [Idle] represents space time; [Row miss] represents that when the processing unit attempts to access data that is not in the current active row, a row miss will occur; PRE represents the precharge command, RP represents the row precharge delay, and the time from the current time to the PRF command can be the time when the next command can be issued as indicated by the clock countdown of the bank; ACT represents the activation command, RCD represents the row to column delay, that is, the time from row activation to when column data can be accessed, CCD represents the column to column delay, RD represents the read command, RL represents the read latency for reading the operand from the memory, that is, the time from issuing the read command to when the operand reaches the processing unit, Proc represents the stage where the processing unit receives and processes the data, and Occu._proc represents the time of occupying the processing unit, that is, the processing unit is processing two operands. As Figure 7a shown, by superimposing the time from the current time to the PRF command, the precharge delay RP, the row to column delay RCD, and the read latency RL, the arrival time T of operands Src1:Bank and Src2:Bank from the current time to reaching the processing unit can be obtained A (BK 0,1,… ).
[0136] For operands from cache units (such as global caches), the arrival time T of the operands at the computing unit A (GB) is determined by the release time of the cache unit, the read latency RL(GB) for reading the operand from the cache unit, and the idle time T of the transmission bus between the computing unit and the cache unit F (BUS). For example, the formula: T A (GB) = max(T F (GB) + RL(GB), T F (BUS)) can be used to determine the arrival time of operands from the cache unit at the computing unit. That is, the operand arrival time is the release time of the global cache plus the larger of the read latency and the bus idle time. Among them, the release time T of the global cache F(GB) generally refers to the time required to wait before the data in the global cache can be used by the next access request. The bus idle time can be understood as the time when the bus can transfer data, or the time from the completion of the previous transfer to the time when the bus can start the next transfer; the operand arrival time is determined by max(T F (GB) + RL(GB), T F (BUS)) because data transfer is limited not only by the read speed of the memory but also by the availability of the bus. If the bus is still busy after the global cache releases and readies the data, then the data transfer must wait for the bus to become idle. If the bus is idle during the time when the global cache releases and readies the data, then the data can be transferred immediately; if the bus is busy, then it must wait for the bus to become idle before transferring the data. Therefore, the arrival time of the operand from the global cache to the computing unit is the release time of the global cache plus the larger of the read latency and the bus idle time (i.e., the time when the bus becomes idle (i.e., the moment)). It should be understood that the arrival time of the operand from the local cache can be calculated in the same way as the arrival time of the operand from the global cache described above, which will not be elaborated here.
[0137] Exemplarily, Figure 7b a schematic diagram of another DRAM access process is shown, as Figure 7b shown. Src1:Bank indicates that operand 1 comes from a bank of the DRAM; Src2:Buf indicates that operand 2 comes from the global cache; [Idle] also represents the idle time; [Row miss] also represents a row miss; PRE represents the precharge command, RP represents the row precharge delay, ACT represents the activation command, RCD represents the row to column delay, RD represents the read command, RL represents the read latency for reading the operand from the bank, buf can represent the read latency for reading from the global cache, Proc represents the stage where the processing unit receives and processes the data (Process), Occu.proc represents the time of occupying the processing unit (Occupancy Processor), that is, the processing unit is processing two operands. Com represents the time when the transfer bus Bus between the processing unit and the cache unit starts to transfer data. By Figure 7bIt can be seen that the idle time of the transmission bus is greater than the sum of the release time of the global cache (i.e., the time from the current moment to sending an RD command to the global cache) and the read latency buf of reading the operand from the global cache. Therefore, the arrival time of Src1:Bank and Src2:Buf at the processing unit from the current time (Current time) can be the idle time of the transmission bus.
[0138] As described above, the operands of the i-th instruction may all come from the storage unit, or one may come from the storage unit and the cache unit. Therefore, in step S15123, the arrival time of the operands of the i-th instruction at the computing unit can include the arrival time of the operands from the storage unit, or the arrival time of the operands from the storage unit and the arrival time of the operands from the cache unit. Thus, the start processing time for the computing unit to start processing the operands of the i-th instruction can be the arrival time of the operands of the i-th instruction at the computing unit plus the maximum value of the idle time T F (PU), that is, the start processing time can be equal to T S =max(T A (GB), T A (BK 0,1,… ), T F (PU)), or the start processing time can also be equal to T S =max(T A (BK 0,1,… ), T F (PU)); where the idle time of the computing unit is the time (i.e., the moment) when the computing unit becomes idle.
[0139] In step S15124, the start processing time T S for the computing unit to start processing the operands of the i-th instruction can be regarded as the latest computing start time, and based on this latest computing start time, DRAM command scheduling is performed in an ASAP manner, and the issue time of the earliest issued DRAM command (such as PRE, ACT, READ, etc.) is used as the issue time of the i-th instruction, which is equivalent to calculating the issue time of the earliest DRAM command based on the start processing time for the computing unit to start processing the operands of the i-th instruction and using it as the issue time of the i-th instruction. Then, the instruction with the earliest issue time can be selected as the target instruction to be issued currently, or the instruction with the earliest start processing time can also be selected as the target instruction to be issued currently.
[0140] In step S1513, after determining the target instruction currently issued, the virtual memory space controller can, based on the issuance time of the target instruction, control the hardware states of each hardware component in the hardware state space to be directly updated to the hardware states corresponding to the issuance time of the target instruction by sending an update request to the hardware state space, which is equivalent to directly adjusting the clock countdown of each hardware component in the hardware state space to the issuance time of the target instruction. This enables the hardware state space in the emulator to adopt an instruction-driven update working mode, which means that the hardware state space is not updated step by step in cycles but is updated when a new instruction is issued. Therefore, the number of simulation iterations is reduced from the number of DRAM cycles required to calculate the instruction sequence to the length of the instruction sequence, which enables the emulator to achieve a 10- to 300-fold reduction in simulation time on different near-memory computing architectures and is beneficial to improving the efficiency of the compilation process.
[0141] Among them, the virtual memory controller deleting the target instruction from the instruction queue represents that the target instruction has been issued. Furthermore, in step S1514, after the target instruction has been issued, the virtual content controller can update the hardware states of the hardware components occupied by the target instruction in the hardware state space according to the hardware occupation duration of the hardware components occupied by the target instruction, which is equivalent to simulating the execution process of the target instruction; specifically, the occupation duration of the target instruction occupying the DRAM can be determined according to the issuance time of the target instruction, the automatic precharge policy of the DRAM, and the number of consecutive DRAM columns read by the instruction, and then the clock countdown of the DRAM can be updated based on the occupation duration of the DRAM, that is, the hardware state indicated by the clock countdown is updated, and the clock countdowns (i.e., hardware states) of other hardware components (such as the computing unit, cache unit, and transmission bus) will also be updated according to their respective occupied occupation durations.
[0142] It should be understood that the above steps S1511 to S1514 can be executed in a loop until all the instructions in the instruction queue have been issued (i.e., the instruction queue is empty), and the hardware states of each hardware component in the hardware state space become idle states (i.e., the hardware components have executed all the instructions in the instruction queue). In this case, the delay obtained from the start of the simulation to the end of the simulation is the total duration required for the complete instruction sequence to be executed on the near-memory computing architecture.
[0143] In practical applications, the emulator can also extend to support workloads of non-neural network models (i.e., operator-level intermediate representations) by modifying the performance model of the computing unit (for example, using other existing emulators in the art (such as Gem5) to simulate near-memory CPUs or other general computing units), adding corresponding IRs, and reusing memory access modeling. The embodiments of the present disclosure do not limit this.
[0144] In step S16, based on the simulation results corresponding to each candidate compilation strategy, a target compilation strategy is determined from multiple candidate compilation strategies, and the compilation result of the target compilation strategy is determined as the target compilation result of the operator-level intermediate representation of the neural network model, so as to deploy the neural network model on the near-memory computing architecture based on the target compilation result.
[0145] In practical applications, the candidate compilation strategy with the best simulation result (i.e., the shortest total duration) can be selected from multiple candidate compilation strategies as the target compilation strategy. Of course, the candidate compilation strategy with the second-best simulation result can also be selected as the target compilation strategy, etc. The embodiments of the present disclosure do not limit this. Furthermore, after obtaining the target compilation result of the operator-level intermediate representation of the neural network model, the neural network model can be deployed on the near-memory computing architecture based on the target compilation result. The embodiments of the present disclosure do not limit the deployment process of the neural network model on the near-memory computing architecture.
[0146] According to the compilation method of the embodiments of the present disclosure, by using the tree-structured abstract model to describe the near-memory computing architecture, any near-memory computing architecture can be uniformly described in time series by using the tree structure. Furthermore, by combining the operator-level intermediate representation of the neural network model, a compilation strategy search space is constructed. Then, multiple candidate compilation strategies are determined by predicting the performance of the compilation strategies in the search space. Then, through the emulator, the accurate simulation results corresponding to each candidate compilation strategy obtained by simulating the execution process of the compilation result on the near-memory computing architecture based on the abstract model are used to determine the final target compilation strategy and target compilation result, which can provide a better target compilation strategy for the deployment of any neural network model on any near-memory computing architecture, so as to realize the deployment of any neural network model on any near-memory computing architecture, and has high generality.
[0147] Based on the compilation method provided by the above embodiments of the present disclosure, the embodiments of the present disclosure also provide Figure 8 A framework schematic diagram of a compilation method shown in Figure 8As shown, the input is a neural network model to be compiled and deployed, as well as the hardware configuration information of the near-memory computing architecture (including the architecture configuration of the near-memory architecture and DRAM timing parameters, etc.). First, the operator analysis of the neural network model is performed through the ONNX framework to obtain a series of operator-level IRs converted by the ONNX framework, and a tree-like architecture abstraction model of the near-memory computing architecture is generated based on a general near-memory architecture abstraction method (i.e., a method of abstracting the near-memory computing architecture using a tree structure); then, based on the operator-level IR and the tree-like architecture abstraction model, a search space for creating a compilation strategy for the operator-level IR is created. The compilation strategy includes an operator partitioning strategy, a mapping strategy, and a scheduling strategy. Among them, a hardware data-guided design space pruning strategy can be used to reduce the size of the search space. Among them, the hardware data, for example, includes the number of rows and columns read from the DRAM, and the number of DRAM row precharges, etc. To further improve the search efficiency, a performance predictor is also designed based on the DRAM timing parameters to find the candidate compilation strategies corresponding to the instruction files (i.e., the files formed by partial instruction sequences) with the top K prediction performances. Then, the compiler can be used to generate K different compilation results based on these K candidate compilation strategies, and the simulator is called to evaluate the accurate performance of these compilation results. Among them, the input of the simulator is the compilation result (i.e., the complete instruction sequence), and the output simulation result represents the total time required for the execution of the complete instruction sequence to be completed. Finally, according to the simulation result, the optimal compilation strategy (i.e., the target compilation strategy) and the compilation result can be output.
[0148] It should be understood that the compilation method proposed in the embodiments of the present disclosure uses the near-memory computing architecture configuration based on the real world, and these configurations can also be flexibly changed to adapt to other hardware architectures. In addition, the compilation strategy search method for the search space proposed in the embodiments of the present disclosure adopts the idea of pruning and hierarchical simulation, and some more efficient search algorithms, such as simulated annealing or genetic algorithms, can also replace the search method in the embodiments of the present disclosure to provide higher search efficiency.
[0149] According to the compilation method of the embodiments of the present disclosure, various types and multi-level near-memory computing architectures can be supported through the near-memory architecture abstraction method, so that the simulator and the compiler can be adapted to different types of near-memory computing architectures; the proposed simulator uses instruction-driven to update the hardware state space, which is about 1.7 times faster than the existing simulator, and at the same time, the simulation error can be guaranteed to be within 10%.
[0150] In addition, by verifying the compilation method of the embodiments of the present disclosure in experiments for single operators and complete models (such as convolutional neural network models and large language models, etc.) on various near-memory computing architectures, compared with the existing compilation methods, using the compilation method of the embodiments of the present disclosure can achieve an acceleration of 1.09 to 1.56 times on single operators and an acceleration of 1.05 to 3.43 times on complete models.
[0151] The text analyzes the differences in compilation optimization strategies caused by the differences between near-memory computing architectures, and accordingly proposes a unified architecture abstraction and a general compilation method based on the architecture abstraction. Moreover, it deeply analyzes the operator splitting and mapping strategy space in near-memory computing architectures and designs a series of effective means for rapid strategy space exploration. Thus, a compilation method that can be applied to any near-memory computing architecture is proposed, realizing a general compilation and deployment toolchain for near-memory computing architectures, which has high generality.
[0152] Figure 9 The block diagram of a compilation system according to an embodiment of the present disclosure is shown, as Figure 9 shown, the device includes:
[0153] An acquisition module 901, configured to acquire the operator-level intermediate representation of a neural network model to be deployed to a near-memory computing architecture and the hardware configuration information of the near-memory computing architecture, where the hardware configuration information is used to indicate the hardware configuration of the near-memory computing architecture;
[0154] An abstraction module 902, configured to generate a tree-like architecture abstraction model of the near-memory computing architecture based on the hardware configuration information of the near-memory computing architecture, where the tree-like architecture abstraction model is used to describe the near-memory computing architecture in a tree structure;
[0155] A construction module 903, configured to construct a search space for compilation strategies based on the operator-level intermediate representation of the neural network model and the tree-like architecture abstraction model of the near-memory computing architecture, where the compilation strategies are used to compile the operator-level intermediate representation;
[0156] A prediction module 904, configured to determine multiple candidate compilation strategies and the compilation results corresponding to each candidate compilation strategy by performing performance prediction on the compilation strategies in the search space, where the compilation results include the complete instruction sequence obtained by compiling the operator-level intermediate representation according to the candidate compilation strategies;
[0157] A simulation module 905, configured to simulate the execution process of the compilation results corresponding to each candidate compilation strategy on the near-memory computing architecture based on the tree architecture abstraction model, so as to obtain the simulation results corresponding to each candidate compilation strategy, where the simulation results represent the total duration required for the compilation results to be executed and completed on the near-memory computing architecture;
[0158] A determination module 906, configured to determine a target compilation strategy from the multiple candidate compilation strategies according to the simulation results corresponding to each candidate compilation strategy, and determine the compilation result of the target compilation strategy as the target compilation result of the operator-level intermediate representation of the neural network model, so as to deploy the neural network model on the near-memory computing architecture based on the target compilation result.
[0159] In a possible implementation manner, the near-memory computing architecture includes multiple levels of hardware structures. One level of hardware structure in the near-memory computing architecture corresponds to one layer of nodes in the tree architecture abstraction model. Each layer in the tree architecture abstraction model includes at least one central node, and each central node is connected to at least one of the following types of child nodes: central nodes of the next lower level, storage nodes, computing nodes, and cache nodes in the layer where the central node belongs; the central node represents the data transmission relationship between two adjacent levels; wherein, the storage node represents the storage unit of the memory in the near-memory computing architecture and includes the attribute information of the storage unit, the computing node represents the computing unit in the near-memory computing architecture and includes the attribute information of the computing unit, and the cache node represents the cache unit in the near-memory computing architecture and includes the attribute information of the cache unit.
[0160] In a possible implementation, a single compilation strategy includes a partitioning strategy and a corresponding mapping strategy. The partitioning strategy represents the partitioning result of partitioning the operator-level intermediate representation along multiple dimensions at multiple levels. The mapping strategy represents the way of mapping the partitioning result of the operator-level intermediate representation to the memory in the near-memory computing architecture. The partitioning result represents multiple sub-intermediate representations obtained by partitioning the operator-level intermediate representation. Among them, constructing the search space of the compilation strategy based on the operator-level intermediate representation of the neural network model and the tree-like architecture abstraction model of the near-memory computing architecture includes: based on the operator-level intermediate representation and the tree-like architecture abstraction model, in the order from the highest level to the lowest level of the memory in the near-memory computing architecture, partitioning the operator-level intermediate representation along multiple dimensions at multiple levels under the satisfaction of the first constraint condition to obtain multiple partitioning strategies. Among them, the first constraint condition includes that the product of the number of partitions along each dimension among multiple dimensions at each level is less than or equal to the number of hardware structures corresponding to each level in the memory, and the partitioning methods of different partitioning strategies are different under the first constraint condition. For each sub-intermediate representation partitioned by each partitioning strategy, partitioning each sub-intermediate representation into multiple data slices under the satisfaction of the second constraint condition, and mapping each data slice to the memory in the near-memory computing architecture under the satisfaction of the third constraint condition to obtain multiple mapping strategies corresponding to each partitioning strategy. Among them, the second constraint condition includes that the product of the size of each data slice and the data bit-width of each data slice is less than or equal to the product of the burst length corresponding to the memory and the column bit-width of the memory. The third constraint condition includes that the number of data slices mapped to the same row in the memory is less than or equal to the ratio of the total number of columns in that row of the memory to the burst length.
[0161] In a possible implementation, determining a plurality of candidate compilation strategies and compilation results corresponding to each candidate compilation strategy by performing performance prediction on the compilation strategies in the search space includes: determining, based on the hardware data corresponding to the near-memory computing architecture, the performance upper limit corresponding to each compilation strategy in the search space, where the performance upper limit represents the upper limit of the time required to complete the execution of the compilation result corresponding to the compilation strategy on the near-memory computing architecture, and the hardware data includes the bandwidth of the memory, the row miss penalty, the number and performance of computing units in the near-memory computing architecture; determining a plurality of initial candidate compilation strategies whose performance upper limits are higher than a preset performance threshold based on the performance upper limits corresponding to each compilation strategy in the search space; using a compiler to perform partial compilation on the operator-level intermediate representation based on each of the initial candidate compilation strategies and a specific instruction set to obtain partial instruction sequences compiled by each of the initial candidate compilation strategies, where the specific instruction set covers instructions executed on various near-memory computing architectures; using a performance predictor to perform performance prediction on the partial instruction sequences compiled by each of the initial candidate compilation strategies to obtain the predicted performance of the partial instruction sequences compiled by each of the initial candidate compilation strategies, where the predicted performance represents the time required to complete the execution of the partial instruction sequences on the near-memory computing architecture; determining the target predicted performance corresponding to each initial candidate compilation strategy based on the predicted performance of the partial instruction sequences compiled by each initial candidate compilation strategy and the number of loop executions corresponding to the partial instruction sequences, where the target predicted performance represents the time required to complete the execution of the complete instruction sequence compiled based on the initial candidate compilation strategy on the near-memory computing architecture; selecting a plurality of candidate compilation strategies from the plurality of initial candidate compilation strategies according to the target predicted performance corresponding to each initial candidate compilation strategy; using the compiler to perform complete compilation on the operator-level intermediate representation based on the plurality of candidate compilation strategies and the specific instruction set respectively to obtain the compilation results corresponding to each candidate compilation strategy.
[0162] In a possible implementation, the determining, based on the hardware data corresponding to the near-memory computing architecture, the performance upper limit corresponding to each compilation strategy in the search space includes: analyzing each compilation strategy in the search space to obtain the analysis result corresponding to each compilation strategy, where the analysis result includes the total access data volume and the number of row misses to the memory under each compilation strategy, and the load calculation amount of the computing unit; determining the performance upper limit corresponding to each compilation strategy in the search space according to the analysis result corresponding to each compilation strategy and the hardware data.
[0163] In a possible implementation, using the performance predictor to perform performance prediction on partial instruction sequences compiled by each initial candidate compilation strategy to obtain the predicted performance of the partial instruction sequences compiled by each initial candidate compilation strategy includes: determining, based on the partial instruction sequences compiled by each initial candidate compilation strategy, the total number of accesses to the memory in the near-memory computing architecture, the number of row misses in continuously accessing the memory, and the amount of data accessed by the computing units in the near-memory computing architecture to the cache unit or register in the partial instruction sequences compiled by each initial candidate compilation strategy; determining the predicted performance of the partial instruction sequences compiled by each initial candidate compilation strategy according to the total number of accesses to the memory in the near-memory computing architecture and the column-to-column latency, the number of row misses in continuously accessing the memory and the row miss penalty, and the amount of data accessed by the computing units in the near-memory computing architecture to the cache unit or register and the bus bandwidth in the partial instruction sequences compiled by each initial candidate compilation strategy.
[0164] In a possible implementation, the simulation module uses a simulator to perform simulation on the execution process of the compilation results corresponding to each candidate compilation strategy in the near-memory computing architecture based on the tree-like architecture abstraction model, and obtains the simulation results corresponding to each candidate compilation strategy; wherein, the simulator includes: a hardware state space, an instruction queue, and a virtual memory controller; wherein, the hardware state space is used to simulate and track the hardware states of each hardware component in the near-memory computing architecture based on the tree-like architecture abstraction model, and the hardware components include a storage unit, a computing unit, a cache unit, and a transmission bus; the instruction queue is used to store the complete instruction sequences compiled based on the candidate compilation strategies and maintain the data dependencies in the complete instruction sequences, the instruction queue includes multiple instruction groups, instructions independent of each other executed by different computing units are located in different instruction groups, each instruction group includes multiple sequentially executed instructions, and instructions in instruction groups without data dependencies can be issued in parallel; the virtual content controller is used to determine the issue time of the instructions in the instruction queue according to the hardware state space and control the update time of the hardware states of each hardware component in the hardware state space.
[0165] In a possible implementation, tracking the hardware states of various hardware components in the near-memory computing architecture includes: tracking the activation states of individual storage units in the memory of the near-memory computing architecture to issue an activation command before an unactivated storage unit is accessed; tracking the currently activated row address in the storage unit to indicate whether a subsequent access causes a row miss and scheduling a precharge command and an activation command to activate the row to be accessed when a row miss occurs; wherein, multiple clock countdowns are used in the hardware state space to respectively track the hardware states of various hardware components, wherein the hardware state of the storage unit includes the time when the next command can be issued for the storage unit, and the hardware states of other hardware components except the storage unit include an idle state or an occupied state; the clock countdowns corresponding to each hardware component are updated according to the hardware occupation duration of the hardware components involved in issuing the instruction.
[0166] In a possible implementation, using the emulator to simulate the execution process of the compilation results corresponding to each candidate compilation strategy in the near-memory computing architecture based on the tree architecture abstraction model to obtain the simulation results corresponding to each candidate compilation strategy includes: for the compilation result corresponding to any candidate compilation strategy, cyclically performing the following simulation operations: the instruction queue determines the set of instructions that can be issued simultaneously based on the data dependencies involved in each instruction group, and sends the set of instructions to the virtual memory controller, and the set of instructions includes the instructions that are currently ranked first in each instruction group without data dependencies; the virtual content controller determines the issue time of each instruction in the set of instructions according to the hardware state space, and selects the instruction with the earliest issue time as the target instruction to be issued currently; the virtual content controller controls the hardware states of various hardware components in the hardware state space to be updated to the hardware states corresponding to the issue time of the target instruction based on the issue time of the target instruction, and deletes the target instruction from the instruction queue; the virtual content controller updates the hardware states of the hardware components occupied by the target instruction in the hardware state space according to the hardware occupation duration of the hardware components occupied by the target instruction; when all the instructions in the instruction queue have been issued and the hardware states of various hardware components in the hardware state space become idle states, the simulation results corresponding to the candidate compilation strategy are obtained.
[0167] In a possible implementation, determining the issue time of each instruction in the instruction set according to the hardware state space includes: for the i-th instruction in the instruction set, when the operand of the i-th instruction comes from a storage unit, based on the hardware state space, determining the activation state of the row where the operand of the i-th instruction is located in the storage unit, and according to the activation state of the row where the operand of the i-th instruction is located, the clock countdown of the storage unit, and the read latency for reading the operand from the storage unit, determining the arrival time of the operand of the i-th instruction at the computing unit; when the operand of the i-th instruction comes from a cache unit, determining the arrival time of the operand of the i-th instruction at the computing unit according to the release time of the cache unit, the read latency for reading the operand from the cache unit, and the idle time of the transmission bus between the cache unit and the computing unit; according to the arrival time of the operand of the i-th instruction at the computing unit and the idle time of the computing unit, determining the start processing time for the computing unit to start processing the operand of the i-th instruction; based on the start processing time for the computing unit to start processing the operand of the i-th instruction, determining the issue time of the i-th instruction.
[0168] In a possible implementation, selecting the instruction with the earliest issue time as the target instruction to be issued currently includes: selecting the instruction with the earliest start processing time as the target instruction to be issued currently.
[0169] According to the compilation system of the embodiments of the present disclosure, by using a tree-structured abstract model to describe a near-memory computing architecture, any near-memory computing architecture can be uniformly described in terms of timing using the tree structure. Furthermore, by combining the operator-level intermediate representation of a neural network model, a compilation strategy search space is constructed. Then, by predicting the performance of the compilation strategies in the search space, a plurality of candidate compilation strategies are determined. Then, through the emulator, based on the abstract model, the exact simulation results corresponding to each candidate compilation strategy for the execution process of the compilation result on the near-memory computing architecture are obtained to determine the final target compilation strategy and target compilation result. It can provide a better target compilation strategy for the deployment of any neural network model on any near-memory computing architecture, thereby realizing the deployment of any neural network model on any near-memory computing architecture, and having high generality.
[0170] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.
[0171] An embodiment of the present disclosure also provides an electronic device, including a memory adopting a near-storage computing architecture, a processor, and a computer program stored on the memory, where the processor executes the computer program to implement the steps of the above method.
[0172] An embodiment of the present disclosure also provides a non-volatile computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0173] An embodiment of the present disclosure also provides a computer program product, including a computer program or a non-volatile computer-readable storage medium carrying the computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.
[0174] Figure 10 The block diagram of an electronic device 1900 according to an embodiment of the present disclosure is shown. For example, the electronic device 1900 may be provided as a server or a terminal device. Referring to Figure 10 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a storage unit 1932 for storing instructions executable by the processing component 1922, such as application programs. Among them, the storage unit 1932 and the processing component 1922 may adopt a near-storage computing architecture, or may be separate memory and processor, and the embodiments of the present disclosure do not limit this. Among them, the application programs stored in the storage unit 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above compilation method.
[0175] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the storage unit 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0176] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the storage unit 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.
[0177] A computer-readable storage medium can be a tangible device that can retain and store programs / instructions used by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a dynamic random access memory (DRAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device, such as a punched card or raised structures in grooves storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as being a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagated through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0178] The computer programs (or computer-readable program instructions) described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0179] A computer program (or computer program instructions) for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit may execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0180] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0181] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. The computer-readable program instructions can also be stored in a computer-readable storage medium, which instructions cause a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0182] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0183] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or by combinations of special-purpose hardware and computer instructions.
[0184] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.
Claims
1. A compiling method, characterized in that: include: Obtaining an operator-level intermediate representation of a neural network model to be deployed to a near-storage computing architecture and hardware configuration information of the near-storage computing architecture, wherein the hardware configuration information is used to indicate a hardware configuration of the near-storage computing architecture; Based on the hardware configuration information of the near storage computing architecture, a tree-like architecture abstract model of the near storage computing architecture is generated, wherein the tree-like architecture abstract model is used to describe the near storage computing architecture in a tree structure; Based on the operator-level intermediate representation of the neural network model and the tree-structure abstract model of the near-storage computing architecture, construct a search space for a compilation strategy, wherein the compilation strategy is used to compile the operator-level intermediate representation; By performing performance prediction on the compilation strategies in the search space, a plurality of candidate compilation strategies and compilation results corresponding to each candidate compilation strategy are determined, wherein the compilation results include a complete instruction sequence obtained by compiling the operator-level intermediate representation according to the candidate compilation strategy; Based on the tree-structure abstract model, simulating the execution process of the compilation results corresponding to each candidate compilation strategy on the near storage computing architecture, and obtaining the simulation results corresponding to each candidate compilation strategy, wherein the simulation results represent the total time required for the compilation results to be executed on the near storage computing architecture; According to the simulation results corresponding to each candidate compilation strategy, a target compilation strategy is determined from the multiple candidate compilation strategies, and the compilation result of the target compilation strategy is determined as the target compilation result of the operator-level intermediate representation of the neural network model, so as to deploy the neural network model on the near storage computing architecture based on the target compilation result.
2. The method according to claim 1, characterized in that The near storage computing architecture includes a hardware structure of multiple levels, and a hardware structure of one level in the near storage computing architecture corresponds to a layer of nodes in the tree-like architecture abstract model. Each layer in the tree-like architecture abstract model includes at least one central node, and each central node is connected to at least one of the following child nodes: a central node of a next level, a storage node, a computing node, and a cache node in the level to which the central node belongs; the central node represents a data transmission relationship between two adjacent levels; Among them, the storage node represents the storage unit of the memory in the near storage computing architecture and includes the attribute information of the storage unit, the computing node represents the computing unit in the near storage computing architecture and includes the attribute information of the computing unit, and the cache node represents the cache unit in the near storage computing architecture and includes the attribute information of the cache unit.
3. The method according to claim 1, characterized in that A single compilation strategy includes a partitioning strategy and a corresponding mapping strategy, wherein the partitioning strategy represents a partitioning result of partitioning the operator-level intermediate representation along multiple dimensions at multiple levels, the mapping strategy represents a manner of mapping the partitioning result of the operator-level intermediate representation to a memory in the near storage computing architecture, and the partitioning result represents a plurality of sub-intermediate representations partitioned from the operator-level intermediate representation; The search space of the compilation strategy is constructed based on the operator-level intermediate representation of the neural network model and the tree-structure abstract model of the near-storage computing architecture, including: Based on the operator-level intermediate representation and the tree-like architecture abstract model, in order from the highest level to the lowest level of the memory in the near storage computing architecture, the operator-level intermediate representation is divided in multiple dimensions along multiple dimensions of the operator-level intermediate representation layer by layer under the first constraint condition to obtain multiple division strategies; wherein the first constraint condition includes that the product of the number of divisions along each dimension in the multiple dimensions at each level is less than or equal to the number of hardware structures corresponding to each level in the memory, and different division strategies have different division methods under the first constraint condition; For each sub-intermediate representation divided by each partitioning strategy, each sub-intermediate representation is divided into multiple data slices while satisfying the second constraint, and each data slice is mapped to the memory of the near storage computing architecture while satisfying the third constraint, so as to obtain multiple mapping strategies corresponding to each partitioning strategy; wherein the second constraint includes: the product of the size of each data slice and the data bit width of each data slice is less than or equal to the product of the burst length corresponding to the memory and the column bit width of the memory; the third constraint includes: the number of data slices mapped to the same row in the memory is less than or equal to the ratio between the total number of columns of the row in the memory and the burst length.
4. The method according to claim 1 or 3, characterized in that: The step of performing performance prediction on the compilation strategies in the search space to determine a plurality of candidate compilation strategies and compilation results corresponding to each candidate compilation strategy includes: Based on the hardware data corresponding to the near storage computing architecture, determining the performance upper limit corresponding to each compilation strategy in the search space, wherein the performance upper limit represents the upper limit of the time required for the compilation result corresponding to the compilation strategy to be executed on the near storage computing architecture, and the hardware data includes the bandwidth of the memory in the near storage computing architecture, the row miss cost, and the number and performance of the computing units; Based on the performance upper limits corresponding to the respective compilation strategies in the search space, determining a plurality of initial candidate compilation strategies whose performance upper limits are higher than a preset performance threshold; Using a compiler to partially compile the operator-level intermediate representation based on the initial candidate compilation strategies and a specific instruction set to obtain a partial instruction sequence compiled by each initial candidate compilation strategy, wherein the specific instruction set covers instructions executed on various near-storage computing architectures; Using a performance predictor to perform performance prediction on a portion of instruction sequences compiled by each initial candidate compilation strategy, to obtain predicted performance of the portion of instruction sequences compiled by each initial candidate compilation strategy, wherein the predicted performance represents the time required for the portion of instruction sequences to be executed on the near storage computing architecture; Determine a target prediction performance corresponding to each initial candidate compilation strategy based on the prediction performance of the partial instruction sequence compiled by each initial candidate compilation strategy and the number of loop executions corresponding to the partial instruction sequence, wherein the target prediction performance represents the time required for the complete instruction sequence compiled based on the initial candidate compilation strategy to be executed on the near storage computing architecture; Selecting a plurality of candidate compilation strategies from the plurality of initial candidate compilation strategies according to target prediction performance corresponding to each initial candidate compilation strategy; The compiler is used to completely compile the operator-level intermediate representation based on the multiple candidate compilation strategies and the specific instruction set to obtain compilation results corresponding to each candidate compilation strategy.
5. The method according to claim 4, characterized in that The determining, based on the hardware data corresponding to the near storage computing architecture, the performance upper limit corresponding to each compilation strategy in the search space includes: By analyzing each compilation strategy in the search space, an analysis result corresponding to each compilation strategy is obtained, wherein the analysis result includes a total amount of access data to the memory and a number of row misses, as well as a load calculation amount of a computing unit; According to the analysis results corresponding to the respective compilation strategies and the hardware data, the performance upper limit corresponding to each compilation strategy in the search space is determined.
6. The method according to claim 4, characterized in that The method of using a performance predictor to perform performance prediction on a portion of instruction sequences compiled by each initial candidate compilation strategy to obtain predicted performance of the portion of instruction sequences compiled by each initial candidate compilation strategy includes: Based on the partial instruction sequences compiled by each initial candidate compilation strategy, determine the total number of accesses to the memory in the near storage computing architecture, the number of row misses of continuous accesses to the memory, and the amount of data accessed by the computing unit in the near storage computing architecture in the partial instruction sequences compiled by each initial candidate compilation strategy; The predicted performance of the partial instruction sequences compiled by each initial candidate compilation strategy is determined based on the total number of accesses to the memory in the near storage computing architecture and the column-to-column delay, the number of row misses and the row miss cost of continuous access to the memory, and the amount of data accessed by the computing unit in the near storage computing architecture to the cache unit or register and the bus bandwidth.
7. The method according to claim 1, characterized in that The method uses a simulator to simulate the execution process of the compilation results corresponding to each candidate compilation strategy in the near storage computing architecture based on the tree-like architecture abstract model, and obtains the simulation results corresponding to each candidate compilation strategy; The simulator includes: a hardware state space, an instruction queue, and a virtual memory controller; the hardware state space is used to simulate and track the hardware state of each hardware component in the near storage computing architecture based on the tree architecture abstract model, and the hardware components include a storage unit, a computing unit, a cache unit, and a transmission bus; The instruction queue is used to store the complete instruction sequence compiled based on the candidate compilation strategy and maintain the data dependency relationship in the complete instruction sequence. The instruction queue includes multiple instruction groups. Independent instructions executed by different computing units are located in different instruction groups. The same instruction group includes multiple instructions to be executed sequentially. Instructions in instruction groups without data dependency can be issued in parallel. The virtual content controller is used to determine the issuing time of the instructions in the instruction queue according to the hardware state space and to update the hardware state of each hardware component in the hardware state space.
8. The method according to claim 7, characterized in that Tracking the hardware status of each hardware component in the near storage computing architecture includes: Tracking the activation status of each storage unit of the memory in the near storage computing architecture to issue an activation command before an inactivated storage unit is accessed; Tracking the currently activated row address in the storage unit to indicate whether a subsequent access triggers a row miss and scheduling a precharge command and an activate command to activate the row to be accessed when a row miss is triggered; Among them, multiple clock countdowns are used in the hardware status space to track the hardware status of each hardware component respectively, wherein the hardware status of the storage unit includes the time when the next command can be issued to the storage unit, and the hardware status of other hardware components except the storage unit includes an idle state or an occupied state; the clock countdown corresponding to each hardware component will be updated according to the hardware occupancy time of the hardware component involved in issuing the instruction.
9. The method according to claim 7 or 8, characterized in that: The using of the simulator to simulate the execution process of the compilation results corresponding to each candidate compilation strategy in the near storage computing architecture based on the tree-like architecture abstract model to obtain the simulation results corresponding to each candidate compilation strategy includes: For the compilation results corresponding to any candidate compilation strategy, the following simulation operations are performed cyclically: The instruction queue determines the instruction set that can be issued simultaneously based on the data dependency relationship involved in each instruction group, and sends the instruction set to the virtual memory controller, wherein the instruction set includes the instruction currently ranked first in each instruction group without data dependency relationship; The virtual content controller determines the issuing time of each instruction in the instruction set according to the hardware state space, and selects the instruction with the earliest issuing time as the target instruction to be issued currently; The virtual content controller controls the hardware state of each hardware component in the hardware state space to be updated to the hardware state corresponding to the issuance time of the target instruction based on the issuance time of the target instruction, and deletes the target instruction from the instruction queue; The virtual content controller updates the hardware state of the hardware component occupied by the target instruction in the hardware state space according to the hardware occupation time of the hardware component occupied by the target instruction; When all instructions in the instruction queue are issued and the hardware state of each hardware component in the hardware state space becomes an idle state, a simulation result corresponding to the candidate compilation strategy is obtained.
10. The method according to claim 9, characterized in that Determining the issuance time of each instruction in the instruction set according to the hardware state space includes: For an i-th instruction in the instruction set, when an operand of the i-th instruction comes from a storage unit, determining, based on the hardware state space, an activation state of a row in which the operand of the i-th instruction is located in the storage unit, and determining an arrival time of the operand of the i-th instruction at a computing unit according to the activation state of the row in which the operand of the i-th instruction is located, a clock countdown of the storage unit, and a read delay of reading the operand from the storage unit; In the case where the operand of the i-th instruction comes from a cache unit, determining the arrival time of the operand of the i-th instruction at the computing unit according to the release time of the cache unit, the read latency of reading the operand from the cache unit, and the idle time of the transmission bus between the cache unit and the computing unit; Determine a start processing time for the computing unit to start processing the operand of the i-th instruction according to an arrival time of the operand of the i-th instruction at the computing unit and an idle time of the computing unit; The issue time of the i-th instruction is determined based on the start processing time of the computing unit starting to process the operand of the i-th instruction.
11. The method according to claim 10, characterized in that The selecting the instruction with the earliest issuance time as the target instruction to be issued currently includes: selecting the instruction with the earliest start processing time as the target instruction to be issued currently.
12. A compilation system, characterized in that: include: An acquisition module, used to acquire an operator-level intermediate representation of a neural network model to be deployed to a near storage computing architecture and hardware configuration information of the near storage computing architecture, wherein the hardware configuration information is used to indicate a hardware configuration of the near storage computing architecture; An abstraction module, configured to generate a tree-structured abstract model of the near storage computing architecture based on hardware configuration information of the near storage computing architecture, wherein the tree-structured abstract model is configured to describe the near storage computing architecture in a tree-structured manner; A construction module, configured to construct a search space of a compilation strategy based on an operator-level intermediate representation of the neural network model and a tree-structure abstract model of the near-storage computing architecture, wherein the compilation strategy is used to compile the operator-level intermediate representation; A prediction module, configured to determine a plurality of candidate compilation strategies and compilation results corresponding to each candidate compilation strategy by performing performance prediction on the compilation strategies in the search space, wherein the compilation results include a complete instruction sequence obtained by compiling the operator-level intermediate representation according to the candidate compilation strategy; A simulation module, for simulating the execution process of the compilation results corresponding to each candidate compilation strategy on the near storage computing architecture based on the tree architecture abstract model, to obtain the simulation results corresponding to each candidate compilation strategy, wherein the simulation results represent the total time required for the compilation results to be executed on the near storage computing architecture; A determination module is used to determine a target compilation strategy from the multiple candidate compilation strategies according to the simulation results corresponding to each candidate compilation strategy, and determine the compilation result of the target compilation strategy as the target compilation result of the operator-level intermediate representation of the neural network model, so as to deploy the neural network model on the near storage computing architecture based on the target compilation result.
13. An electronic device comprising a memory and a processor, and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 11.
14. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Joint compiling method and system for heterogeneous hardware architecture
CN110968320A
Neural network compiling method for storage and calculation integrated platform
CN112465108A
Land space planning and compiling collaborative design platform based on spatial data subdivision
CN114328789A
Compiling method of multilayer intermediate representation based on field programmable logic gate array
CN118689486A
Joint compilation optimization method and device for heterogeneous computing and medium
CN118819543A
Cited By
Post quantum cryptography task simulation system and method
CN121173461A
Storage unit type selection method and device, electronic equipment and computer storage medium
CN121354611A