A DNN General Deployment Method and System Based on a Heterogeneous System
Through fine-grained splitting and mathematical optimization methods, the generality and efficiency problems of DNN deployment on heterogeneous systems are solved, efficient utilization of hardware resources and energy efficiency improvement are achieved, and the application of various application scenarios is adapted to various application scenarios.
Patent Information
- Application Number
- CN202510266459.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The existing DNN deployment acceleration solutions on heterogeneous systems lack universality and efficiency, resulting in low hardware resource utilization and inability to meet the needs of large-scale parameter calculations, especially in terms of latency and power consumption.
By splitting the deep neural network into a sublayer in fine-grained manner, combining the hardware parameters and communication link modeling of heterogeneous systems, monitoring delay and power consumption in real time, using mathematical planning optimization algorithms for offline iterative optimization, generating a final allocation plan, and optimizing the task allocation and storage management of the computing core.
It improves the utilization rate of hardware resources, improves system efficiency and energy efficiency, adapts to different application needs, and reduces development costs and time.
Smart Images

Figure CN119759595B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of artificial intelligence hardware acceleration, and relates to a general DNN deployment method and system based on a heterogeneous system. Background Art
[0002] Currently, with the continuous development of DNN, the scale of its parameters has shown an explosive growth. Taking some advanced image recognition models as an example, the number of parameters can reach billions or even tens of billions. This large-scale parameter growth has brought many severe challenges to the system. In terms of power consumption, in order to process such a huge amount of parameter calculations, hardware devices need to consume a large amount of electrical energy. The server clusters used to run DNN models in data centers consume an astonishing amount of electricity every year, which not only increases the operating cost but also places higher demands on energy supply. In terms of latency, a large number of computing tasks prolong the data processing time. In application scenarios with extremely high real-time requirements, such as autonomous driving and real-time video surveillance, longer latency may lead to serious consequences and the inability to make accurate decisions in a timely manner. In terms of memory resources, huge parameters need to occupy a large amount of memory space, which makes memory management extremely complex and prone to problems such as memory shortage or memory fragmentation.
[0003] To address these challenges, the DNN accelerator architecture has been continuously evolving. The early single-core designs have been gradually phased out, and dedicated architectures have emerged. In recent years, heterogeneous designs have become the focus of research and attention. A heterogeneous system usually consists of multiple different types of computing cores, such as CPUs, GPUs, FPGAs, ASICs, etc. Each core has its unique performance advantages and applicable scenarios. By combining different cores, their respective advantages can be fully utilized to improve the overall system performance. For example, GPUs are good at handling large-scale parallel computing tasks and are suitable for convolutional layer calculations in DNN; FPGAs have reconfigurability and can be customized according to different algorithm requirements; ASICs have extremely high efficiency and low power consumption in specific computing tasks.
[0004] In a heterogeneous multi-core architecture, in order to improve the core utilization rate, it is crucial to design an effective layer scheduling method. However, the existing layer scheduling strategies have many limitations. On the one hand, most existing scheduling strategies are specific to a certain hardware architecture and lack generality and universality. For example, some scheduling strategies designed for a specific GPU architecture cannot be effectively applied in other types of GPUs or heterogeneous systems containing multiple computing cores. This requires developers to redesign the scheduling strategy when facing different hardware platforms, increasing the development cost and time. On the other hand, there is also much room for improvement in the performance of existing scheduling strategies. They often cannot fully consider complex resource distributions, data communication delays, and collaborative work between different computing cores in a heterogeneous system, resulting in low system resource utilization and the overall performance not reaching the optimal level.
[0005] In summary, in the current deployment and acceleration of DNN on heterogeneous systems, there is an urgent need for a general and efficient solution to overcome the deficiencies of the existing technologies and meet the growing demands of artificial intelligence applications.
[0006] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present application, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0007] To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary is not a comprehensive review, nor is it intended to identify key / important elements or delineate the scope of protection of these embodiments. Instead, it serves as a prelude to the detailed description that follows.
[0008] Embodiments of the present disclosure provide a general DNN deployment method and system based on a heterogeneous system. During each DNN operation iteration, the latency and power consumption of the allocation scheme are monitored, and through mathematical modeling, the allocation scheme is optimized in real time, maximizing the utilization of hardware resources, improving hardware utilization, and enhancing system efficiency and energy efficiency.
[0009] In some embodiments, the method includes:
[0010] Fine-grained splitting of the deep neural network to be deployed, dividing it into multiple sub-layers as the minimum unit for heterogeneous system scheduling, and maintaining the data dependency relationship between sub-layers during the splitting process;
[0011] Based on the hardware parameters of the heterogeneous system, mapping the sub-layers to the computing cores of the heterogeneous system for preliminary mapping allocation;
[0012] Modeling the communication links of the heterogeneous system to form a preliminary allocation scheme, and estimating the latency and power consumption of each communication link;
[0013] According to the preliminary allocation scheme, monitor the communication latency and power consumption during the data transmission process in real time, and perform offline iterative optimization on the allocation scheme through a mathematical programming optimization algorithm to generate a final optimized allocation scheme;
[0014] Based on the final optimized allocation scheme, control the computing cores of the heterogeneous system to execute sub-layer computing tasks and manage the tensor life cycle of the storage module.
[0015] Preferably, during the fine-grained splitting of the deep neural network, the input-output dependency relationship between sub-layers is maintained through a data dependency relationship matrix.
[0016] Preferably, the latency of the communication link The calculation formula is as follows;
[0017] ,
[0018] wherein, represents the size of the communication data packet in the calculation core array, represents the data bit width of the dual - end calculation core interface, represents the energy consumption.
[0019] Preferably, the mathematical programming optimization algorithm performs offline iterative optimization on the allocation scheme, with minimizing the total delay and total power consumption as the objective function to optimize the allocation variables;
[0020] In the optimization process, the following constraint conditions are imposed: each sub - layer is only allocated to one calculation core; the allocation scheme needs to satisfy the hierarchical dependence of the data - dependence matrix; the weight storage capacity of each calculation core does not exceed the upper limit of the capacity of the directly connected storage module; the communication link delay of the special path does not exceed the preset threshold;
[0021] Repeat the optimization until the preset number of iterations is reached or the performance index is satisfied.
[0022] Preferably, the specific manner of the data transmission process is as follows:
[0023] According to the allocation scheme, transmit the tensor corresponding to the sub - layer to the target calculation core;
[0024] If the target storage module has sufficient space, set the active variables of all involved CLs to 1. After the transmission is completed, accumulate the delays and power consumptions of all active CLs to obtain the estimated total delay and power consumption; if the target storage module has insufficient space, evict the existing tensors according to the priority. The priority is calculated based on the following factors: whether the calculation core directly connected to the storage module has completed the calculation task involving the tensor; the size of the tensor and the time interval since the last use;
[0025] If there is still insufficient space after eviction, send a refresh signal to the off - chip storage, re - divide the tensor set and repeat the transmission process until the tensor transmission is completed and the estimated delay and power consumption are obtained. At this time, the estimated total delay and power consumption are the sum of the delays and power consumptions of all active CLs and the delays and power consumptions of the tensor refresh signal.
[0026] Preferably, when the storage space is insufficient in the eviction of existing tensors according to the priority, the tensors with the lowest priority and the smallest volume are evicted first.
[0027] In some embodiments, the system includes the following modules:
[0028] Allocation Manager: Used to maintain allocation variables, which are used to indicate the allocation status of sub - layers on target cores. When the set of allocation variables changes, an enable signal and the label of the new target core are sent to the storage section that stores the set of tensors required by the storage sub - layer, the packet header is regenerated, and data transmission is activated;
[0029] Transmission Manager: Maintains the power consumption and bit - width of each CL in terms of CL. Monitors the CL status through active variables and packet sizes, estimates the transmission delay and records it;
[0030] Storage Manager: Tracks the complete life cycle of each tensor in memory, performs memory load and eviction operations. Before the target core executes operations, it checks whether the set of tensors is stored on the section directly connected to the target core. If not, it determines whether the storage space is sufficient. When the space is insufficient, eviction operations are performed. If the space is still insufficient after all tensors are evicted, a tensor refresh signal is sent to off - chip storage;
[0031] Planning and Optimization Module: A planning and optimization accelerator based on RISCV that generates task assignments for mapping sub - layers to computing cores, that is, the set of allocation variables. It uses a mathematical planning optimizer to perform offline optimization of heterogeneous system resource deployment, and the optimization objective function is the total power consumption and total delay.
[0032] Preferably, the planning and optimization module has a built - in hash lookup table for maintaining the mapping relationship between the sub - layer topology and the computing core, and constrains the optimization process through a data - dependency matrix.
[0033] Preferably, the transmission manager monitors the handshake feedback signals of each computing core through active variables, dynamically updates the active state of the communication link. When the handshake feedback signal is pulled high, the active variable signal of the corresponding CL is pulled high, and the packet size data is requested from the core that sent the handshake feedback signal, that is, the current communication receiver, and recorded.
[0034] Preferably, the planning and optimization module uses a Gurobi optimizer.
[0035] A DNN general deployment method and system based on a heterogeneous system provided by an embodiment of the present disclosure can achieve the following technical effects:
[0036] After the layer allocation is completed in the present invention, the scheduling process is further refined. Each layer is divided into finer - grained sub - layers (SL, Sublayer). Although this division increases the difficulty of scheduling, it can output the data required for the next - layer operation faster, improving the operating efficiency of the system. For application scenarios with relatively low requirements for real - time performance and accuracy, multiple layers can be fused and regarded as a whole for operation scheduling. This flexible scheduling method can better adapt to different application requirements.
[0037] Aiming at the problem that the existing dedicated solutions for DNN allocation and deployment on heterogeneous systems lack generalizability, in each iteration of DNN operation, the present invention monitors the latency and power consumption of the allocation scheme and uses mathematical modeling to optimize the allocation scheme in real time, and then applies the optimized scheme to the heterogeneous system. The entire optimization process is completely offline and does not require additional operations on the host side, greatly improving the autonomy and operation efficiency of the system.
[0038] The above general description and the following description are only exemplary and explanatory and are not intended to limit the present application. Brief Description of the Drawings
[0039] One or more embodiments are exemplarily illustrated by corresponding drawings. These exemplary illustrations and the drawings do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation, and wherein:
[0040] Figure 1 is a schematic diagram of the method flow of the present invention;
[0041] Figure 2 is a schematic diagram of the fine-grained layer splitting of DNN;
[0042] Figure 3 is a schematic diagram of the topological storage structure of the heterogeneous system computing array;
[0043] Figure 4 is a schematic diagram of the many-core topology of the heterogeneous system computing array topology;
[0044] Figure 5 is a schematic diagram of the module structure of the heterogeneous system computing array;
[0045] Figure 6 is a schematic diagram of the tensor transmission process. Detailed Description of the Embodiments
[0046] In order to be able to understand the characteristics and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure will be described in detail below with reference to the drawings. The attached drawings are for reference and illustration purposes only and are not intended to limit the embodiments of the present disclosure. In the following technical description, for the sake of explanation, numerous details are provided to give a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be shown in a simplified manner to simplify the drawings.
[0047] In the description, claims, and above-mentioned accompanying drawings of the embodiments of the present disclosure, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so as to implement the embodiments of the present disclosure described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion.
[0048] Unless otherwise specified, the term "plurality" means two or more.
[0049] A DNN general deployment method based on a heterogeneous system. The core of this method lies in the resource evaluation and scheduling of the hardware part, with a focus on the communication cost and transmission delay of the data link. In the software part, first, the DNN to be deployed is split into fine-grained layers, which are split into a set of multiple sub-layers (SLs). Each SL can be scheduled independently and is the smallest unit of scheduling for the entire heterogeneous system. During the splitting process, the data dependency relationship is strictly maintained to ensure the correctness of the calculation. At the same time, according to the hardware parameters of each core, all SLs are mapped to each core to complete the preliminary mapping allocation. In the hardware part, each module works together to achieve efficient scheduling evaluation.
[0050] Specifically, as Figure 1 shown, it includes:
[0051] S1: Split the deep neural network to be deployed into fine-grained layers, which are divided into multiple sub-layers as the smallest unit of scheduling for the heterogeneous system, and maintain the data dependency relationship between each sub-layer during the splitting process.
[0052] S2: Based on the hardware parameters of the heterogeneous system, map the sub-layers to the computing cores of the heterogeneous system for preliminary mapping allocation.
[0053] S3: Model the communication links of the heterogeneous system to form a preliminary allocation plan, and estimate the delay and power consumption of each communication link.
[0054] S4: According to the preliminary allocation plan, monitor the communication delay and power consumption during the data transmission process in real time, and perform offline iterative optimization on the allocation plan through a mathematical programming optimization algorithm to generate a final optimized allocation plan.
[0055] S5: Based on the final optimized allocation plan, control the computing cores of the heterogeneous system to execute the sub-layer calculation tasks and manage the tensor life cycle of the storage module.
[0056] As a refinement of the above embodiment, during the process of splitting the deep neural network into fine-grained layers, the input-output dependency relationship between each sub-layer is maintained through a data dependency relationship matrix.
[0057] As a refinement of the above embodiments, as Figure 2 shown, each transmission link between cores is called a CL (communication link). The bandwidth performance of the CL can be predetermined when forming a heterogeneous system, that is, the data bit width of the dual-core computing interface, and the latency of the communication link The calculation formula is as follows;
[0058] ,
[0059] where, represents the size of the communication data packet in the computing core array, represents the data bit width of the dual-core computing interface, represents the energy consumption (process, length, clock frequency). Through the manufacturing process, physical link length, and clock frequency for estimation, the power consumption of each CL can be estimated during system initialization and maintained in the transmission manager in the form of a lookup table.
[0060] As a refinement of the above embodiments, the mathematical programming optimization algorithm performs offline iterative optimization on the allocation scheme, with minimizing the total latency and total power consumption as the objective function to optimize the allocation variables;
[0061] The following constraints are imposed during the optimization process: Each sublayer is only allocated to one computing core; the allocation scheme needs to satisfy the hierarchical dependencies of the data dependency matrix; the weight storage capacity of each computing core does not exceed the upper limit of the capacity of the storage module directly connected to it; the communication link latency of the special path does not exceed the preset threshold;
[0062] Repeat the optimization until the preset number of iterations is reached or the performance index is satisfied.
[0063] As a refinement of the above embodiments, the specific manner of the data transmission process is as follows:
[0064] According to the allocation scheme, transmit the tensor corresponding to the sublayer to the target computing core;
[0065] If the target storage module has sufficient space, set the active variables of all involved CLs to 1. After the transmission is completed, accumulate the latency and power consumption of all active CLs to obtain the estimated total latency and power consumption; if the target storage module has insufficient space, evict the existing tensors according to the priority, and the priority is calculated based on the following factors: whether the computing core directly connected to the storage module has completed the computing task involving the tensor; the size of the tensor and the time interval since the last use;
[0066] If there is still insufficient space after eviction, a refresh signal is sent to off-chip storage, the tensor set is re-partitioned, and the transmission process is repeated until the tensor transmission is completed and the estimated latency and power consumption are obtained. The estimated total latency and power consumption at this time are the sum of the latency and power consumption of all active CLs and the latency and power consumption of the tensor refresh signal.
[0067] As a refinement of the above embodiment, when the storage space is insufficient during the eviction of existing tensors according to priority, the tensors with the lowest priority and the smallest volume are preferentially evicted.
[0068] Embodiment 2
[0069] A DNN general deployment system based on a heterogeneous system, comprising the following modules:
[0070] Allocation manager: used to maintain allocation variables, which are used to indicate the allocation of sub-layers on target cores. When the set of allocation variables changes, an enable signal and the label of the new destination core are sent to the storage section storing the tensor set required by the storage sub-layer, the packet header is regenerated, and data transmission is activated.
[0071] Transmission manager: maintains the power consumption and bit width of each CL in terms of CL, monitors the CL status through active variables and packet sizes, estimates the transmission latency and records it.
[0072] Storage manager: tracks the complete life cycle of each tensor in memory, performs memory load and eviction operations. Before the target core executes an operation, it checks whether the tensor set is stored on the section directly connected to the target core. If not, it determines whether the storage space is sufficient. When the space is insufficient, eviction operations are performed. If the space is still insufficient after all tensors are evicted, a tensor refresh signal is sent to off-chip storage.
[0073] Planning and optimization module: a planning and optimization accelerator based on RISCV, generates task allocations for mapping sub-layers to computing cores, that is, the set of allocation variables, and uses a mathematical programming optimizer to perform offline optimization on the resource deployment of the heterogeneous system. The optimization objective function is the total power consumption and the total latency.
[0074] As a refinement of the above embodiment, Figure 5 The allocation manager in it is the allocation management module of the entire heterogeneous system, mainly maintaining allocation variables , (indicating the allocation situation on core j. If it is 1, then core j is responsible for executing the computing task of, indicating the th sub-layer). When a is allocated to core j, the input tensor set Tinput and the weight tensor set Tweight together form the required tensor set When the set changes, it indicates that the allocation mapping relationship has changed. The allocation manager sends an enable signal and the label of the new destination core j to the storage storage segment, regenerates the packet header, and activates data transmission. The specific data path is determined by the routing policy and the on-chip network.
[0075] As a refinement of the above embodiment, Figure 5 the transmission manager in it is an adapted hardware module designed for the above scheduling evaluation scheme. The transmission manager maintains the power consumption and bit width of each CL in terms of CL, and monitors the CL status through the active variable act and the packet size packet_size. The active variable act is modified by monitoring the handshake feedback Response signal of each computing core. When the Response signal is pulled high, it indicates that the inter-core communication has completed the handshake and starts, then the act signal of the corresponding CL is pulled high, and the packet size data is requested from the core that sent the Response signal, that is, the receiver of the current communication, and recorded. Subsequently, the transmission manager estimates the transmission delay and records it.
[0076] As a refinement of the above embodiment, Figure 5 the storage manager in it is another additionally designed hardware module facing the storage side, aiming to track the complete life cycle of each tensor in the memory and perform memory loading and eviction operations. Before core j executes the relevant operations, it needs to be transferred to the target core. Before starting, as shown in the overall architecture of the heterogeneous system in Figure 2 , the core and the storage segment are not fully connected structures. The storage manager first checks whether it is stored on the segment directly connected to core j, that is, the input / output value storage IOMem_j and the weight value storage WeightMem_j. Then the cost is only the storage access consumption e_memor and the access delay t_memor. Otherwise, the storage manager judges whether the storage space is sufficient to accommodate the upcoming tensor set.
[0077] If the space is sufficient, the tensor is transferred to the target memory. The transfer process is controlled by the allocation manager and follows the preset transmission policy. At the same time, the transmission manager sets the active variable act of all involved CLs to 1. When the transmission work of Ti is completed, the transmission manager accumulates the delays and power consumptions of all active CLs to obtain the estimated total delay and power consumption:
[0078] ,
[0079] ,
[0080] wherein, represents the total delay, Indicates the latency on the active CLs involved. Indicates the memory access latency. Indicates the total power consumption. Indicates the power consumption on the active CLs involved. Indicates the power consumption of the memory access latency. The power consumption and latency of each CL are obtained through a lookup table.
[0081] If there is insufficient space, an eviction operation needs to be performed on the storage space, and one or more low-priority tensors will be evicted. The priority of the tensors is maintained in the storage manager and is affected by system design, size, and the interval of recent use. Specifically, the storage manager checks whether the computing cores directly connected to the current storage segment have completed all the calculations involving the tensor. The more cores that have not completed, the higher the priority of the tensor. When an eviction operation needs to be carried out, the storage manager retrieves all the tensors with the lowest priority and preferentially evicts the smallest tensors to minimize the negative impact of tensor unloading, and then rechecks whether the space is sufficient. If the remaining space still cannot accommodate the upcoming tensors after all tensors have been evicted, the storage manager sends a tensor refresh signal T_flush to the off-chip storage, requesting to re-partition the tensor transfer packets to obtain a smaller set of tensors. Repeat the above steps until the tensor transfer is completed and the estimated latency and power consumption are obtained. The estimated total latency and power consumption at this time are:
[0082] ,
[0083] ,
[0084] Among them, Indicates the latency on the active CLs involved when there is insufficient space. Indicates the time required for sending the tensor refresh signal and reorganizing the tensor packets. Indicates the power consumption on the active CLs involved when there is insufficient space. Indicates the power consumption required for sending the tensor refresh signal and reorganizing the tensor packets.
[0085] As a refinement of the above embodiment, Figure 5 The planning optimization module in is a RISC-V-based planning optimization accelerator. By the highly modular and customizable characteristics of RISC-V ISA, the structure is streamlined. This module generates a task assignment that maps SL to the computing cores, that is Set. The SL granularity topology of the entire DNN and its mapping relationship with the computing cores are maintained through a built-in hash lookup table, and the data dependency matrix Matrix_dependency among individual SLs is reflected. This module mainly deploys a mathematical programming optimizer, and Gurobi is adopted in this solution to perform offline optimization on the heterogeneous system resource deployment.
[0086] The objective function for optimization is the total power consumption and the total latency, with the total latency being the main optimization objective. At the same time, the optimization tasks need to meet the following constraints:
[0087] (1) , indicating that one SL can only be assigned to one computing core;
[0088] (2) The assignment of SLs to cores must satisfy the data dependencies between different levels and needs to follow the results of the lookup table;
[0089] (3) The total weight stored on each core shall not exceed the weight storage upper limit of the board directly connected to the core;
[0090] (4) For special paths, the data link latency needs to meet the path latency requirements.
[0091] The planning and optimization module executes the task optimization task to generate the optimized Set. Based on the new set for assignment, and estimate the total latency and total power consumption of the optimized system. Repeat this step until the optimization iteration times reach the upper limit or meet the preset performance requirements. Thus, the optimization deployment solution of the entire DNN on the heterogeneous system is completed.
[0092] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure, enabling those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, process, and other changes. The embodiments only represent possible variations. Unless explicitly required, the individual components and functions are optional, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terms used in this application are only for describing the embodiments and do not limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to also include the plural forms. Similarly, as used in this application, the term "and / or" refers to any and all possible combinations including one or more of the associated listed items. Additionally, when used in this application, the term "comprise" and its variants "comprises" and / or "comprising" etc. mean the presence of the stated features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or groups of these. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, or device comprising the element. Herein, what each embodiment focuses on can be the differences from other embodiments, and the same or similar parts among the embodiments can be referred to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method parts disclosed in the embodiments, the relevant parts can refer to the description of the method parts.
[0093] Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner can depend on the specific application and design constraints of the technical solution. The skilled person can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the embodiments of the present disclosure. The skilled person can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0094] In the embodiments disclosed herein, the disclosed methods, products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units can be merely a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Additionally, the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms. The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to implement this embodiment. Additionally, in the embodiments of the present disclosure, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0095] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the block can occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, which can depend on the functions involved. In the descriptions corresponding to the flowcharts and block diagrams in the accompanying drawings, the operations or steps corresponding to different blocks can also occur in a different order than that disclosed in the description. Sometimes, there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, which can depend on the functions involved. Each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
Claims
1. A general DNN deployment method based on a heterogeneous system, characterized in that, Including: Perform fine-grained splitting on the deep neural network to be deployed, divide it into multiple sub-layers, which serve as the smallest units for heterogeneous system scheduling, and maintain the data dependency relationships among the sub-layers during the splitting process; Based on the hardware parameters of the heterogeneous system, map the sub-layers to the computing cores of the heterogeneous system for preliminary mapping allocation; Model the communication links of the heterogeneous system to form a preliminary allocation plan, and estimate the latency and power consumption of each communication link; According to the preliminary allocation plan, monitor the communication latency and power consumption during the data transmission process in real time, and perform offline iterative optimization on the allocation plan through a mathematical programming optimization algorithm to generate a final optimized allocation plan; Based on the final optimized allocation plan, control the computing cores of the heterogeneous system to execute the sub-layer computing tasks and manage the tensor life cycle of the storage module; The specific method of the data transmission process is as follows: According to the allocation plan, transmit the tensors corresponding to the sub-layers to the target computing core; If the target storage module has sufficient space, set the active variables of all involved CLs to 1. After the transmission is completed, accumulate the latency and power consumption of all active CLs to obtain the estimated total latency and power consumption. If the target storage module has insufficient space, evict the existing tensors according to the priority. The priority is calculated based on the following factors: whether the computing core directly connected to the storage module has completed the computing task involving this tensor; the size of the tensor and the time interval since the last use; If there is still insufficient space after eviction, send a refresh signal to the off-chip storage, re-partition the tensor set and repeat the transmission process until the tensor transmission is completed and the estimated latency and power consumption are obtained. At this time, the estimated total latency and power consumption are the accumulated latency and power consumption of all active CLs and the latency and power consumption of the tensor refresh signal.
2. The DNN general deployment method based on a heterogeneous system according to claim 1, wherein During the fine-grained splitting process of the deep neural network, maintain the input-output dependency relationships among the sub-layers through a data dependency relationship matrix.
3. The DNN general deployment method based on a heterogeneous system according to claim 1, wherein The delay of the communication link The calculation formula is as follows; , Among them, represents the size of communication data packets in the computing core array, represents the data bit width of the dual-end computing core interface, represents the energy consumption.
4. The DNN general deployment method based on a heterogeneous system according to claim 1, wherein The mathematical programming optimization algorithm performs offline iterative optimization on the allocation plan, with minimizing the total latency and total power consumption as the objective function to optimize the allocation variables; Apply the following constraint conditions during the optimization process: each sub-layer is only allocated to one computing core; the allocation plan needs to satisfy the hierarchical dependencies of the data dependency relationship matrix; the weight storage capacity of each computing core does not exceed the upper limit of the capacity of the directly connected storage module; the communication link latency of the special path does not exceed the preset threshold; Repeat the optimization until the preset number of iterations is reached or the performance index is satisfied.
5. The DNN general deployment method based on a heterogeneous system according to claim 1, characterized in that When the storage space is insufficient during the eviction of existing tensors according to the priority, preferentially evict the tensors with the lowest priority and the smallest volume.
6. A DNN general deployment system based on a heterogeneous system using the method according to any one of claims 1-5, characterized in that, Including the following modules: Allocation Manager: used to maintain the allocation variables, which are used to indicate the allocation status of the sub-layers on the target cores. When the set of allocation variables changes, send an enable signal and the label of the new destination core to the storage section storing the tensor set required by the sub-layer, regenerate the packet header, and activate the data transmission; Transmission Manager: maintain the power consumption and bit width of each CL in terms of CL, monitor the CL status through active variables and packet sizes, estimate the transmission latency and record it; Storage Manager: Tracks the complete life cycle of each in-memory tensor, performs memory load and eviction operations. Before the operations are executed on the target core, it checks whether the tensor set is stored on the board directly connected to the target core. If not, it determines whether the storage space is sufficient. When the space is insufficient, it performs the eviction operation. If the space is still insufficient after all tensors are evicted, it sends a tensor refresh signal to the off-chip storage. Planning Optimization Module: A planning optimization accelerator based on RISCV that generates task assignments for mapping sub-layers to computing cores, that is, assigns a set of variables, and uses a mathematical programming optimizer to perform offline optimization of heterogeneous system resource deployment. The optimization objective function is the total power consumption and the total latency.
7. The DNN general deployment system based on a heterogeneous system according to claim 6, characterized in that, The planning optimization module has a built-in hash lookup table for maintaining the mapping relationship between the sub-layer topology and the computing core, and constrains the optimization process through a data dependency matrix.
8. The DNN general deployment system based on a heterogeneous system according to claim 6, characterized in that, The transmission manager monitors the handshake feedback signals of each computing core through active variables, dynamically updates the active state of the communication link. When the handshake feedback signal is pulled high, it pulls high the active variable signal of the corresponding CL, and requests the packet size data from the core that sent the handshake feedback signal, that is, the current communication receiver, and records it.
9. The DNN general deployment system based on a heterogeneous system according to claim 6, characterized in that The planning optimization module uses a Gurobi optimizer.
Citation Information
Patent Citations
Simplified convolutional neural network-oriented low-cost accelerator architecture and processing method thereof
CN111191774A
Anti-radiation low-delay neural network reasoning acceleration chip
CN117474061A