Operator fusion method based on Shenwei processor

By performing operator fusion on Shenwei processor and preloading parameters to local data memory, the problems of insufficient parallel processing capabilities and low memory bandwidth utilization in the prior art are solved, and more efficient computing and performance improvements are achieved.

CN119759583BActive Publication Date: 2025-05-23青岛国实科技集团有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510244893.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-23
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The existing technology fails to fully consider the multi-core architecture of Shenwei processors, resulting in the failure to fully utilize parallel processing capabilities, affecting the execution speed of the algorithm. In addition, the existing operator fusion method has insufficient optimization of the memory access mode during the fusion process, resulting in a low memory bandwidth utilization rate, which further restricts the improvement of algorithm performance.

Method used

An operator fusion method based on Shenwei processor is provided. By obtaining the operators of the calculation graph in the deep learning model, analyzing their data access mode, determining the fusion operators that meet the fusion conditions based on the preset fusion judgment mechanism, and preloading their parameters into the local data memory of Shenwei processor, reducing the frequent transmission of data between the processor and the external memory.

Benefits of technology

Make full use of the computing resources of Shenwei processors to improve parallel computing efficiency, optimize memory access mode, reduce memory bandwidth bottlenecks, and significantly improve the execution speed and overall performance of the algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119759583B_ABST
    Figure CN119759583B_ABST
Patent Text Reader

Abstract

The present invention relates to an operator fusion method based on a Shenwei processor, which belongs to the technical field of operator fusion. The method comprises obtaining operators of a calculation graph in a deep learning model, analyzing data access modes of the operators, determining operators to be fused that meet fusion conditions according to a preset fusion judgment mechanism, and obtaining feature graphs of the operators to be fused, performing data processing on the feature graphs using the operators to be fused, pre-loading parameters of the operators to be fused into a local data memory of the Shenwei processor, judging the size of the total data flow of the feature graph and the capacity of the local data memory, dividing data blocks of the feature graphs according to the judgment results, completing operations of the operators to be fused using the data blocks, and outputting the data blocks after the operations, thereby reducing the frequent transmission of data between the processor and the external memory, making full use of the computing resources of the Shenwei processor, improving the efficiency of parallel computing, optimizing the memory access mode, and reducing the bottleneck of the memory bandwidth.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and particularly to an operator fusion method based on the Shenwei processor. Background Art

[0002] In the field of deep learning, operator fusion is an effective means to improve computing efficiency. Operator fusion reduces the number of data transfers between memory and computing units by combining multiple computing operations, reduces computing latency, and thus improves the execution speed of the algorithm. Due to the increasing computing complexity and data scale of deep learning algorithms, higher requirements are also placed on computing devices.

[0003] The Shenwei processor is a high-performance processor with significant advantages in floating-point computing, parallel processing, etc. The Shenwei processor adopts a many-core architecture, has a large number of computing cores and high parallel capabilities, providing strong support for the acceleration of deep learning algorithms.

[0004] However, existing operator fusion methods are mainly designed for general-purpose processors such as CPUs and GPUs, and there are certain limitations when executing deep learning tasks, such as low energy efficiency ratio and insufficient parallel processing capabilities. The characteristics of the Shenwei processor are not fully utilized, resulting in poor operator fusion effects when executing deep learning algorithms on the Shenwei processor. In addition, existing operator fusion technologies do not fully consider the many-core architecture of the Shenwei processor, resulting in insufficient utilization of parallel processing capabilities and affecting the execution speed of the algorithm. Moreover, in the fusion process of existing operator fusion methods, the memory access pattern is not optimized enough, resulting in low memory bandwidth utilization, further restricting the improvement of algorithm performance. Summary of the Invention

[0005] Aiming at the deficiencies in the related technologies, the purpose of the present invention is to provide an operator fusion method based on the Shenwei processor to solve the technical problems that the existing technology fails to fully consider the many-core architecture of the Shenwei processor, resulting in insufficient utilization of parallel processing capabilities and affecting the execution speed of the algorithm, and that in the fusion process of existing operator fusion methods, the memory access pattern is not optimized enough, resulting in low memory bandwidth utilization, further restricting the improvement of algorithm performance.

[0006] The present invention provides an operator fusion method based on the Shenwei processor, including the following steps:

[0007] Data acquisition step: Acquire the operators of the computation graph in the deep learning model, analyze the data access patterns of the operators, determine the operators to be fused that meet the fusion conditions according to a preset fusion determination mechanism, and acquire the feature maps of the operators to be fused;

[0008] Data loading step: the feature graph is processed using the operator to be fused, wherein the parameters of the operator to be fused are pre-loaded into the local data memory of the Shenwei processor;

[0009] Data operation step: determine the total data flow of the feature map and the size of the capacity of the local data storage device, divide the data blocks of the feature map according to the determination result, use the data blocks to complete the operation of the operator to be fused, and output the data blocks after the operation.

[0010] The embodiment of the present invention reduces the frequent transmission of data between the processor and the external memory by fusing operators that meet the fusion conditions and pre-loading the parameters into the local data memory of the Shenwei processor, which can fully utilize the computing resources of the Shenwei processor, improve the efficiency of parallel computing, optimize the memory access mode, and reduce the bottleneck of memory bandwidth.

[0011] In some embodiments of the present invention, the data calculation step is specifically:

[0012] The spatial area of ​​the local data memory is divided into a calculation area, a cache area and a communication area according to the working set size of the operator to be fused. The calculation area is used to store data required for calculation, the cache area is used to store temporary data, and the communication area is used for data exchange buffer.

[0013] The embodiment of the present invention divides the spatial area of ​​the local data storage into a computing area, a cache area and a communication area, thereby achieving efficient management and utilization of memory resources, reducing data access conflicts and memory bandwidth pressure, and thus significantly improving computing efficiency and overall performance.

[0014] In some embodiments of the present invention, the process of determining the total data flow of the feature graph and the size of the capacity of the local data storage is specifically as follows:

[0015] If the total data flow of the feature map exceeds the capacity of the local data storage, the optimal block size is dynamically calculated, and the operation of the operator to be fused is completed in sequence using multiple data blocks in a block pipeline manner according to the calculated optimal block, and the data blocks after the operation are output;

[0016] Otherwise, the data blocks in the feature map are continuously loaded into the local data memory to complete the calculation of the operator to be fused, and the calculated data blocks are output.

[0017] The embodiment of the present invention dynamically determines the relationship between the total data flow of the feature map and the capacity of the local data storage device, and flexibly selects the optimal block size or continuous loading method. When the feature map data flow exceeds the capacity of the local data storage device, a block pipeline method is adopted to divide the data into optimal blocks suitable for the storage capacity to avoid memory overflow and make full use of computing resources. When the feature map data flow does not exceed the capacity, a continuous loading method is adopted to reduce the overhead of data segmentation and loading, improve computing efficiency, and maximize performance and resource utilization.

[0018] In some embodiments of the present invention, the process of using a block pipeline to sequentially complete the operation of the operator to be fused using multiple data blocks according to the calculated optimal block is specifically as follows:

[0019] Two buffers are divided in the calculation area of ​​the local data memory, and a DMA request is sent to the DMA controller according to the interface of the Shenwei processor. After the data required for calculation in the buffer is obtained according to the DMA request, the operation of the operator to be fused is completed.

[0020] The embodiment of the present invention divides two buffers in the calculation area of ​​the local data storage. When one buffer is used for calculation, the other buffer pre-fetches the next block of data through the DMA controller, thereby realizing the parallelization of calculation and data loading, reducing the waiting time, and directly obtaining the data required for calculation through the DMA controller, avoiding the frequent intervention of the CPU, and improving the data transmission efficiency. Multiple data blocks are calculated in turn through the buffers, forming an efficient pipeline processing mode, which significantly improves the calculation throughput and overall performance.

[0021] In some embodiments of the present invention, the calculation model for dynamically calculating the optimal block size is:

[0022]

[0023] in, is the optimal block size; is the maximum value of the preset block; is the available space of the local data memory; is the number of operators to be fused.

[0024] The embodiment of the present invention dynamically calculates the optimal block size, combines the available space of the local data storage and the number of operators to be fused, and ensures that the block size can fully utilize storage resources and meet computing requirements.

[0025] In some embodiments of the present invention, the data loading step specifically includes:

[0026] The data access mode is determined according to the architecture mode of the Shenwei processor, and the method of loading the parameters of the operator to be fused into the local data memory is selected according to the data access mode, wherein:

[0027] When the data access mode is continuous access, a DMA request is sent to a DMA controller according to an interface of the Shenwei processor, and the DMA controller continuously loads the data blocks in the feature map into a local data memory according to the DMA request;

[0028] When the data access mode is stride access, the data blocks of the feature map are rearranged in the main memory, the discrete data are combined into continuous data blocks, and the data blocks are continuously loaded into the local data memory by sending a DMA request to the DMA controller;

[0029] When the data access mode is random access, after analyzing and identifying data in the data block whose access times exceed a preset number, the data is moved to adjacent positions to form continuous data blocks, and then a DMA request is sent to the DMA controller to continuously load the data blocks into the local data memory.

[0030] In the embodiment of the present invention, when the data access mode is continuous access, continuous data blocks are directly loaded into the local data memory through the DMA controller, making full use of the efficient data transmission capability of DMA and reducing CPU intervention; during stride access, discrete data are grouped into continuous data blocks through data rearrangement and then loaded through DMA, avoiding the inefficiency of stride access and improving data transmission efficiency; during random access, high-frequency access data is analyzed and reorganized into continuous data blocks, thereby optimizing the data loading efficiency in the random access mode, reducing memory access latency, and being able to dynamically select the data loading method according to the data access mode to ensure that data is efficiently loaded into the local data memory.

[0031] In some embodiments of the present invention, the process of determining the operators to be fused that meet the fusion conditions according to the preset fusion determination mechanism is specifically as follows:

[0032] Analyze the data dependency between operators to confirm whether there are operator pairs with direct data dependency and no circular dependency, and whether the output of an operator is only used by one subsequent operator;

[0033] Evaluate the computing characteristics of the operators, confirm whether the computing density between operators is similar, determine whether the access to the operators meets the alignment requirements preset by the Shenwei processor, and whether the total resource demand after the fusion of the operators exceeds the capacity limit of a single CPE;

[0034] Based on a preset benefit indicator, a quantitative analysis is performed to evaluate whether the benefit indicator after the operator fusion meets the preset benefit indicator, and the benefit indicator includes the number of data transfer times, data reuse rate and calculation density.

[0035] The embodiment of the present invention avoids data conflicts or calculation errors that may be introduced after fusion by confirming whether there is direct data dependency and no circular dependency between operators, and whether the output of an operator is only used by one subsequent operator. By evaluating the computing density, alignment requirements and resource requirements of the operators, it is ensured that the fused operators can run efficiently on the Shenwei processor to avoid resource overload or performance degradation. The benefits after fusion are quantitatively evaluated through preset benefit indicators to ensure that the fusion operation can significantly improve performance, reduce data transfer overhead and improve computing efficiency.

[0036] In some embodiments of the present invention, the data acquisition step includes:

[0037] Classifying and analyzing the operators according to the computational characteristics of the operators, and classifying the operators in the deep learning model into computationally intensive operators, communication-intensive operators, and hybrid operators;

[0038] The computation-intensive operator completes the computation within a single CPE, the communication-intensive operator completes the computation collaboratively between multiple CPEs by optimizing the communication path, and the hybrid operator completes the computation collaboratively between multiple CPEs.

[0039] The embodiment of the present invention divides operator types. Computation-intensive operators complete calculations within a single CPE, fully utilizing the computing power of a single computing core and avoiding communication overhead. Communication-intensive operators collaboratively complete calculations between multiple CPEs by optimizing communication paths, reducing communication delays and improving data transmission efficiency. Hybrid operators collaboratively complete calculations between multiple CPEs, balancing computing and communication requirements and maximizing overall performance. Operators are classified into computationally intensive, communication-intensive and hybrid types, and different computing strategies are adopted for different types of operators to optimize resource allocation and computing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the specific embodiments of the present invention will be described in detail below in conjunction with the accompanying drawings. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0041] Figure 1 A structural diagram of a Shenwei processor provided in an embodiment of the present invention;

[0042] Figure 2 A memory structure diagram of a Shenwei processor provided in an embodiment of the present invention;

[0043] Figure 3 A GPU memory structure diagram provided for an embodiment of the present invention;

[0044] Figure 4 A flowchart of an operator fusion method based on a Shenwei processor provided in an embodiment of the present invention;

[0045] Figure 5 A flowchart of a data operation step provided by an embodiment of the present invention;

[0046] Figure 6 A flowchart of a data loading step provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0047] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. Based on the embodiments provided in the present application, all other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of the present application.

[0048] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0049] The operator fusion solutions in existing technologies are mainly designed for GPU architecture, relying on shared memory and cache mechanisms, and are unable to fully utilize the hardware advantages of the Shenwei processor.

[0050] As attached Figure 1 -Attached Figure 2 As shown, the Sunway processor of the embodiment of the present invention is a 4×4 CPE grid structure architecture design without shared memory, cache, and each CPE is equipped with an independent 256KB local data memory. Among them, CPE (Computing Processing Element) is a computing processing unit, which is one of the core computing units in the Sunway processor architecture and is used to perform specific computing tasks.

[0051] The Shenwei processor adopts a multi-core design. Each computing core group contains a master core and multiple slave cores. The master core is responsible for controlling and coordinating the work of the slave cores and performing complex tasks. The slave cores focus on performing specific computing tasks and are usually larger in number to achieve high parallelism.

[0052] Among them, the architecture module of Shenwei processor:

[0053] Registers are used to store temporary data and instructions and provide fast access.

[0054] LDM (Local Data Memory) is used for high-speed data storage and access from the core.

[0055] Dcache (data cache) is used to store frequently accessed data to reduce the latency of accessing main memory.

[0056] The second-level cache (secondary cache) provides a larger cache capacity for storing more data and further improving access speed.

[0057] Main memory, or primary storage, has a larger capacity but is slower to access and is used to store all data and instructions.

[0058] The Shenwei processor contains multiple core groups, each core group has 6 master cores and 64 slave cores. Data is loaded from the main memory to the secondary cache, then to the data cache, and finally to the register or local data memory for use by the master core and slave core.

[0059] As attached Figure 3 As shown, the GPU memory architecture includes:

[0060] GPU core, the computing unit of the GPU, is used to process graphics and computing tasks in parallel.

[0061] GPU video memory, a high-speed memory dedicated to the GPU, used to store textures, frame buffers, and other graphics data.

[0062] System memory, CPU main memory, GPU accesses system memory through PCIe bus.

[0063] When the GPU processes data, the data is transferred from the system memory to the GPU video memory, and then to the shared memory and registers of the GPU core for use by the computing unit.

[0064] There are significant differences between the Shenwei processor and the GPU. The Shenwei processor uses a unified memory architecture, and its 6 main cores and 384 slave cores share the same main memory. This design eliminates the need for each core to frequently switch and move data between different memory areas when accessing and processing data, improving computing efficiency and simplifying memory management. In traditional GPU architectures, there is a separation between CPU memory and video memory, and the interaction of data between the two requires an additional transmission process, which increases the delay and complexity of data processing to a certain extent, affecting the smoothness and efficiency of the overall operation.

[0065] The memory bandwidth of the Shenwei processor is seriously inferior to that of the GPU. The aggregate memory bandwidth of the Shenwei processor is 300GB / s, while the GPU memory bandwidth exceeds 3000GB / s (H100). For example, NVIDIA's H100 GPU memory bandwidth can reach more than 3TB / s, which is much higher than the memory bandwidth of the Shenwei processor. At the same time, GPU processors generally have multiple levels of hardware cache, such as level 1 cache, level 2 cache, and level 3 cache, which can effectively reduce memory access latency, improve data access efficiency, and make programming more convenient. However, the Shenwei processor needs to control the LDM (local data memory) through software to achieve efficient computing, which increases the complexity of programming and the requirements for developers.

[0066] The Sunway processor has a fast computing speed, but in actual operation, the memory loading speed is relatively slow, which causes the processor to be idle while waiting for data to be loaded. In addition, slow memory access has become a bottleneck restricting the performance of the Sunway processor, so the use of operator fusion can greatly improve performance. Operator fusion combines multiple operations into one operation, reducing the number of memory accesses and data handling overhead, thereby improving overall computing efficiency.

[0067] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0068] The technical solution of the present invention is described in detail below in conjunction with specific embodiments and the accompanying drawings.

[0069] As attached Figure 4 As shown, the present invention provides an operator fusion method based on the Shenwei processor, comprising the following steps:

[0070] Data acquisition step S1: acquiring operators of the computational graph in the deep learning model, analyzing the data access mode of the operators, determining the operators to be fused that meet the fusion conditions according to the preset fusion judgment mechanism, and acquiring the feature graph of the operators to be fused;

[0071] Data loading step S2: the feature graph is processed using the operator to be fused, wherein the parameters of the operator to be fused are pre-loaded into the local data memory of the Shenwei processor;

[0072] Data operation step S3: Determine the total data flow of the feature map and the capacity of the local data storage device, divide the data blocks of the feature map according to the determination result, use the data blocks to complete the operation of the operator to be fused, and output the data blocks after the operation.

[0073] Based on the above method, by fusing operators that meet the fusion conditions and pre-loading the parameters into the local data storage of the Shenwei processor, the frequent transmission of data between the processor and the external memory is reduced, the computing resources of the Shenwei processor can be fully utilized, the parallel computing efficiency is improved, the memory access mode is optimized, and the bottleneck of memory bandwidth is reduced.

[0074] Combined with Figure 5 As shown, the data operation step S3 is specifically as follows:

[0075] According to the working set size of the operator to be fused, the spatial area of ​​the local data memory is divided into a computing area, a cache area and a communication area. The computing area is used to store the data required for calculation, the cache area is used to store temporary data, and the communication area is used for data exchange buffering;

[0076] If the total data flow of the feature map exceeds the capacity of the local data storage, the optimal block size is dynamically calculated, and the block pipeline method is used to use multiple data blocks to complete the operation of the operator to be fused in sequence according to the calculated optimal block, and the data block after the operation is output;

[0077] Otherwise, the data blocks in the feature map are continuously loaded into the local data memory to complete the calculation of the operator to be fused, and the calculated data blocks are output;

[0078] Among them, the calculation model for dynamically calculating the optimal block size is:

[0079]

[0080] in, is the optimal block size; is the maximum value of the preset block; is the available space of the local data memory; is the number of operators to be fused.

[0081] Optionally, a local data memory with a space capacity of 256KB is dynamically divided into a calculation area, a cache area and a communication area. When the operation of the input feature map is matrix multiplication, the scale of the input matrix cabinet is determined. At this time, if the scale of the input matrix is ​​32×32, that is, the total data flow is less than 80% of the capacity of the local data memory, the data blocks in the feature map are continuously loaded into the local data memory to complete the calculation of the operator to be fused; otherwise, for large-scale matrix operations, that is, the total data flow of the input matrix is ​​greater than 80% of the capacity of the local data memory, the optimal block size is dynamically calculated, and a block pipeline method is adopted according to the calculated optimal block. When a data block is processed in the calculation area, the next data is loaded in the prefetch area through the DMA controller at the same time. The algorithm uses multiple data blocks to complete the operation of the operator to be fused in sequence, realizes the effective overlap of calculation and communication, and dynamically adjusts the prefetch depth according to the operation situation. On the basis of the LRU cache replacement strategy, it integrates data reuse awareness and access mode optimization. That is, the LRU cache replacement strategy is used as the basic strategy to give priority to replacing the least recently used data block. The reuse frequency of data blocks is dynamically analyzed through data reuse awareness, and data blocks with high reuse are retained first. Combined with access mode optimization, the replacement strategy is adjusted according to the data access mode: for continuous access mode, continuous data blocks are retained first; for random access mode, combined with reuse awareness, high-frequency access data blocks are retained first, which further improves the utilization efficiency of local data storage.

[0082] Furthermore, the process of using a block pipeline to sequentially complete the operation of the operator to be fused using multiple data blocks according to the calculated optimal block is as follows:

[0083] Two buffers are divided in the calculation area of ​​the local data memory. A DMA request is sent to the DMA controller according to the interface of the Shenwei processor. After obtaining the data required for calculation in the buffer according to the DMA request, the operation of the operator to be fused is completed. When one buffer is used for calculation, the other buffer pre-fetches the next block of data through the DMA controller. The dual buffer design realizes the pipeline parallelism of calculation and data loading.

[0084] Combined with Figure 6 As shown, the data loading step S2 specifically includes:

[0085] The data access mode is determined according to the architecture mode of the Shenwei processor, and the method of loading the parameters of the operator to be fused into the local data memory is selected according to the data access mode, wherein:

[0086] When the data access mode is continuous access, a DMA request is sent to the DMA controller according to the interface of the Shenwei processor, and the DMA controller continuously loads the data blocks in the feature map into the local data memory according to the DMA request;

[0087] When the data access mode is stride access, the data blocks of the feature map are rearranged in the main memory, and after the discrete data are formed into continuous data blocks, the data blocks are continuously loaded into the local data memory by sending DMA requests to the DMA controller;

[0088] When the data access mode is random access, after analyzing and identifying the data in the data block that has been accessed more than a preset number of times, the data is moved to adjacent positions to form continuous data blocks, and then DMA requests are sent to the DMA controller to continuously load the data blocks into the local data memory; by dynamically selecting the data loading method according to the data access mode, it is ensured that the data is efficiently loaded into the local data memory.

[0089] Based on the hardware characteristics of the Shenwei processor, in terms of point-to-point transmission, by optimizing the data exchange path between adjacent CPEs and taking measures to avoid bandwidth competition, efficient local data transmission is achieved; in terms of broadcast transmission, a hierarchical broadcast mechanism is designed based on a 4×4 grid structure, and bandwidth sharing is optimized through a multi-level tree distribution method; in terms of data pipeline, effective overlap of computing and communication is achieved through task segmentation and scheduling, and bandwidth utilization is optimized through measures such as data alignment and merging, and memory access conflict avoidance.

[0090] Furthermore, when the input feature graph of the operator fusion is a matrix multiplication and addition operation (d=a×b+c), the data access pattern is first analyzed. According to the data access pattern analysis, four data loads (a, b, tmp, c) and two data stores (tmp, d) are required. After the optimization of the operator fusion, (a, b, c) are loaded into the local data memory at the same time and the final result is directly calculated in the local data memory, so as to optimize the data access to three loads (a, b, c) and one storage (d). For large-scale calculations that exceed the capacity of the local data memory, the formula for dynamically calculating the optimal block size is used to realize the direct accumulation of block calculation results in the local data memory, which significantly improves the computing efficiency.

[0091] In some embodiments of the present invention, the process of determining the operators to be fused that meet the fusion conditions according to the preset fusion determination mechanism is specifically as follows:

[0092] Analyze the data dependency between operators to confirm whether there are operator pairs with direct data dependency and no circular dependency, and whether the output of an operator is only used by one subsequent operator;

[0093] Evaluate the computing characteristics of the operators, confirm whether the computing density between operators is similar, determine whether the access to the operators meets the preset alignment requirements of the Shenwei processor, and whether the total resource demand after the operator fusion exceeds the capacity limit of a single CPE; optionally, the preset alignment requirement is that the Shenwei processor requires 4B alignment for access data before sending a request to the DMA controller through the interface.

[0094] Based on a preset benefit indicator, a quantitative analysis is performed to evaluate whether the benefit indicator after operator fusion meets the preset benefit indicator. The benefit indicator includes the number of data transfers, data reuse rate and computing density.

[0095] The computational density is the amount of computation completed by the operator per unit time, reflecting the computational intensity of the operator. It is calculated by the ratio of the number of all floating-point operations completed by the operator to the time required for the operator to complete the computational task.

[0096] Optionally, the operators are analyzed according to the computational graph of the deep learning model. When continuous Element-wise operation operators encounter an operator chain of Add+ReLU+BatchNorm, they are preferentially marked as potential fusion objects because they have similar data access patterns and lower computational density. The data dependencies between potential fusion pairs are verified to ensure that the fusion will not cause computational errors or deadlocks. The computational complexity, memory usage, and data access characteristics of the operators are analyzed to evaluate whether the resource requirements after fusion meet the hardware constraints.

[0097] Furthermore, when the operators to be fused are convolutional layers, batch normalization layers, and ReLU activation layers, the convolution kernel parameters and batch normalization parameters are pre-loaded into the local data storage device. When processing the input feature map, a block pipeline method is adopted. Each data block completes the convolution calculation, batch normalization, and ReLU activation in turn to directly obtain the final output, avoiding repeated reading and writing of intermediate results.

[0098] Furthermore, when the input feature map of the operator fusion is an addition operation, the traditional steps without fusion of the operator are to read data A and data B, write data D, then read data D and data C, and finally write data E, that is, the number of data transfers is 6 times, the number of data reads is 4 times, and the number of data writes is 2 times;

[0099] However, through the operator fusion method of the present invention, it is only necessary to read data A, data B and data C, and then write data E, that is, the number of data transfers is 4 times, the number of data reads is 3 times, and the number of data writes is 1 time. Through the traditional method without operator fusion, the total calculation time is 147.635ms, and after the operator is fused based on the Shenwei processor, the total calculation time is 39.496ms, and the performance is improved by 3.7 times. Through operator fusion, the number of data transfers is reduced, and the impact of slow memory access speed on performance is alleviated, thereby significantly improving the overall performance.

[0100] In some embodiments of the present invention, the data acquisition step S1 further includes:

[0101] Classify and analyze operators according to their computational characteristics, and divide operators in deep learning models into computationally intensive operators, communication-intensive operators, and hybrid operators.

[0102] Among them, computation-intensive operators complete calculations within a single CPE, communication-intensive operators complete calculations collaboratively between multiple CPEs by optimizing communication paths, and hybrid operators complete calculations collaboratively between multiple CPEs.

[0103] Furthermore, compute-intensive operators refer to operators whose amount of calculation is much greater than the amount of data transmitted, such as convolution and matrix multiplication. Compute-intensive operators are characterized by high computational density and large demand for computing resources. By loading the data required for calculation into the local data storage of the CPE, access to the main memory is reduced, and through block computing and data pipelining technology, the capacity and bandwidth of the local data storage are fully utilized.

[0104] Communication-intensive operators refer to operators whose data transmission volume is much larger than the computation volume, such as reduction and broadcast. Communication-intensive operators are characterized by large demands for communication bandwidth and low computation density. By optimizing the data flow path at the 4×4 grid level, reducing the communication overhead across core groups, and using a multi-level tree distribution mechanism, the bandwidth utilization of broadcast and reduction operations is optimized.

[0105] Hybrid operators refer to operators that require both a large amount of computation and frequent data exchange, such as batch normalization and attention mechanism. Hybrid operators are characterized by high computation and communication requirements, and need to balance computation and communication overhead. They reduce memory access conflicts by adopting data reorganization and block processing strategies, and improve data transmission efficiency through DMA batch transfer mechanism.

[0106] It should be noted that the above is a reference method for the operator fusion method based on the Shenwei processor, and the present invention is not limited to this.

[0107] Embodiments of the present invention fuse operators that meet the fusion conditions and pre-load parameters into the local data memory of the Sunway processor, reducing the frequent transmission of data between the processor and external memory, making full use of the computing resources of the Sunway processor, improving the parallel computing efficiency, optimizing the memory access pattern, reducing the bottleneck of memory bandwidth, and solving the technical problems in the prior art that the characteristics of the Sunway processor cannot be fully utilized, resulting in poor operator fusion effect when executing deep learning algorithms on the Sunway processor, and the existing operator fusion technology does not fully consider the many-core architecture of the Sunway processor, resulting in the failure to fully utilize the parallel processing ability and affecting the execution speed of the algorithm. In the existing operator fusion method, the memory access pattern is not optimized enough during the fusion process, resulting in low memory bandwidth utilization, further restricting the improvement of algorithm performance.

[0108] Finally, it should be noted that the embodiments in this specification are described in a progressive manner, and the key points of each embodiment are the differences from other embodiments. The same and similar parts among the embodiments can be referred to each other.

[0109] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that it is still possible to modify the specific implementation manners of the present invention or perform equivalent replacements for some technical features without departing from the spirit of the technical solutions of the present invention, and they should all be covered within the scope of the technical solutions claimed by the present invention.

Claims

1. An operator fusion method based on the Shenwei processor, characterized in that: The steps include: Data acquisition step: acquiring operators of the computational graph in the deep learning model, analyzing the data access mode of the operators, determining operators to be fused that meet the fusion conditions according to a preset fusion determination mechanism, and acquiring feature graphs of the operators to be fused; Data loading step: the feature graph is processed using the operator to be fused, wherein the parameters of the operator to be fused are pre-loaded into the local data memory of the Shenwei processor; Data operation step: judging the total data flow of the feature graph and the capacity of the local data storage, dividing the data blocks of the feature graph according to the judgment result, using the data blocks to complete the operation of the operator to be fused, and outputting the data blocks after the operation; The process of determining the total data flow of the feature graph and the capacity of the local data storage is specifically as follows: If the total data flow of the feature map exceeds the capacity of the local data storage, the optimal block size is dynamically calculated, and the operation of the operator to be fused is completed in sequence using multiple data blocks in a block pipeline manner according to the calculated optimal block, and the data blocks after the operation are output; Otherwise, the data blocks in the feature map are continuously loaded into the local data memory to complete the calculation of the operator to be fused, and the calculated data blocks are output; The data loading step specifically includes: The data access mode is determined according to the architecture mode of the Shenwei processor, and the method of loading the parameters of the operator to be fused into the local data memory is selected according to the data access mode, wherein: When the data access mode is continuous access, a DMA request is sent to a DMA controller according to an interface of the Shenwei processor, and the DMA controller continuously loads the data blocks in the feature map into a local data memory according to the DMA request; When the data access mode is stride access, the data blocks of the feature map are rearranged in the main memory, the discrete data are combined into continuous data blocks, and the data blocks are continuously loaded into the local data memory by sending a DMA request to the DMA controller; When the data access mode is random access, after analyzing and identifying data in the data block whose access times exceed a preset number, the data is moved to adjacent positions to form continuous data blocks, and then a DMA request is sent to the DMA controller to continuously load the data blocks into the local data memory.

2. The operator fusion method based on the Shenwei processor according to claim 1 is characterized in that: The data calculation steps are specifically as follows: The spatial area of ​​the local data memory is divided into a calculation area, a cache area and a communication area according to the working set size of the operator to be fused. The calculation area is used to store data required for calculation, the cache area is used to store temporary data, and the communication area is used for data exchange buffer.

3. The operator fusion method based on the Shenwei processor according to claim 1 is characterized in that: The specific process of using a block pipeline to sequentially complete the operation of the operator to be fused using multiple data blocks according to the optimal block obtained by calculation is as follows: Two buffers are divided in the calculation area of ​​the local data memory, and a DMA request is sent to the DMA controller according to the interface of the Shenwei processor. After the data required for calculation in the buffer is obtained according to the DMA request, the operation of the operator to be fused is completed.

4. The operator fusion method based on the Shenwei processor according to claim 1, characterized in that: The calculation model for dynamically calculating the optimal block size is: in, is the optimal block size; is the maximum value of the preset block; is the available space of the local data memory; is the number of operators to be fused.

5. The operator fusion method based on the Shenwei processor according to any one of claims 1 to 4, characterized in that: The specific process of determining the operators to be fused that meet the fusion conditions according to the preset fusion judgment mechanism is as follows: Analyze the data dependency between operators, confirm whether there are operator pairs with direct data dependency and no circular dependency, and confirm whether the output of an operator is only used by one subsequent operator.

6. The operator fusion method based on the Shenwei processor according to claim 5 is characterized in that: The process of determining the operators to be fused that meet the fusion conditions according to the preset fusion judgment mechanism also includes: Evaluate the computing characteristics of the operators, confirm whether the computing density between operators is similar, determine whether the access to the operators meets the preset alignment requirements of the Shenwei processor, and whether the total resource demand after the operator fusion exceeds the capacity limit of a single CPE.

7. The operator fusion method based on the Shenwei processor according to claim 6 is characterized in that: The operator fusion determination step also includes: Based on a preset benefit indicator, a quantitative analysis is performed to evaluate whether the benefit indicator after the operator fusion meets the preset benefit indicator, and the benefit indicator includes the number of data transfer times, data reuse rate and calculation density.

8. The operator fusion method based on the Shenwei processor according to claim 6, characterized in that: The data acquisition step comprises: Classifying and analyzing the operators according to the computational characteristics of the operators, and classifying the operators in the deep learning model into computationally intensive operators, communication-intensive operators, and hybrid operators; The computation-intensive operator completes the computation within a single CPE, the communication-intensive operator completes the computation collaboratively between multiple CPEs by optimizing the communication path, and the hybrid operator completes the computation collaboratively between multiple CPEs.

Citation Information

Patent Citations

  • Circular buffer for input and output of tensor computations

    US20240281393A1

  • Deep learning memory optimization method for microcontroller

    WO2025000821A1