A Template Calculation Optimization Method, Device and Equipment on a Multi-Core DSP
By dividing the data grid into data blocks on multi-core DSP and using vector computing units for template calculations, combined with data transmission of the three-buffer mechanism, the problem of multi-core DSP in template calculation optimization is solved, and efficient template computing performance is achieved.
Patent Information
- Application Number
- CN202510257886.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-03-05
AI Technical Summary
Due to its unique hardware architecture, multi-core DSP is difficult to achieve optimal performance optimization for template computing applications, and lacks specially designed micro-cores to make full use of hardware features.
By dividing the input data grid into multiple data blocks and assigning a DSP core to each data block, the DSP vector calculation unit is used for secondary calculations, and data transmission is carried out using a three-buffer mechanism to achieve efficient template calculations.
High-performance template computing on the multi-core DSP platform is realized, making full use of DSP's vector computing unit and data mobile unit, improving processing efficiency and resource utilization, and overcoming the performance optimization obstacles caused by hardware architecture differences.
Smart Images

Figure CN119739673B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of computing optimization for multi-core DSP (Digital Signal Processor), and particularly to a method, device, and equipment for optimizing stencil calculations on a multi-core DSP. Background Art
[0002] Stencil calculation is a common calculation mode in high-performance computing (HPC) applications and is used in fields such as image processing, convolutional neural networks, and solving partial differential equations. These calculations are usually the performance bottleneck in many scientific applications because they are memory-bound and require efficient data reuse and access patterns.
[0003] When an application needs to perform a stencil calculation during operation, it can send a stencil calculation execution request to the processor. The processor that receives the stencil calculation execution request can traverse each or specified data points of the data to be executed for the stencil calculation, and perform calculation operations on each data point and the neighbor data points corresponding to each data point. During the traversal process, the same calculation operation is performed on the data point and its neighbor data points each time. Among them, the calculation domain can be a data matrix, such as a two-dimensional grid, a three-dimensional grid, etc.
[0004] Traditional optimization methods mainly focus on CPU and GPU architectures. These methods enhance performance by improving data locality and parallelism, reducing communication overhead, etc. However, there is no consensus on the best optimization method for stencil calculations on emerging HPC processors, especially multi-core digital signal processors (DSPs). The architecture of DSPs is unique, with very long instruction words (VLIWs) and vector units, which provides opportunities to utilize instruction-level and data-level parallelism, but also brings challenges to the performance optimization of stencil calculation-based application programs due to the differences in their hardware architectures.
[0005] In summary, multi-core DSPs are an ideal choice for executing stencil calculation-based tasks due to their VLIWs and vector units. Designing an efficient stencil calculation program for multi-core DSPs includes two parts: microkernel design and off-chip and on-chip data transfer. Among them, efficient microkernel design assumes how to make full use of the VLIWs and vector units of DSPs when the data is on the chip. However, there is currently a lack of microkernels specifically designed for stencil calculation-based application programs to make full use of these hardware characteristics, resulting in the performance of this program not reaching the best. Summary of the Invention
[0006] Based on this, in view of the above technical problems, it is necessary to provide a method, device, and equipment for optimizing stencil calculations on a multi-core DSP that can improve the development energy efficiency of multi-core DSP processors and stencil calculation-based application programs.
[0007] An optimization method for template calculation on a multi-core DSP, the method comprising:
[0008] Dividing the input data grid into a plurality of data blocks, and allocating a DSP core to each data block. Each data block contains a plurality of grid points.
[0009] At the current time step, all DSP cores have completed the first-level calculation on the allocated data blocks, and further divide the data blocks into a plurality of sub-blocks. Each sub-block contains a plurality of adjacent grid points.
[0010] In the next time step, load the sub-blocks into the vector on-chip memory of the DSP, and use the vector calculation unit of the DSP to perform a second-level calculation on the sub-blocks to obtain new values for each grid point.
[0011] Pack the calculated new values into new sub-blocks and store them back in the vector on-chip memory. The new sub-blocks are transferred from the vector on-chip memory to the off-chip memory for loading and storage through the data movement unit of the DSP using a triple-buffer mechanism.
[0012] An optimization device for template calculation on a multi-core DSP, the device comprising:
[0013] A data block allocation module, configured to divide the input data grid into a plurality of data blocks, and allocate a DSP core to each data block. Each data block contains a plurality of grid points.
[0014] A data block splitting module, configured to, at the current time step, when all DSP cores have completed the first-level calculation on the allocated data blocks, further divide the data blocks into a plurality of sub-blocks. Each sub-block contains a plurality of adjacent grid points.
[0015] A template calculation module, configured to, in the next time step, load the sub-blocks into the vector on-chip memory of the DSP, and use the vector calculation unit of the DSP to perform a second-level calculation on the sub-blocks to obtain new values for each grid point.
[0016] A data transfer module, configured to pack the calculated new values into new sub-blocks and store them back in the vector on-chip memory. The new sub-blocks are transferred from the vector on-chip memory to the off-chip memory for loading and storage through the data movement unit of the DSP using a triple-buffer mechanism.
[0017] A computer device, comprising a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:
[0018] Dividing the input data grid into a plurality of data blocks, and allocating a DSP core to each data block. Each data block contains a plurality of grid points.
[0019] At the current time step, all DSP cores have completed the first-level calculation on the allocated data blocks, and further divide the data blocks into multiple sub-blocks, where each sub-block contains multiple adjacent grid points.
[0020] At the next time step, load the sub-blocks into the vector on-chip memory of the DSP, and use the vector calculation unit of the DSP to perform the second-level calculation on the sub-blocks to obtain the new values of each grid point.
[0021] Pack the calculated new values into new sub-blocks and store them back in the vector on-chip memory. The new sub-blocks are transferred from the vector on-chip memory to the off-chip memory for loading and storage through the data movement unit of the DSP using a triple-buffer mechanism.
[0022] The above template calculation optimization method, device, and equipment on a multi-core DSP. First, the input data grid is divided into multiple data blocks, and a DSP core is allocated to each data block, making full use of multi-core parallelism and initiating the first step of parallel computing. This enables the first-level calculation to advance synchronously in multiple threads, laying the foundation for overall high efficiency. At the current time step, all DSP cores are synchronously activated to complete the first-level calculation of the allocated data blocks. The subdivision of the data blocks in the first-level calculation is closely connected to subsequent steps. The fine sub-block division provides an appropriate data granularity for vector on-chip memory loading and second-level calculation, allowing the vector calculation unit of the DSP to be fully utilized, accurately exploiting data-level parallelism, making the calculation granularity finer, and laying the foundation for subsequent in-depth processing. Without reasonable data division in the early stage, the vector calculation unit cannot operate efficiently, resulting in idle or underutilized hardware resources. Subsequently, at the next time step, the sub-blocks obtained by further dividing the data blocks are loaded into the vector on-chip memory of the DSP, and the second-level calculation is carried out with the help of the unique vector calculation unit of the DSP, fully exploiting data-level parallelism and accurately calculating the new values of each grid point. The subsequent steps are closely connected. The fine sub-block division provides an appropriate data granularity for vector on-chip memory loading and second-level calculation, allowing the vector calculation unit of the DSP to be fully utilized, accurately exploiting data-level parallelism. Moreover, after the calculated new values are packed into new sub-blocks and stored back in the vector on-chip memory, a triple-buffer mechanism is adopted to efficiently transfer data between the vector on-chip memory and off-chip memory using the data movement unit of the DSP. The packing after calculating the new values and the triple-buffer mechanism are also interlocked. The packing regularizes the data for using the triple-buffer mechanism, enabling smooth transmission between storage levels through the data movement unit. The triple-buffer mechanism cleverly balances the data loading and storage requirements, avoiding high latency caused by frequent reading and writing of off-chip memory and ensuring that data can be updated and stored back in a timely manner, reducing communication overhead. This comprehensive design from data division, calculation division to storage transmission adapts to the unique hardware architecture of the DSP's very long instruction word (VLIW) and vector unit, overcomes the performance optimization obstacles brought to the template calculation application program due to architecture differences, and finally achieves the goal of high-performance template calculation on the multi-core DSP platform, opening up a new path for template calculation optimization on emerging HPC processors. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic flowchart of a template calculation optimization method on a multi-core DSP in an embodiment;
[0024] Figure 2 It is a schematic flowchart of a microkernel calculation in an embodiment;
[0025] Figure 3 It is a parallel calculation strategy diagram of a microkernel on a multi-core DSP processor in an embodiment;
[0026] Figure 4 Schematic diagram of the microkernel triple-buffer mechanism in one embodiment;
[0027] Figure 5 Block diagram of the microkernel structure in one embodiment, where Figure 5 (a) is the shape information of the 2D5P template, Figure 5 (b) is the schematic diagram of local calculation, Figure 5 (c) is the schematic diagram of data reuse, Figure 5 (d) is the register allocation;
[0028] Figure 6 Schematic diagram of the microkernel calculation execution process in one embodiment;
[0029] Figure 7 Block diagram of the structure of the template calculation optimization device on a multi-core DSP in one embodiment;
[0030] Figure 8 Internal structure diagram of a computer device in one embodiment. Detailed implementation manners
[0031] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0032] In one embodiment, as Figure 1 shown, a method for optimizing template calculation on a multi-core DSP is provided, including the following steps:
[0033] Step 102: Divide the input data grid into several data blocks, and allocate a DSP core to each data block.
[0034] A data block contains several grid points.
[0035] Step 104: At the current time step, all DSP cores complete the first-level calculation on the allocated data blocks, and further divide the data blocks into multiple sub-blocks, where each sub-block contains multiple adjacent grid points.
[0036] Step 106: At the next time step, load the sub-blocks into the vector on-chip memory of the DSP, and use the vector calculation unit of the DSP to perform the second-level calculation on the sub-blocks to obtain new values for each grid point.
[0037] Step 108: Pack the calculated new values into new sub-blocks and store them back in the vector on-chip memory. The new sub-blocks are transferred from the vector on-chip memory to the off-chip memory for loading and storage through the data movement unit of the DSP using the triple-buffer mechanism.
[0038] In the above template calculation optimization method on a multi-core DSP, first, the input data grid is divided into multiple data blocks, and a DSP core is allocated to each data block, making full use of multi-core parallelism and starting the first step of parallel computing. This enables the first-level calculation to advance synchronously in multiple threads, laying the foundation for overall high efficiency. At the current time step, all DSP cores are synchronously activated to complete the first-level calculation of the allocated data blocks. The subdivision of the data blocks in the first-level calculation is closely related to the subsequent steps. The fine sub-block division provides an appropriate data granularity for the vector on-chip memory loading and the second-level calculation, allowing the vector calculation unit of the DSP to be fully utilized, accurately mining data-level parallelism, making the calculation granularity finer, and laying the foundation for subsequent in-depth processing. Without reasonable data division in the early stage, the vector calculation unit cannot operate efficiently, resulting in idle or underutilized hardware resources. Subsequently, at the next time step, the sub-blocks obtained by further dividing the data blocks are loaded into the vector on-chip memory of the DSP, and the second-level calculation is carried out by means of the unique vector calculation unit of the DSP, fully mining data-level parallelism and accurately calculating the new values of each grid point. The subsequent steps are closely related. The fine sub-block division provides an appropriate data granularity for the vector on-chip memory loading and the second-level calculation, allowing the vector calculation unit of the DSP to be fully utilized and accurately mining data-level parallelism. Moreover, after the calculated new values are packed into new sub-blocks and stored back in the vector on-chip memory, a triple-buffer mechanism is adopted, and the data movement unit of the DSP is used to efficiently transfer data between the vector on-chip memory and the off-chip memory. The packing and the triple-buffer mechanism after calculating the new values are also closely linked. The packing regularizes the data for using the triple-buffer mechanism, and the data is smoothly transferred between the storage levels through the data movement unit. The triple-buffer mechanism cleverly balances the data loading and storage requirements, avoiding the high latency caused by frequent reading and writing of the off-chip memory and ensuring that the data can be updated and stored back in a timely manner, reducing the communication overhead. This comprehensive design from data division, calculation division to storage transmission adapts to the unique hardware architecture of the DSP's very long instruction word (VLIW) and vector unit, overcomes the performance optimization obstacles brought by the architecture differences to the template calculation application program, and finally achieves the goal of realizing high-performance template calculation on the multi-core DSP platform, opening up a new path for the template calculation optimization on emerging HPC processors.
[0039] In one embodiment, the parameter information of the multi-core DSP is obtained, where the parameter information includes: the information of the data grid, the template mode information, and the DSP parameter information. The information of the data grid includes: the grid size and the grid dimension. The template mode information includes: the data dependence range, the number of data dependence points, and the data dependence shape. The DSP parameter information: the size of the global shared memory, the number of floating-point multiply-accumulate units, the number of vector registers, the memory bandwidth, and the number of DSP cores.
[0040] In one embodiment, the shape of the microkernel is determined according to the parameter information, and the input data grid is divided into several data blocks composed of the same number of data columns and data rows according to the shape of the microkernel. A DSP core is allocated to each data block.
[0041] In one embodiment, at the current time step, the DSP cores of the allocated data blocks select the instructions corresponding to the first-level calculation mode according to the DSP parameter information, the template mode information, and the microkernel shape, and pack the instructions corresponding to the first-level calculation mode into a first-level instruction packet to ensure that multiple multiply-accumulate instructions can be executed within each clock cycle. The allocated data blocks are subjected to first-level calculation in a vectorized manner according to the multiply-accumulate instructions, and the data blocks are further divided into multiple sub-blocks. In the sub-blocks, multiple adjacent grid points are combined into vectors and processed in parallel by the vector calculation unit.
[0042] In one embodiment, at the next time step, the sub-blocks are loaded into the vector on-chip memory of the DSP by using a triple-buffer mechanism, and the memory access path information is determined based on the grid parameter information, the template mode information, and the DSP parameter information corresponding to the multi-core DSP chip. According to the path information, each sub-block undergoes three stages of loading, calculation, and storage through the first-level triple-buffer, so that the sub-blocks are completely processed at the first level by the vector processing unit. The second-level triple-buffer is used to perform three stages of second-level calculation of loading, calculation, and storage step by step between the main memory and the vector on-chip memory for the sub-blocks that have been processed at the first level through the vector calculation unit of the DSP, and new values of each grid point are obtained. The triple-buffer mechanism includes: a first-level triple-buffer and a second-level triple-buffer. The first-level triple-buffer occurs between the vector on-chip memory and the vector register of the DSP and is used to realize the parallel execution of data loading and calculation inside the sub-blocks. The second-level triple-buffer occurs between the vector on-chip memory and the off-chip memory of the DSP and is used to realize the parallel execution of data loading, calculation, and storage between the sub-blocks.
[0043] In one embodiment, according to the path information, each sub-block undergoes three stages of loading, calculation, and storage through the first-level triple-buffer. In the loading stage, the sub-block is loaded from the vector on-chip memory of the DSP into the vector register of the vector processing unit. In the calculation stage, the floating-point multiply-add operation is performed using the multiply-add unit of the vector register of the vector processing unit. In the storage stage, the sub-block is stored back from the vector register to the vector on-chip memory. Through the instruction set pipeline, the operations of calculating the current sub-block, loading the next sub-block, and storing the previous sub-block are performed within the same time segment, so that all sub-blocks are completely processed at the first level by the vector processing unit.
[0044] In one embodiment, a secondary three-buffer is utilized to perform a three-stage secondary calculation of loading, calculating, and storing on a sub-block that has been processed at the first level between the main memory and the vector on-chip memory through the vector calculation unit of the DSP. In the loading stage, the sub-block is transferred from the main memory to the vector on-chip memory of the DSP. In the calculating stage, all operations of the first-level three-buffer are executed. In the storing stage, the sub-block is stored back from the vector on-chip memory to the main memory. Through cyclic operations, when the sub-block in the first buffer is used for calculation, the next sub-block is loaded into the second buffer of the vector on-chip memory until, after the calculation is completed, new values for each grid point are obtained.
[0045] In one embodiment, the data transfer operations of the sub-block to or from the on-chip are optimized using interleaved single-word and double-word instructions. The size of the sub-block is selected according to the on-chip storage capacity of the DSP, the characteristics of the vector calculation unit, and the data reuse strategy. The loading and storing operations of the sub-block use the DMA engine for asynchronous data transfer. The calculation of the sub-block is performed using a microkernel to utilize the vector calculation unit of the DSP and the optimized instruction pipeline technology. The shape of the microkernel is designed according to the number of vector registers of the DSP, the instruction packet filling constraint, and the instruction latency constraint of the vector calculation unit. The newly calculated values are packed into new sub-blocks and stored back into the vector on-chip memory. All the new sub-blocks are stored in the third buffer using the three-buffer mechanism and are written back from the vector on-chip memory to the main memory through the data movement unit of the DSP, completing the loading and storing of the input data.
[0046] In one embodiment, as Figure 2 shown, a microkernel calculation process is provided, which specifically includes the following:
[0047] S1. Parameter information acquisition: First, necessary multi-core DSP chip parameters are collected, including the size and dimension of the input grid, the data dependence range of the template pattern, the number of points and shape, as well as the global shared memory size of the DSP, the vector on-chip cache size, the number of floating-point multiply-accumulate units, the number of vector registers, the main memory bandwidth, the global shared memory bandwidth, and the number of multi-core DSP cores. These parameters are the basis for the design and optimization of the microkernel.
[0048] S2. Microkernel shape determination: According to the collected parameter information, the shape of the microkernel is determined, including the configuration of data columns and rows. The shape of the microkernel directly affects its execution efficiency on the DSP, so it needs to be optimized based on the hardware characteristics of the DSP and the template calculation requirements.
[0049] S3. Micro-kernel computing strategy formulation: Formulate the computing strategy of the micro-kernel, including the optimization of the instruction pipeline and the configuration of the data reuse strategy. Design the instruction pipeline of the micro-kernel to ensure that multiple FMAC instructions can be executed within each cycle, and configure the data reuse strategy to reduce the number of accesses to external memory and improve the computing efficiency.
[0050] S4. Micro-kernel execution and result output: Execute the micro-kernel computing, use the vector processing unit of the DSP to process data in parallel, and store the result back to the on-chip memory after the computing is completed, and then output the computing result. This step includes data loading, computing execution, result storage and exception handling to ensure the correctness and synchronization of the computing.
[0051] In an alternative embodiment, the DSP core includes an instruction scheduling unit, a scalar processing unit, a vector processing unit and a DMA engine, wherein the vector processing unit includes a vector on-chip cache and a number of vector processing elements, and each vector processing element includes a number of 64-bit vector registers and a number of floating-point multiply-accumulate units. This hardware configuration provides a physical basis for the efficient execution of the micro-kernel.
[0052] It should be noted that by implementing the template computing optimization method on the multi-core DSP described in the present invention, the processing efficiency and resource utilization rate of the multi-core DSP chip can be significantly improved, especially when executing computationally intensive template computing tasks.
[0053] In one of the embodiments, as Figure 3 shown, a parallel computing strategy of the micro-kernel on the multi-core DSP processor is provided, which specifically includes the following contents:
[0054] Micro-kernel initialization: The process starts with the initialization of the micro-kernel, including configuring the DSP parameters required for the micro-kernel, such as the size of the global shared memory, the size of the vector on-chip cache, the number of floating-point multiply-accumulate units, etc.
[0055] Data block allocation and micro-kernel execution: Divide the input grid into multiple data blocks, and each data block is assigned to a DSP core. Each DSP core loads its corresponding data block into the vector register and executes the micro-kernel computing. This step reflects thread-level parallelization, and each DSP core acts as an execution thread to process the data block assigned to it in parallel.
[0056] Vector-level parallelization: In data-level parallelization, use the SIMD instructions of the DSP core to form vectors from consecutive grid points and perform parallel computing in the micro-kernel. This step optimizes the vector processing of data and improves the computing efficiency.
[0057] Result Summarization and Output: After the micro-kernel calculation is completed, the results of each DSP core need to be summarized and output. This step includes storing the calculation results from the vector registers back to the on-chip memory and integrating the results of each core to form the final output grid.
[0058] In one embodiment, as Figure 4 shown, a micro-kernel triple buffering mechanism is provided. Based on the multi-core DSP chip grid parameter information, the template pattern information, and the DSP parameter information, the memory access path information is determined, and the two-level triple buffering mechanism overlaps the calculation and communication tasks.
[0059] In an alternative embodiment, the first-level triple buffer (LV1) decomposes the tile data into smaller sub-tiles, and each sub-tile can be fully processed by the VPU. The sub-tiles are divided into three stages: loading, computing, and storing to achieve data overlap. In the loading stage, a sub-tile of size is loaded from the AM into the vector register of the VPU (vector processing unit). In the computing stage, the FMAC unit of the VPU is used to perform floating-point multiply-add operations. In the storing stage, a sub-tile of size is stored back from the vector register to the tile in the AM. Through the instruction set pipeline, the operations of computing the current sub-tile, loading the next sub-tile, and storing the previous sub-tile are performed within the same time segment. There are delays in the prefetch and the last storing operations, but the overall design enhances parallelism and overlap by minimizing the overhead of loading and storing, thereby improving the overall performance.
[0060] In an alternative embodiment, the second-level triple buffer (LV2) divides the tile into three stages: loading, computing, and storing to achieve data overlap. In the loading stage, a tile of size is transferred from the DDR to the AM. In the computing stage, all operations of LV1 are performed. In the storing stage, a tile of size is stored back from the AM (vector on-chip cache) to the DDR (main memory). Through cyclic operations, when the tile in the first buffer is used for computing, the next tile is loaded into the second buffer of the AM; after the computing is completed, the data is stored into the third buffer and written back from the AM to the DDR. This level also achieves effective overlap of loading, computing, and storing, thereby improving the efficiency of data processing.
[0061] It is worth noting that in this step, through the two-level triple buffering mechanism, the memory bandwidth of the FT-M7032 platform is effectively utilized, the data loading, computing, and storing operations are decoupled, and overlap between them is achieved, thereby improving the efficiency of template calculation.
[0062] In one embodiment, as Figure 5 shown, a microkernel structure block diagram is provided, including a vector register, a VFMAC unit, and an instruction pipeline. Based on the multi-core DSP chip grid parameter information, the template pattern information, and the DSP parameter information, the microkernel shape information is determined.
[0063] It should be noted that the shape of the microkernel refers to the data scale processed each time the microkernel is called. The microkernel shape information includes data column information and data row information, which are constructed by the following inequality model:
[0064] (1) Vector register constraint: Ensure that the common points in the same sub-tile only need to be loaded into the vector register once. For a two-dimensional k-point template, the new value of each point depends on k points, and there are 2d common points among the k points. Therefore, in each sub-tile, there are a total of points, where M REG -1 adjacent points share common points. To ensure that these common points only need to be loaded once, the values of M REG and N REG need to satisfy the following inequality:
[0065] ;
[0066] (2) Instruction packet filling constraint: Ensure that 3 FMAC instructions can be issued each cycle. For a two-dimensional k-point template, each point requires k floating-point multiply-add operations. Therefore, each sub-tile needs to perform floating-point multiply-add operations. To ensure that 3 FMAC instructions can be issued each cycle, needs to be a multiple of 3 and needs to satisfy the following formula:
[0067] ;
[0068] (3) Instruction delay constraint - 1: Avoid pipeline stalls caused by data dependencies and ensure that the delay between instructions is minimized. Due to the data dependency relationship of FMAC instructions, it is necessary to ensure that the delay between instructions is minimized. For example, after the first FMAC instruction completes the write operation, the second FMAC instruction can start the read operation. Therefore, the values of M REG and N REG need to satisfy the following inequality:
[0069] ;
[0070] (4)Instruction Latency Constraint - 2: Optimize data transfer instructions and ensure that prefetch and store operations are completed within two iterations. The VLDW / VLDDW instructions load vectors into the AM, each instruction having a latency of 9 cycles; the VFMULAD instruction performs vector multiplication, each instruction having a latency of 6 cycles; the VSTDW instruction stores the computed data back to the AM, this instruction having a latency of 4 cycles. Therefore, it is necessary to ensure that prefetch and store operations are completed within two iterations, satisfying the following inequality:
[0071] ;
[0072] In an alternative embodiment, according to the above constraints, the shape of the microkernel is determined. For example, for a 2D5P template, the shape of the microkernel can be determined to be 4x4, as shown respectively in Figure 5 (a) Two - dimensional five - point template shape information (2D5P stencil), Figure 5 (b) Local computation, Figure 5 (c) Data reuse, Figure 5 (d) Register allocation.
[0073] It should be noted that in this step, by carefully designing and implementing the microkernel, the vector computing unit of the DSP is effectively utilized, improving the efficiency of template calculation. The design of the microkernel takes into account vector register constraints, instruction packet filling constraints, and instruction latency constraints, and minimizes latency through instruction pipeline design, thereby achieving high - performance template calculation.
[0074] In one of the embodiments, as Figure 6 shown, a microkernel computing execution process is provided, specifically including microkernel initialization, data loading, computing execution, result storage, and exception handling:
[0075] 1. Obtain parameter information
[0076] The present invention first needs to obtain the following parameter information for subsequent optimization design and calculation:
[0077] Grid parameter information: Obtain grid parameter information from the input grid, including the size (MxN) and dimension (2D or 3D) of the grid. For example, for a two - dimensional grid, the number of rows M and the number of columns N of the grid need to be obtained; for a three - dimensional grid, the three - dimensional size MxNxL of the grid needs to be obtained.
[0078] Template pattern information: Obtain template parameter information from the template pattern, including the range of data dependencies, the number of data - dependent points (k), and the shape of data dependencies (box - shaped or star - shaped). For example, for a 2D5P template, the range of data dependencies is a 3x3 window, the number of data - dependent points is 5 points, and the shape of data dependencies is box - shaped.
[0079] DSP parameter information: Obtain DSP parameter information from the DSP chip, including the size of the global shared memory (GSM), the size of the vector on-chip cache (AM), the number of floating-point multiply-accumulate units (VFMAC), the number of vector registers, the bandwidth of the main memory (DDR), the bandwidth of the GSM, and the number of multi-core DSP cores. For example, the parameter information of the FT-M7032 DSP chip includes: GSM = 8MB, AM = 768KB, VFMAC = 3, vector registers = 64, DDR bandwidth = 340.8 GBits / s, GSM bandwidth = 85.2 GBits / s, core number = 8.
[0080] 2. Determine the shape of the microkernel
[0081] The microkernel is the core module of template computing, and determining its shape is crucial for computing performance. The present invention uses the following steps to determine the shape of the microkernel:
[0082] Microkernel structure: The microkernel is tailored for template computing and can process data blocks of a fixed size (e.g., 16x16). The design of the microkernel takes into account data dependencies and computing patterns to ensure that each computing step can fully utilize the vector processing unit of the DSP.
[0083] Template parameter analysis: According to the template parameter information, determine the range and number of points of data dependencies. The design parameters of the microkernel include the number of data points, the number of vectors, and the data dependency distance. These parameters can be adjusted according to different template computing requirements to achieve optimal performance. This information determines the amount of data that the microkernel needs to load. For example, for a 2D5P template, the microkernel needs to load data for 5 points (i.e., a 3x3 window).
[0084] Analysis of the number of vector registers and VFMAC: When designing the microkernel, the use of vector registers is fully considered to achieve data reuse in the registers and reduce the number of memory accesses. First, according to the DSP parameter information, determine the number of vector registers and the number of VFMAC. This information determines the number of data columns and data rows that the microkernel can process. For example, the FT-M7032 DSP chip has 64 vector registers and 3 VFMAC, so the number of data columns and data rows of the microkernel should be 64 and 16 respectively.
[0085] Data reuse analysis: Next, analyze the data dependency relationship and determine the data reuse strategy. Data reuse is a key factor in improving computing efficiency. For example, multiple sub-blocks can be combined into a larger data block to reduce the number of data loads. For example, 4 sub-blocks of 2x2 can be combined into a 4x4 data block, which can reduce the number of data loads and improve data locality.
[0086] Cache Capacity Analysis: Then, analyze the capacity of the AM to determine the shape information of the microkernel so that it can store the data of the entire sub-block. For example, the AM capacity of the FT-M7032 DSP chip is 768KB, so the shape of the microkernel should be able to store a 64x16 data block, that is, 1024KB.
[0087] Comprehensive Analysis: Finally, based on the above analysis results, determine the shape information of the microkernel, such as the number of data columns and the number of data rows. For example, for the FT-M7032 DSP chip and the 2D5P template, the shape of the microkernel can be 64x16.
[0088] 3. Instruction Pipeline Optimization
[0089] Instruction Selection and Packing: Select the instructions that are most suitable for the microkernel computing mode and pack them into instruction packets to ensure that multiple FMAC (fused multiply-accumulate instructions) instructions can be executed within each clock cycle.
[0090] Latency Hiding: Through a carefully designed instruction pipeline, hide the latency caused by data dependencies. For example, by pre-computing and storing intermediate results to avoid repeated data access in subsequent calculations.
[0091] 4. Data Reuse Strategy
[0092] Vector Register Reuse: Implement multiple reuses of data in vector registers within the microkernel to reduce access to external memory.
[0093] Cache Optimization: Utilize the cache mechanism of the DSP processor, such as L1 and L2 caches, to further reduce access to external memory.
[0094] 5. Microkernel Computing
[0095] Microkernel Initialization: Before the microkernel starts computing, it is necessary to initialize the necessary parameters and registers first. This includes loading template parameters, configuring vector registers, and setting the working mode of the VFMAC unit.
[0096] Data Loading: The microkernel loads the required sub-block data from the on-chip memory (AM) into the vector registers. This step utilizes the DMA engine to achieve asynchronous data transfer to reduce the computing waiting time.
[0097] Computing Execution: Once the data is loaded into the vector registers, the microkernel starts to execute the computing. During the computing process, the microkernel utilizes multiple floating-point multiply-accumulate units (VFMAC) in the vector processing unit (VPU) of the DSP to process the data in parallel. The computing process of the microkernel follows a specific algorithm logic, and according to the requirements of the template calculation, performs predetermined mathematical operations on each grid point and its neighboring points.
[0098] Intermediate result processing: The intermediate results generated during the calculation can be temporarily stored in the vector register for subsequent calculations, which can reduce the number of accesses to external storage and improve the calculation efficiency.
[0099] Result storage: After the calculation is completed, the microkernel stores the result data back to the on-chip memory (AM). This step also uses the DMA engine to achieve asynchronous data transfer so that the microkernel can perform other calculation tasks simultaneously.
[0100] Exception handling and synchronization: During the microkernel calculation, possible exception situations need to be handled and the synchronization of the calculation needs to be ensured. This includes ensuring that all dependent data has been loaded and that the calculation results are correctly stored.
[0101] 6. Performance model construction
[0102] Performance prediction: Construct a performance model to predict the performance under different microkernel sizes, instruction pipeline configurations, and triple-buffer mechanism settings.
[0103] Parameter adjustment: Adjust the parameters of the microkernel according to the performance test results to optimize the performance.
[0104] It should be noted that the sub - blocks adopt an overlapping block strategy. Each sub - block contains a boundary data area, which is used to ensure that the data of grid points inside the sub - block is independent of the data of other sub - blocks during the calculation process and to ensure the correctness of the calculation. The size of the sub - block is selected according to the on - chip storage capacity of the digital signal processor, the characteristics of the vector calculation unit, and the data reuse strategy, so as to maximize data locality, reduce the number of data transmissions, and make full use of the vector calculation unit of the DSP. The calculation of the sub - block is carried out in a vectorized manner, forming vectors from multiple adjacent grid points and performing parallel processing by the vector calculation unit, so as to make full use of the instruction - level parallelism and data - level parallelism of the DSPs and improve the calculation efficiency. The loading and storage operations of the sub - block adopt a triple - buffer mechanism, decoupling the loading, calculation, and storage operations, allowing them to overlap with the calculation tasks, so as to effectively utilize the DMA engine of the DSP for asynchronous data transmission and increase the degree of calculation - communication overlap. The loading and storage operations of the sub - block use the DMA engine for asynchronous data transmission to improve the data transmission efficiency and reduce the data transmission delay. The loading and storage operations of the sub - block adopt a two - level triple - buffer mechanism, decoupling the loading and storage operations, allowing them to overlap with the calculation tasks, so as to increase the degree of calculation - communication overlap and further accelerate the template calculation. The calculation of the sub - block is carried out using a specially designed micro - kernel to make full use of the vector calculation unit of the digital signal processor, and by optimizing the instruction pipeline, reducing the instruction execution delay, and improving the calculation efficiency. The data reuse strategy includes combining multiple sub - blocks into a larger data block to reduce the number of data loadings, increase data locality, reduce the number of data transmissions, and improve the calculation efficiency. The loading and storage operations of the sub - block adopt data pre - fetching and post - fetching techniques to reduce the data transmission delay and improve the data transmission efficiency.
[0105] It should be understood that although Figure 1 - Figure 2 , Figure 6 the steps in the flowchart of Figure 1 - Figure 2 , Figure 6 are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order limit for the execution of these steps, and these steps can be executed in other orders. Moreover,
[0106] in one embodiment, as Figure 7As shown, a template calculation optimization device on a multi-core DSP is provided, including: a data block allocation module 702, a data block splitting module 704, a template calculation module 706, and a data transfer module 708, where:
[0107] The data block allocation module 702 is configured to divide the input data grid into several data blocks and allocate a DSP core to each data block. Each data block contains several grid points.
[0108] The data block splitting module 704 is configured to, at the current time step, when all DSP cores have completed the first-level calculation on the allocated data blocks, further divide the data blocks into multiple sub-blocks, and each sub-block contains multiple adjacent grid points.
[0109] The template calculation module 706 is configured to, at the next time step, load the sub-blocks into the vector on-chip memory of the DSP and use the vector calculation unit of the DSP to perform the second-level calculation on the sub-blocks to obtain the new values of each grid point.
[0110] The data transfer module 708 is configured to pack the calculated new values into new sub-blocks and store them back into the vector on-chip memory. The new sub-blocks are transferred to the off-chip memory for loading and storage through the data movement unit of the DSP using a triple-buffer mechanism.
[0111] For the specific limitations of a template calculation optimization device on a multi-core DSP, reference can be made to the limitations of a template calculation optimization method on a multi-core DSP in the above text, which will not be elaborated here. Each module in the above template calculation optimization device on a multi-core DSP can be implemented in whole or in part through software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each of the above modules.
[0112] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 8As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a method for optimizing template calculation on a multi-core DSP. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the outer shell of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0113] Those skilled in the art can understand that Figure 7 - Figure 8 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0114] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the following steps are implemented:
[0115] Divide the input data grid into several data blocks, and allocate a DSP core to each data block. Each data block contains several grid points.
[0116] At the current time step, all DSP cores have completed the first-level calculation of the allocated data blocks, and further divide the data blocks into multiple sub-blocks. Each sub-block contains multiple adjacent grid points.
[0117] At the next time step, load the sub-blocks into the vector on-chip memory of the DSP, and use the vector calculation unit of the DSP to perform a second-level calculation on the sub-blocks to obtain new values for each grid point.
[0118] Pack the calculated new values into new sub-blocks and store them back in the vector on-chip memory. The new sub-blocks are transferred from the vector on-chip memory to the off-chip memory for loading and storage through the data movement unit of the DSP using a triple-buffer mechanism.
[0119] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0120] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0121] The above-described embodiments merely represent several implementation manners of this application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of this application, several modifications and improvements can still be made, and these all belong to the protection scope of this application. Therefore, the protection scope of this application should be subject to the appended claims.
Claims
1. A template calculation optimization method on a multi-core DSP, characterized in that: The method comprises: Acquire parameter information of the multi-core DSP, wherein the parameter information includes: information of the data grid, template mode information and DSP parameter information; Divide the input data grid into a plurality of data blocks, and allocate a DSP core to each of the data blocks; the data block includes a plurality of grid points; At the current time step, all of the DSP cores complete the first-level calculation of the allocated data block, and further divide the data block into a plurality of sub-blocks, each of which contains a plurality of adjacent grid points; At the next time step, the sub-block is loaded into the vector on-chip memory of the DSP, and the vector computing unit of the DSP is used to perform secondary calculation on the sub-block to obtain a new value of each grid point; The sub-blocks are loaded into the vector on-chip memory of the DSP by using a triple buffer mechanism, and the memory access path information is determined based on the grid parameter information corresponding to the multi-core DSP chip, the template mode information and the DSP parameter information; According to the path information, each of the sub-blocks is loaded, calculated and stored in three stages through a first-level three-buffer, so that the sub-block is completely processed by a vector processing unit at the first level; Using the secondary triple buffer, the vector computing unit of the DSP performs three stages of secondary computing of loading, computing and storing the sub-blocks that have been processed at the primary level between the main memory and the vector on-chip memory, so as to obtain a new value of each grid point; The three-buffer mechanism includes: a primary three-buffer and a secondary three-buffer; The first level three buffering occurs between the vector on-chip memory and the vector register of the DSP, and is used to realize the parallel execution of data loading and calculation within the sub-block; The secondary triple buffer occurs between the vector on-chip memory and the off-chip memory of the DSP, and is used to implement parallel execution of data loading, calculation and storage between sub-blocks; The calculated new values are packaged into new sub-blocks and stored back in the vector on-chip memory. The new sub-blocks are transferred from the vector on-chip memory to the off-chip memory through the data movement unit of the DSP using a triple buffer mechanism for loading and storage.
2. The method according to claim 1, characterized in that The information of the data grid includes: grid size and grid dimension; The template pattern information includes: data dependency range, data dependency point number and data dependency shape; The DSP parameter information includes: the size of global shared memory, the number of floating-point multiplication and addition units, the number of vector registers, the memory bandwidth, and the number of DSP cores.
3. The method according to claim 2, characterized in that Divide the input data grid into several data blocks and assign a DSP core to each of the data blocks, including: The microkernel shape is determined according to the parameter information, the input data grid is divided into a plurality of data blocks with the same number of data columns and data rows according to the microkernel shape, and a DSP core is allocated to each of the data blocks.
4. The method according to any one of claims 1 to 3, characterized in that: At the current time step, all the DSP cores complete the first-level calculation of the allocated data block, and further divide the data block into a plurality of sub-blocks, each of which contains a plurality of adjacent grid points, including: At the current time step, the DSP core to which the data block has been allocated selects instructions corresponding to the primary computing mode according to DSP parameter information, template mode information and microkernel shape, and packages the instructions corresponding to the primary computing mode into a primary instruction package to ensure that a plurality of multiplication and accumulation instructions can be executed in each clock cycle; According to the multiply-accumulate instruction, the allocated data block is subjected to primary calculation in a vectorized manner, and the data block is further divided into a plurality of sub-blocks. In the sub-blocks, a plurality of adjacent grid points are formed into vectors, which are processed in parallel by a vector calculation unit.
5. The method according to claim 4, characterized in that The method further comprises: performing three stages of loading, calculating and storing each sub-block through a first-level three-buffer according to the path information, so that the sub-block is completely processed by a vector processing unit at the first level, including: According to the path information, each of the sub-blocks is loaded, calculated and stored in three stages through the first-level three buffers. In the loading stage, the sub-block is loaded from the vector on-chip memory of the DSP to the vector register of the vector processing unit; in the calculation stage, the multiplication and addition operation unit of the vector register of the vector processing unit is used to perform floating-point multiplication and addition operations; in the storage stage, the sub-block is stored from the vector register back to the vector on-chip memory; Through the instruction set pipeline, the calculation of the current sub-block, the loading of the next sub-block and the storage of the previous sub-block are performed in the same time segment, so that all the sub-blocks are completely processed by the vector processing unit at the first level.
6. The method according to claim 4, characterized in that The sub-blocks that have been processed at the first level are loaded, calculated, and stored in three stages between the main memory and the vector on-chip memory by using the second-level triple buffer through the vector calculation unit of the DSP, and a new value of each grid point is obtained, including: The sub-blocks that have been processed at the first level are loaded, calculated and stored in three stages between the main memory and the vector on-chip memory by using the second-level triple buffer through the vector calculation unit of the DSP. In the loading stage, the sub-blocks are transferred from the main memory to the vector on-chip memory of the DSP; in the calculation stage, all operations of the first-level triple buffer are performed; in the storage stage, the sub-blocks are stored from the vector on-chip memory back to the main memory; Through a loop operation, when the sub-block in the first buffer is used for calculation, the next sub-block is loaded into the second buffer of the vector on-chip memory until the calculation is completed and a new value of each grid point is obtained.
7. The method according to any one of claims 4 to 6, characterized in that: The data transmission of the sub-blocks onto the chip or off the chip is optimized by using interleaved single-word and double-word instructions; The size of the sub-block is selected according to the on-chip storage capacity of the DSP, the characteristics of the vector computing unit and the data reuse strategy; The loading and storing operations of the sub-blocks adopt a DMA engine to perform asynchronous data transfer; The calculation of the sub-block is performed by using a micro-kernel to utilize the vector computing unit and optimized instruction pipeline technology of the DSP; The shape of the microkernel is designed according to the number of vector registers of the DSP, the instruction packet filling constraint and the instruction delay constraint of the vector computing unit; The calculated new value is packaged into a new sub-block and stored back in the vector on-chip memory. The new sub-block is transferred from the vector on-chip memory to the off-chip memory through the data movement unit of the DSP using a triple buffer mechanism for loading and storage, including: The calculated new values are packaged into new sub-blocks and stored back in the vector on-chip memory. All the new sub-blocks are stored in the third buffer using a three-buffer mechanism and written back from the vector on-chip memory to the main memory through the data movement unit of the DSP to complete the loading and storage of input data.
8. A template calculation optimization device on a multi-core DSP, characterized in that: The device comprises: A data block allocation module is used to obtain parameter information of a multi-core DSP, wherein the parameter information includes: information of a data grid, template mode information, and DSP parameter information; divide the input data grid into a plurality of data blocks, and allocate a DSP core to each of the data blocks; the data block contains a plurality of grid points; A data block segmentation module, used for, at the current time step, all the DSP cores complete the first-level calculation of the allocated data block, and further divide the data block into a plurality of sub-blocks, wherein the sub-blocks include a plurality of adjacent grid points; The template calculation module is used to load the sub-block into the vector on-chip memory of the DSP at the next time step, and use the vector calculation unit of the DSP to perform secondary calculation on the sub-block to obtain a new value of each grid point; use a three-buffer mechanism to load the sub-block into the vector on-chip memory of the DSP, determine the memory access path information based on the grid parameter information corresponding to the multi-core DSP chip, the template mode information and the DSP parameter information; perform three stages of loading, calculating and storing for each sub-block through the first-level three-buffer according to the path information, so that the sub-block is completely processed by the vector processing unit at the first level The method comprises the following steps: using a secondary three-buffer to load, calculate and store three stages of secondary calculations between the main memory and the vector on-chip memory of the sub-block that has been processed at the first level through the vector calculation unit of the DSP, so as to obtain a new value of each grid point; the three-buffer mechanism comprises: a primary three-buffer and a secondary three-buffer; the primary three-buffer occurs between the vector on-chip memory and the vector register of the DSP, and is used to realize the parallel execution of data loading and calculation within the sub-block; the secondary three-buffer occurs between the vector on-chip memory and the off-chip memory of the DSP, and is used to realize the parallel execution of data loading, calculation and storage between sub-blocks; The data handling module is used to pack the calculated new values into new sub-blocks and store them back into the vector on-chip memory. The new sub-blocks are transferred from the vector on-chip memory to the off-chip memory through the data movement unit of the DSP using a triple buffer mechanism for loading and storage.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.