Thread block data prefetching method and device, electronic equipment and storage medium

By creating data prefetch thread warps in a general parallel computing architecture, the problems of low efficiency and accuracy of cross-thread block data prefetching are solved, achieving more efficient data prefetching and improving processor performance.

CN120780618APending Publication Date: 2025-10-14YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510825991.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-14

AI Technical Summary

Technical Problem

In general parallel computing architectures, the efficiency and accuracy of data prefetching across thread blocks are low, which limits the application of prefetching technology and makes it impossible to effectively improve the computing performance of the processor.

Method used

By creating data prefetch warps, calculating the data starting base address of each warp in the next thread block, and synchronously calculating the data starting base address of the next thread block when executing the warp of the current thread block, data prefetching across thread block boundaries is achieved.

Benefits of technology

The efficiency and accuracy of data prefetching are improved, the application scope of prefetching technology in general parallel computing architecture is enhanced, and the computing performance of the processor and the computing efficiency of data operation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780618A_ABST
    Figure CN120780618A_ABST
Patent Text Reader

Abstract

The invention provides a thread block data prefetching method and device, electronic equipment and a storage medium, and relates to the technical field of computers.The thread block data prefetching method comprises the steps that a data prefetching thread bundle corresponding to a next thread block is created; the data prefetching thread bundle is used for calculating a data initial base address of each thread bundle in the next thread block; executing each thread bundle in the current thread block and executing the data prefetching thread bundle, and determining a data initial base address of each thread bundle in the next thread block; and obtaining operation data corresponding to the next thread block based on the data initial base address of each thread bundle in the next thread block. According to the method and the device provided by the invention, the data pre-fetching of the cross-thread block boundary is realized, and the efficiency and the accuracy of the data pre-fetching are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a thread block data prefetching method, device, electronic device and storage medium. Background Art

[0002] General-purpose parallel computing architectures rely on the scheduling of dozens of concurrent warps to achieve high parallelism. Although various current warp scheduling strategies can quickly context switch between available warps to mask the performance loss of long-latency memory operations, long-latency memory operations remain one of the most critical performance bottlenecks in high-performance computing.

[0003] To address the performance degradation caused by long-latency memory operations, researchers applied prefetching technology to general-purpose parallel computing architectures. In general-purpose parallel computing architectures, there are regular strides between the computational data of threads. The base address and stride calculated from the current thread warp can be used to issue prefetch requests for subsequent threads. Since scheduling is performed at the granularity of thread warps in general-purpose parallel computing architectures, data prefetching between threads is essentially data prefetching between thread warps. The starting base address of a thread block is a complex function of the thread index and the thread block index, and also depends on the upper-level programmer's habit of allocating thread block data. Therefore, the starting base address accessed by the first thread in a thread block is difficult to predict. As a result, when prefetching across thread blocks, the prefetch address does not match the required fetch, and the prefetch is invalid, which limits the application of prefetching technology in general-purpose parallel computing architectures.

[0004] Therefore, how to improve the efficiency and accuracy of data prefetching in a general parallel computing architecture has become a technical problem that needs to be solved urgently in the industry. Summary of the Invention

[0005] The present invention provides a thread block data prefetching method, device, electronic device and storage medium, which are used to solve the technical problem of how to improve the efficiency and accuracy of data prefetching in a general parallel computing architecture.

[0006] The present invention provides a thread block data prefetching method, comprising: Creating a data prefetch warp corresponding to a next thread block; the data prefetch warp is used to calculate a data starting base address of each warp in the next thread block; Executing each thread warp in the current thread block and executing the data prefetch thread warp, and determining a data starting base address of each thread warp in the next thread block; Based on the data starting base addresses of each thread warp in the next thread block, operation data corresponding to the next thread block is obtained.

[0007] In some embodiments, creating a data prefetch warp corresponding to the next thread block includes: determining a current thread block; splitting the current thread block to determine each thread bundle of the current thread block; based on thread block index information and thread block size information of the next thread block, creating a data prefetch thread bundle corresponding to the next thread block; writing each thread bundle of the current thread block and the data prefetch thread bundle into a thread bundle queue; the thread bundle queue is used to schedule execution of each thread bundle.

[0008] In some embodiments, the execution of each thread bundle in the current thread block and the execution of the data prefetch thread bundle to determine the data start base address of each thread bundle in the next thread block comprises: scheduling execution of each thread bundle in the thread bundle queue, so that the data prefetch thread bundle determines the data start base address of each thread bundle in the next thread block based on the start base address of the next thread block and the thread bundle address step.

[0009] In some embodiments, the thread bundle queue includes an active thread bundle table and a suspended thread bundle table; the active thread bundle table is used to store state information of an executed thread bundle; the suspended thread bundle table is used to store state information of a suspended thread bundle; The scheduling execution of each thread bundle in the thread bundle queue comprises: scheduling execution of a current thread bundle in the active thread bundle table; In the case where the data response waiting duration corresponding to the current thread bundle is greater than a preset duration, transferring the current thread bundle from the active thread bundle table to the suspended thread bundle table; In the case where the data corresponding to the current thread bundle is received, transferring the current thread bundle from the suspended thread bundle table to the active thread bundle table; In the case where the execution of the current thread bundle ends, removing the current thread bundle from the active thread bundle table.

[0010] In some embodiments, the obtaining of the operation data corresponding to the next thread block based on the data start base address of each thread bundle in the next thread block comprises: determining the data start base address of a current thread bundle in the next thread block; based on the data start base address of the current thread bundle, obtaining operation data corresponding to the current thread bundle; caching the operation data corresponding to the current thread bundle, and switching the current thread bundle to a next thread bundle.

[0011] In some embodiments, the method further comprises: acquire a data access request of the next thread block; compare a first data address in the data access request with a second data address of the operation data; determine that the operation data corresponding to the next thread block is pre-fetched incorrectly in a case that the first data address is inconsistent with the second data address.

[0012] In some embodiments, the determining that the operation data corresponding to the next thread block is pre-fetched incorrectly comprises: increase a count value of a counter in a case that the operation data corresponding to the next thread block is pre-fetched incorrectly; the counter is used to record a number of times that the operation data corresponding to each thread block is pre-fetched incorrectly; empty a cache space of the operation data corresponding to each thread block in a case that the count value is greater than a preset threshold.

[0013] The application provides a thread block data pre-fetching device, comprising: a creating module, used for creating a data pre-fetching thread bundle corresponding to a next thread block; the data pre-fetching thread bundle is used to calculate data start base addresses of each thread bundle in the next thread block; an executing module, used for executing each thread bundle in a current thread block and executing the data pre-fetching thread bundle, and determining the data start base addresses of each thread bundle in the next thread block; a pre-fetching module, used for acquiring operation data corresponding to the next thread block based on the data start base addresses of each thread bundle in the next thread block.

[0014] The application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor; the processor executes the computer program to implement the thread block data pre-fetching method.

[0015] The application provides a non-transitory computer readable storage medium, which stores a computer program; the computer program is executed by a processor to implement the thread block data pre-fetching method.

[0016] The thread block data prefetching method, device, electronic device and storage medium provided by the present invention create a data prefetching thread bundle corresponding to the next thread block; execute each thread bundle in the current thread block and execute the data prefetching thread bundle to determine the data starting base address of each thread bundle in the next thread block; based on the data starting base address of each thread bundle in the next thread block, obtain the operation data corresponding to the next thread block; since the data prefetching thread bundle is created, the data starting base address of each thread bundle in the next thread block is synchronously calculated during the execution of each thread bundle of the current thread block, thereby realizing data prefetching across thread block boundaries, effectively improving the application scope of prefetching technology in general parallel computing architecture, improving the efficiency and accuracy of data prefetching, further improving the computing performance of the processor by prefetching data across thread block boundaries, and improving the computing efficiency of operation data. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.

[0018] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 It is a schematic diagram of the general parallel computing architecture provided by the present invention.

[0020] Figure 2 It is a schematic diagram of the computing task scheduling granularity provided by the present invention.

[0021] Figure 3 This is a schematic diagram of cross-thread block boundary address prediction provided by the present invention.

[0022] Figure 4 It is a flowchart of the thread block data prefetching method provided by the present invention.

[0023] Figure 5 It is a structural diagram of the thread block data prefetching device provided by the present invention.

[0024] Figure 6 It is a schematic diagram of the hardware architecture of the data prefetching method provided by the present invention.

[0025] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0027] It should be noted that the terms "first," "second," and the like in the present invention are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps, units, or modules is not necessarily limited to those steps, units, or modules explicitly listed, but may include other steps, units, or modules that are not explicitly listed or that are inherent to these processes, methods, products, or apparatuses.

[0028] ‌ Figure 1 is a schematic diagram of the general parallel computing architecture provided by the present invention, such as Figure 1 As shown, a processor adopting a general parallel computing architecture (general parallel computing architecture processor) can be a graphics processing unit (GPU), a general-purpose graphics processing unit (GPGPU), a domain-specific architecture (DSA), etc.

[0029] Structurally, a general-purpose parallel computing architecture consists of at least a streaming multiprocessor (SM), an interconnection network, a level 2 cache (L2 cache), and dynamic random access memory (DRAM). DRAM is primarily used as the main memory (main memory) of a computer system.

[0030] A streaming multiprocessor is an arithmetic unit primarily composed of an instruction pipeline and a level-1 cache (L1 cache). The instruction pipeline is key to the computing functionality of the streaming multiprocessor. During the execution phase of the instruction pipeline, threads within a warp are assigned to various streaming processors for execution. A streaming processor (SP) implements functions such as the arithmetic logic unit (ALU), special function unit (SFU), and floating-point unit (FPU). The number of streaming processors determines the degree of parallelism within a warp. When the number of streaming processors matches the number of threads in a warp, the degree of parallelism is maximized, allowing all threads in the warp to execute in full parallelism. The streaming multiprocessor also includes a warp scheduling module for warp scheduling and branch control; a register file for caching data and intermediate results during computation; and shared memory for data exchange between warps and threads. The L1 cache, connected to the interconnect network, can be used to prefetch data and instructions from the L2 cache, further reducing pipeline memory access time.

[0031] The interconnection network connects the L2 cache with the L1 cache of each streaming multiprocessor to achieve efficient data transmission. The bandwidth upper limit of the interconnection network determines the memory access efficiency and speed of the entire processor and is one of the performance bottlenecks limiting general-purpose parallel computing processors.

[0032] The L2 cache pre-fetches data in the memory by merging memory accesses to reduce pipeline stalls caused by data memory access and improve computing efficiency. Every two or more L2 caches reuse a memory controller to achieve access to the memory (DRAM).

[0033] A general-purpose parallel computing architecture processor acts as a coprocessor, handling computationally intensive tasks that the central processing unit (CPU) is not good at. The CPU issues computing tasks, which are usually scheduled in the form of thread blocks in a general-purpose parallel computing architecture processor.

[0034] Figure 2 This is a schematic diagram of the computing task scheduling granularity provided by the present invention, such as Figure 2As shown, each computing task consists of multiple parallel functions (kernels), each of which can be divided into several thread blocks. A thread block is the smallest scheduling granularity visible to upper-level programmers. Thread blocks are assigned to various streaming multiprocessors for computation. Within the streaming multiprocessor, the instruction pipeline divides thread blocks into several warps. Each warp contains several threads. The instruction pipeline schedules tasks using warps as the smallest granularity.

[0035] The current memory wall problem (the significant gap between processor performance and memory access performance) is one of the main obstacles hindering the improvement of high-performance processor computing performance. Memory data read and write speeds significantly lag behind computing speed. In general parallel computing architectures, computing efficiency is improved through high-concurrency operations across a large number of parallel threads, but the memory wall problem persists. Long-latency data load / store instructions for memory access can significantly block instruction pipeline execution. Although various warp scheduling strategies can mask latency by suspending long-latency warps, sudden L1 cache misses can still cause memory system congestion, significantly increasing the queuing delay of suspended warps.

[0036] The application of prefetching technology in general-purpose parallel computing architectures enables prefetching computational data for warps within the same thread block, significantly improving computational efficiency. However, for cross-thread block data prefetching, there is a significant discrepancy between ideal and actual prediction conditions.

[0037] Figure 3 Schematic diagram of cross-thread block boundary address prediction provided by the present invention, such as Figure 3 As shown in the figure, under ideal prediction conditions, each thread block (TB) has the same size. The storage addresses of the data corresponding to the thread blocks are continuous, and the stride of the starting base addresses of adjacent thread blocks (TB_stride, the difference between the addresses of the first thread data storage in adjacent thread blocks) is the same. The starting base address (Base Address) is the starting point of data storage and is used to determine the location of data in memory. Each thread block has a starting base address, which is the starting point for all threads in the thread block to access memory. It can be expressed as: TB_stride = thread_stride * TB_thread_num.

[0038] Among them, TB_stride is the stride of the starting base addresses of adjacent thread blocks, that is, the thread block address stride; thread_stride is the thread address stride, which is bound to the calculation accuracy of the general parallel computing architecture and is a fixed value; TB_thread_num is the number of threads in the thread block.

[0039] The starting base address of a thread block can be calculated by a simple linear function, denoted as: TB_base_addr = Global_base_addr + block_id * TB_stride.

[0040] wherein TB_base_addr is the starting base address of a thread block; Global_base_addr is the global base address, which is the storage address of the data of the first thread of the first thread block in the computing task; and block_id is the thread block index.

[0041] After the thread block is divided into a plurality of thread bundles, the starting base address of a thread bundle is calculated as follows: TB_warp_base_addr = TB_base_addr + warp_id * warp_stride.

[0042] wherein TB_warp_base_addr is the starting base address of a thread bundle; warp_id is the thread bundle index; and warp_stride is the thread bundle address stride. The thread bundle address stride refers to the interval of the memory addresses accessed by the first threads in adjacent thread bundles. In other words, it is the difference between the starting base addresses of adjacent thread bundles.

[0043] The thread bundle address stride can be denoted as: warp_stride = thread_stride * warp_thread_num.

[0044] wherein warp_thread_num is the number of threads in a thread bundle.

[0045] The storage addresses of the data of the threads in a thread bundle are calculated as follows: warp_thread_addr = TB_warp_base_addr + thread_id * thread_stride.

[0046] wherein warp_thread_addr is the storage address of the data of a thread; thread_id is the thread index; and thread_stride is the thread address stride. The thread address stride refers to the interval of the memory addresses accessed by adjacent threads in the same thread bundle. In other words, it is the difference between the memory addresses accessed by each thread in the thread bundle.

[0047] However, in practical general-purpose parallel computing architectures, thread block scheduling policies typically do not assign adjacent thread blocks to the same streaming multiprocessor when allocating thread blocks. Furthermore, thread block sizes are not necessarily equal, resulting in inconsistent strides for data access addresses even for adjacent thread blocks. Furthermore, the upper-level programming environment's allocation of thread block input data depends on programming habits. These factors result in the thread block's starting base address not being a simple linear function of the thread block index and address stride. Therefore, predicting the starting base address of the first thread in a thread block cannot be achieved using a simple prediction algorithm. This results in a mismatch between the prefetched address and the requested address when prefetching data across thread blocks, rendering the prefetch ineffective and limiting the application of prefetching technology in general-purpose parallel computing architectures.

[0048] In order to solve the above technical problems, Figure 4 FIG. 1 is a flow chart of the thread block data prefetching method provided by the present invention, as shown in FIG. Figure 4 As shown, the method includes step 410 , step 420 and step 430 .

[0049] Step 410: Create a data prefetch warp corresponding to the next thread block; the data prefetch warp is used to calculate the data starting base address of each warp in the next thread block.

[0050] Specifically, the thread block data prefetching method provided in an embodiment of the present invention is implemented by a thread block data prefetching device. This device can be implemented in software, such as a thread block data prefetching program running in a streaming multiprocessor, or in hardware, such as a processor that executes the thread block data prefetching method.

[0051] During the execution of the current computing task, the current computing task can be divided into multiple thread blocks. The current thread block is the thread block that needs to be scheduled for execution. The next thread block is the thread block that is scheduled for execution after the current thread block.

[0052] The warp scheduling unit in the streaming multiprocessor's pipeline management module can split the current thread block into multiple warps and add these warps to the warp queue for scheduling at the warp granularity. Furthermore, a prefetch scheduling unit can be configured in the streaming multiprocessor's pipeline management module to create a data prefetch warp corresponding to the next thread block. The data prefetch warp is primarily used to calculate the data starting base address (TB_warp_base_addr) for each warp in the next thread block. The data starting base address refers to the starting base address of the computation data in memory.

[0053] Step 420: Execute each warp in the current thread block and execute data prefetch warps to determine the data starting base address of each warp in the next thread block.

[0054] Specifically, the data prefetch warp corresponding to the next thread block and each warp of the current thread block can be added to the warp queue, so that the instruction pipeline in the streaming multiprocessor can synchronously execute each warp in the current thread block and the data prefetch warp.

[0055] After the data prefetch thread warp enters the instruction pipeline, the stream processor calculates the data starting base address of each thread warp in the next thread block according to the thread block starting base address calculation method and the thread warp starting base address calculation method.

[0056] Step 430: Based on the data starting base addresses of each thread warp in the next thread block, obtain the operation data corresponding to the next thread block.

[0057] Specifically, in general parallel computing architectures, warps serve as the smallest scheduling unit. Each warp size after each thread block is split is equal, meaning the address stride for data access across warps is the same. In other words, the computational data corresponding to each warp can be retrieved based on the data starting base address of each warp.

[0058] According to the operation data corresponding to each thread warp in the next thread block, the operation data corresponding to the next thread block can be obtained.

[0059] The thread block data prefetching method provided by the embodiment of the present invention creates a data prefetching thread warp corresponding to the next thread block; executes each thread warp in the current thread block and executes the data prefetching thread warp to determine the data starting base address of each thread warp in the next thread block; obtains the operation data corresponding to the next thread block based on the data starting base address of each thread warp in the next thread block; since the data prefetching thread warp is created, the data starting base address of each thread warp in the next thread block is synchronously calculated during the execution of each thread warp of the current thread block, thereby realizing data prefetching across thread block boundaries, effectively improving the application scope of prefetching technology in general parallel computing architecture, improving the efficiency and accuracy of data prefetching, further improving the computing performance of the processor by prefetching data across thread block boundaries, and improving the computing efficiency of operation data.

[0060] It should be noted that each embodiment of the present invention can be freely combined, the order can be changed, or it can be executed separately, and does not need to rely on or depend on a fixed execution order.

[0061] In some embodiments, creating a data prefetch warp corresponding to a next thread block includes: Determine the current thread block; Split the current thread block and determine the thread warps of the current thread block; Creating a data prefetch warp corresponding to the next thread block based on the thread block index information and thread block size information of the next thread block; Each warp of the current thread block and the data prefetch warp are written into a warp queue; the warp queue is used to schedule the execution of each warp.

[0062] Specifically, the process of creating the data prefetch warp corresponding to the next thread block is synchronized with the process of splitting the current thread block.

[0063] The warp scheduling unit divides the current thread block into its own warps. The prefetch scheduling unit creates a data prefetch warp for the next thread block. This data prefetch warp includes the thread block index and size of the next thread block, which can be used to calculate the starting base address of the thread block.

[0064] Each warp of the current thread block and the data prefetch warp are written into the warp queue, so that the warp scheduling unit can schedule each warp to enter the instruction pipeline for execution according to the warp queue.

[0065] The thread block data prefetching method provided by the embodiment of the present invention synchronously performs the segmentation of the current thread block and the creation of the data prefetching warp, fully utilizes the computing performance of the processor, and improves the efficiency of data prefetching.

[0066] In some embodiments, executing each warp in a current thread block and executing a data prefetch warp, and determining a data starting base address for each warp in a next thread block includes: The warps in the warp queue are scheduled for execution, so that the data prefetch warp determines the data starting base address of each warp in the next thread block based on the starting base address of the next thread block and the warp address stride.

[0067] Specifically, the instruction pipeline schedules execution of each warp (including warps in the current thread block and data prefetch warps) based on the warp queue. When executing the data prefetch warp, the data starting base address of each warp in the next thread block can be calculated based on the starting base address of the next thread block and the warp address stride. For example, the starting base address of a warp is calculated as follows: TB_warp_base_addr = TB_base_addr + warp_id * warp_stride.

[0068] Among them, TB_warp_base_addr is the starting base address of the thread warp; TB_base_addr is the starting base address of the thread block; warp_id is the thread warp index; warp_stride is the thread warp address stride.

[0069] The thread block data prefetching method provided by the embodiment of the present application fully utilizes the computing performance of the processor and improves the efficiency of data prefetching.

[0070] In some embodiments, the thread bundle queue comprises an active thread bundle table and a suspended thread bundle table; the active thread bundle table is configured to store state information of the executed thread bundle; and the suspended thread bundle table is configured to store state information of the suspended thread bundle. Scheduling the execution of each thread bundle in the thread bundle queue comprises: Scheduling the execution of the current thread bundle in the active thread bundle table; In the case that the data response waiting duration corresponding to the current thread bundle is greater than the preset duration, the current thread bundle is transferred from the active thread bundle table to the suspended thread bundle table; In the case that the data corresponding to the current thread bundle is received, the current thread bundle is transferred from the suspended thread bundle table to the active thread bundle table; In the case that the execution of the current thread bundle is completed, the current thread bundle is removed from the active thread bundle table.

[0071] Specifically, in the stream pipeline management module in the stream multi-processor, the thread bundle queue can specifically comprise an active thread bundle table and a suspended thread bundle table. The active thread bundle table is configured to store state information of the thread bundle being executed in the instruction pipeline; and the suspended thread bundle table is configured to store state information of the thread bundle being suspended in the instruction pipeline.

[0072] When the thread bundle scheduling unit schedules the execution of each thread bundle in the thread bundle queue, the execution of the current thread bundle is scheduled according to the active thread bundle table.

[0073] If the data response waiting duration corresponding to the current thread bundle is greater than the preset duration, it indicates that the current thread bundle performs a long-latency memory operation, and during the waiting period for the data response, the current thread bundle can be transferred from the active thread bundle table to the suspended thread bundle table. The preset duration can be set according to actual conditions.

[0074] In the case that the data corresponding to the current thread bundle is received, the current thread bundle is transferred from the suspended thread bundle table to the active thread bundle table. After the current thread bundle processes the operation data and the execution is completed, the current thread bundle is removed from the active thread bundle table.

[0075] The thread block data prefetching method provided by the embodiment of the present application manages the execution scheduling of each thread bundle by setting the active thread bundle table and the suspended thread bundle table, thereby improving the computing efficiency of the operation data.

[0076] In some embodiments, the operation data corresponding to the next thread block is obtained based on the data start base address of each thread bundle in the next thread block, comprising: Determine the data starting base address of the current thread warp in the next thread block; Based on the data starting base address of the current thread warp, obtain the operation data corresponding to the current thread warp; Cache the computation data corresponding to the current warp and switch the current warp to the next warp.

[0077] Specifically, based on the data starting base address of the current warp in the next thread block, the computational data corresponding to the current warp can be retrieved from the L2 cache or memory. The cache scheduling unit in the L1 cache of the streaming multiprocessor can cache the computational data corresponding to the current warp in the data cache unit, waiting to be called by the stream processor for computation.

[0078] After caching the operation data corresponding to the current thread warp, the current thread warp is switched to the next thread warp, and the operation data corresponding to the next thread warp is continuously scheduled and cached.

[0079] The thread block data prefetching method provided by the embodiment of the present invention completes data prefetching in sequence according to the order of each thread warp, effectively expanding the application scope of the prefetching technology in a general parallel computing architecture and improving the efficiency and accuracy of data prefetching.

[0080] In some embodiments, the method further comprises: Get the data access request of the next thread block; Comparing the first data address in the data access request with the second data address of the operation data; When the first data address is inconsistent with the second data address, it is determined that prefetching of operation data corresponding to the next thread block is incorrect.

[0081] Specifically, when executing the warp of the next thread block, the cache scheduling unit of the L1 cache receives the data access request from the next thread block and then compares the first data address in the data access request with the second data address of the computation data. The first data address is the data storage address specified by the thread block when initiating the memory access request. The second data address is the memory address where the computation data is actually stored.

[0082] If the first data address and the second data address are consistent, the data prefetch for the next thread block is correct. If the first data address and the second data address are inconsistent, the data prefetch for the next thread block is incorrect. In this case, action is required, such as reloading the correct data from memory or notifying the system or program that the data prefetch failed and requires further processing.

[0083] The thread block data prefetching method provided by the embodiment of the present invention compares the first data address in the data memory access request with the second data address of the operation data to determine whether the operation data prefetching is erroneous, thereby improving the efficiency and accuracy of data prefetching.

[0084] In some embodiments, determining that prefetching of computation data corresponding to the next thread block fails includes: In the case where the prefetching of the operation data corresponding to the next thread block fails, the count value of the counter is increased; the counter is used to record the number of prefetching errors of the operation data corresponding to each thread block; When the count value is greater than a preset threshold, the cache space of the operation data corresponding to each thread block is cleared.

[0085] Specifically, a counter may be set to record the number of errors in prefetching computing data corresponding to each thread block.

[0086] In the case that the prefetching of the operation data corresponding to the next thread block fails, the count value of the counter may be increased, for example, the count value may be increased by one.

[0087] When the count value is greater than a preset threshold, an error signal is sent to the prefetch scheduling unit of the pipeline management module. The error signal is used to control the prefetch scheduling unit to clear the cache space of the operation data corresponding to each thread block.

[0088] If the preset threshold is set too high, prefetched data with errors will occupy cache space, resulting in low cache utilization. If the preset threshold is set too low, frequent flushing will occur, consuming significant processor time and system resources, negatively impacting overall system performance. Therefore, it is important to set the preset threshold appropriately based on actual conditions.

[0089] The thread block data prefetching method provided by an embodiment of the present invention sets a counter to record the number of errors in the prefetching of computing data corresponding to each thread block. When the count value is greater than a preset threshold, the cache space of the computing data corresponding to each thread block is cleared, thereby improving the utilization of the cache space.

[0090] The following describes an apparatus provided by an embodiment of the present invention. The apparatus described below and the method described above can refer to each other.

[0091] Figure 5 Schematic diagram of the structure of the thread block data prefetching device provided by the present invention. Figure 5 As shown, the device includes: A creation module 510 is configured to create a data prefetch warp corresponding to the next thread block; the data prefetch warp is used to calculate the data starting base address of each warp in the next thread block; An execution module 520 is configured to execute each warp in the current thread block and perform data prefetching on the warp, and determine a data starting base address for each warp in the next thread block; The prefetch module 530 is configured to obtain operation data corresponding to the next thread block based on the data starting base address of each thread warp in the next thread block.

[0092] The thread block data prefetching device provided by the embodiment of the present invention creates a data prefetching thread warp corresponding to the next thread block; executes each thread warp in the current thread block and executes the data prefetching thread warp to determine the data starting base address of each thread warp in the next thread block; obtains the operation data corresponding to the next thread block based on the data starting base address of each thread warp in the next thread block; since the data prefetching thread warp is created, the data starting base address of each thread warp in the next thread block is synchronously calculated during the execution of each thread warp of the current thread block, thereby realizing data prefetching across thread block boundaries, effectively improving the application scope of prefetching technology in general parallel computing architecture, improving the efficiency and accuracy of data prefetching, and further improving the computing performance of the processor by prefetching data across thread block boundaries, thereby improving the computing efficiency of operation data.

[0093] Figure 6 Schematic diagram of the hardware architecture of the data prefetching method provided by the present invention, such as Figure 6 As shown, the streaming multiprocessor includes an instruction pipeline, a pipeline management module and a first-level cache (L1 cache).

[0094] The instruction pipeline implements the operations of hardware instruction fetching, decoding, issuing, executing and writing back in the processor, and is the core module of the general parallel computing architecture.

[0095] The pipeline management module mainly includes a warp scheduling unit, a warp queue and a prefetch scheduling unit.

[0096] The warp scheduling unit is responsible for dividing received thread blocks into warps. Computational tasks are calculated and scheduled at the warp granularity in the instruction pipeline. The warp scheduling unit schedules the divided warps, as well as the data prefetch warps created by the prefetch scheduling unit, into the instruction pipeline for computation. When a warp is waiting for a data response from a long-latency memory access operation, it suspends it in a suspended warp table and prioritizes scheduling warps that are not performing memory access operations into the instruction pipeline. This conceals long-latency memory access operations and improves computational efficiency. When a data response arrives for a long-latency memory access operation, it is retrieved from the suspended warp table and scheduled back into the instruction pipeline for subsequent computation.

[0097] The thread bundle queue is used to maintain the thread bundle state information of the thread bundle being scheduled in the instruction pipeline, including an active thread bundle table and a suspended thread bundle table. The active thread bundle table stores the thread bundle information being executed in the current instruction pipeline, and the suspended thread bundle table stores the thread bundle information being suspended in the current instruction pipeline. When a thread bundle executes a long-latency memory access operation, the thread bundle scheduling unit moves it from the active thread bundle table to the suspended thread bundle table during its waiting for data response, and moves it from the suspended thread bundle table to the active thread bundle table after the data response arrives. When a thread bundle completes all operations and calculations, it is removed from the active thread bundle table.

[0098] The prefetch scheduling unit implements the creation of a data prefetch thread bundle and the cooperative control of the data prefetch unit in the L1 cache. When the thread bundle scheduling unit schedules the thread bundle of the current thread block into the instruction pipeline for operation, the prefetch scheduling unit creates a data prefetch thread bundle bound to the next thread block and writes the data prefetch thread bundle into the active thread bundle table for the thread bundle scheduling unit to schedule into the instruction pipeline for synchronous execution with the thread bundle of the current thread block. The role of the prefetch thread bundle is to calculate the data starting base address of the next thread block and calculate the base addresses of the thread bundles in the thread block based on the data starting base address. Then the prefetch scheduling unit sends the data starting base addresses of the thread bundles in the next thread block to the data prefetch unit of the L1 cache for data prefetching. Meanwhile, when the prefetch error counter of the L1 cache reaches a set threshold, the prefetch scheduling unit controls to empty the prefetch data of the L1 cache for the next prefetching.

[0099] The L1 cache is composed of four parts: a prefetch error counter, a data prefetch unit, a cache scheduling unit, and a data cache unit.

[0100] The prefetch error counter is used to count the number of mismatches between the prefetch data and the actual required data, so as to avoid the waste of prefetch space and time caused by the mismatch between the prefetch data and the actual data. Each time the actual memory data does not match the prefetch data, the counter is incremented, and when the counter reaches a set threshold (defined by the user), an error signal is sent to the prefetch scheduling unit, and the prefetch scheduling unit empties the prefetch cache, re-prefetches, and clears the prefetch error counter.

[0101] The data prefetch unit implements data prefetching. After the data prefetch unit receives the thread bundle base address sent by the prefetch scheduling unit, the data storage addresses of the threads in the thread bundle are calculated, and a memory access request is initiated to the L2 cache, and the returned response data is put into the data cache unit. When the prefetch scheduling unit initiates to empty the prefetch cache, the data prefetch unit controls to empty the prefetch data in the data cache unit.

[0102] The cache scheduling unit implements the control and scheduling of traditional cache mechanisms, including Miss Status Handling Register (MSHR), merged memory access, cache replacement and other functions.

[0103] The data cache unit is used to store prefetched and cached data. When the instruction pipeline initiates a memory access request, it first enters the data cache unit to check whether the data at the access address has been cached or prefetched. If so, the data is retrieved from the data cache unit and returned to the instruction pipeline. If not, the cache miss process begins, and the cache scheduling unit merges the memory access and waits for the L2 cache response.

[0104] Based on the above hardware architecture, the specific steps of the thread block data prefetching method provided by the present invention may include: 1) The host computer sends computing tasks and configuration information to the general-purpose parallel computing processor and starts the operation; 2) The general-purpose parallel computing processor divides the issued computing tasks into thread blocks and assigns the divided thread blocks to each streaming multiprocessor (SM); 3) After the SM receives the thread block, the warp scheduling unit of the pipeline management module splits the first thread block TB0 into multiple warps and writes them into the active warp table of the warp queue, scheduling them at the warp granularity; 4) The prefetch scheduling unit creates a prefetch warp bound to the next thread block TB1. This warp contains all parameters involved in the thread block base address calculation, such as the thread block index and thread block size of thread block TB1. 5) The prefetch scheduling unit writes the prefetched thread warp into the active thread warp table, and the thread warp scheduling unit schedules it into the instruction pipeline for execution in the next cycle; 6) After the prefetched thread warp enters the instruction pipeline, it calculates the starting base address TB_base_addr of thread block TB1 according to the thread block base address calculation method (which can be controlled by the user through instructions, and the calculation algorithm and calculation process are not limited); 7) In the next instruction cycle, the prefetch thread warp calculates the starting address of each thread warp in thread block TB1 based on the starting base address TB_base_addr of thread block TB1 and the warp address stride warp_stride (in the general parallel computing architecture, the warp is the smallest scheduling unit. The warp sizes after each thread block is split are equal, that is, the address stride of data access across warps is the same. After TB0 calculation is completed, subsequent thread blocks do not need to repeat the calculation. The warp address stride is equal to the number of threads in the warp multiplied by the thread address stride: warp_stride = thread_stride * warp_thread_num. The thread address stride thread_stride is bound to the calculation accuracy of the general parallel computing architecture and is a fixed value). TB_warp_base_addr = TB_base_addr + warp_id * warp_stride; 8) After the prefetch warp calculates the base address of each warp in thread block TB1, the prefetch scheduling unit removes the prefetch warp from the active warp table and releases the prefetch warp resources; 9) The prefetch scheduling unit sends the calculated thread warp base address to the data prefetch unit of the L1 cache and initiates a prefetch request; 10) The data prefetch unit of the L1 cache calculates the thread data address warp_thread_addr = TB_warp_base_addr + thread_id * thread_stride according to the received thread warp base address; 11) The data prefetch unit sends the thread data address to the cache scheduling unit. After receiving the thread data address, the cache scheduling unit combines it with the missing address in the MSHR and initiates a merged memory access to the L2 cache, waiting for the data response from the L2 cache. 12) The L2 cache processes the memory access request initiated by the L1 cache. If the cache hits, the response data is returned to the L1 cache. If the cache misses, it initiates a memory access to the DRAM and returns the response data to the L1 cache after receiving it. 13) After receiving the response data, the L1 cache writes the data into the data cache unit and sends a prefetch completion signal to the data prefetch module; 14) After receiving the prefetch completion signal, the data prefetch module calculates the thread address from the base address of the next thread warp and sends it to the cache scheduling unit, completing data prefetching in sequence according to steps 11) to 14); 15) When the instruction pipeline performs data memory access, the cache scheduling module of the L1 cache receives the memory access request and compares the memory access address with the address corresponding to the cached data in the data cache unit; 16) If the cache hits, the corresponding data is returned to the instruction pipeline. If the cache misses, the access address is stored in MSHR and waits for the merged access. At the same time, the value of the prefetch error counter is increased by one. 17) When the value of the prefetch error counter reaches the set threshold (user-defined), an error signal is sent to the prefetch scheduling unit of the pipeline management module; 18) After receiving the error signal, the prefetch scheduling unit clears the prefetch data cached in the data cache unit and clears the prefetch error counter, and then sends the next prefetch request and prefetch address to the data prefetch unit for the next data prefetch.

[0105] The thread block data prefetching method provided by the present invention realizes data prefetching across thread block boundaries by creating a data prefetching thread warp, cooperates with the thread warp scheduling strategy, uses the prefetching thread warp to realize the calculation of the data base address of the next thread block and the thread warp memory access address, rather than a fixed prefetching algorithm, which has wider application and greater flexibility, effectively improves the application scope of the prefetching technology in the general parallel computing architecture, realizes efficient and accurate data prefetching, and effectively controls the waste of prefetching resources and costs through a prefetching error counter, further improves the computing performance of the processor and improves computing efficiency through cross-boundary data prefetching.

[0106] Figure 7 Schematic diagram of the structure of the electronic device provided by the present invention, such as Figure 7 As shown, the electronic device may include: a processor (Processor) 710, a communication interface (Communications Interface) 720, a memory (Memory) 730 and a communication bus (Communications Bus) 740, wherein the processor 710, the communication interface 720, and the memory 730 communicate with each other via the communication bus 740. The processor 710 may call the logic commands in the memory 730 to execute the method described in the above embodiments, for example: Create a data prefetch thread bundle corresponding to the next thread block; the data prefetch thread bundle is used to calculate the data starting base address of each thread bundle in the next thread block; execute each thread bundle in the current thread block and execute the data prefetch thread bundle to determine the data starting base address of each thread bundle in the next thread block; based on the data starting base address of each thread bundle in the next thread block, obtain the operation data corresponding to the next thread block.

[0107] Furthermore, the logical commands in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several commands for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.

[0108] The processor in the electronic device provided by the embodiment of the present invention can call the logic instructions in the memory to implement the above method. Its specific implementation method is consistent with the implementation method of the above method and can achieve the same beneficial effects, which will not be repeated here.

[0109] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the method provided in the above embodiments is implemented.

[0110] Its specific implementation is consistent with the aforementioned method implementation and can achieve the same beneficial effects, so it will not be repeated here.

[0111] An embodiment of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, the method described above is implemented.

[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0113] Through the above description of the embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A thread block data prefetching method, characterized in that: include: Create a data prefetch warp corresponding to the next thread block; The data prefetch warp is used to calculate the data starting base address of each warp in the next thread block; Executing each thread warp in the current thread block and executing the data prefetch thread warp, and determining a data starting base address of each thread warp in the next thread block; Based on the data starting base addresses of each thread warp in the next thread block, operation data corresponding to the next thread block is obtained.

2. The thread block data prefetching method according to claim 1, wherein: The step of creating a data prefetch warp corresponding to the next thread block includes: Determine the current thread block; Splitting the current thread block to determine thread warps of the current thread block; Creating a data prefetch warp corresponding to the next thread block based on the thread block index information and the thread block size information of the next thread block; Each warp of the current thread block and the data prefetch warp are written into a warp queue; the warp queue is used to schedule execution of each warp.

3. The thread block data prefetching method according to claim 2, wherein: The executing each warp in the current thread block and executing the data prefetch warp, and determining the data starting base address of each warp in the next thread block, includes: The warps in the warp queue are scheduled for execution, so that the data prefetch warp determines the data starting base address of each warp in the next thread block based on the starting base address of the next thread block and the warp address stride.

4. The thread block data prefetching method according to claim 3, wherein: The warp queue includes an active warp table and a suspended warp table; the active warp table is used to store status information of executed warps; The suspended warp table is used to store state information of suspended warps; The scheduling and executing of each warp in the warp queue includes: Scheduling execution of the current warp in the active warp table; When the data response waiting time corresponding to the current warp is longer than a preset time, the current warp is transferred from the active warp table to the suspended warp table; When receiving data corresponding to the current warp, transferring the current warp from the suspended warp table to the active warp table; When the execution of the current warp is completed, the current warp is removed from the active warp table.

5. The thread block data prefetching method according to claim 1, wherein: The acquiring operation data corresponding to the next thread block based on the data starting base address of each thread warp in the next thread block includes: Determine a data starting base address of the current thread warp in the next thread block; Obtaining operation data corresponding to the current thread warp based on the data starting base address of the current thread warp; Cache the operation data corresponding to the current warp, and switch the current warp to the next warp.

6. The thread block data prefetching method according to any one of claims 1 to 5, characterized in that: The method further comprises: Obtaining a data memory access request of the next thread block; Comparing the first data address in the data access request with the second data address of the operation data; When the first data address is inconsistent with the second data address, it is determined that prefetching of the operation data corresponding to the next thread block is incorrect.

7. The thread block data prefetching method according to claim 6, wherein: The determining that the prefetching of the operation data corresponding to the next thread block is an error includes: In the case where the prefetching of the operation data corresponding to the next thread block fails, increasing the count value of the counter; the counter is used to record the number of prefetching errors of the operation data corresponding to each thread block; When the count value is greater than a preset threshold, the cache space of the operation data corresponding to each thread block is cleared.

8. A thread block data prefetching device, characterized in that: include: A creation module is used to create a data prefetch thread warp corresponding to the next thread block; The data prefetch warp is used to calculate the data starting base address of each warp in the next thread block; an execution module, configured to execute each thread warp in a current thread block and execute the data prefetch thread warp, and determine a data starting base address of each thread warp in the next thread block; The prefetch module is configured to obtain operation data corresponding to the next thread block based on the data starting base address of each thread warp in the next thread block.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the thread block data prefetching method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the thread block data prefetching method according to any one of claims 1 to 7 is implemented.

Citation Information

Cited By

  • Artificial intelligence chip and operation method thereof

    CN121116908A