A Prefetch Optimization Method, Optimization System and Storage Medium for GPGPU

By calculating the step distance between thread bundles and thread blocks, the data prefetching method of GPGPU is optimized, and the problem of large memory access delay in the prior art is solved and the computing speed is improved.

CN119415445BActive Publication Date: 2025-06-24SHANDONG INSPUR SCI RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510031641.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-06-24
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

In the prior art, the data prefetching method of GPGPU is based on a simple prediction model and fails to fully consider the dynamic characteristics of thread execution and the spatial locality of data access, resulting in low cache hit rate and large memory access latency, which limits the computing speed of GPGPU.

Method used

By calculating the step distance between the thread bundle and the thread block, prefetching the future memory access address, and generating priority according to the state of the thread bundle, optimizing the sending of data prefetch requests.

Benefits of technology

It improves the memory access efficiency of GPGPU, reduces memory access latency, and improves the overall computing speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119415445B_ABST
    Figure CN119415445B_ABST
Patent Text Reader

Abstract

The present invention relates to a prefetch optimization method, optimization system and storage medium for a GPGPU, belonging to the field of GPGPU, and includes: S1. Calculate the step distance between warps based on the addresses of two consecutive warps in a thread block, and calculate the prefetch addresses of the remaining warps in the thread block; calculate the step distance between thread blocks based on the addresses of two consecutive thread blocks in the same task, and calculate the prefetch addresses of the remaining thread blocks in the task; S2. Assign stage attributes to warps according to the calculation status of the step distance between warps and the step distance between thread blocks; S3. Grab the prefetch addresses of warps based on the stage attributes of warps and send prefetch requests; S4. Set the status of warps based on the emission status of prefetch requests; S5. Generate the priorities of warps based on the status of warps. Through the provided step calculation unit, the prefetch queue or wake-up history queue can be quickly calculated, improving the performance and efficiency of the GPGPU.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of medical data processing, and particularly relates to a prefetch optimization method, an optimization system and a storage medium for GPGPU. Background Art

[0002] The existing data prefetch technology has significant limitations in the application of GPGPU: traditional data prefetch methods often rely on simple prediction models and fail to fully consider the dynamic characteristics of thread execution and the spatial locality of data access in GPGPU. This results in a low cache hit rate during execution, and memory access latency becomes a key factor restricting performance improvement. Especially when dealing with large-scale parallel tasks, resource competition and cache misses between thread blocks are particularly prominent. When the memory access latency is too high, the computing units will be idle while waiting for data to be transferred from memory. At the same time, when the data loading stage cannot provide data in a timely manner due to high latency, subsequent computing and storage stages will be affected, and the smoothness of the entire pipeline is disrupted, reducing the number of computing tasks completed per unit time. Summary of the Invention

[0003] In view of the problems in the prior art, the present invention provides a prefetch optimization method, an optimization system and a storage medium for GPGPU, which solve the problem that the existing data prefetch method often uses a simple prediction model, resulting in a large memory access latency and thus reducing the overall computing speed of GPGPU.

[0004] The technical solution adopted by the present invention is as follows:

[0005] In a first aspect, the present application provides a prefetch optimization method for GPGPU, the method comprising the following steps:

[0006] Step S1: Calculate the step distance between warps based on the addresses of two consecutive warps in a thread block, and calculate the prefetch addresses of the remaining warps in the thread block based on the step distance between warps;

[0007] Calculate the step distance between thread blocks based on the addresses of two consecutive thread blocks in the same task, and calculate the prefetch addresses of the remaining thread blocks in the task based on the step distance between thread blocks;

[0008] Store the prefetch addresses of the remaining warps and the prefetch addresses of the remaining thread blocks into a step list;

[0009] Step S2: Assign stage attributes to warps according to the calculation status of the step distance between warps and the step distance between thread blocks;

[0010] Step S3: Fetch the prefetch address of the warp in the step list based on the stage attribute of the warp, generate a corresponding prefetch request, and send the prefetch request to the data cache;

[0011] Step S4: Set the state of the warp based on the emission state of the prefetch request;

[0012] Step S5: Generate the priority of the warp based on the state of the warp.

[0013] Preferably, in Step S1, calculate the step distance between tasks based on the addresses of two consecutive tasks, and calculate the prefetch address of the subsequent tasks in this loop task based on the step distance between tasks.

[0014] Preferably, in Step S2, the prefetch stage attributes include an invalid stage, an initial stage, a semi-active stage, and an active stage, and the prefetch stage attributes are updated in real time according to the calculation status;

[0015] Among them, the invalid stage means that the warp is in an idle state and has not taken effect; the initial stage means that the warp has taken effect, but the available step distance has not been calculated; the semi-active stage means that only the step distance between warps has been calculated in this task, but the step distance between thread blocks has not been calculated; the active stage means that the step distance between warps and the step distance between thread blocks have been calculated in this task.

[0016] Preferably, Step S3 includes the following steps:

[0017] Step S3-1: Perform call sorting according to the prefetch stage attribute;

[0018] Step S3-2: Fetch the warp prefetch address queue in sequence based on the call sorting in Step S3-1, and generate a corresponding prefetch request to be sent to the data cache.

[0019] Preferably, in Step S5, the warp states include an idle state, a ready state, a completed state, and an ended state. The warp states in the step queue unit are recorded in real time, and the priority of the warp is adjusted in real time according to the warp state;

[0020] Among them, the warp state of the warp that has not sent a prefetch request is the idle state; the state of the warp when it is ready to send a request is the ready state; the state of the warp that has sent a prefetch request is the completed state; the state of the warp that has performed a memory access operation according to the prefetch request is the ended state.

[0021] Preferably, after the calculation task is completed, store the prefetch address of the warp in the step list fetched in Step S3 into the cache list.

[0022] In a second aspect, the present application provides an optimization system for a GPGPU, including:

[0023] A step calculation unit, configured to calculate and store the step distance and prefetch address queue required for step prefetching in each mode;

[0024] A step queue unit, configured to store the calculated prefetch address queue;

[0025] A prefetch stage register, configured to assign a stage attribute to a warp according to the calculation status of the inter-warp step distance and the inter-thread-block step distance: an invalid stage, an initial stage, a semi-active stage, and an active stage;

[0026] A warp status register, configured to generate the priority of a warp based on the status of the warp: an idle state, a ready state, a completed state, and an ended state;

[0027] A request emission unit, configured to grab the prefetch address queue of an entry in the step queue unit and emit a corresponding prefetch request to the data cache;

[0028] A pre-scheduling unit, configured to read the status and stage of an entry in the step queue unit and output the scheduling priority in real time to the scheduling module.

[0029] Preferably, a polling scheduling policy is reserved in the scheduling module.

[0030] Preferably, the optimization system further includes a history cache unit, which is configured to cache the executed address sequence and memory access pattern prefetch information for subsequent invocation of the cached prefetch information when performing the same task, and the cache queue applies a first-in-first-out replacement policy.

[0031] In a third aspect, the present application provides a storage medium storing program instructions, and when the program instructions are running, they execute the prefetch optimization method of the GPGPU as described in the first aspect.

[0032] It can be seen from the above technical solutions that the present invention has the following advantages: By setting the step calculation unit, it is possible to predict future memory access requirements using the step prefetch technology, and in combination with the historical wake-up mechanism, quickly wake up the prefetch address queue that has been executed, improving the performance and efficiency of the GPGPU. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0034] Figure 1 It is a flowchart of the method in the embodiment of the specific implementation manner of the present invention.

[0035] Figure 2 This is the overall architecture diagram of the method and the optimized system in the specific embodiment of the present invention.

[0036] Figure 3 This is the schematic diagram of the step distance in the specific embodiment of the present invention.

[0037] Figure 4 This is the schematic diagram of the prefetch address calculation in the specific embodiment of the present invention. Specific Embodiments

[0038] In the following detailed description, various embodiments of the present disclosure will be described more fully. The present disclosure can have various embodiments and adjustments and changes can be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents and / or alternative solutions falling within the spirit and scope of the various embodiments of the present disclosure.

[0039] In view of the problems in the prior art, the present invention provides a load / store unit of a GPGPU, a GPGPU prefetch optimization method and a storage medium, which solve the problem that the data prefetch method in the prior art often leads to a large memory access latency due to a simple prediction model, thereby reducing the overall computing speed of the GPGPU.

[0040] As Figure 1 shown in Figure 3 This embodiment provides a GPGPU prefetch optimization method, and the method includes the following steps:

[0041] Step S1: Calculate the step distance between warps based on the addresses of two consecutive warps in a thread block, and calculate the prefetch addresses of the remaining warps in the thread block based on the step distance between warps;

[0042] Calculate the step distance between thread blocks based on the addresses of two consecutive thread blocks in the same task, and calculate the prefetch addresses of the remaining thread blocks in the task based on the step distance between thread blocks;

[0043] Store the prefetch addresses of the remaining warps and the prefetch addresses of the remaining thread blocks into the step list;

[0044] Based on that the step distance between warps is obtained from the address difference between two adjacent warps in the same thread block, as shown in ① in Figure 3 ; based on that the step distance between thread blocks is obtained from the address difference between two adjacent thread blocks in the same task, as shown in ② in Figure 3 ; based on that the step distance between tasks is obtained from the address difference between adjacent tasks when executing a loop, as shown inFigure 3 As shown in ③, since not every task is a loop task and shows regularity, the step distance between tasks does not always exist. The calculated step distance is added to the address being executed to obtain the prefetch address, which is sent to the request emission unit, and it emits the request to the data cache to achieve data prefetching;

[0045] Step S2: Assign a stage attribute to the warp according to the calculation status of the step distance between warps and the step distance between thread blocks;

[0046] Step S3: Grab the warp prefetch address in the step list based on the stage attribute of the warp, generate the corresponding prefetch request, and send the prefetch request to the data cache;

[0047] Step S4: Set the status of the warp based on the emission status of the prefetch request;

[0048] Step S5: Generate the priority of the warp based on the status of the warp.

[0049] In this embodiment, in step S1, the step distance between tasks is calculated based on the addresses of two consecutive tasks, and the prefetch address of the subsequent task in this loop task is calculated based on the step distance between tasks.

[0050] In this embodiment, in step S2, the stage attributes of prefetching include an invalid stage, an initial stage, a semi-active stage, and an active stage, and the prefetch stage attributes are updated in real time according to the calculation status;

[0051] The invalid stage means that the entry is still idle and has not taken effect; the initial stage means that the entry has taken effect but the available step distance has not been calculated; the semi-active stage means that only the step distance between warps has been calculated for the current entry, but the step distance between thread blocks has not been calculated; the active stage means that the step distance between thread blocks has been calculated for the current entry. It should be noted that the step distance between tasks does not always exist, so the judgment of the prefetch stage does not consider the step distance between tasks.

[0052] In this embodiment, step S3 includes the following steps:

[0053] Step S3-1: Perform a call sorting according to the prefetch stage attribute;

[0054] Step S3-2: Grab the warp prefetch address queue in sequence based on the call sorting in step S3-1, and generate the corresponding prefetch request and send it to the data cache.

[0055] In this embodiment, in step S5, the warp status includes an idle state, a ready state, a completed state, and an ended state, and the warp status register records the warp status in the step queue unit in real time;

[0056] Adjust the priority of the warp in real time according to the warp state.

[0057] In this embodiment, after the computing task is completed, the prefetch address of the warp in the step list grabbed in step S3 is stored in the cache list.

[0058] As Figure 2 shown, this embodiment also provides an optimization system for GPGPU, including: a step calculation unit, a step queue unit, a request emission unit, a history cache unit, and a pre-scheduling unit;

[0059] The step calculation unit calculates and stores the step distance required for step prefetch in each mode, which determines the address difference during prefetch. The calculated address difference is added to the current address to obtain a prefetch address queue, and this queue is sent to the step queue unit.

[0060] The step queue unit stores the calculated prefetch address queue for the request emission unit to grab. At the same time, the warp state register and prefetch stage register maintained in this unit can record the prefetch state of the warp and the prefetch stage of the current entry;

[0061] Among them, the prefetch stage register is used to assign a stage attribute to the warp according to the calculation status of the step distance between warps and the step distance between thread blocks: invalid stage, initial stage, semi-active stage, and active stage;

[0062] The warp state register is used to generate the priority of the warp based on the state of the warp: idle state, ready state, completed state, and ended state.

[0063] The request emission unit grabs an entry from the step queue unit and emits a corresponding prefetch request to the data cache. Among them, the prefetch stage register of the step queue unit is read for judging the priority of entry grabbing.

[0064] The history cache unit caches the prefetch information such as the executed address sequence and memory access pattern. When the same task is executed, the cached prefetch information can be quickly called to quickly activate a new prefetch queue. In this embodiment, the history cache unit applies the first-in first-out replacement policy.

[0065] The pre-scheduling unit reads the state and stage in the step queue unit and outputs the scheduling priority in real time according to the read information, and sends it to the scheduling module to achieve the priority collaborative scheduling based on prefetch.

[0066] In this embodiment, a polling scheduling policy is retained in the scheduling module. When prefetching is not enabled or disabled, the scheduling module schedules warps using the basic polling scheduling policy.

[0067] After reading the warp status register in the step queue unit, the pre-scheduling unit adjusts the priority of each warp in real time according to the status of each warp: Since all requests have been issued and prefetching can be performed, the completed status has the highest priority; when a warp is ready to issue a request, it is in the ready state and has a higher priority; when a warp has not issued a request yet, it is in the idle state and has a lower priority; when a warp has performed a memory access operation and no longer needs prefetching, it is in the end state and has the lowest priority.

[0068] Among them, when a warp is in the idle state, prefetching is not activated and the warp does not issue a prefetch request; when a warp is in the ready state, the warp has prepared the required information and issues a prefetch request to activate the prefetch operation; when a warp successfully issues all prefetch requests, it is marked as the completed state; when a memory access instruction arrives in the optimization system, the memory access operation starts and prefetching is no longer performed, and the warp is marked as the end state.

[0069] For further explanation of the present invention, as Figure 4 shown, in this task, each thread block contains 4 warps, and the addresses are as shown.

[0070] Before prefetching starts, since the step distance has not been calculated yet, the stage in the prefetch stage register is in the idle state;

[0071] When the task starts to execute, at this time the step calculation unit still does not have the condition to calculate the step distance, so the prefetch stage changes to the initial stage and waits for the calculation result of the calculation unit;

[0072] When executing to warp 1, the step calculation unit can calculate the step distance between warps and send the result to the step queue unit. The prefetch stage changes to semi-activated, and prefetching of the warp addresses can start: The calculation unit takes the current address as a reference, adds the step distance to get the prefetch address queue of future warps, and writes it into the step queue unit. In the example Figure 3 shown, the step distance between warps = 0x1020 - 0x1000 = 0x0020.

[0073] When all the warps in thread block 0 have finished execution and the warp 0 in thread block 1 starts to execute, the step calculation unit can calculate the step distance between thread blocks and write the result into the step queue unit. After obtaining the step distance between thread blocks, the step queue unit will change the prefetch stage to the activation stage, even if the step distance between tasks has not been obtained yet. Meanwhile, the step calculation unit can calculate the prefetch address queue of future thread blocks to complete the prediction of the addresses of future thread blocks and the warps inside them. In the example Figure 3 shown, the step distance between thread blocks = 0x1080 - 0x1000 = 0x0080.

[0074] Until task 0 has finished execution and starts to loop to execute task 1, the step calculation unit obtains the address of warp 0 in thread block 0 in task 1, calculates the step distance between tasks, writes it into the step queue unit, and can calculate all the addresses of future tasks. In the example Figure 3 shown, the step distance between tasks = 0x3000 - 0x1000 = 0x2000.

[0075] After entering the semi - activation state from the prefetch stage, the request emission unit can grab the entries in the step queue unit and generate prefetch requests. When all requests have been emitted, the warp status register in the step queue unit will set the status of this warp to completed. Until the memory access instruction enters the optimization system and no longer prefetches, the warp status is marked as the end state.

[0076] According to the warp status, the pre - scheduling unit will generate the scheduling priority order of warps in real - time and send the priorities to the scheduling module, so that the scheduling module can perform cooperative scheduling on warps, giving priority to scheduling the warps that have completed prefetching to improve the memory access efficiency and further optimize the computing efficiency of the GPGPU.

[0077] When a task ends and switches to another task, the prefetch address queue of the executed task will be sent to the history cache unit. Until the next time this task is executed again, the history cache unit will resend this queue to the step queue unit to achieve the re - activation of the queue, directly perform prefetching based on the queue, and save the time of recalculation.

[0078] In this embodiment, a storage medium is also provided, storing program instructions which, when running, execute the prefetch optimization method of the GPGPU as described above.

[0079] It can be understood that the systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email transceiver device, a game console, a tablet computer, a wearable device, or a combination of any several of these devices.

[0080] In a typical configuration, a computer includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0081] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.

[0082] The computer-readable medium includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of the computer's storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, disk storage, quantum memory, graphene-based storage media, or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0083] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0084] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising said element.

[0085] The specific embodiments of this specification are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0086] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the" and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0087] It should be understood that although the terms first, second, third, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "upon" or "in response to determining".

[0088] The above are only the preferred embodiments of one or more embodiments of this specification and are not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.

Claims

1. A GPGPU prefetch optimization method, characterized in that: The method comprises the following steps: Step S1, calculating the step distance between the thread warps based on the addresses of two consecutive thread warps in the thread block, and calculating the pre-fetch addresses of the remaining thread warps in the thread block based on the step distance between the thread warps; Calculating a step distance between thread blocks based on addresses of two consecutive thread blocks in the same task, and calculating prefetch addresses of remaining thread blocks in the task based on the step distance between the thread blocks; storing the prefetch addresses of the remaining thread warps and the prefetch addresses of the remaining thread blocks in a step list; Step S2: assigning a phase attribute to the thread warp according to the calculation status of the step distance between thread warps and the step distance between thread blocks; The prefetch stage attributes include invalid stage, initial stage, semi-activated stage and activated stage, and the prefetch stage attributes are updated in real time according to the calculation status; The invalid phase means that the thread warp is in an idle state and has not taken effect; the initial phase means that the thread warp has taken effect, but the usable stepping distance has not been calculated; the semi-activated phase means that only the stepping distance between thread warps has been calculated in the task, but the stepping distance between thread blocks has not been calculated; the activated phase means that the stepping distance between thread warps and the stepping distance between thread blocks have been calculated in the task; Step S3: based on the phase attribute of the thread warp, the thread warp prefetch address in the step list is captured, a corresponding prefetch request is generated, and the prefetch request is sent to the data cache; Step S4, setting the state of the warp based on the emission state of the prefetch request; Step S5: generating the priority of the thread warp based on the state of the thread warp; The thread warp states include idle state, ready state, completed state, and ended state. The thread warp states in the step queue unit are recorded in real time and the thread warp priorities are adjusted in real time according to the thread warp states. The thread warp state of the thread warp that has not issued a prefetch request is in the idle state; the thread warp state when preparing to issue a request is in the ready state; the thread warp state that has issued a prefetch request is in the completed state; the thread warp state that has executed a memory access operation according to the prefetch request is in the ended state.

2. The GPGPU prefetching optimization method according to claim 1, characterized in that: In step S1 , a step distance between tasks is calculated based on the addresses of two consecutive tasks, and pre-fetch addresses of two subsequent identical consecutive tasks are calculated based on the step distance between tasks.

3. The GPGPU prefetching optimization method according to claim 2, characterized in that: Step S3 includes the following steps: Step S3-1, sorting calls according to pre-fetching stage attributes; Step S3-2: based on the call order of step S3-1, the thread warp prefetch address queue is captured in sequence, and a corresponding prefetch request is generated and sent to the data cache.

4. The GPGPU prefetching optimization method according to claim 1, characterized in that: When the computing task is completed, the thread warp prefetch addresses in the step list captured in step S3 are stored in the cache list.

5. A GPGPU optimization system, characterized in that: include: A step calculation unit, used for calculating and storing the step distance and prefetch address queue required for step prefetching in each mode; A step queue unit, used for storing the calculated pre-fetch address queue; A prefetch stage register, used to assign stage attributes to the warp according to the calculation status of the inter-warp step distance and the inter-thread block step distance: inactive stage, initial stage, semi-active stage and active stage; The invalid phase means that the thread warp is in an idle state and has not taken effect; the initial phase means that the thread warp has taken effect, but the usable stepping distance has not been calculated; the semi-activated phase means that only the stepping distance between thread warps has been calculated in the task, but the stepping distance between thread blocks has not been calculated; the activated phase means that the stepping distance between thread warps and the stepping distance between thread blocks have been calculated in the task; A warp state register for generating warp priorities based on the warp states: idle, ready, completed, and terminated; The thread warp state of the thread warp that has not issued a prefetch request is an idle state; the thread warp state when preparing to issue a request is a ready state; the thread warp state that has issued a prefetch request is a completed state; the thread warp state that has executed a memory access operation according to the prefetch request is an ended state; a request transmitting unit, for fetching a prefetch address queue of an entry in the step queue unit, and transmitting a corresponding prefetch request to the data cache; The pre-scheduling unit is used to read the status and stage of the entries in the step queue unit, and output the scheduling priority in real time to the scheduling module.

6. The GPGPU optimization system according to claim 5, characterized in that: The polling scheduling strategy is retained in the scheduling module.

7. The GPGPU optimization system according to claim 5, characterized in that: The optimization system also includes a history cache unit, which is used to cache the pre-fetch information of the executed address sequence and memory access mode, and to call the cached pre-fetch information when the same task is executed subsequently. The cache queue applies a first-in-first-out replacement strategy.

8. A storage medium, characterized in that: Program instructions are stored, and when the program instructions are run, the GPGPU prefetch optimization method according to any one of claims 1 to 4 is executed.

Citation Information

Patent Citations

  • Register resource dynamic scheduling system

    CN118377536A

  • Alignment of Cache Fetch Return Data Relative to a Thread

    US20090063818A1