CGRA hardware architecture for multi-task fusion execution and operation method
By introducing data storage modules, memory access monitor modules and tiles into the CGRA hardware architecture, using the global control signal Ctrl to generate a gated clock, and coordinating the preservation and restoration of checkpoint registers, the problem of task blocking caused by memory access misses in multi-task fusion execution is solved, and efficient multi-tasking performance is achieved.
Patent Information
- Application Number
- CN202510715575.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-19
AI Technical Summary
During the multi-task fusion execution process, memory access miss in any task will hinder the execution of other tasks, and lead to the propagation of computing errors and waste of energy.
A CGRA hardware architecture designed for multi-task converged execution is employed, comprising a data storage module, a memory access monitor module, and several tiles. A global control signal, Ctrl, generates a gated clock to coordinate the saving and restoring of checkpoint registers, enabling converged execution of tasks.
Even if a memory access miss occurs in any task, it will not hinder the execution of other tasks, avoiding the propagation of computational errors and waste of power, and improving multi-tasking performance.
Smart Images

Figure CN120670368A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a CGRA hardware architecture and an operating method for multi-task fusion execution, belonging to the technical field of computer multi-task processing. Background Art
[0002] Coarse-grained reconfigurable architecture (CGRA) is a reconfigurable computing platform that combines a programmable tile (or PE) array with a runtime adaptive interconnect, designed to natively support the concurrent execution of heterogeneous tasks. Its architectural advantage in multi-tasking scenarios comes from spatial isolation, that is, the ability to statically or dynamically partition the tile array into independent regions, each of which performs different tasks. In real-world applications such as autonomous systems or smart sensors, the computing patterns and latency requirements of tasks such as perception, decision-making, and communication vary. CGRA ensures performance by isolating physical resources to eliminate interference between tasks. This hardware-enforced isolation also reduces synchronization overhead, making it energy-efficient in multi-tasking environments with strict power constraints.
[0003] Current research on CGRA multi-task resource allocation is divided into static and dynamic methods. Static methods analyze tasks offline to estimate computing resource requirements, spatially partition the tile array according to task complexity, and generate a configuration through the CGRA mapper to isolate resources for concurrent execution. Although simple, this fixed partitioning often leads to resource fragmentation. Dynamic methods monitor input data changes, adjust spatial partitioning at runtime, and reallocate the tile array to rebalance the task pipeline. However, both paradigms focus mainly on fine-grained spatial optimization and ignore resource allocation in the temporal dimension. Specifically, they fail to exploit idle clock cycles caused by data dependencies or pipeline bubbles. In this case, although tiles are spatially allocated to specific tasks, they are not fully utilized in the temporal dimension.
[0004] Inspired by the latest advances in GPU kernel fusion technology, we found that fusing multiple task kernels to execute concurrently by sharing the entire CGRA tile array can significantly improve computing resource utilization in the time dimension. Unlike traditional spatial partitioning that assigns different tiles to a single task, this fused execution mode achieves higher multi-tasking performance by allowing all tasks to share the entire tile array in the time dimension. However, this approach brings new challenges to the CGRA hardware architecture: the sharing of the tile array by multiple tasks in the time dimension is driven by unified configuration information, which actually creates interdependencies for multiple tasks: if a task encounters a memory access miss, it will not be able to obtain the next configuration information, thereby forcing the execution of other tasks to stop. On the contrary, if the next configuration information is forced to be obtained to allow the remaining tasks to execute normally, the calculation of the memory access miss task will result in incorrect output and waste a lot of power. Summary of the Invention
[0005] In order to solve the problem that during the multi-task fusion execution, memory access loss of any task hinders the execution of other tasks and causes the propagation of calculation errors of the task and wastes power, the present invention proposes a CGRA hardware architecture and operation method for multi-task fusion execution.
[0006] The technical solution adopted by the present invention to solve the above problems is: the CGRA hardware architecture proposed by the present invention for multi-task fusion execution includes:
[0007] Data storage module, memory access monitor module and several tiles;
[0008] The data storage module is used to store memory access hit / miss signals;
[0009] The memory access monitor module is used to process the memory access hit / miss signal and generate a global control signal Ctrl;
[0010] The multiple tiles are used to generate gated clocks according to the global control signal Ctrl, and coordinate the save / restore of the checkpoint register TaskCkptReg, the activation of the clock gating gClk, and the coverage of the NOPs instruction in each clock cycle to achieve the fusion execution of CGRA multi-tasks.
[0011] Furthermore, the data storage module is a hierarchical data storage structure including a first-level scratchpad memory, a second-level data cache, and a third-level external memory;
[0012] The first-level scratch pad memory consists of n banks, each of which is used to store memory access hit / miss signals;
[0013] The second-level data cache is used to store some data required for multi-tasking execution;
[0014] The third level external memory is used to store all the data required for multi-tasking execution.
[0015] Furthermore, the memory access monitor module includes n access status registers, and each bank in the first-level scratch memory corresponds to a unique access status register;
[0016] The access status register includes multiple flag bits for receiving memory access hit / miss signals, wherein when the flag bit is 1, it indicates a memory access hit, and when the flag bit is 0, it indicates a memory access miss.
[0017] Furthermore, the memory access monitor module generates a global control signal Ctrl by performing a bitwise AND operation on all access status registers, where the bit width of the global control signal Ctrl is m, and each bit is 0 or 1, where m is the number of concurrent tasks currently being executed, 0 indicates that the corresponding task has a memory access loss, and 1 indicates that the corresponding task has a memory access hit.
[0018] Furthermore, several tiles are arranged in a rectangular shape, each tile includes a functional unit and m checkpoint registers added to the local register, the checkpoint registers are written using sequential logic and read using combinational logic, and store the operands of the previous cycle and the configuration memory read address;
[0019] A method for operating a CGRA hardware architecture for multi-task fusion execution, comprising:
[0020] Step 1: Enter the number of computing tasks and combine multiple computing tasks to obtain x multi-task scenarios;
[0021] Step 2: The first-level scratch memory sequentially calls data from the second-level data cache and the third-level external memory, obtains the memory access status of the specific task of the uniquely corresponding access status register through each bank in the first-level scratch memory, and performs a bitwise AND operation on all access status registers through the memory access monitor module to generate a global control signal Ctrl;
[0022] Step 3: Input the global control signal Ctrl into the tile, generate a gated clock by combining it with the main clock through AND gate logic, control the update of local registers and checkpoint registers through the gated clock, compare the current configuration memory read address with the checkpoint value of the checkpoint register, generate a retry signal, and use the retry signal to determine whether the functional unit receives the original configuration Config or the NOPs instruction overwrite, and select the number of tiles used for execution and multi-tasking to perform CGRA multi-tasking fusion execution.
[0023] Furthermore, during the fusion execution of CGRA multi-tasks in step 3, if one of the tasks loses memory access, in the current round, the execution operation in the corresponding clock cycle is replaced with NOPs empty operation, the clock of the local register continues to be shut down, and other tasks trigger the checkpoint save mechanism on the rising edge of the corresponding clock cycle, and perform memory access and execution operations normally. In the next round, the current configuration memory read address and the checkpoint value of the checkpoint register are compared to generate a retry signal, set the retry signal to 1, and retry the load operation. The load hits, and in subsequent clock cycles, all tasks have memory access hits, resume normal execution, and disable the gated clock.
[0024] The beneficial effects of the present invention are:
[0025] 1. The present invention utilizes physically isolated spatial partitioning to avoid interference between tasks and implements CGRA multi-task fusion execution. Even if any task encounters memory access loss, it will not hinder the execution of other tasks, thus avoiding the problem of computational error propagation and energy waste caused by the task.
[0026] 2. Even when faced with highly complex tasks, the present invention maintains a utilization rate between 60% and 75%, always showing the lowest throughput and resource utilization. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 A structural block diagram of the CGRA hardware architecture for multi-task fusion execution provided by the present invention;
[0028] Figure 2 A schematic diagram of a flow chart of an operating method of a CGRA hardware architecture for multi-task fusion execution provided by the present invention;
[0029] Figure 3 A schematic diagram comparing the throughput and utilization of the CGRA hardware architecture provided by the present invention and the prior art;
[0030] Figure 4 This is a schematic diagram of the dual-task scenario execution of the CGRA hardware architecture provided by the present invention. DETAILED DESCRIPTION
[0031] Specific implementation method 1: Combination Figure 1 This embodiment is described as follows. Figure 1 As shown, the structure of a CGRA hardware architecture for multi-task fusion execution described in this embodiment includes:
[0032] Data storage module, memory access monitor module and several tiles;
[0033] Data storage module The storage module is a hierarchical data storage structure, including the first level note memory (1 st levelscratchpad memory), second-level data cache (2 st level data cache) and third-level external memory (3 st level external memory), the first-level scratch pad memory contains multiple banks, each bank is used to store memory access hit / miss signals.
[0034] The memory access monitor module uses n access status registers to process memory hit / miss signals from the first-level scratch memory and generate a global control signal, Ctrl. Each bank in the memory access monitor corresponds to an access status register (AccessStatusReg), where each bit corresponds to the memory access status of a specific task. A bit set to 1 indicates a memory hit, while a bit set to 0 indicates a memory miss. Because a task can access multiple banks, the monitor performs a bitwise AND operation on all access status registers to generate the global control signal, Ctrl. The global control signal, Ctrl, has a bit width of m, with each bit set to 0 or 1, where m is the number of concurrent tasks currently executing. A value of 0 indicates a memory miss for the corresponding task, and a value of 1 indicates a memory hit for the corresponding task. For example, in a dual-task scenario (a two-bit Ctrl signal), "00" indicates that both tasks experienced a memory miss, while "11" indicates that both tasks experienced a memory hit. Notably, the bit width of the Ctrl signal varies with the number of concurrent tasks, ensuring accurate status tracking across varying workloads.
[0035] Several Tiles such as Figure 1 As shown, it is arranged in a rectangular shape. The CGRA structure itself includes routing xbar, funcxbar, and config memory. Config memory is used to read out an instruction in each clock cycle to control the functions of Funcunits, routing xbar, and func xbar. Routing xbar is used to transfer data from other tiles to func units for calculation, or to tiles in other directions (to N, NW, etc.). Registers in the corresponding direction are used to assist in the transfer operation (maintain timing). Func xbar is used to transfer the calculation results of func unit to tiles in other directions (to N, NW, etc.). Among them, routing refers to routing, func refers to functional unit, config memory refers to configuration memory, xbar refers to arbitrator, and tile refers to the basic unit module that can independently execute computing tasks.
[0036] This implementation adds a checkpoint register on this basis. Each local register (Reg) adds a task-specific checkpoint register. The number of checkpoint registers is m (TaskCkptReg[0] and TaskCkptReg[1] in the case of dual tasks) to preserve the calculation state. These registers are written using sequential logic and read using combinational logic to store the operand (OPRD) of the previous cycle and the configuration memory read address (CfgAddr). The gated clock (gClk) is generated by combining the main clock and the Ctrl signal through AND gate logic to control the update of local and checkpoint registers. The retry signal Retry is generated by comparing the current CfgAddr with the checkpoint value. It determines whether the functional unit Func Units receives the original configuration Config or is overwritten by the NOPs instruction. At the same time, the operand is selected from the ongoing calculation or the checkpoint register to realize the fusion execution of CGRA multi-tasks.
[0037] Specific implementation method 2: Figure 2 As shown, the operation method of the CGRA hardware architecture for multi-task fusion execution described in this embodiment includes:
[0038] S1: Input the number of computing tasks;
[0039] Enter the number of computing tasks and combine multiple computing tasks in a staggered manner to obtain x multi-task scenarios;
[0040] S2: Perform data call and generate memory access hit / miss signal through the first-level scratch memory;
[0041] The first-level scratch memory calls data from the second-level data cache and the third-level external memory in turn. If the data stored in the first-level scratch memory does not meet the current multi-tasking requirements, it calls data from the second-level data cache. If the data in the second-level data cache still does not meet the current multi-tasking requirements, it calls data from the third-level external memory.
[0042] S3: Obtain the memory access status of the specific task of the uniquely corresponding access status register through each bank in the first-level scratch memory, and perform a bitwise AND operation on all access status registers through the memory access monitor module to generate a global control signal Ctrl;
[0043] S4: Input the global control signal Ctrl to Tile, select the number of tiles used for execution and multi-tasking, and perform CGRA multi-task fusion execution;
[0044] The global control signal Ctrl is input into the tile, and the gated clock is generated by combining the master clock with the AND gate logic. The gated clock is used to control the update of the local register and the checkpoint register. The current configuration memory read address is compared with the checkpoint value of the checkpoint register to generate a retry signal. The retry signal determines whether the functional unit receives the original configuration Config or the NOPs instruction overwrite, and selects the number of tiles used for execution and multi-tasking. The data transmission direction of the tile task combination is as follows Figure 1 As shown by the gray connection lines in the Tlie matrix, the transmission direction can be E (east), S (south), W (west), N (north) or a combination of E, W and S, N. During the fusion execution of CGRA multi-tasks, if one of the tasks loses memory access, in the current round, the execution operation in the corresponding clock cycle is replaced by NOPs empty operation, the clock of the local register continues to be shut down, and other tasks trigger the checkpoint save mechanism on the rising edge of the corresponding clock cycle, and perform memory access and execution operations normally. In the next round, the current configuration memory read address is compared with the checkpoint value of the checkpoint register, a retry signal is generated, the retry signal is set to 1, and a retry load operation is performed. The load hits, and in subsequent clock cycles, all tasks have memory access hits, resume normal execution, disable the gated clock, and perform fusion execution of CGRA multi-tasks.
[0045] In summary, the present invention utilizes physically isolated spatial partitioning to avoid interference between tasks. Even if any task has a memory access loss, it will not hinder the execution of other tasks, thus avoiding the problem of propagation of computational errors and waste of electricity caused by the task. Specific implementation method three
[0047] Combine Figure 3 This embodiment is described. To verify the technical effect of the present invention, this embodiment is verified as follows:
[0048] The test tasks used in this implementation are shown in Table 1. A storage recorder is an electronic device capable of real-time acquisition and storage of electrical signals, physical quantities, and various other types of data. It supports multi-channel simultaneous recording for long-term monitoring and signal analysis, and is widely used in industrial testing and scientific research. The storage recorder is equipped with extensive waveform calculation capabilities, allowing users to flexibly add multiple calculation channels for real-time analysis of collected signals. Sixteen calculation tasks were selected from the user manual of a commercial MR-6000 storage recorder (see Table 1 for details). 48 multitask scenarios were constructed by staggered combinations of these 16 tasks. In a 4×4 CGRA, the combination of two adjacent tasks produces 16 multitask scenarios. In a 5×5 CGRA, the combination of three adjacent tasks produces 16 multitask scenarios. In a 6×6 CGRA, the combination of four adjacent tasks produces 16 multitask scenarios. The total number of nodes and edges after the combination is shown in Table 2.
[0049] Table 1
[0050]
[0051]
[0052]
[0053] In this implementation, the comparison targets include PartitionBase, FuseBase, and FexMo. PartitionBase is a traditional CGRA multitasking resource allocation method that uses spatial partitioning. FuseBase uses temporal fusion but does not utilize memory access monitors, tile-level checkpoints, or energy management. FexMo represents the present invention. The following will evaluate FexMo's multitasking throughput, tile utilization, and dynamic power savings.
[0054] like Figure 3As shown, the three comparison methods, PartitionBase, FuseBase, and FexMo, exhibit significant differences at 4×4, 5×5, and 6×6 CGRA sizes, but consistent behavioral patterns are observed across all scenarios. PartitionBase, as the base spatial partitioning strategy, consistently exhibits the lowest throughput and resource utilization, although its performance remains stable—utilization remains between 60% and 75% regardless of task complexity. This stability stems from its physically isolated spatial partitioning execution model, which avoids interference between tasks but completely sacrifices the potential for resource reuse. While FuseBase demonstrates theoretical parallelism advantages in low-contention scenarios, its performance fluctuates significantly. Its utilization plummets to 30%-50% in highly complex or data-dependent task combinations, significantly lower than PartitionBase. This "performance reversal" exposes a fundamental flaw in its design: conventional temporal fusion binds all tasks to a single configuration, and any local blocking caused by memory misses freezes global resources. Notably, this anomaly becomes even more pronounced in the 6×6 CGRA. For example, in a four-task scenario involving filters and FFTs, FuseBase's utilization is approximately 15% lower than PartitionBase's. This is because FuseBase cannot dynamically decouple tasks that generate memory misses. Congestion caused by tasks with high memory miss rates spreads rapidly, revealing a key paradox in conventional temporal fusion methods.
[0055] In contrast, FexMo demonstrates a clear advantage. Its checkpoint mechanism ensures that tasks can be continuously executed even during stalls caused by memory misses. For example, in a task combination involving filters and FFTs, FexMo maintains a stable utilization rate of 75%-85%, up to 20% higher than PartitionBase. Crucially, this advantage scales with the size of the CGRA - in a 6×6 four-task scenario, FexMo's throughput is an average of 1.94× higher than the baseline, while FuseBase's throughput is only 1.05×, as Tile is only 59.3%. This scalability validates FexMo's design philosophy: achieving temporal resource reuse through checkpoint saving and recovery, rather than freezing resources through global blocking.
[0056] Table 3
[0057]
[0058]
[0059] As shown in Table 3, experimental data for 4×4, 5×5, and 6×6 CGRA sizes reveals a consistent trend in dynamic power savings under FexMo's clock gating mechanism, with performance modulated by task characteristics and CGRA size. Across all scenarios, dynamic power savings ranged from 18.4% to 36.4%, demonstrating FexMo's adaptability to varying multi-task workloads. High-complexity tasks involving frequent memory operations, such as FFT and RMS, consistently achieved the highest savings (30-38%) due to extended tile idle cycles during memory misses, resulting in long clock gating times. In contrast, low-complexity tasks such as MAX and MIN exhibited modest savings (18-23%) due to overlapping execution that limited tile idle cycles. Hybrid tasks, such as filter paired with FFT, fell in the middle range (23-29%). Notably, larger CGRAs exhibited slightly reduced dynamic power savings, as the increased number of tiles and resource reuse reduced clock gating opportunities, but absolute energy efficiency improved due to higher throughput.
[0060] Example
[0061] like Figure 4 As shown, this embodiment demonstrates the dual-task execution scenario of the CGRA hardware architecture for multi-task fusion execution proposed by the present invention. Two tasks share one CGRA Tile in a time-fused manner, and the initiation interval is 4, as follows:
[0062] (1) During clock cycle 0, Task 0 encounters a memory miss. The checkpoint save mechanism is triggered on the rising edge of clock cycle 1, and OPRD and CfgAddr = 0 are saved to TaskCkptReg[0]. At the same time, the Ctrl signal is updated to "01", activating the clock gating of TaskCkptReg[0].
[0063] (2) Task 1 performs memory access normally in clock cycle 1
[0064] (3) The pre-planned addition operation “+” of task 0 in clock cycle 2 is replaced by NOPs, and the clock of the local register continues to be shut down to prevent error propagation and reduce dynamic power consumption.
[0065] (4) The multiplication operation “×” of Task 1 in clock cycle 3 is executed normally and is not hindered by the memory access miss of Task 0.
[0066] (5) When the initiation interval is 4, clock cycle 4 repeats the Config of clock cycle 0, causing CfgAddr to match the CfgAddr stored in TaskCkptReg[0], so the retry signal Retry is set to 1. FuncUnits uses the OPRD stored in TaskCkptReg[0] as the operand to retry the load operation, and the load hits.
[0067] (6) In the subsequent clock cycle (for example, clock cycle 5), Ctrl is updated to "11", indicating that both tasks have memory access hits, and normal execution is resumed, disabling the gated clock.
[0068] In summary, through this precise clock cycle-level control, the present invention achieves CGRA multi-task fusion execution and improves multi-task performance. During multi-task fusion execution, if any task encounters a memory miss, it will neither hinder the execution of other tasks nor cause computational errors in that task to propagate and waste energy.
[0069] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any technician familiar with the present profession can make some changes or modifications to equivalent embodiments of equivalent changes using the technical content disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modification, equivalent replacement and improvement of the above embodiments made according to the technical essence of the present invention, within the spirit and principles of the present invention, without departing from the content of the technical solution of the present invention, shall still fall within the scope of protection of the technical solution of the present invention.
Claims
1. A CGRA hardware architecture for multi-task fusion execution, characterized by: include: Data storage module, memory access monitor module and several tiles; The data storage module is used to store memory access hit / miss signals; The memory access monitor module is used to process the memory access hit / miss signal and generate a global control signal Ctrl; The multiple tiles are used to generate gated clocks according to the global control signal Ctrl, and coordinate the save / restore of the checkpoint register TaskCkptReg, the activation of the clock gating gClk, and the coverage of the NOPs instruction in each clock cycle to achieve the fusion execution of CGRA multi-tasks.
2. The CGRA hardware architecture for multi-task fusion execution according to claim 1, characterized in that: The data storage module is a hierarchical data storage structure, including a first-level note memory, a second-level data cache, and a third-level external memory; The first-level scratch pad memory includes n banks, each bank being used to store a memory access hit / miss signal; The second-level data cache is used to store part of the data required for multi-tasking execution; The third-level external memory is used to store all data required for multi-tasking execution.
3. The CGRA hardware architecture for multi-task fusion execution according to claim 1, characterized in that: The memory access monitor module includes n access status registers, each bank in the first-level scratch pad memory corresponds to a unique access status register; The access status register includes a plurality of flag bits for receiving a memory access hit / miss signal, wherein when the flag bit is 1, it indicates a memory access hit, and when the flag bit is 0, it indicates a memory access miss.
4. The CGRA hardware architecture for multi-task fusion execution according to claim 3, characterized in that: The memory access monitor module generates a global control signal Ctrl by performing a bitwise AND operation on all access status registers, wherein the bit width of the global control signal Ctrl is m, and each bit is 0 or 1, wherein m is the number of concurrent tasks currently being executed, 0 indicates that the corresponding task has a memory access loss, and 1 indicates that the corresponding task has a memory access hit.
5. The CGRA hardware architecture for multi-task fusion execution according to claim 1, characterized in that: Several tiles are arranged in a rectangular shape. Each tile includes a routing xbar, a func xbar, a config memory, a local register, and a functional unit. The local register includes m checkpoint registers, which are written using sequential logic and read using combinational logic, and store operands of the previous cycle and configuration memory read addresses; Config memory is used to read out an instruction in each clock cycle to control Func units, routing xbar, and func xbar; Routing xbar is used to transfer data from other tiles to func units for calculation, or to other tiles. Registers in the corresponding direction are used to assist in the transfer operation, and the timing remains unchanged during the transfer process. Func xbar is used to transfer the calculation results of func unit to other tiles; The functional unit performs the corresponding computing operations under the control of the original configuration Config, or does not perform any operations under the control of the NOPs instruction, and performs the fusion execution of CGRA multi-tasks.
6. A method for operating a CGRA hardware architecture for multi-task fusion execution, applied to the CGRA hardware architecture for multi-task fusion execution according to any one of claims 1 to 5, characterized in that: include: Step 1: Enter the number of computing tasks; Step 2: The data storage module generates a global control signal Ctrl by calling data; Step 3: Input the global control signal Ctrl into the tile, generate a gated clock by combining it with the main clock through AND gate logic, control the update of local registers and checkpoint registers through the gated clock, compare the current configuration memory read address with the checkpoint value of the checkpoint register, generate a retry signal, and use the retry signal to determine whether the functional unit receives the original configuration Config or the NOPs instruction overwrite, and select the number of tiles used for execution and multi-tasking to perform CGRA multi-tasking fusion execution.
7. The method for operating a CGRA hardware architecture for multi-task fusion execution according to claim 6, characterized in that: During the fusion execution of CGRA multi-tasks in step 3, all tasks share the same tile. If a memory access loss occurs in one of the tasks, the checkpoint save mechanism is triggered at the rising edge of the corresponding clock cycle to save the execution progress of the task. All subsequent operations performed on the tile are temporarily replaced with NOPs to avoid error propagation. At the same time, the clock of the local register is shut down to save power consumption. Other tasks using the tile can perform memory access and execution operations normally. In each clock cycle, the current configuration memory read address is compared with the checkpoint value of the checkpoint register. If they are the same, the retry signal is set to 1, so that the task that lost the memory access retry the memory access operation. If the memory access hits, all operations performed by the task on the tile are restored to normal, the gated clock is disabled, and the task continues to execute.