Cache mode dynamic switching-based AI calculation acceleration method and system

By dynamically switching modes in the processor's last-level cache and utilizing cache blocks to perform AI calculations, the problem of large area overhead and low utilization of AI accelerators in edge computing devices is solved, achieving efficient and stable AI computing acceleration.

CN121349913APending Publication Date: 2026-01-16SHANGHAI JIAOTONG UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511419949.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-30
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

In existing edge computing devices, AI accelerators have large area overhead and low utilization, making it difficult to achieve dynamic switching and efficient scheduling of cache resources without compromising multi-level cache consistency, and lacking system-level support.

Method used

By implementing dynamic mode switching in the processor's last-level cache, AI computations are performed using cache blocks, including fast refresh, status marking, and isolation mechanisms. Combined with direct physical address lookup and scheduling, this enables efficient execution of AI computation tasks.

Benefits of technology

It significantly reduces chip costs, increases computing power per unit area, enables efficient and seamless switching between general computing and AI acceleration tasks, ensures system stability and multi-level cache consistency, and supports a variety of AI operators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121349913A_ABST
    Figure CN121349913A_ABST
Patent Text Reader

Abstract

The invention discloses an AI calculation acceleration method and system based on cache mode dynamic switching, and the method comprises the steps: determining a to-be-switched target cache block in a last-stage cache after receiving an AI calculation acceleration request; based on the physical address, the data is quickly refreshed to guarantee the consistency; marking the AI calculation mode as an AI calculation mode and shielding the AI calculation mode in an allocation strategy; executing in-memory calculation by utilizing a storage array and a vector operation unit which are integrated; and after the task is completed, clearing the mark and recovering the available state. The system comprises a cache switching controller, a matrix multiplication scheduler, a data preprocessing module, a cache isolation module and an instruction interface and software scheduling unit. By multiplexing the last level of cache and dynamically switching the working mode of the last level of cache, the AI calculation is efficiently completed while the system stability is ensured, and the computing power per unit area of a chip is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of integrated circuit and computer architecture technology, and in particular to an AI computing acceleration method and system based on dynamic switching of cache modes. Background Technology

[0002] As artificial intelligence applications are increasingly deployed in resource-constrained scenarios such as mobile devices and edge computing, these scenarios place high demands on the AI ​​inference capabilities of processors while simultaneously imposing strict limitations on chip area and power consumption. Matrix multiplication, as a core operation in deep learning models, is of paramount importance in terms of computational efficiency.

[0003] To accelerate matrix operations, existing technologies primarily employ two approaches: one is to use vector units on general-purpose processors (CPUs) or graphics processing units (GPUs), which have lower energy efficiency and area efficiency; the other is to introduce dedicated AI accelerators, such as hardware units based on systolic arrays. While dedicated accelerators, exemplified by Gemmini, can provide high performance, they introduce independent, large-capacity on-chip memory and computing arrays, resulting in significant additional chip area overhead. This greatly increases costs in edge devices, and the dedicated hardware remains idle when performing non-AI tasks, leading to low utilization.

[0004] In recent years, the academic community has proposed the concept of "in-cache computation," aiming to utilize the processor's inherent cache hierarchy for computation to avoid wasting area. For example, existing research has attempted to embed simple computational logic within cache arrays. However, integrating these studies into modern processors faces significant challenges: First, how to safely and quickly switch some cache resources from data caching mode to computation mode without compromising multi-level cache coherency; second, how to efficiently schedule the data flow of AI computations in this mode; and finally, how to design a complete set of hardware and software interfaces so that this function can be easily invoked by system software.

[0005] Therefore, there is an urgent need in this field for an AI computing acceleration solution that can be seamlessly integrated with existing processor architectures, enable dynamic reuse of cache resources, and have full system-level support. Summary of the Invention

[0006] In view of the above-mentioned shortcomings of the existing technology, the present invention provides an AI computing acceleration method and system based on dynamic switching of cache mode. The method and system reuse the last level cache of the processor and realize the dynamic switching of its working mode, so as to efficiently complete the AI ​​computing task while ensuring system stability and significantly improve the computing power per unit area of ​​the chip.

[0007] To achieve the above objectives, a first aspect of the present invention provides an AI computing acceleration method based on dynamic switching of cache modes, the AI ​​computing acceleration method based on dynamic switching of cache modes comprising:

[0008] S1: After receiving an acceleration request for an AI computing task, determine the target cache block to be switched to AI computing mode;

[0009] S2: Perform a fast refresh operation on the target cache block to write back the valid data stored therein to the main memory or the lower-level cache to ensure cache consistency;

[0010] S3: Mark the state of the target cache block as AI computing mode, and mask the target cache block in the cache allocation strategy;

[0011] S4: Utilize the storage array in the cache block in AI computing mode and its integrated vector operation unit to perform in-memory computation of AI operators;

[0012] S5: After the AI ​​computing task is completed, clear the AI ​​computing mode mark of the target cache block and restore its available state in the cache allocation strategy.

[0013] In some embodiments of the first aspect of this application, step S1, determining the target cache block to be switched to AI computing mode, includes:

[0014] Parse the acceleration request to obtain the resource requirement parameters of the AI ​​computing task;

[0015] Based on the resource requirement parameters, the required cache space size is estimated, and the cache space size is mapped to the specific cache organization structure of the last-level cache.

[0016] Based on the current cache occupancy status, select the specific physical location of the target cache block and generate its physical address list;

[0017] Verify the availability of the target cache block and finally confirm the target cache block.

[0018] In some embodiments of the first aspect of this application, the resource requirement parameters include computation operator type, data size, and computation precision;

[0019] The basic unit of cache allocation in the cache organization structure is a path or a storage block;

[0020] The target cache block is located in the processor's last-level cache.

[0021] In some embodiments of the first aspect of this application, step S2, performing a fast refresh operation on the target cache block, includes:

[0022] The scheduler initiates refresh requests directly based on the physical address;

[0023] The mode switching controller directly queries the cache directory to accurately locate the cache line storing valid data in the target cache block;

[0024] Perform targeted write-back and invalidation operations on the cached lines of the valid data;

[0025] After the refresh operation is completed, it is confirmed that all storage units in the target cache block are in an invalid state.

[0026] In some embodiments of the first aspect of this application, step S3, which involves marking the state of the target cache block as AI computation mode and masking it in the cache allocation strategy, includes:

[0027] Mark the target cache block as AI computing mode in a dedicated status register or status lookup table;

[0028] The allocation and replacement strategy of the cache controller is dynamically modified to exclude the target cache block from the cache allocation scope of ordinary programs;

[0029] The isolation state shall be maintained until a release command is received.

[0030] In some embodiments of the first aspect of this application, step S4, which involves performing in-memory computation of AI operators using the storage array in the cache block in AI computation mode and its integrated vector operation unit, includes:

[0031] Load the data to be computed into the target cache block and perform data preprocessing according to the in-memory computing requirements;

[0032] Based on the computational task requirements, selectively enable the vector operation units integrated in the target cache block;

[0033] The control vector operation unit processes data in the storage array in parallel and temporarily stores the operation results in a designated storage area of ​​the same cache block;

[0034] For computation tasks exceeding the cache capacity, a block or shard strategy is used to repeatedly execute the data loading, computation, and temporary storage steps until the task is completed.

[0035] In some embodiments of the first aspect of this application, the data preprocessing includes data expansion, transpose, and bit-width conversion operations;

[0036] The vector operation unit includes a multiplier and an adder, used to perform multiplication and addition operations.

[0037] In some embodiments of the first aspect of this application, step S5, which involves clearing the AI ​​computation mode tag of the target cache block and restoring its available state in the cache allocation strategy, includes:

[0038] Upon receiving the AI ​​computing task completion signal, send a release command containing the target cache block identifier to the cache switching controller;

[0039] Clear the AI ​​computation mode flag of the target cache block in the dedicated status register or status lookup table;

[0040] Restore the availability of the target cache block in the cache allocation strategy, so that it can be included in the cache allocation scope of normal programs;

[0041] A release completion confirmation signal is returned to the scheduler, and the target cache block resumes normal cache function.

[0042] In some embodiments of the first aspect of this application, the AI ​​computing acceleration method based on dynamic switching of cache mode initiates an acceleration request for AI computing tasks through an extended instruction set, which is compatible with the RISC-V open-source architecture.

[0043] To achieve the above objectives, a second aspect of the present invention provides an AI computing acceleration system based on dynamic switching of cache modes, the AI ​​computing acceleration system based on dynamic switching of cache modes comprising:

[0044] The cache switching controller is used to directly query the cache directory based on the physical address and perform a fast refresh operation on the target cache block to maintain cache consistency, while managing the status flags of the target cache block;

[0045] The matrix multiplication scheduler is used to receive acceleration requests for AI computing tasks, parse resource requirements, determine the physical address of the target cache block, and schedule the in-memory computing process.

[0046] The data preprocessing module is used to adapt AI operators, including convolution unrolling, data transposition and bit width conversion, and achieves efficient data transfer through the direct memory access module;

[0047] The cache isolation module is used to isolate and restore target cache blocks by using status registers and dynamically modifying cache allocation strategies.

[0048] The instruction interface and software scheduling unit provide hardware instruction interfaces and system-level scheduling functions to coordinate AI computations with routine tasks.

[0049] The advantages of this invention are as follows: First, by reusing the processor's inherent last-level cache, only about 10% of the additional area is needed for embedding computing logic and control units. Compared to introducing a separate accelerator, this significantly reduces chip manufacturing costs and increases computing power per unit area, making it particularly suitable for cost-sensitive edge computing scenarios. Second, the innovative direct refresh and query mechanism based on physical addresses minimizes the overhead of cache mode switching, enabling the system to quickly respond to intermittent AI computing tasks and achieving efficient and seamless switching between general computing and AI acceleration tasks. Third, through hardware-level state marking and allocation strategy modification, strong isolation is achieved between the AI ​​computing area and the ordinary cache area, fundamentally avoiding data access conflicts, strictly maintaining the consistency of multi-level caches, and ensuring the stable operation of the entire system. Fourth, configurable data mapping and scheduling strategies can support a variety of AI operators. This architecture is easily scalable to multi-core processor environments, providing a core foundation for building efficient heterogeneous computing systems. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 This is a flowchart illustrating an AI computing acceleration method based on dynamic switching of cache modes as described in this invention.

[0052] Figure 2 This is a schematic diagram of the architecture of the on-chip computing system under cache and a schematic diagram of the structure of the last-level cache of the present invention;

[0053] Figure 3 This is a schematic diagram of the memory structure topology for near-memory computation within the cache of the present invention;

[0054] Figure 4 This is a schematic diagram of the structure and control flow of the cache mode switching control module of the present invention;

[0055] Figure 5 This is a schematic diagram illustrating the cache allocation and release process for AI computation according to the present invention.

[0056] Figure 6 This is a schematic diagram illustrating the configurable data flow state when the computing module of the present invention is turned on or off.

[0057] Figure 7 This is a flowchart illustrating the process of implementing single matrix block or slice operation using a configurable architecture according to the present invention.

[0058] Figure 8 This is a schematic diagram of the structure of an AI computing acceleration system based on dynamic switching of cache mode as described in this invention. Detailed Implementation

[0059] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0060] Figure 1 This diagram illustrates a flowchart of an AI computing acceleration method based on dynamic switching of caching modes according to the present invention. Figure 1 As shown, the method includes the following steps:

[0061] Step S1: After receiving the acceleration request for the AI ​​computing task, determine the target cache block to be switched to AI computing mode.

[0062] like Figure 2 and Figure 3 As shown, after receiving an acceleration request for an AI computing task, determining the target cache block to be switched to AI computing mode is the foundation for subsequent mode switching and accelerated computing. This determination process specifically includes the following sub-steps:

[0063] Step S1.1: Parse the acceleration request and assess resource requirements.

[0064] First, the processor core initiates an acceleration request by calling extended instructions specifically designed for AI acceleration (such as the open-source RISC-V extended instruction set). The matrix multiplication scheduler receives and parses the acceleration instructions from the processor core, extracting key parameters, including but not limited to the type of computational operator to be executed (such as matrix multiplication, convolution, etc.), data size (such as the size of the matrix), and computational precision requirements. Based on these parameters, the scheduler estimates the cache space required to complete the AI ​​task.

[0065] For example, for a large matrix multiplication operation, it is necessary to take into account the storage requirements of the left matrix, the right matrix, and the result matrix, and combine the block computing strategy to determine the required cache capacity.

[0066] Step S1.2: Map resource requirements to specific cache organization structures.

[0067] Based on the calculated cache space requirements, the scheduler, referring to the specific runtime parameters of the processor's last-level cache (such as total capacity, set associativity, number of memory banks, etc.), transforms the abstract capacity requirements into specific, operable cache resource units.

[0068] In this embodiment of the invention, the most basic allocation unit in a cache allocation process is the path. Therefore, the output of this step is to specify how many paths of cache need to be allocated, or in which memory banks a specific number of computing units should be allocated.

[0069] For example, in a 1024KB total capacity, 4-way set-associative cache, if an estimated 256KB of dedicated space is needed, it may be necessary to specify one of the caches as the target.

[0070] Step S1.3: Select the specific physical location of the target cache block and generate its address.

[0071] Based on the resource allocation scheme determined in the previous step (such as the specified number of paths), and combined with the current occupancy status of the cache, the scheduler accurately selects the target cache block to be switched.

[0072] In the process of selecting a target cache block, in order to maximize performance and reduce the impact on normal programs, paths that are currently idle or have low occupancy rates can be prioritized.

[0073] Once selected, the scheduler directly calculates the physical addresses corresponding to these target cache blocks, rather than virtual addresses. This design avoids the huge lookup overhead and latency caused by the virtual-to-physical address translation in traditional cache refresh processes, laying the foundation for subsequent fast refreshes.

[0074] Step S1.4: Verify the availability of the target cache block and finally confirm it.

[0075] First, the scheduler submits a list of the physical addresses of the selected target cache blocks to the cache switching controller. The cache switching controller then queries the cache directory to verify the current status of these cache blocks, such as whether they are in an "invalid" state or whether they are "valid" but can be flushed.

[0076] The scheduler then operates based on this state: if all target blocks are available or can be released immediately through a fast refresh mechanism, the selection is confirmed as valid and the target cache blocks are officially determined; if there are conflicts, fine-tuning and re-verification may be required.

[0077] Finally, the scheduler outputs the determined physical address range of the target cache block, and the information of the target cache block is recorded in a dedicated status register or lookup table for use in subsequent cache mode switching operations.

[0078] Step S2: Perform a fast refresh operation on the target cache block to ensure cache consistency.

[0079] like Figure 4As shown, a fast refresh operation is performed on the target cache block, writing back the valid data stored therein to main memory or a lower-level cache to ensure cache consistency. In this invention, the fast refresh mechanism overcomes the problems of high latency and low efficiency of traditional refresh strategies. Its specific implementation includes the following sub-steps:

[0080] Step S2.1: The scheduler directly initiates a refresh request based on the physical address.

[0081] Unlike traditional methods where the processor core issues a virtual address refresh request, in this step, the matrix multiplication scheduler, which has already determined the physical address of the target cache block, directly issues a refresh instruction to the cache switching controller, along with the physical address range of the target cache block.

[0082] In this invention, this design skips the process of converting virtual addresses to physical addresses, avoiding the huge and inefficient query space caused by a physical cache block potentially corresponding to multiple virtual addresses, and greatly reducing the scope of operation.

[0083] Step S2.2: The mode switching controller directly queries the cache directory to accurately locate valid data.

[0084] After receiving the list of physical addresses, the mode switching controller does not go through the traditional cache controller path, but directly connects to and queries the cache directory. The cache directory is a hardware module that records the status of all cache lines (such as valid, invalid, dirty).

[0085] By querying, the controller can quickly and accurately identify which cache lines corresponding to the target physical address currently store valid data (especially "dirty" data, i.e. data that has been modified but not written back to main memory).

[0086] Step S2.3: Perform targeted write-back and invalidation operations only on valid rows.

[0087] Based on the precise results returned by the cache catalog, the mode switching controller initiates targeted operations only on those cache lines marked as valid:

[0088] First, force the valid data in that row to be written back to main memory, or write it to the lower-level cache according to the cache consistency protocol;

[0089] Subsequently, the cache line is immediately marked as invalid. This selective refresh mechanism avoids the traditional practice of indiscriminately refreshing the entire cache block, greatly reducing unnecessary data movement and thus significantly reducing the time overhead of refresh operations and the system bus bandwidth usage.

[0090] Step S2.4: Confirm that the refresh is complete and the status is synchronized.

[0091] Once all target cache lines marked as valid have completed the write-back and invalidation operations, the cache switching controller will receive a completion confirmation signal.

[0092] At this point, all storage units in the target cache block are in a known and consistent invalid state, which removes the obstacle to safely and losslessly switching it to AI computing mode. This direct lookup and selective refresh process based on physical addresses is the core guarantee for achieving efficient and low-latency dynamic switching of cache modes.

[0093] Step S3: Mark the target cache block's state as AI computing mode and mask the target cache block in the cache allocation strategy.

[0094] After completing the fast refresh operation on the target cache block, the system needs to officially activate it as an AI computing resource and ensure its isolation from the processor's ordinary tasks. This process is achieved through the following sub-steps:

[0095] Step S3.1: Mark the target cache block as AI computing mode in the status register.

[0096] The system sets up a dedicated status register or status lookup table for the last-level cache. In this step, the cache switching controller marks the physical location information of the refreshed target cache block in the status register and sets its status to AI computing mode or isolated.

[0097] In this embodiment of the invention, the register serves as the decision-making basis for the cache controller, providing timely and accurate status information for its subsequent cache allocation behavior.

[0098] Step S3.2: Dynamically modify the cache allocation strategy to block the target cache block.

[0099] To achieve physical isolation, the system needs to modify the original allocation and replacement strategy of the cache controller (such as LRU, random replacement, etc.). Specifically, the cache isolation module will dynamically adjust the selection range of the allocation strategy based on the information in the status register.

[0100] For example, in an 8-way set-associative cache, if the 2nd and 3rd ways are marked as AI computing mode, the isolation module will limit the effective selection range of the allocation algorithm to the 0th, 1st, 4th, 5th, 6th, and 7th ways by updating the policy-related configuration registers or lookup tables.

[0101] In this way, when a normal program requests the allocation of a new cache line, the cache controller will automatically skip the path in AI computing mode and only select from the remaining available paths, thereby ensuring in hardware logic that isolated cache blocks will not be accidentally overwritten by normal data.

[0102] Step S3.3: Achieve continuous isolation until the mode is lifted.

[0103] After completing the above marking and configuration, the target cache block enters the protected AI computing mode. In subsequent system operation, the cache controller will continue to follow the modified allocation strategy, making the shielded cache block invisible to ordinary programs and dedicated to AI computing tasks.

[0104] This mechanism effectively prevents resource competition and data corruption between ordinary processes and AI-accelerated tasks, ensuring the stability and security of system operation.

[0105] The isolation state will only be lifted when the AI ​​computation task is completed and the scheduler explicitly issues a release instruction. The status flag of the target cache block will be cleared, and the allocation policy will be restored to its original range, making it available to ordinary programs again.

[0106] Step S4: Utilize the storage array in the cache block in AI computing mode and its integrated vector operation unit to perform in-memory computation of AI operators.

[0107] After the target cache block is successfully isolated and marked for AI computing mode, the system can utilize its hardware resources to perform efficient in-memory computations. This process, led by the matrix multiplication scheduler, coordinates the work of the memory array and integrated logic units, and specifically includes the following sub-steps:

[0108] Step S4.1: Load and lay out the data.

[0109] The scheduler uses a custom direct memory access module to move the matrix or vector data to be computed from main memory or other cache levels to the target cache block already in AI computing mode. The data layout is optimized to match the data flow characteristics of in-memory computation.

[0110] For example, specific blocks of the left and right matrices in matrix multiplication can be placed in different memory locations or different computation units.

[0111] like Figure 7 As shown, the left matrix is ​​placed in a buffer area that serves as a global input buffer, while the right matrix is ​​pre-deployed in a storage array containing activation operation units. The data preprocessing module performs data preprocessing operations such as expansion, transposition, or bit-width conversion on the data at this stage to adapt to the requirements of specific AI operators.

[0112] Step S4.2: Configure the computing unit and start the operation.

[0113] The scheduler selectively enables the vector operation units (including multipliers and adders) integrated beneath the target cache block storage array via control signals, based on the computation task.

[0114] like Figure 6 As shown, this configuration is highly flexible: for different subarrays within the same computing unit, certain subarrays can be designated for performing multiply-accumulate operations, while adjacent subarrays can be temporarily used to cache intermediate results. This structure allows data to flow over extremely short distances between storage and computing units, achieving true in-memory computation and significantly reducing the energy and latency overhead of data movement.

[0115] Step S4.3: Perform parallel vector operations and temporarily store the results.

[0116] Under the control of the scheduler, the computation process officially begins. The activated vector operation units perform operations on multiple data from the memory array in parallel.

[0117] For example, when performing block matrix multiplication, the arithmetic unit simultaneously reads a row and a column (or a set of data) from the arrays storing the left and right matrix blocks, performing multiplication and accumulation operations. The resulting partial sum or final result is written directly back to a specified memory array within the same cache block that is not currently being used for computation, according to the scheduling policy, rather than being written back to processor registers or further memory. This method of performing computations and temporarily storing results within the cache is key to achieving high energy efficiency.

[0118] Step S4.4: Coordinate the data flow to complete the entire computation.

[0119] For computation tasks larger than the cache capacity, the scheduler will execute the above data loading, computation, and temporary storage steps in a loop according to the block or shard strategy until the entire operator computation is completed.

[0120] The entire process is precisely controlled by a hardware scheduler, ensuring that computing units and storage bandwidth are fully utilized, thereby efficiently accelerating AI operators such as matrix multiplication or vector operations.

[0121] Step S5: After the AI ​​computation task is completed, clear the AI ​​computation mode flag of the target cache block and restore its available state in the cache allocation strategy.

[0122] like Figure 5 As shown, after the AI ​​computing task is successfully completed, the system needs to safely and efficiently release the occupied cache resources so that they can return to their normal cache function. This process is achieved through the following sub-steps:

[0123] Step S5.1: Receive the release command and start the mode recovery process.

[0124] Once the matrix multiplication scheduler confirms that the AI ​​computation task has been completed, it sends a release command to the cache switching controller.

[0125] The instruction contains identification information of the target cache block that needs to be freed, such as its physical address range or specific road number.

[0126] Step S5.2: Clear the status flags of the AI ​​computing mode.

[0127] Upon receiving a release command, the cache switching controller accesses a dedicated status register or status lookup table to find the entry corresponding to the target cache block and changes its status from AI computing mode or isolated mode to available mode.

[0128] Step S5.3: Restore the availability of the target cache block in the cache allocation strategy.

[0129] The controller then notifies the cache isolation module, which will dynamically modify the cache controller's allocation and replacement strategy and cancel the previously set blocking.

[0130] Specifically, for example in a set-associative cache, the cache isolation module updates the threshold in the configuration register or lookup table that determines the allocation range, bringing previously excluded specific paths back into the valid selection range of cache allocation strategies (such as random replacement or LRU strategy).

[0131] At this point, when the cache controller allocates new cache lines for regular program data, it will reconsider these cache blocks that have just been released.

[0132] Step S5.4: Complete the release and confirm that the resources are ready.

[0133] After the above state clearing and policy update operations are completed, the cache switching controller will return an acknowledgment signal to the scheduler. The entire release process does not involve any data refresh or migration operations, because the intermediate or final results generated by AI calculations are already considered temporary data and do not need to be written back to main memory.

[0134] Therefore, resource release is almost instantaneous, without introducing additional performance overhead. The target cache block then fully recovers its normal function and can be immediately used for caching subsequent ordinary program data, achieving seamless and dynamic sharing of cache resources between general computing and AI acceleration tasks.

[0135] Figure 8 A schematic diagram of the structure of an AI computing acceleration system based on dynamic switching of caching modes, according to the present invention, is shown. Figure 8 As shown, the system includes a cache switching controller 101, a matrix multiplication scheduler 102, a data preprocessing module 103, a cache isolation module 104, and an instruction interface and software scheduling unit 105.

[0136] The cache mode switching controller 101 is the core hardware module that ensures cache consistency and enables dynamic mode switching. It is directly connected to the cache directory and the refresh unit. Its core function is to maintain data consistency in the cache system during mode switching. By receiving the physical address of the target cache block transmitted by the matrix multiplication scheduler, it directly queries the cache directory to obtain the status (valid or invalid) of the cache block corresponding to that address. For valid data, it triggers a precise refresh operation to write the data back to main memory or the lower-level cache.

[0137] Meanwhile, it marks the AI ​​operation mode of the cache block by controlling the status register and adjusts the threshold of the cache allocation strategy by using a lookup table to achieve the masking and restoration of the target cache block, ensuring that mode switching is both safe and efficient and does not compromise system cache consistency.

[0138] The matrix multiplication scheduler 102 is a computation scheduling core located in the last level cache of the processor, responsible for connecting the processor with the computational resources in the cache. It receives AI acceleration requests from the processor, parses key parameters such as the computation scale and data precision, and, in conjunction with the number of physical cache paths specified by the user, directly generates the physical address sequence of the target cache block (without virtual-physical address mapping conversion) and passes it to the cache mode switching controller.

[0139] After the mode switch is completed, it formulates a scheduling strategy based on the cache hardware structure (such as Bank and Mat layout), schedules the cache blocks in AI computing mode and the integrated vector operation units to perform operations such as matrix multiplication, and at the same time manages the storage and accumulation of intermediate results, coordinates the working rhythm of the data preprocessing module and the operation unit, and maximizes the utilization of computing resources in the cache.

[0140] Among them, the data preprocessing module 103 is a key supporting module for achieving compatibility with multiple types of AI operators, mainly responsible for operator adaptation and efficient data transfer. It has the ability to be compatible with various AI model operators, and can convert convolution operators into matrix multiplication form through convolution unrolling. At the same time, it supports preprocessing operations such as data transposition and bit serialization, ensuring that different types of AI operators can adapt to the data flow requirements of in-cache computation.

[0141] In addition, this module integrates a customized direct memory access module, which can bypass the conventional data path and realize high-speed data transfer between the main memory and the cache computing area. It can accurately move the preprocessed input data and weight data to the corresponding cache storage area, providing a data foundation for subsequent in-memory computing.

[0142] Among them, the cache isolation module 104 is a crucial functional module that ensures that AI operations and regular tasks do not interfere with each other. It achieves the isolation of cache resources through hardware mechanisms and policy adjustments. It relies on a specially set status register to record the operating mode (regular operation, AI operation) of each cache block, providing a basis for the cache controller's allocation decisions.

[0143] In AI computing mode, this module modifies the cache allocation and eviction policy to exclude cache blocks in AI mode from the resource allocation scope of regular tasks, ensuring that the cache controller does not allocate these cache blocks to ordinary program data. Simultaneously, by coordinating with the cache mode switching controller, it restores the normal allocation permissions of cache blocks after computation is complete, achieving secure isolation and efficient reuse of cache resources.

[0144] The instruction interface and software scheduling unit 105 is a key unit for realizing full-stack hardware and software collaboration, and includes two parts: hardware interface and software scheduling logic.

[0145] The instruction interface is compatible with the open-source RISC-V instruction set, which establishes an instruction interaction channel between the processor and the cache computing module, supporting the transmission and status feedback of instructions such as accelerated requests, mode switching, and resource release.

[0146] The software scheduling unit provides system-level scheduling functions, encapsulating the details of underlying hardware operations. It is responsible for triggering cache mode switching, allocating computing tasks, and coordinating the resource usage timing between AI operations and regular processes. This unit, through a full-stack implementation, can effectively verify the feasibility of in-cache computation and the actual system performance, thereby reducing the compatibility risks between module design and actual system deployment.

[0147] The advantages of this invention are as follows: First, by reusing the processor's inherent last-level cache, only about 10% of the additional area is needed for embedding computing logic and control units. Compared to introducing a separate accelerator, this significantly reduces chip manufacturing costs and increases computing power per unit area, making it particularly suitable for cost-sensitive edge computing scenarios. Second, the innovative direct refresh and query mechanism based on physical addresses minimizes the overhead of cache mode switching, enabling the system to quickly respond to intermittent AI computing tasks and achieving efficient and seamless switching between general computing and AI acceleration tasks. Third, through hardware-level state marking and allocation strategy modification, strong isolation is achieved between the AI ​​computing area and the ordinary cache area, fundamentally avoiding data access conflicts, strictly maintaining the consistency of multi-level caches, and ensuring the stable operation of the entire system. Fourth, configurable data mapping and scheduling strategies can support a variety of AI operators. This architecture is easily scalable to multi-core processor environments, providing a core foundation for building efficient heterogeneous computing systems.

[0148] In summary, this invention provides a full-stack design from hardware architecture and control logic to software interface, effectively reducing the risks from module design to system integration, accelerating the practical application and deployment of the technology, and has broad industrial application value.

[0149] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An AI computing acceleration method based on dynamic switching of cache modes, characterized in that, The AI computing acceleration method based on dynamic switching of cache modes comprises: S1: after receiving an acceleration request of an AI computing task, determining a target cache block to be switched to an AI computing mode; S2: performing a quick flush operation on the target cache block to write back valid data stored therein to a main memory or a lower-level cache to ensure cache consistency; S3: marking the state of the target cache block as the AI computing mode and shielding the target cache block in a cache allocation strategy; S4: using a storage array in the cache block in the AI computing mode and a vector operation unit integrated therein to perform in-memory computing of an AI operator; S5: after completion of the AI computing task, clearing the AI computing mode mark of the target cache block and restoring its available state in the cache allocation strategy.

2. The AI computing acceleration method based on dynamic switching of cache modes according to claim 1, characterized in that, In step S1, the determination of the target cache block to be switched to the AI computing mode comprises: parsing the acceleration request to obtain resource requirement parameters of the AI computing task; based on the resource requirement parameters, estimating a required cache space size and mapping the cache space size to a specific cache organization structure of a last-level cache; selecting a specific physical location of the target cache block according to a current occupancy state of the cache and generating a physical address list thereof; verifying the availability of the target cache block and finally confirming the target cache block.

3. The AI computing acceleration method based on dynamic switching of cache modes according to claim 2, characterized in that, The resource requirement parameters include a computing operator type, a data size and a computing precision; the basic unit of cache allocation in the cache organization structure is a way or a storage bank; the target cache block is located in the last-level cache of a processor.

4. The AI computing acceleration method based on dynamic switching of cache modes according to claim 1, characterized in that, In step S2, the quick flush operation performed on the target cache block comprises: a scheduler directly initiates a flush request based on a physical address; a mode switching controller directly queries a cache directory to accurately locate cache lines storing valid data in the target cache block; performing targeted write-back and invalidation operations on the cache lines of the valid data; after completion of the flush operation, confirming that all storage units in the target cache block are in an invalid state.

5. The AI computing acceleration method based on dynamic switching of cache modes according to claim 1, characterized in that, In step S3, the marking of the state of the target cache block as the AI computing mode and the shielding in the cache allocation strategy comprise: marking the target cache block as the AI computing mode in a dedicated state register or a state lookup table; dynamically modifying allocation and replacement strategies of a cache controller to exclude the target cache block from the cache allocation range of ordinary programs; continuously maintaining the isolation state until a release instruction is received.

6. The AI computing acceleration method based on dynamic switching of cache modes according to claim 1, characterized in that, In step S4, the in-memory computing of the AI operator using the storage array in the cache block in the AI computing mode and the vector operation unit integrated therein comprises: loading data to be computed to the target cache block and performing data preprocessing according to in-memory computing requirements; selectively enabling the vector operation unit integrated in the target cache block according to computing task requirements; controlling the vector operation unit to process data in the storage array in parallel and temporarily storing the operation results in a designated storage area of the same cache block; for a computing task exceeding the cache capacity, adopting a block or slice strategy to cyclically perform data loading, operation and temporary storage steps until the task is completed.

7. The AI computing acceleration method based on dynamic switching of cache modes according to claim 6, characterized in that, The data preprocessing includes data unfolding, transposition and bit width conversion operations; The vector operation unit includes a multiplier and an adder for performing multiplication and addition operations.

8. The AI computing acceleration method based on dynamic switching of cache modes according to claim 1, characterized in that, In step S5, the clearing of the AI computing mode flag of the target cache block and the restoration of its availability in the cache allocation strategy include: Receiving an AI computing task completion signal, sending a release instruction containing the target cache block identifier to the cache switch controller; Clearing the AI computing mode flag of the target cache block in the dedicated state register or state lookup table; Restoring the availability of the target cache block in the cache allocation strategy, so that it is included in the cache allocation range of the general program again; Returning a release completion confirmation signal to the scheduler, and the target cache block restores the general cache function.

9. The AI computing acceleration method based on dynamic switching of cache modes according to any one of claims 1 to 8, characterized in that, The AI computing acceleration method based on dynamic switching of cache modes initiates an acceleration request of an AI computing task through an extended instruction set, and the extended instruction set is compatible with the RISC-V open source architecture.

10. An AI computing acceleration system based on dynamic switching of cache modes, the system comprising: a memory controller configured to: determine a cache mode for a memory device; and send a command to the memory device to switch to the determined cache mode. The AI computing acceleration system based on dynamic switching of cache modes includes: A cache switch controller for directly querying a cache directory based on a physical address and performing a quick refresh operation on a target cache block to maintain cache consistency while managing the state flag of the target cache block; A matrix multiplication scheduler for receiving an acceleration request of an AI computing task, analyzing resource requirements, determining the physical address of a target cache block, and scheduling an in-memory computing process; A data preprocessing module for adapting AI operators, including convolution unfolding, data transposition and bit width conversion, and achieving efficient data transfer through a direct memory access module; A cache isolation module for isolating and restoring a target cache block by modifying the cache allocation strategy through a state register and dynamically; An instruction interface and software scheduling unit for providing a hardware instruction interface and a system-level scheduling function to coordinate AI computing and regular tasks.

Citation Information

Patent Citations

  • In-memory computing AI accelerator design architecture based on RISC-V architecture and control method

    CN119201836A