Cache and operation method thereof

By introducing a hit-miss check unit and a reference table into the AI ​​chip, the problem of selecting the cache line in use during cache line replacement is solved, achieving safer and more efficient cache operations.

CN121070871AActive Publication Date: 2025-12-05SHANGHAI BIREN TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511613670.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2025-12-05
Estimated Expiration
2045-11-06

AI Technical Summary

Technical Problem

In artificial intelligence chips, how can we ensure that cache lines not being used are not selected when replacing cache lines to avoid data loss or performance degradation?

Method used

A hit-miss check unit and a reference table are introduced to determine whether to execute a replacement request by checking the reference status of the cache line, ensuring that the replacement is only executed when the target cache line is either out of service or not in use.

Benefits of technology

It improves the security and efficiency of cache replacement, avoids unnecessary data loss, and enhances system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121070871A_ABST
    Figure CN121070871A_ABST
Patent Text Reader

Abstract

The invention provides a cache and an operation method thereof. The cache includes a reference table and a hit-miss check unit. The hit-miss check unit is coupled to the reference table. In response to one of the plurality of computing cores sending a replacement request to the cache, the hit-miss check unit checks a corresponding reference bit corresponding to a target cache line of the replacement request in the reference table. In response to the corresponding reference bit corresponding to the target cache line in the reference table indicating that the target cache line is quit, the hit-miss check unit executes the replacement request. In response to the fact that the corresponding reference bit corresponding to the target cache line in the reference table indicates that the target cache line is busy, the hit-miss check unit does not temporarily execute the replacement request until the corresponding reference bit corresponding to the target cache line of the replacement request in the reference table indicates that the target cache line is quit.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence (AI) chips, and in particular to a cache and an operation method thereof. BACKGROUND

[0002] In the architecture of an artificial intelligence (AI) chip, a graphics processing unit (GPU), a General-purpose GPU (GPGPU), or the like, the replacement of a cache line is crucial. The replacement refers to that when the cache is full and new data needs to be loaded, the system selects a cache line that is not used by any program at the moment. How to ensure that the selected cache line is not used by any program is one of the many technical issues in the field. SUMMARY

[0003] The present application is directed to a cache and an operation method thereof to safely perform a replacement request.

[0004] In an embodiment according to the present application, the cache includes a referring table and a hit-miss check unit. The hit-miss check unit is coupled to the referring table. In response to a replacement request being sent by one of the plurality of compute cores to the cache, the hit-miss check unit checks a corresponding reference bit in the referring table corresponding to a target cache line of the replacement request. In response to the corresponding reference bit in the referring table corresponding to the target cache line indicating that the target cache line is retired, the hit-miss check unit performs the replacement request. In response to the corresponding reference bit in the referring table corresponding to the target cache line indicating that the target cache line is busy, the hit-miss check unit does not perform the replacement request until the corresponding reference bit in the referring table corresponding to the target cache line of the replacement request indicates that the target cache line is retired.

[0005] In an embodiment according to the present application, the operation method comprises: in response to one of the plurality of computing cores sending a replacement request to the cache, checking, by a hit-miss check unit of the cache, a corresponding reference bit in a reference table of the cache corresponding to a target cache line of the replacement request; in response to the corresponding reference bit in the reference table corresponding to the target cache line indicating that the target cache line is retired, executing, by the hit-miss check unit, the replacement request; and in response to the corresponding reference bit in the reference table corresponding to the target cache line indicating that the target cache line is busy, temporarily not executing, by the hit-miss check unit, the replacement request until the corresponding reference bit in the reference table corresponding to the target cache line of the replacement request indicates that the target cache line is retired.

[0006] Based on the above, each cache line of the cache is configured with a dedicated reference bit. Each reference bit in the reference table is used to mark the reference state (or busy state) of a corresponding cache line. The reference state refers to whether the cache line is referenced (used) by any program. When the hit-miss check unit receives a replacement request, the hit-miss check unit checks a corresponding reference bit in the reference table corresponding to a target cache line of the replacement request. The corresponding reference bit can ensure that the target cache line is busy (referenced) or retired (not referenced). Therefore, the cache can safely execute the replacement request. BRIEF DESCRIPTION OF DRAWINGS

[0007] Figure 1 is a circuit block diagram of an artificial intelligence (AI) chip according to an embodiment of the present application.

[0008] Figure 2 is a flowchart diagram of an operation method of an artificial intelligence chip according to an embodiment of the present application.

[0009] Figure 3 is a diagram of a reference table according to an embodiment of the present application.

[0010] Figure 4 is a circuit block diagram of a management unit according to an embodiment of the present application.

[0011] DETAILED DESCRIPTION

[0012] 100: AI chip,

[0013] 110: computing core,

[0014] 120: cache,

[0015] 121: hit-miss check unit,

[0016] 122: tag array,

[0017] 123: operation engine,

[0018] 124: cache line array,

[0019] 125: reference table,

[0020] 126: management unit,

[0021] 130: main memory,

[0022] 310_1: 1st group reference map,

[0023] 310_2: 2nd group reference map,

[0024] 310_c: cth group reference map,

[0025] 410: exit engine,

[0026] 420: arbiter,

[0027] ref_1: 1st reference bit,

[0028] ref_2: 2nd reference bit,

[0029] ref_m: mth reference bit. DETAILED DESCRIPTION

[0030] Reference will now be made in detail to exemplary embodiments of the present application, examples of which are illustrated in the accompanying drawings. Wherever possible, the same reference numbers are used in the drawings and the description to refer to the same or like parts.

[0031] The term "coupled" or "connected" used in the detailed description section of the present application, including in the claims, can refer to either a direct or indirect connection. For example, if a first device is coupled or connected to a second device, it can be directly connected to the second device or it can be indirectly connected to the second device through one or more other devices or some connection means. The terms "first", "second", and the like used in the detailed description section of the present application, including in the claims, refer to names of elements rather than the order of the elements, and are used to distinguish between different embodiments or aspects of the present application. In addition, the same reference numbers are used in the drawings and the description to represent the same or similar parts. The same reference numbers or the same terms used in different embodiments can be referred to each other according to the related description. It should be understood that the features of the following embodiments can be combined with each other. For example, the features of the second embodiment can be combined with the features of the first embodiment. Those skilled in the art can select appropriate combinations of features according to actual design requirements.

[0032] An operation device such as an artificial intelligence (AI) chip can provide tremendous computing power. The tremendous computing power of an AI chip is derived from a large number of hardware cores inside. An AI chip usually contains multiple programmable processors, for example, a Stream Processor Cluster (SPC). Each programmable processor usually contains multiple compute units (CUs, or compute cores), and each compute core usually contains multiple execution units (EUs, or execution cores). An execution core includes at least one of a tensor core (Tcore), an integer (INT) core, a floating point (FP) core, and a vector core (Vcore), for example. By programming the compute cores, an AI chip can support general-purpose computing, scientific computing, and neural network computing. The compute cores of an AI chip usually access data in a main memory through a cache, such as a last level cache (LLC). The following embodiments will illustrate implementation examples of a cache.

[0033] Figure 1 is a circuit block diagram of an AI chip 100 according to an embodiment of the present application. In Figure 1 In the embodiment shown, the AI chip 100 includes multiple compute cores 110, a cache 120, and a main memory 130. The number of compute cores 110 can be determined according to actual design and application. A compute core 110 is also referred to as a compute unit (CU). Each compute core 110 usually contains multiple execution units (EUs, or execution cores) and a shared memory. Different execution cores in the same compute core can exchange data with each other through the shared memory. The cache 120 is coupled between the compute cores 110 and the main memory 130. The cache 120 can be a last level cache (LLC) of the AI chip 100 or other caches. The compute cores 110 access data in the main memory 130 through the cache 120.

[0034] In the embodiment shown, the AI chip 100 includes multiple compute cores 110, a cache 120, and a main memory 130. The number of compute cores 110 can be determined according to actual design and application. A compute core 110 is also referred to as a compute unit (CU). Each compute core 110 usually contains multiple execution units (EUs, or execution cores) and a shared memory. Different execution cores in the same compute core can exchange data with each other through the shared memory. The cache 120 is coupled between the compute cores 110 and the main memory 130. The cache 120 can be a last level cache (LLC) of the AI chip 100 or other caches. The compute cores 110 access data in the main memory 130 through the cache 120. Figure 1In the illustrated embodiment, the cache 120 includes a hit-miss check unit 121, a tag array 122, an operation engine 123, a cache line array 124, a referring table 125, and a management unit 126. The hit-miss check unit 121 is coupled to the tag array 122, the operation engine 123, the referring table 125, and the management unit 126, while the operation engine 123 is coupled to the cache line array 124. Depending on different designs, in some embodiments, the implementation of at least one of the hit-miss check unit 121, the operation engine 123, and the management unit 126 can be in the form of hardware circuit. In other embodiments, the implementation of at least one of the hit-miss check unit 121, the operation engine 123, and the management unit 126 can be in the form of a combination of more than one of hardware, firmware, and software (i.e., program).

[0035] In terms of hardware, at least one of the hit-miss check unit 121, the operation engine 123, and the management unit 126 can be implemented as logic circuits on an integrated circuit. For example, the functions of at least one of the hit-miss check unit 121, the operation engine 123, and the management unit 126 can be implemented as various logic blocks, modules, and circuits in one or more hardware controllers, microcontrollers, hardware processors, microprocessors, application-specific integrated circuits (ASICs), digital signal processors (DSPs), field programmable gate arrays (FPGAs), central processing units (CPUs), or other processing units. The functions of at least one of the hit-miss check unit 121, the operation engine 123, and the management unit 126 can be implemented as hardware circuits, such as various logic blocks, modules, and circuits in an integrated circuit, using a hardware description language (e.g., Verilog HDL or VHDL) or other suitable programming language.

[0036] In software or firmware form, the functions of at least one of the hit-miss check unit 121, the operation engine 123, and the management unit 126 can be implemented as programming codes. For example, at least one of the hit-miss check unit 121, the operation engine 123, and the management unit 126 is implemented by using general programming languages (e.g., C, C++, or assembly language) or other suitable programming languages. The programming codes can be recorded / stored in a "non-transitory machine-readable storage medium". In some embodiments, the non-transitory machine-readable storage medium includes, for example, semiconductor memories and / or storage devices. An electronic device (e.g., a CPU, a hardware controller, a microcontroller, a hardware processor, or a microprocessor) can read and execute the programming codes from the non-transitory machine-readable storage medium, thereby implementing the functions of at least one of the hit-miss check unit 121, the operation engine 123, and the management unit 126.

[0037] The tag array 122 includes a plurality of tag sets, and each tag set includes a plurality of tags. The cache line array 124 includes a plurality of cache lines. The plurality of cache lines of the cache line array 124 one-to-one correspond to the plurality of tags of the tag array 122. Each tag is used to store corresponding cache line information. The hit-miss check unit 121 fetches a corresponding tag set from the tag array 122 based on a set address, and then compares a tag address field of each tag in the corresponding tag set with a tag address carried by an access request. If the tag address field of a tag matches the tag address carried by the access request, the hit-miss check unit 121 determines "hit". If the tag address field of each tag does not match the tag address carried by the access request, the hit-miss check unit 121 determines "miss".

[0038] The operation engine 123 is coupled to the hit-miss check unit 121 and the cache line array 124. In response to an access request sent by one of the plurality of compute cores 110 to the cache 120, the hit-miss check unit 121 checks whether the access request is a hit. In response to the hit-miss check unit 121 determining that the access request is a hit, the operation engine 123 accesses a cache line corresponding to the access request in the plurality of cache lines of the cache line array 124. In response to the hit-miss check unit 121 determining that the access request is a miss, the operation engine 123 accesses the main memory 130.

[0039] Figure 2 This is a flowchart illustrating an operation method of an artificial intelligence chip according to an embodiment of the present invention. Please refer to... Figure 1 and Figure 2 In response to one of the multiple computing cores 110 sending a replacement request to the cache 120, the hit-miss check unit 121 checks the corresponding reference bit in the reference table 125 for the target cache line of the replacement request (step S210). The specific structure of the reference table 125 can be determined according to the actual design and application.

[0040] Figure 3 This is a schematic diagram illustrated according to an embodiment of the present invention, referencing Table 125. Figure 3 The referenced Table 125 shown can be used as... Figure 1 The example shown is one of many implementation examples referenced in Table 125. In Figure 3 In the illustrated embodiment, reference table 125 includes multiple set referring maps, such as reference map 310_1, reference map 310_2, ..., reference map 310_c. Reference maps 310_1 to 310_c in reference table 125 correspond one-to-one with multiple cache line groups of cache line array 124. Each of the multiple cache line groups includes multiple cache lines. Each reference map 310_1 to 310_c includes multiple reference bits, such as reference bit ref_1, reference bit ref_2, ..., reference bit ref_m. Reference bits ref_1 to ref_m in reference maps 310_1 to 310_c correspond one-to-one with different cache lines of cache line array 124.

[0041] In reference table 125, the first reference bit ref_1 to the m-th reference bit ref_m of reference diagrams 310_1 to 310_c (group 1) are used to indicate the reference status (or busy status) of different corresponding cache lines. The reference status refers to whether a cache line is referenced (used) by any program. For example (but not limited to this), suppose a cache line of cache line array 124 corresponds to the first reference bit ref_1 of reference diagram 310_1. When a cache line is referenced (used) by one or more programs, the first reference bit ref_1 of reference diagram 310_1 is set to logical "truth", such as the logical value "1". Conversely, when a cache line is not referenced (used) by any program, the first reference bit ref_1 of reference diagram 310_1 is reset to logical "false", such as the logical value "0".

[0042] Referring to Figure 1 , Figure 2 and Figure 3 , in response to the corresponding reference bit in the reference table 125 corresponding to the target cache line indicating that the target cache line is retired (i.e., not referenced), the hit-miss check unit 121 performs the replacement request (step S220). In response to the corresponding reference bit in the reference table 125 corresponding to the target cache line indicating that the target cache line is busy (referenced), the hit-miss check unit 121 temporarily does not perform the replacement request until the corresponding reference bit in the reference table 125 corresponding to the target cache line of the replacement request indicates that the target cache line is retired (step S230).

[0043] The management unit 126 is configured to manage the reference table 125. The management unit 126 is coupled to the hit-miss check unit 121 and the operation engine 123. In response to the hit-miss check unit 121 performing a replacement request, the management unit 126 sets the corresponding reference bit in the reference table 125 corresponding to the target cache line of the replacement request. In response to one of the plurality of compute cores 110 sending an access request to the cache, the management unit 126 sets the reference bit in the reference table 125 corresponding to the target cache line of the access request.

[0044] In response to the operation engine 123 completing a current access request, the management unit 126 checks whether the target cache line of the access request to be performed by the operation engine 123 is the same as the target cache line of the current access request. In response to the target cache line of any one of the access requests to be performed by the operation engine 123 being the same as the target cache line of the current access request, the management unit 126 maintains the set state of the reference bit in the reference table 125 corresponding to the target cache line of the access request. In response to the target cache line of all of the access requests to be performed by the operation engine 123 being different from the target cache line of the current access request, the management unit 126 resets the reference bit in the reference table 125 corresponding to the target cache line of the access request.

[0045] In summary, each cache line of the cache line array 124 is configured with a dedicated reference bit. Each reference bit in the reference table 125 is configured to indicate the reference state (or busy state) of a corresponding cache line. The reference state refers to whether the cache line is referenced (used) by any program. When the hit-miss check unit 121 receives a replacement request, the hit-miss check unit 121 checks a corresponding reference bit in the reference table 125 corresponding to the target cache line of the replacement request. The corresponding reference bit in the reference table 125 can ensure that the target cache line is busy (referenced) or retired (not referenced). Therefore, the cache 120 can safely perform the replacement request.

[0046] Figure 4 This is a circuit block diagram of the management unit 126 according to an embodiment of the present invention. Figure 4 The management unit 126 shown can be used as Figure 1 This is one of many implementation examples of the management unit 126 shown. Figure 4 The hit-miss check unit 121, operation engine 123, reference table 125, and management unit 126 shown can be referenced. Figures 1 to 3 The relevant explanations. In Figure 4 In the illustrated embodiment, management unit 126 includes an exit engine 410 and an arbitrator 420. The exit engine 410 is coupled to the operation engine 123. The arbitrator 420 manages the reference table 125. The arbitrator 420 is coupled to the exit engine 410 and a hit-miss check unit 121. In response to the hit-miss check unit 121 executing a replacement request, the arbitrator 420 sets the corresponding reference bit in the reference table 125 corresponding to the target cache line of the replacement request. In response to one of the compute cores 110 sending an access request to the cache 120, the arbitrator 420 sets the reference bit in the reference table 125 corresponding to the target cache line of the access request.

[0047] In response to the operation engine 123 completing the current access request, the exit engine 410 checks whether the target cache line of the pending access requests in the operation engine 123 is the same as the target cache line of the current access request. If the check result of the exit engine 410 indicates that the target cache line of any pending access request in the operation engine 123 is the same as the target cache line of the current access request, the arbitrator 420 maintains the set state of the reference bit corresponding to the target cache line of the access request in the reference table 125. If the check result of the exit engine 410 indicates that the target cache lines of all pending access requests in the operation engine 123 are different from the target cache line of the current access request, the arbitrator 420 resets the reference bit corresponding to the target cache line of the access request in the reference table 125.

[0048] Sometimes, the set command issued by the hit-miss check unit 121 to the arbitrator 420 and the reset command issued by the exit engine 410 to the arbitrator 420 may arrive at the arbitrator 420 simultaneously. In this case, the arbitration strategy of the arbitrator 420 includes that the set command issued by the hit-miss check unit 121 to the arbitrator 420 takes precedence over the reset command issued by the exit engine 410 to the arbitrator 420.

[0049] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A cache, characterized by, The cache comprises: a reference table; and a hit-miss check unit coupled to the reference table, wherein in response to one of the plurality of compute cores sending a replacement request to the cache, the hit-miss check unit checks a corresponding reference bit in the reference table corresponding to a target cache line of the replacement request; in response to the corresponding reference bit in the reference table corresponding to the target cache line indicating that the target cache line is evicted, the hit-miss check unit performs the replacement request; and in response to the corresponding reference bit in the reference table corresponding to the target cache line indicating that the target cache line is busy, the hit-miss check unit temporarily does not perform the replacement request until the corresponding reference bit in the reference table corresponding to the target cache line of the replacement request indicates that the target cache line is evicted.

2. The cache of claim 1, wherein, The cache further comprises: a cache line array comprising a plurality of cache line groups, wherein each of the plurality of cache line groups comprises a plurality of cache lines, the plurality of cache lines of the cache line array one-to-one corresponding to different reference bits of the reference table; and an operation engine coupled to the hit-miss check unit and the cache line array, wherein in response to one of the plurality of compute cores sending an access request to the cache, the hit-miss check unit checks whether the access request is a hit; in response to the hit-miss check unit determining that the access request is a hit, the operation engine accesses a cache line corresponding to the access request among the plurality of cache lines of the cache line array; and in response to the hit-miss check unit determining that the access request is a miss, the operation engine accesses a main memory.

3. The cache of claim 2, wherein, The reference table comprises a plurality of group reference maps, the plurality of group reference maps one-to-one corresponding to the plurality of cache line groups of the cache line array, each of the plurality of group reference maps comprising a plurality of reference bits, the plurality of reference bits one-to-one corresponding to the plurality of cache lines of the cache line array.

4. The cache of claim 3, wherein, The cache further comprises: a management unit to manage the reference table, wherein the management unit is coupled to the hit-miss check unit and the operation engine, and in response to the hit-miss check unit performing the replacement request, the management unit sets the corresponding reference bit in the reference table corresponding to the target cache line of the replacement request.

5. The cache of claim 4, wherein in response to one of the plurality of compute cores sending the access request to the cache, the management unit sets a reference bit in the reference table corresponding to a target cache line of the access request.

6. The cache of claim 4, wherein in response to the operation engine completing a current access request, the management unit checks whether a target cache line of a pending access request of the operation engine is the same as a target cache line of the current access request; and in response to the target cache line of the pending access request of the operation engine being the same as the target cache line of the current access request, the management unit sets the reference bit in the reference table corresponding to the target cache line of the pending access request. in response to any one of the pending access requests of the operation engine having the same target cache line as the target cache line of the current access request, the management unit maintaining a set state of a reference bit in the reference table corresponding to the target cache line of the access request; and in response to all of the pending access requests of the operation engine having different target cache lines than the target cache line of the current access request, the management unit resetting the reference bit in the reference table corresponding to the target cache line of the access request.

7. The cache of claim 4, wherein, The management unit comprises: an exit engine coupled to the operation engine; and an arbiter to manage the reference table, wherein the arbiter is coupled to the exit engine and the hit-miss check unit, wherein, in response to the hit-miss check unit executing the replacement request, the arbiter setting the corresponding reference bit in the reference table corresponding to the target cache line of the replacement request.

8. The cache of claim 7, wherein, in response to one of the plurality of compute cores sending the access request to the cache, the arbiter setting the reference bit in the reference table corresponding to the target cache line of the access request.

9. The cache of claim 7, wherein, in response to the operation engine completing a current access request, the exit engine checking whether the target cache line of any one of the pending access requests of the operation engine is the same as the target cache line of the current access request; in response to the exit engine checking result indicating that the target cache line of any one of the pending access requests of the operation engine is the same as the target cache line of the current access request, the arbiter maintaining the set state of the reference bit in the reference table corresponding to the target cache line of the access request; and in response to the exit engine checking result indicating that the target cache line of all of the pending access requests of the operation engine are different from the target cache line of the current access request, the arbiter resetting the reference bit in the reference table corresponding to the target cache line of the access request.

10. The cache of claim 7, wherein, The hit-miss check unit has a higher priority than the exit engine to issue a set command to the arbiter.

11. A method of operating a cache, the method comprising: The operation method comprises: in response to one of the plurality of compute cores sending the access request to the cache, the hit-miss check unit checking a corresponding reference bit in a reference table of the cache corresponding to a target cache line of the access request, wherein the hit-miss check unit is coupled to the reference table; in response to the corresponding reference bit in the reference table corresponding to the target cache line indicating that the target cache line is an exit, the hit-miss check unit executing the replacement request; and in response to the corresponding reference bit in the reference table corresponding to the target cache line indicating that the target cache line is an exit, the hit-miss check unit executing the replacement request; and in response to the corresponding reference bit in the reference table corresponding to the target cache line of the replacement request indicating that the target cache line is busy, temporarily not performing the replacement request by the hit-miss check unit until the corresponding reference bit in the reference table corresponding to the target cache line of the replacement request indicates that the target cache line is retired.

12. The method of claim 11, wherein, The cache line array of the cache comprises a plurality of cache line groups, each of the plurality of cache line groups comprises a plurality of cache lines, the plurality of cache lines of the cache line array correspond to different reference bits of the reference table one-to-one, and the operation method further comprises: in response to one of the plurality of computing cores sending an access request to the cache, checking by the hit-miss check unit whether the access request is a hit; in response to the hit-miss check unit determining that the access request is a hit, accessing by the operation engine of the cache a cache line corresponding to the access request in the plurality of cache lines of the cache line array, wherein the operation engine is coupled to the hit-miss check unit and the cache line array; and in response to the hit-miss check unit determining that the access request is a miss, accessing by the operation engine a main memory.

13. The method of operation of claim 12, wherein, The reference table comprises a plurality of group reference maps, the plurality of group reference maps correspond to the plurality of cache line groups of the cache line array one-to-one, each of the plurality of group reference maps comprises a plurality of reference bits, and the plurality of reference bits correspond to the plurality of cache lines of the cache line array one-to-one.

14. The method of claim 13, wherein, The operation method further comprises: managing by a management unit of the cache the reference table, wherein the management unit is coupled to the hit-miss check unit and the operation engine; and in response to the hit-miss check unit performing the replacement request, setting by the management unit the corresponding reference bit in the reference table corresponding to the target cache line of the replacement request.

15. The method of operation of claim 14, wherein, The operation method further comprises: in response to one of the plurality of computing cores sending the access request to the cache, setting by the management unit a reference bit in the reference table corresponding to the target cache line of the access request.

16. The method of claim 14, wherein, The operation method further comprises: in response to the operation engine completing a current access request, checking by the management unit whether the target cache line of a to-be-executed access request of the operation engine is the same as the target cache line of the current access request; in response to the target cache line of any one to-be-executed access request of the operation engine being the same as the target cache line of the current access request, maintaining by the management unit the set state of the reference bit in the reference table corresponding to the target cache line of the access request; and in response to the target cache lines of all to-be-executed access requests of the operation engine being different from the target cache line of the current access request, resetting by the management unit the reference bit in the reference table corresponding to the target cache line of the access request.

17. The method of claim 14, wherein, The operation method further comprises: the reference table is managed by an arbiter of the management unit, wherein the arbiter is coupled to an exit engine of the management unit and the hit-miss check unit, and the exit engine is coupled to the operation engine; and in response to the hit-miss check unit performing the replacement request, setting, by the arbiter, the corresponding reference bit in the reference table corresponding to the target cache line of the replacement request.

18. The method of operation of claim 17, wherein, The operation method further comprises: in response to one of the plurality of compute cores sending the access request to the cache, setting, by the arbiter, a reference bit in the reference table corresponding to the target cache line of the access request.

19. The method of claim 17, wherein, The operation method further comprises: in response to the operation engine completing a current access request, checking, by the exit engine, whether a target cache line of a pending access request of the operation engine is the same as a target cache line of the current access request; in response to a result of the checking by the exit engine indicating that the target cache line of any one of the pending access requests of the operation engine is the same as the target cache line of the current access request, maintaining, by the arbiter, a set state of a reference bit in the reference table corresponding to the target cache line of the access request; and in response to the result of the checking by the exit engine indicating that the target cache line of all of the pending access requests of the operation engine is different from the target cache line of the current access request, resetting, by the arbiter, the reference bit in the reference table corresponding to the target cache line of the access request.

20. The operating method according to claim 17, characterized in that, The hit-miss check unit has a higher priority than the exit engine in issuing a set command to the arbiter.

Citation Information

Patent Citations

  • Method and apparatus for secure context switching in a system including a processor and cached virtual memory

    CN101438290A

  • Metaphysical address space for holding lossy metadata in hardware

    CN101770429A

  • Management method of computer cache system

    CN102999443A

  • Last-stage cache based on linked list structure and supporting dynamic partition granularity access

    CN117785737A

  • Cache of artificial intelligence chip and operation method thereof

    CN120803969A