Method for supporting cache coherency based on virtual address for artificial intelligence processor with large capacity on-chip memory and apparatus using same

By setting non-overlapping external memory address areas in the artificial intelligence processor and providing virtual addresses, the problem of difficult to maintain consistency between multiple caches is solved, efficient cache consistency and virtual address support are achieved, and the performance and efficiency of the processor are improved.

CN120225997APending Publication Date: 2025-06-27ELECTRONICS & TELECOMM RES INST

Patent Information

Application Number
CN202380081612.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-07-07
Filing Date
2023-11-01
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The prior art is difficult to effectively maintain cache consistency between multiple caches, especially in artificial intelligence processors with large capacity on-chip memory, and it is difficult to support virtual addresses and quickly call software running on multiple processor cores.

Method used

By setting an external memory address area that does not overlap each other in an artificial intelligence processor and providing a virtual address, multiple processor cores can access multiple caches using virtual addresses to achieve cache consistency.

Benefits of technology

It effectively maintains cache consistency between multiple caches, supports virtual addresses, improves processor performance and efficiency, and can quickly call software running on multiple processor cores.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120225997A_ABST
    Figure CN120225997A_ABST
Patent Text Reader

Abstract

A method for supporting cache coherency based on a virtual address for an artificial intelligence processor having a large capacity on-chip memory and an apparatus using the same are disclosed. A method for supporting cache coherency according to an embodiment of the present invention comprises the following steps performed by an artificial intelligence processor comprising a plurality of processor cores and a plurality of caches: setting non-overlapping external memory address regions in each of the plurality of caches; and providing a virtual address for the plurality of processor cores to access the plurality of caches.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to a method for supporting cache coherence based on virtual addresses for an artificial intelligence processor having a large-capacity on-chip memory and an apparatus for the method, and more particularly, to a structure for supporting cache coherence and virtual addresses in a cache integrated with a processor core, the processor core forming an artificial intelligence processor having a large-capacity on-chip memory.

[0002] This application claims the benefit of Korean Patent Application No. 10-2022-0163064, filed on November 29, 2022, and Korean Patent Application No. 10-2023-0088228, filed on July 7, 2023, the entire disclosures of which are incorporated herein by reference. Background Art

[0003] Recently, in response to the demand for high-performance artificial intelligence processors, as Figure 1 shown, many processor cores (access node controllers: ANCs) have been integrated into a single die, and caches for respective processor cores have been installed in the die such that the number of caches is the same as the number of processor cores. In this case, the number of cases where many processor cores process different multiple data based on the same instruction configuration increases.

[0004] In addition, maintaining cache coherence between multiple caches using only existing directory-based cache coherence designs may lead to excessive performance degradation.

[0005] Therefore, a new scheme for maintaining cache coherence between multiple caches is needed. Summary of the Invention

[0006] Technical Problem An object of the present disclosure is to provide a method capable of maintaining cache coherence between multiple caches in an artificial intelligence processor having a large-capacity on-chip memory.

[0007] Another object of the present disclosure is to provide a structure for supporting virtual addresses that allows multiple processor cores to process different multiple data based on the same instruction configuration.

[0008] Still another object of the present disclosure is to quickly call software running on multiple processor cores.

[0009] Technical Solution A method for supporting cache coherence based on virtual addresses according to the present disclosure to achieve the above object includes: for an artificial intelligence processor including a plurality of processor cores and a plurality of caches, setting non-overlapping external memory address regions for the corresponding plurality of caches; and providing virtual addresses, wherein the plurality of processor cores use the virtual addresses to access the plurality of caches.

[0010] Here, the setting step may include setting the external memory address region by considering the programs running on the plurality of processor cores and the access distances between the plurality of processor cores and the plurality of caches.

[0011] Here, the setting step may further include setting the external memory address region such that the data required to run each program is stored in the cache with the shortest access distance to each of the plurality of processor cores.

[0012] Here, the setting step may further include setting the external memory address region based on the memory mapping information related to the plurality of caches through a bus existing between the plurality of processor cores and the plurality of caches.

[0013] Here, the method may further include detecting, by the artificial intelligence processor, the programs running on the plurality of processor cores.

[0014] Here, the virtual address may be provided as corresponding to an exclusive address such that the plurality of processor cores can access the same data simultaneously.

[0015] Here, the bus may be configured to obtain mapping information between the physical address and the virtual address corresponding to the plurality of caches based on a translation lookaside buffer (TLB) that can be set by the system controller of the artificial intelligence processor, and provide the virtual address to the plurality of processor cores considering the mapping information.

[0016] Here, the external memory address region may correspond to the address region of a high bandwidth memory (HBM) located outside the artificial intelligence processor.

[0017] Here, for each piece of data to be stored in the plurality of caches, considering data dependencies to determine the cache in which the corresponding data will be stored, and the cache may be determined such that when the data dependencies are high, the corresponding data is stored in the cache with a relatively short cache distance corresponding to the distance between the caches.

[0018] In addition, an artificial intelligence processor according to the present disclosure includes: a plurality of processor cores; a plurality of caches corresponding to the plurality of processor cores; a system controller (SCP); and a bus located between the plurality of processor cores and the plurality of caches, wherein the bus is configured to set non-overlapping external memory address regions for the corresponding plurality of caches and provide virtual addresses, and wherein the plurality of processor cores access the plurality of caches using the virtual addresses.

[0019] Here, the bus may be configured to set the external memory address regions in consideration of the programs running on the plurality of processor cores and the access distances between the plurality of processor cores and the plurality of caches.

[0020] Here, the bus may be configured to set the external memory address regions such that the data required to run each program is stored in the cache with the shortest access distance to each of the plurality of processor cores.

[0021] Here, the bus may be configured to cause the bus existing between the plurality of processor cores and the plurality of caches to set the external memory address regions based on the memory mapping information related to the plurality of caches.

[0022] Here, the bus may be configured to detect the programs running on the plurality of processor cores.

[0023] Here, the virtual addresses may be provided as corresponding to exclusive addresses such that the plurality of processor cores can access the same data simultaneously.

[0024] Here, the bus may be configured to obtain the mapping information between the physical addresses and the virtual addresses corresponding to the plurality of caches based on a translation lookaside buffer (TLB) that can be set by the system controller of the artificial intelligence processor, and provide the virtual addresses to the plurality of processor cores in consideration of the mapping information.

[0025] Here, the external memory address regions may correspond to the address regions of a high-bandwidth memory (HBM) located outside the artificial intelligence processor.

[0026] Here, for each piece of data to be stored in the plurality of caches, the cache in which the corresponding data is to be stored may be determined in consideration of data dependencies, and the cache may be determined such that when the data dependencies are high, the corresponding data is stored in the cache with a relatively short cache distance corresponding to the distance between the caches.

[0027] Advantageous Effects According to the present disclosure, a method capable of maintaining cache coherence between a plurality of caches in an artificial intelligence processor having a large-capacity on-chip memory can be provided.

[0028] Further, the present disclosure may provide a structure for supporting virtual addresses, which allows multiple processor cores to process different multiple pieces of data based on the same instruction configuration.

[0029] In addition, the present disclosure can quickly call software running on multiple processor cores. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a diagram showing the structure of an artificial intelligence processor composed of 128 processor cores; Figure 2 is a flowchart of operations of a method for supporting cache coherence based on virtual addresses for an artificial intelligence processor having a large-capacity on-chip memory according to an embodiment of the present disclosure; Figure 3 and Figure 4 is a diagram showing an example of the structure of an artificial intelligence processor for supporting cache coherence and virtual addresses according to the present disclosure; and Figure 5 is a diagram showing an example of a mapping process for supporting virtual addresses according to the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] The present disclosure will be described in detail below with reference to the drawings. Repeated descriptions and descriptions of known functions and configurations that are considered to unnecessarily obscure the gist of the present disclosure will be omitted. Embodiments of the present disclosure are intended to fully describe the present disclosure to persons having ordinary knowledge in the art to which the present disclosure pertains. Therefore, the shapes, sizes, etc. of components in the drawings may be exaggerated to make the description clearer.

[0032] In this specification, each phrase such as "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" may include any one of the items listed together in the corresponding phrase or all possible combinations thereof.

[0033] Hereinafter, preferred embodiments of the present disclosure will be described in detail with reference to the drawings.

[0034] Figure 2 is a flowchart of operations of a method for supporting cache coherence based on virtual addresses for an artificial intelligence processor having a large-capacity on-chip memory according to an embodiment of the present disclosure.

[0035] Referring to Figure 2, in a method for supporting cache coherence based on virtual addresses for an artificial intelligence processor with a large-capacity on-chip memory according to an embodiment of the present disclosure, in step S210, an artificial intelligence processor including a plurality of processor cores and a plurality of caches sets non-overlapping external memory address regions for the corresponding plurality of caches.

[0036] Here, the external memory address region may correspond to an address region of a high-bandwidth memory (HBM) located outside the artificial intelligence processor. For example, the external memory address region may correspond to an HBM memory pseudo-channel region.

[0037] Here, an artificial intelligence processor according to an embodiment of the present disclosure may include a plurality of processor cores, a plurality of caches, a system controller, and a bus located between the plurality of processor cores and the plurality of caches.

[0038] Therefore, the caches connected to each processor core may be configured such that by storing only specific pseudo-channel regions in the external memory address region, there is only one copy stored in the HBM memory in the artificial intelligence processor chip. That is, only one copy is stored in each cache, and thus all caches in the artificial intelligence processor may be configured at one level. The non-overlapping HBM memory regions in this way are stored in the corresponding caches, thereby naturally supporting the cache coherence mechanism and preventing performance degradation caused by the processing of the cache coherence mechanism.

[0039] In this case, the external memory address region may be set considering the programs running on the plurality of processor cores and the access distances between the plurality of processor cores and the plurality of caches.

[0040] Here, the external memory address region may be set such that the data required to run each program is stored in the cache with the shortest access distance to each of the plurality of processor cores.

[0041] Therefore, each processor core according to the present disclosure can access the cache closest to it and then use the data required to run the program.

[0042] For example, each of the plurality of processor cores can access the cache closer to it (= the cache with a shorter access distance) or access the cache farther from it (= the cache with a longer access distance), so as to read the data stored at the required address and write data to it.

[0043] Here, the bus existing between the plurality of processor cores and the plurality of caches may set the external memory address region based on the memory mapping information related to the plurality of caches.

[0044] For example, a bus according to an embodiment of the present disclosure may set HBM memory regions stored in respective caches, and may manage the regions set in this way in a form corresponding to configurable memory mapping information by the bus.

[0045] In addition, although not shown in Figure 2 , in a method for supporting cache coherence based on virtual addresses for an artificial intelligence processor having a large-capacity on-chip memory according to an embodiment of the present disclosure, the artificial intelligence processor detects programs running on multiple processor cores.

[0046] In addition, in a method for supporting cache coherence based on virtual addresses for an artificial intelligence processor having a large-capacity on-chip memory according to an embodiment of the present disclosure, the artificial intelligence processor including multiple processor cores and multiple caches may provide virtual addresses, where the multiple processor cores use the virtual addresses to access the multiple caches.

[0047] Here, the virtual address may be provided as corresponding to an exclusive address, such that the multiple processor cores may access the same data simultaneously.

[0048] Here, the bus may obtain mapping information between physical addresses and virtual addresses corresponding to the multiple caches based on a translation lookaside buffer (TBL) that can be set by a system controller of the artificial intelligence processor, and may provide the virtual addresses to the multiple processor cores in consideration of the mapping information.

[0049] For example, a bus of an artificial intelligence processor according to an embodiment of the present disclosure may obtain mapping information from a virtual address set by the system controller to a physical address through a translation lookaside buffer (TLB) in a state where the translation lookaside buffer (TLB) that can be set by the system controller is inserted between the processor core and the cache. For example, as Figure 5 shown, mapping information between the virtual address (VAx) of each processor core (ANCx) and the physical address (PAx) of the cache may be obtained.

[0050] Here, physical address regions converted by respective translation lookaside buffers (TLBs) may be exclusive of each other.

[0051] In this case, for each piece of data to be stored in the multiple caches, the cache in which the corresponding data is to be stored may be determined in consideration of data dependencies, where when the data dependencies are high, the corresponding data may be stored in a cache with a relatively short cache distance corresponding to the distance between the caches.

[0052] For example, for Figure 3The cache 0 300-1 shown in [description] can pre-compute the cache distances to cache 1 300-2 and cache 300-3, and can store the access times depending on the computed cache distances respectively. In this way, for cache 1 300-2 and cache 3 300-3, the corresponding cache distances and access times can also be computed and stored. Thereafter, the data dependencies of multiple pieces of data to be stored in the corresponding caches can be checked. It can be determined that the data with relatively high data dependencies will be stored in the cache with a relatively short cache distance, so that other caches can access the data quickly.

[0053] By a method for supporting cache coherence based on virtual addresses for an artificial intelligence processor with a large-capacity on-chip memory, a method capable of maintaining cache coherence among multiple caches in an artificial intelligence processor with a large-capacity on-chip memory can be provided.

[0054] In addition, a structure for supporting virtual addresses can be provided, which allows multiple processor cores to process different multiple pieces of data based on the same instruction configuration.

[0055] In addition, software running on multiple processor cores can be called quickly.

[0056] Figure 3 and Figure 4 is a diagram showing an example of the structure of an artificial intelligence processor for supporting cache coherence and virtual addresses according to the present disclosure.

[0057] First, referring to Figure 3 , an artificial intelligence processor for supporting cache coherence and virtual addresses according to the present disclosure may include multiple processor cores 320-1 to 320-3, multiple caches 300-1 to 300-3, a system controller 310, and a bus 300 disposed between the multiple processor cores and the multiple caches.

[0058] Here, although a structure including three processor cores and three caches is shown in Figure 3 , this structure may actually include more processor cores and more caches.

[0059] The bus 300 may be disposed between the multiple processor cores 320-1 to 320-3 and the multiple caches 300-1 to 300-3.

[0060] Hereinafter, the process of supporting cache coherence based on virtual addresses through the bus 300 will be described in detail.

[0061] The bus 300 sets non-overlapping external memory address regions for the corresponding multiple caches.

[0062] Here, the external memory address region may correspond to the address region of a high-bandwidth memory (HBM) located outside the artificial intelligence processor. For example, the external memory address region may correspond to the HBM memory pseudo-channel region. In Figure 3 this, these regions are shown as corresponding to the HBM memory pseudo-channel controllers (PCx) 340-1 to 340-3.

[0063] Accordingly, the caches connected to the respective processor cores may be configured such that by storing only specific pseudo-channel regions in the external memory address region, only one copy of the data stored in the HBM memory exists in the artificial intelligence processor chip. That is, only one copy is stored in each cache, and thus all the caches in the artificial intelligence processor may be configured at one level. The HBM memory regions that do not overlap with each other in this way are stored in the corresponding caches, thereby naturally supporting the cache coherence mechanism and preventing performance degradation caused by the processing of the cache coherence mechanism.

[0064] In this case, the external memory address region may be set in consideration of the programs running on the multiple processor cores and the access distances between the multiple processor cores and the multiple caches.

[0065] Here, the external memory address region may be set such that the data required to run each program is stored in the cache with the shortest access distance to each of the multiple processor cores.

[0066] Accordingly, each processor core according to the present disclosure may access the cache closest to it and then use the data required to run the program.

[0067] For example, each of the multiple processor cores may access the cache closer to it (= the cache with a shorter access distance) or access the cache farther from it (= the cache with a longer access distance) to read the data stored at the required address and write data to it.

[0068] Here, the bus existing between the multiple processor cores and the multiple caches may set the external memory address region based on the memory mapping information related to the multiple caches.

[0069] For example, the bus according to an embodiment of the present disclosure may set the HBM memory regions stored by the respective caches and may manage the regions set in this way in a form corresponding to the configurable memory mapping information.

[0070] In addition, the bus 300 detects the programs running on the multiple processor cores.

[0071] In addition, the bus 300 can provide virtual addresses, where multiple processor cores use these virtual addresses to access multiple caches.

[0072] Here, the virtual addresses can be provided as corresponding to exclusive addresses, such that multiple processor cores can access the same data simultaneously.

[0073] Here, the bus can obtain mapping information between physical addresses and virtual addresses corresponding to multiple caches based on a translation lookaside buffer (TBL) that can be set by the system controller of the artificial intelligence processor, and can provide the virtual addresses to multiple processor cores considering the mapping information.

[0074] For example, referring to Figure 4 , the bus 400 of the artificial intelligence processor according to an embodiment of the present disclosure can obtain mapping information from a virtual address set by the system controller to a physical address through translation lookaside buffers (TLBs) 410 to 430 in a state where the translation lookaside buffers (TLBs) set by the system controller are inserted between the processor core and the cache. For example, as Figure 5 shown, mapping information between the virtual address (VAx) for each processor core (ANCx) and the physical address (PAx) of the cache can be obtained.

[0075] Here, the physical address regions converted by each translation lookaside buffer (TLB) can be exclusive of each other.

[0076] In this case, for each piece of data to be stored in multiple caches, the cache in which the corresponding data will be stored can be determined considering data dependencies. When the data dependencies are high, the corresponding data can be stored in a cache with a relatively short cache distance corresponding to the distance between the caches.

[0077] Through the artificial intelligence processor, a method capable of maintaining cache coherence between multiple caches in an artificial intelligence processor having a large-capacity on-chip memory can be provided.

[0078] In addition, a structure for supporting virtual addresses can be provided, and this structure allows multiple processor cores to process different multiple pieces of data based on the same instruction configuration.

[0079] In addition, software running on multiple processor cores can be called quickly.

[0080] As described above, in the method for supporting cache coherence based on virtual addresses for an artificial intelligence processor having a large-capacity on-chip memory according to the present disclosure and the device for the method, the configurations and solutions in the above embodiments are unrestrictedly applied, and some or all of the above embodiments can be selectively combined and configured, so that various modifications can be made.

Claims

1. A method for supporting cache coherence based on virtual addresses, comprising: Performing, by an artificial intelligence processor including a plurality of processor cores and a plurality of caches, the following operations: Setting non-overlapping external memory address regions for the corresponding plurality of caches; And Providing virtual addresses, wherein the plurality of processor cores access the plurality of caches using the virtual addresses.

2. The method according to claim 1, wherein, The setting step includes: Setting the external memory address regions in consideration of programs running on the plurality of processor cores and access distances between the plurality of processor cores and the plurality of caches.

3. The method according to claim 2, wherein The setting step further includes: Setting the external memory address regions such that data required to run each program is stored in the cache with the shortest access distance to each of the plurality of processor cores.

4. The method according to claim 3, wherein, The setting step further includes: Setting the external memory address regions based on memory mapping information related to the plurality of caches through a bus existing between the plurality of processor cores and the plurality of caches.

5. The method according to claim 2, further comprising: Detecting, by the artificial intelligence processor, the programs running on the plurality of processor cores.

6. The method according to claim 1, wherein The virtual addresses are provided as corresponding to exclusive addresses such that the plurality of processor cores can access the same data simultaneously.

7. The method according to claim 4, wherein The bus is configured to obtain mapping information between physical addresses corresponding to the plurality of caches and the virtual addresses based on a table bypass buffer TBL that can be set by a system controller of the artificial intelligence processor, and provide the virtual addresses to the plurality of processor cores in consideration of the mapping information.

8. The method according to claim 1, wherein The external memory address regions correspond to address regions of a high-bandwidth memory HBM located outside the artificial intelligence processor.

9. The method according to claim 1, wherein, For each piece of data to be stored in the plurality of caches, determine the cache in which the corresponding data will be stored in consideration of data dependencies, and the cache is determined such that when the data dependencies are high, the corresponding data is stored in the cache with a relatively short cache distance corresponding to the distance between caches.

10. An artificial intelligence processor, comprising: A plurality of processor cores; A plurality of caches corresponding to the plurality of processor cores; A system controller SCP; And A bus located between the plurality of processor cores and the plurality of caches, wherein the bus is configured to set non-overlapping external memory address regions for the corresponding plurality of caches and provide virtual addresses, wherein the plurality of processor cores access the plurality of caches using the virtual addresses.

11. The artificial intelligence processor according to claim 10, wherein, The bus is configured to set the external memory address regions in consideration of programs running on the plurality of processor cores and access distances between the plurality of processor cores and the plurality of caches.

12. The artificial intelligence processor according to claim 11, wherein, The bus is configured to set the external memory address regions such that data required to run each program is stored in the cache with the shortest access distance to each of the plurality of processor cores.

13. The artificial intelligence processor according to claim 12, wherein, The bus is configured such that the bus existing between the multiple processor cores and the multiple caches sets the external memory address region based on memory mapping information related to the multiple caches.

14. The artificial intelligence processor according to claim 11, wherein, The bus is configured to detect the program running on the multiple processor cores.

15. The artificial intelligence processor according to claim 10, wherein, The virtual address is provided to correspond to an exclusive address, enabling the multiple processor cores to access the same data simultaneously.

16. The artificial intelligence processor according to claim 13, wherein, The bus is configured to obtain mapping information between the physical address corresponding to the multiple caches and the virtual address based on a table bypass buffer TBL that can be set by the system controller of the artificial intelligence processor, and provide the virtual address to the multiple processor cores in consideration of the mapping information.

17. The artificial intelligence processor according to claim 10, wherein, The external memory address region corresponds to the address region of a high-bandwidth memory HBM located outside the artificial intelligence processor.

18. The artificial intelligence processor according to claim 10, wherein, For each piece of data to be stored in the multiple caches, the cache in which the corresponding data is to be stored is determined considering data dependencies, and the cache is determined such that when the data dependencies are high, the corresponding data is stored in a cache with a relatively short cache distance corresponding to the distance between caches.

Citation Information

Patent Citations

  • Photosensitive resin composition, photosensitive resin layer using the same and color filter

    KR1020230088228A

Cited By

  • Memory allocation method and device, chip, electronic equipment and storage medium

    CN121979808A