Cache system, cache processing method, electronic device and storage medium
By dividing the cache area according to the multi-level cache configuration information of the processing core in the many-core system, a flexible multi-level cache system is constructed, which solves the problem of mismatch between processor computing speed and memory access speed, and improves data access efficiency and system performance.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2026-04-02
AI Technical Summary
In many-core systems, the processor's computing speed and memory access speed are mismatched, affecting computing performance and system power consumption. The fixed design of existing multi-level cache systems limits system performance and flexibility.
Based on the multi-level cache configuration information of the processing cores, the first cache area is divided from the target storage areas of multiple processing cores to form a multi-level cache system. Different levels of cache areas can be flexibly configured to meet the task requirements of the processing cores and improve computing and caching performance.
By flexibly configuring a multi-level caching system, data access speed and efficiency can be improved, hardware resource consumption can be saved, and overall computing and system performance can be enhanced.
Smart Images

Figure CN2025122180_02042026_PF_FP_ABST
Abstract
Description
Cache system, cache processing method, electronic device and storage medium TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, and particularly relates to a cache system, a cache processing method, an electronic device and a storage medium. BACKGROUND
[0002] A many-core system can process large-scale data and complex tasks in parallel, and provides efficient computing power for application fields such as scientific computing, artificial intelligence and big data analysis, but the computing speed of the processor and the memory access speed may not match, affecting the computing performance. SUMMARY
[0003] The present disclosure provides a cache system, a cache processing method, an electronic device and a storage medium.
[0004] In a first aspect, the present disclosure provides a cache system, comprising a plurality of processing cores and a configuration module, the processing core comprising a target storage area:
[0005] The configuration module is configured to determine multi-level cache configuration information of a first processing core, and divide a first cache area from the target storage area of the plurality of processing cores according to the multi-level cache configuration information, the first cache area divided from the plurality of processing cores constituting a multi-level cache system of the first processing core, and the processing core where the first cache area of different levels is located being related to the position information of the first processing core, the first processing core being any one of the plurality of processing cores.
[0006] The first processing core is configured to, when receiving a data access request, search and acquire data indicated in the data access request from the multi-level cache system of the first processing core.
[0007] In a second aspect, the present disclosure provides a cache processing method applied to the cache system in the embodiments of the present disclosure, and the method comprises:
[0008] Determining multi-level cache configuration information of a first processing core;
[0009] Dividing a first cache area from the target storage area of the plurality of processing cores according to the multi-level cache configuration information, the first cache area divided from the plurality of processing cores constituting a multi-level cache system of the first processing core, and the processing core where the first cache area of different levels is located being related to the position information of the first processing core.
[0010] In a third aspect, the present disclosure provides an electronic device, comprising: a plurality of processing cores; and an on-chip network configured to interact data between the plurality of processing cores and external data; wherein one or more instructions are stored in one or more of the processing cores, and the one or more instructions are executed by the one or more processing cores to enable the one or more processing cores to perform the cache processing method in the embodiments of the present disclosure.
[0011] In a fourth aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the cache processing method in the embodiments of the present disclosure.
[0012] In a fifth aspect, the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, implements the cache processing method in the embodiments of the present disclosure.
[0013] According to the multi-level cache configuration information of the first processing core, the first cache area is divided from the target storage area of the plurality of processing cores, so that the plurality of different levels of first cache areas divided constitute the multi-level cache system of the first processing core. In this way, the multi-level cache system of the first processing core can be flexibly configured according to the multi-level cache configuration information of the first processing core. Different processing cores can more specifically and flexibly configure the multi-level cache system, better meet the needs of different processing cores, thereby improving the overall computing and caching performance. Moreover, the multi-level cache system is configured based on the target storage area of the processing core itself, without the need for additional design of other hardware resources, saving hardware resource consumption. When the first processing core receives a data access request, the data access speed and efficiency can be improved by searching and obtaining from the multi-level cache system of the first processing core.
[0014] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0015] The accompanying drawings are included to provide a further understanding of the present disclosure and constitute a part of the specification, which together with the embodiments of the present disclosure are used to explain the present disclosure and do not constitute a limitation of the present disclosure. The above and other features and advantages will become more apparent to those skilled in the art from the detailed description of the specific example embodiments, with reference to the accompanying drawings, in which:
[0016] FIG. 1 is a schematic diagram of a hardware architecture of a many-core system according to an embodiment of the present disclosure;
[0017] FIG. 2 is a schematic diagram of a hardware architecture of another many-core system according to an embodiment of the present disclosure;
[0018] FIG. 3 is a schematic diagram of a cache system according to an embodiment of the present disclosure;
[0019] FIG. 4 is a schematic diagram of a target storage area according to an embodiment of the present disclosure;
[0020] FIG. 5 is a multi-level cache system according to an embodiment of the present disclosure;
[0021] FIG. 6 is a schematic diagram of a multi-cache area and a cache page table according to an embodiment of the present disclosure;
[0022] FIG. 7 is a schematic diagram of a multi-level cache system according to an embodiment of the present disclosure;
[0023] FIG. 8 is a schematic diagram of a multi-level cache system according to an embodiment of the present disclosure;
[0024] FIG. 9 is a schematic diagram of a multi-level cache system according to an embodiment of the present disclosure;
[0025] FIG. 10 is a flowchart of a cache processing method according to an embodiment of the present disclosure;
[0026] FIG. 11 is a block diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the exemplary embodiments of the present disclosure are described below in conjunction with the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered only as exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of clarity and conciseness, the description below omits the description of well-known functions and structures.
[0028] In the case of no conflict, each embodiment of the present disclosure and each feature in the embodiments can be combined with each other.
[0029] As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items.
[0030] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms "comprises" and / or "comprising," when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. "Coupled" or "connected" or similar terms are not restricted to physical or mechanical connections or associations, but can also include electrical connections, whether direct or indirect.
[0031] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art. It will be further understood that terms, such as those defined in commonly used dictionaries, should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted in an overly literal or overly formal sense unless expressly so defined herein.
[0032] For the ease of understanding, the concepts involved in the embodiments of the present disclosure are briefly described first.
[0033] Many-core system: A many-core system is a high-performance computing architecture that integrates multiple processor cores on a single chip to improve performance and energy efficiency through parallel computing. A processing core is an independent computing unit within a processor, which can execute its own instruction stream and has its own register set and other hardware resources. With the development of multi-core technology, a processor can include multiple processing cores, so that a processor can execute multiple tasks or threads simultaneously.
[0034] Multi-level cache system: The speed difference between the processor and the main memory can be alleviated by providing different levels of storage, thereby improving the overall system performance. In a multi-level cache system, the cache area can be divided into multiple levels according to access speed, memory size, etc. For example, in one possible embodiment, the multi-level cache system includes: Level 1 cache (L1 cache): usually a cache layer close to the processor, with faster access speed, which can be used to store data and instructions to be executed, etc., and due to cost and physical space limitations, the size of L1 cache is usually relatively small; Level 2 cache (L2 cache): L2 cache usually has a larger capacity than L1 cache, but the access speed is relatively slower than L1 cache; Level 3 cache (L3 cache): L3 cache has a larger capacity than L2 cache and L1 cache, but the access speed is relatively slow, and L3 cache can also be used to store shared data between multiple processing cores.
[0035] The many-core system can process large-scale data and complex tasks in parallel, but the computing speed of the processor and the memory access speed do not match, which affects the computing performance and system power consumption, and in the related technology, a multi-level cache technology can be used, and in the hardware design, the positions, sizes and link relations of the levels of cache are fixed in the processor, that is, the designed multi-level cache system is part of the chip hardware logic circuit, and is provided for all processing cores to use, which requires additional division of hardware resources from the processor, and the execution tasks of different processing cores in the many-core system can be different, and different tasks can have different requirements for the multi-level cache system, and the fixed multi-level cache system can also limit the performance of the many-core system.
[0036] According to the cache system and the cache processing method provided in the embodiments of the present disclosure, the first cache area can be divided from the target storage area of the plurality of processing cores according to the multi-level cache configuration information of the first processing core, so that the plurality of different levels of first cache areas divided, that is, the multi-level cache system of the first processing core, so that the multi-level cache system of the processing core can be flexibly configured according to the multi-level cache configuration information, and a multi-level cache system more suitable for the task executed by the processing core can also be obtained, which is more targeted and flexible, thereby improving the overall performance, and the multi-level cache system is configured based on the resource configuration of the target storage area of the processing core itself, without occupying other hardware resources in the processor, saving hardware resource consumption, and when the first processing core receives a data access request, the multi-level cache system of the first processing core can be searched and obtained, improving the data access speed and efficiency.
[0037] According to the cache system and the cache processing method provided in the embodiments of the present disclosure, the cache system and the cache processing method can be applied to the many-core system, the many-core system includes a plurality of processing cores, and based on the hardware architecture of the processing core of the many-core system, the multi-level cache system of the processing core can be flexibly configured, and the cache processing method can be executed by the many-core system.
[0038] FIG. 1 is a schematic diagram of a hardware architecture of a many-core system provided by an embodiment of the present disclosure.
[0039] As shown in FIG. 1, the many-core system includes a plurality of processing cores, and each processing core has a private target storage area, for example, the target storage area is a static random access memory (SRAM), and the plurality of processing cores can communicate through an on-chip network.
[0040] Based on the many-core system shown in FIG. 1, in the embodiments of the present disclosure, a corresponding multi-level cache system can be configured according to the requirements of the processing core, different processing cores perform different computing tasks in parallel, and specifically, the present disclosure provides a cache system, which comprises: a plurality of processing cores and a configuration module, the processing core comprises a target storage area, the plurality of processing cores are processing cores on a many-core chip, wherein:
[0041] 1) the configuration module is configured to determine the multi-level cache configuration information of the first processing core, and divide the first cache area from the target storage area of the plurality of processing cores according to the multi-level cache configuration information, the plurality of processing cores correspond to the divided first cache area, which constitutes the multi-level cache system of the first processing core, and the processing cores at different levels of the first cache area are related to the position information of the first processing core, and the first processing core is any one of the plurality of processing cores.
[0042] In the embodiments of the present disclosure, the target storage area is located in the private memory of the processing core, for example, as shown in FIG. 1, the target storage area can be the private SRAM of the processing core.
[0043] Or the target storage area is located in the local shared storage area or the global shared storage area of the plurality of processing cores, for example, as shown in FIG. 2, which is a hardware architecture diagram of another many-core system in the embodiments of the present disclosure, as shown in FIG. 2, the plurality of processing cores can also have a local shared storage area and a global shared storage area, and when dividing, the first cache area can also be divided from the local shared storage area or the global shared storage area according to the multi-level cache configuration information, for example, the first cache in the multi-level cache system of the processing core 0 can be divided from the local shared storage area corresponding to the processing core 0.
[0044] Among them, for determining the multi-level cache configuration information of the first processing core, the present disclosure provides several possible embodiments:
[0045] In one possible embodiment, the configuration module is configured to determine the multi-level cache configuration information of the first processing core according to the task information of the task executed by the first processing core, and the storage resource information of the target storage area in the plurality of processing cores, and the position information between the first processing core and the plurality of processing cores.
[0046] In the embodiments of the present disclosure, the task information can include the data volume, the computing speed and other information associated with the task, and the position information between the first processing core and the plurality of processing cores represents the hardware position relationship, the distance between the positions and other information between the processing cores in the many-core system.
[0047] In this way, in the embodiments of the present disclosure, the multi-level cache system of the processing core can be dynamically configured according to the requirements of the task executed by the processing core, which is more flexible and can better meet the task requirements of each processing core, thereby improving the system performance.
[0048] In a possible embodiment, the configuration module is configured to: acquire an input configuration instruction, and determine the multi-level cache configuration information of the first processing core indicated in the configuration instruction.
[0049] That is, in the embodiments of the present disclosure, the multi-level cache configuration information of the plurality of processing cores can also be pre-configured by the user. Specifically, the configuration can also be performed in combination with the execution task demand of the processing core, the hardware position layout of the processing core in the many-core system, the storage resource information of the target storage area in the processing core, and the like, and the configuration is comprehensively considered, and then, for example, in the compilation phase, the configuration instruction containing the multi-level cache configuration information of the first processing core can be input, that is, the multi-level cache system of the first processing core can be configured according to the multi-level cache configuration information.
[0050] Further, in the embodiments of the present disclosure, the multi-level cache configuration information of the first processing core can also be determined in combination with the above two embodiments, for example, the multi-level cache system of the first processing core can be statically determined by the compiler according to the configuration instruction in the compilation phase, and then the multi-level cache system of the first processing core can be configured, which can be used as the initial design state of the multi-level cache system of the first processing core, and then in the running phase, the multi-level cache system of the first processing core can be dynamically determined and dynamically adjusted according to the task information of the first processing core, the current storage resource information of the target storage area of the processing core, and the like, so that the flexibility and performance of the design of the multi-level cache system of different processing cores can be further improved.
[0051] In the embodiments of the present disclosure, the multi-level cache configuration information includes the number of levels of the multi-level cache system, the positions of the plurality of level cache areas, and the sizes of the plurality of level cache areas, and when the first cache area is divided from the target storage area of the plurality of processing cores according to the multi-level cache configuration information, the configuration module is configured to: according to the number of levels, the processing cores corresponding to the positions of the plurality of level cache areas, and the sizes of the plurality of level cache areas, divide the first cache area from the target storage area of the processing core corresponding to the position of the plurality of level cache areas according to the sizes of the plurality of level cache areas, and determine the linking information between the plurality of first cache areas.
[0052] The plurality of first cache areas are the cache areas of different levels in the multi-level cache system, and the first cache areas of the same level can be located in the target storage area of one or more processing cores, and the first cache areas of different levels can also be located in the target storage area of the same processing core or different processing cores.
[0053] For example, referring to FIG. 3, a schematic diagram of a cache system in an embodiment of the present disclosure is shown. As shown in FIG. 3, the first processing core is processing core 0, the number of levels of the multi-level cache system of the processing core 0 is four, and the first-level cache area is located in the target storage area of the processing core 0. The second-level cache area is divided into two, and is located in the target storage area of the processing core 3 and the processing core 4. The third-level cache area is divided into three, and is located in the target storage area of the processing core 5, the processing core 6, and the processing core 7. The fourth-level cache area is located in the target storage area of the processing core 8. The link information between the cache areas of different levels is the link relationship indicated by the arrows in FIG. 3.
[0054] 2) The first processing core is configured to, when receiving a data access request, search and acquire data indicated in the data access request from the multi-level cache system of the first processing core.
[0055] Specifically, the present disclosure provides a possible embodiment. According to the data address in the data access request, the first-level cache area included in the multi-level cache system of the first processing core is queried. If the query is successful, the data indicated in the data access request is acquired from the first-level cache area. If the query is unsuccessful, according to the link information corresponding to the first-level cache area, the next-level cache area linked with the first-level cache area is queried. The target cache area that is queried is determined, and the data indicated in the data access request is acquired from the target cache area.
[0056] For example, as shown in FIG. 3, the data access request of the processing core 0 first accesses the first-level cache area in the multi-level cache system of the processing core 0, that is, the L1 cache deployed in the processing core 0. If the query in the L1 cache is unsuccessful, the link information in the L1 cache is queried. It is determined that the next-level L2 cache having the link relationship is located in the processing core 3 and the processing core 4. The processing core 0 simultaneously sends the data address of the data to be queried to the processing core 3 and the processing core 4. If one of the processing core 3 and the processing core 4 queries successfully, the queried data and the query hit information are returned to the processing core 0. Then, the processing core 0 sends the termination search information to the processing core 3 and the processing core 4. If neither of the processing core 3 and the processing core 4 queries successfully, the next-level cache area can be continuously accessed according to the above steps.
[0057] In the embodiments of the present disclosure, a flexible multi-level cache system configuration is provided for a many-core system. The first cache area can be divided from the target storage area of the plurality of processing cores according to the multi-level cache configuration information of the first processing core. The plurality of processing cores correspond to the divided first cache area, and the first processing core is configured with a multi-level cache system. In this way, the required multi-level cache system can be flexibly configured according to the needs of different processing cores, which not only improves the performance and flexibility of the many-core system, but also improves the performance of the multi-level cache system. In this way, when the first processing core receives a data access request, the corresponding data can be found and obtained from the multi-level cache system of the first processing core, and the data acquisition efficiency can be improved.
[0058] The cache system in the embodiments of the present disclosure is further described below.
[0059] In the embodiments of the present disclosure, the first cache area divided from the target storage area of the processing core can be further divided to further subdivide the storage of different information and improve the data storage and reading efficiency. In a possible embodiment, the first cache area includes a link area and a cache data area. The cache data area is used to store target data of a task executed by the first processing core. The link area is used to store link information of a next-level cache area corresponding to a current-level cache area. The link information includes a position of the next-level cache area and a data address range of target data stored in the next-level cache area.
[0060] In addition, in order to identify the data address of the target data stored in the cache area, the cache system in the embodiments of the present disclosure is further configured with a tag area. In a possible embodiment, the tag area is located in the first cache area of the first processing core. The tag area is used to store the data address of the target data stored in the corresponding first cache area.
[0061] In a possible embodiment, the tag area is located in a shared area of the plurality of processing cores. The tag area is used to store the data address of the target data stored in the first cache area of the plurality of processing cores associated with the shared area.
[0062] In the embodiments of the present disclosure, the tag area and the cache data area can be configured in the same processing core to improve the query efficiency. In addition, the tag area and the cache data area can also be configured in different processing cores. In this way, the cache data areas of different processing cores can share one tag area, so that the storage caused by the tag can be saved. However, it should be noted that in the case that the plurality of processing cores share the tag area, the position of the tag area corresponding to the cache data area needs to be recorded in the link area in the cache area of the processing core, so that the data address of the target data stored in the cache data area can be accurately found.
[0063] Further, the target storage area in the embodiments of the present disclosure can be divided into the first cache area and a random access area, and the random access area is used to store the data allocated statically at compile time.
[0064] For example, referring to FIG. 4, which is a schematic diagram of the division of a target storage area in the embodiments of the present disclosure, and taking the case that the tag area is located in the first cache area of the processing core as an example, as shown in FIG. 4, the target storage area can be divided into the first cache area and the random access area, the first cache area can be further divided into the link area, the tag area and the cache data area, the tag area is used to store the data address, the cache data area is used to store the target data, and the link area is used to store the link information of the next level cache area corresponding to the cache area.
[0065] In this way, in the embodiments of the present disclosure, the target storage area of the processing core is divided into the cache area and the random access area, for the data stored in the random access area, the processing core can directly access the data in the random access area through the address, and the access to the cache area can be performed through the Cache mechanism, and the data in the cache area can be missing, at this time, if the data is not queried from the multi-level cache system of the processing core, the data can be queried from the random access area or the main memory outside the processing core, and it can be ensured that the required data can be obtained.
[0066] It should be noted that the configuration of the multi-level cache system in the embodiments of the present disclosure is not limited, and several possible embodiments are further provided.
[0067] In one possible embodiment, the cache area at the same level includes a plurality of sub-cache areas, and the plurality of sub-cache areas are located in the target storage areas of different processing cores, the link information between the first cache areas at different levels in the multi-level cache system is in a tree structure or an inverse tree structure, and the link information of the next level cache area stored in the link information is the link information of the sub-cache area in the next level cache area which has a connection edge in the tree structure or the inverse tree structure.
[0068] In the embodiments of the present disclosure, for example, the remaining storage resources in the target storage area of one processing core can be insufficient, and the access speeds of the cache areas based on different levels are different, and the position distribution has a near-far requirement, so the cache area of the same level can be split into multiple sub-cache areas, and respectively deployed to the target storage areas of different processing cores, so as to improve the configuration flexibility of the multi-level cache system. However, there can be many link relationships between two adjacent levels, and the data query can be too complex, so in the case that the cache area of the same level includes multiple sub-cache areas, the multi-level cache system in the embodiments of the present disclosure can adopt a tree structure or an inverse tree structure, so that only the link information of some sub-cache areas in the next level needs to be stored in the link area in the processing core, instead of the link information of all sub-cache areas in the next level, so as to improve the data query and access efficiency.
[0069] For example, referring to FIG. 5, a multi-level cache system in a tree structure or an inverse tree structure in the embodiments of the present disclosure is shown. As shown in FIG. 5, assuming that the L2 cache is distributed on 2 processing cores, the L3 cache is distributed on 4 processing cores, and the L4 cache is distributed on 2 processing cores, there are 8 cache connection edges between L2 and L3, and the cache search signal explosion can occur. Therefore, the multi-level cache system in a tree structure or an inverse tree structure is adopted, and the link area of the processing core only needs to record the link information of the connected sub-cache areas in the next level cache area, instead of all sub-cache areas corresponding to the next level cache area. For example, in the tree structure shown in FIG. 7, the link area of the second-level sub-cache area 1 corresponding to the L2 cache only needs to store the link information of the third-level sub-cache area 1 and the third-level sub-cache area 2 corresponding to the L3 cache. As shown in FIG. 7, the inverse tree structure can also be adopted, that is, one-to-many. At this time, the link area of one sub-cache area only needs to record the link information of the sub-cache areas having the connection relationship in the next level cache area.
[0070] In a possible embodiment, the target storage area of the processing core includes one or more first cache areas, and the plurality of first cache areas represent the cache areas of different levels in the multi-level cache system of the first processing core. Alternatively, the target storage area of the processing core further includes a second cache area in the multi-level cache system of another processing core. The link area is further used to store a cache page table, and the cache page table is used to record the core identifier of the processing core to which the plurality of cache areas included in the target storage area of the processing core belong, and the level identifier.
[0071] For example, referring to FIG. 6, a schematic diagram of a plurality of cache regions and a cache page table in an embodiment of the present disclosure is shown. As shown in FIG. 6, a target storage region of a processing core can be divided into a plurality of cache regions and random access regions, and the plurality of cache regions can include cache region 1, cache region 2, and cache region 3. The plurality of cache regions can be in a multi-level cache system centered on different processing cores, or can be different levels of a multi-level cache system of a processing core. For example, cache region 1 is a second-level cache region in a multi-level cache system of processing core 0, cache region 2 is a third-level cache region in a multi-level cache system of processing core 0, and cache region 3 is a second-level cache region in a multi-level cache system of processing core 1.
[0072] Further, in an embodiment of the present disclosure, a cache page table can be used in a processing core to record which level of a multi-level cache region of a processing core a plurality of cache regions correspond to. In this way, using a cache page table to record is more clear and simple, and when data access is queried, the correct cache region can be quickly indicated for querying, thereby improving efficiency.
[0073] In a possible embodiment, in an embodiment of the present disclosure, the multi-level cache system can also be configured based on the hardware architecture of a plurality of processing cores in a many-core system. Specifically, the plurality of processing cores are divided into an m*n grid according to position information, and at least one processing core is included in the grid. The processing cores in the same grid share the same level of cache region. A configuration module is configured to:
[0074] 1) According to the divided m*n grid, taking a first grid including a first processing core as the center, a first cache region is divided from a target storage region of the processing core included in the first grid to serve as a first-level cache region. 2) According to a preset adjacent routing rule, a second grid corresponding to the adjacent routing rule is determined, and a first cache region is divided from a target storage region of the processing core included in the second grid to serve as a second-level cache region. The first grid is continuously taken as the center until the number of levels is satisfied, and a multi-level cache system of the first processing core is obtained. The multi-level cache system at least includes a first-level cache region and a second-level cache region.
[0075] The target storage region of the processing core where the first-level cache region and the second-level cache region are located satisfies a storage resource condition. For example, if a certain processing core used to divide a second-level cache region is determined according to the adjacent routing rule, and the target storage region of the processing core has less storage resources, or the to-be-stored data of the random access region of the processing core itself exceeds a certain threshold, it can be insufficient to divide a cache region of appropriate size. Therefore, the processing core can not be used to divide a cache region, and the random access region in the processing core can be used normally.
[0076] The adjacent routing rules can be adjacent grids in up, down, left and right directions or adjacent grids in diagonal directions, and the disclosure embodiments are not limited in this way. It can be understood that the smaller the number of levels in the multi-level cache system of the first processing core, the closer the grid of the processing core to the first grid.
[0077] For example, referring to FIG. 7, which is a schematic diagram of a multi-level cache system in a grid division in the disclosure embodiments, a plurality of processing cores in a many-core system are divided into an m*n grid, as shown in FIG. 7, each box represents a processing core, and in each grid, two processing cores are taken as an example, G i,j represents all the processing cores in the i-th row and the j-th column of the grid, G i,j All the processing cores in G 3,2 can share the cache area of the same level, for example, in FIG. 7, G 3,2 is taken as the center, and the two processing cores in G i,j share the L1 cache.
[0078] And, G i,j is taken as the center, and the target storage area of the processing cores in G i,j is divided into a first-level cache area, and G i-1,j has a one-hop distance from the grid, and can be taken as a second grid that meets the adjacent routing rules, for example, the second grid includes G i+1,j , G i,j-1 , G i,j+1 As shown in FIG. 7, in order to more clearly describe the cache areas of different levels, different gray colors of boxes are used to distinguish the cache areas of different levels in FIG. 7, for example, in FIG. 7, G 3,2 The second grid corresponding to the grid is an adjacent grid in up, down, left and right directions, and a second-level cache area is divided from the target storage area of the processing cores included in the second grid, and similarly, G i,j is a third grid that is adjacent to the second grid, for example, as shown in FIG. 7, G 3,2 The third grid corresponding to the grid can be determined, for example, the third grid includes G i-1,j-1 , G i+1,j-1 , G i-1,j+1 , G i-1,j+1 , and a third-level cache area is divided from the target storage area of the processing cores included in the third grid, and further, a multi-level cache system centered on G i,j can be obtained, and the processing cores in G i,j can correspond to a same set of multi-level cache systems.
[0079] For another example, referring to FIG. 8, which is a schematic diagram of another multi-level cache system in a grid division in the disclosure embodiments, for example, G i,j is taken as the center, G i,jThe first cache area is divided in the target storage area of the middle processing core. If the second network meeting the adjacent routing rule includes grids in the up, down, left, right and diagonal positions, and the determined second grid includes G i-1,j , G i+1,j , G i,j-1 , G i,j+1 , G i-1,j-1 , G i+1,j-1 , G i-1,j+1 , G i-1,j+1 , and the determined third grid includes G i-2,j-1 , G i-1,j , G i-2,j+1 , G i+2,j-1 , G i+2,j , G i+2,j+1 .
[0080] In addition, in the embodiment of the disclosure, when determining the other grids meeting the adjacent routing rule corresponding to the first grid, if it is determined that the storage resources of the target storage area of the processing core in the other grids meeting the adjacent routing rule do not meet the requirements, or the to-be-stored data of the random storage area in the other grids itself exceeds a certain threshold, the other grids meeting the adjacent routing rule can not be used for cache division. For example, referring to FIG. 9, which is a schematic diagram of another grid-shaped multi-level cache system in the embodiment of the disclosure, G 3,2 is taken as the center, if G 3,1 and G 4,3 do not meet the storage resource conditions, they can not be used, and the second cache area or the third cache area is divided from other grids meeting the adjacent routing rule.
[0081] In this way, in the embodiment of the disclosure, the configuration of the multi-level cache system can be performed based on the hardware deployment architecture of the multiple processing cores themselves in the many-core system, without occupying additional hardware resources, and the performance of the multi-level cache system is improved by comprehensively considering the position and storage resources and other factors.
[0082] Based on the cache system in the embodiment of the disclosure, FIG. 10 is a flowchart of a cache processing method provided by the embodiment of the disclosure. Referring to FIG. 10, the method includes the following steps.
[0083] S1010: determining the multi-level cache configuration information of a first processing core.
[0084] S1020: dividing a first cache area from the target storage areas of multiple processing cores according to the multi-level cache configuration information, and the multiple processing cores corresponding to the divided first cache area constitute a multi-level cache system of the first processing core, and the processing cores where the cache areas of different levels are located are related to the position information of the first processing core.
[0085] Further, when the first processing core accesses data during task execution, the first processing core can preferentially query from the multi-level cache system of the first processing core to improve data access efficiency. Specifically, the present disclosure also provides a possible implementation, which includes: receiving a data access request, the data access request including at least a data address, the data access request being received by the first processing core; and according to the data address, searching and obtaining the data indicated in the data access request from the multi-level cache system of the first processing core.
[0086] It should be noted that the cache processing method in the embodiments of the present disclosure involves configuration and multi-level cache system, which is the same as the description of the cache system in the above embodiments, and will not be repeated here.
[0087] In this way, in the embodiments of the present disclosure, the flexibility of the many-core system can be fully utilized, and the corresponding multi-level cache system can be dynamically and flexibly configured for the processing core, which better meets the computing and data access requirements of the processing core itself, improves the system performance, and further improves the data access efficiency and the performance of the processing core in executing tasks when querying and obtaining data from the corresponding multi-level cache system.
[0088] The execution subject of the cache processing method in the embodiments of the present disclosure can be hardware execution or execution through a processor running computer executable code, and the embodiments of the present disclosure do not limit the execution subject.
[0089] It can be understood that the above-mentioned various method embodiments of the present disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to the limited space, the present disclosure will not be repeated. Those skilled in the art can understand that the specific execution order of each step in the above method should be determined according to its function and possible internal logic.
[0090] In addition, the present disclosure also provides an electronic device, a computer readable storage medium, and a computer program product, all of which can be used to implement any of the cache processing methods provided by the present disclosure. The corresponding technical solutions and descriptions are referred to the corresponding description in the method part, and will not be repeated.
[0091] FIG. 11 is a block diagram of an electronic device according to an embodiment of the present disclosure.
[0092] Referring to FIG. 11, the present disclosure provides an electronic device, which includes a plurality of processing cores 1101 and a network on chip 1102, wherein the plurality of processing cores 1101 are connected with the network on chip 1102, and the network on chip 1102 is used to interact data between the plurality of processing cores and external data.
[0093] The one or more instructions are stored in the one or more processing cores 1101 and are executed by the one or more processing cores 1101 to enable the one or more processing cores 1101 to perform the cache processing method described above.
[0094] In some embodiments, the electronic device can be a brain-like chip. Since the brain-like chip can use vectorized computing, it needs to load parameters such as weight information of a neural network model through an external memory, for example, a double data rate (DDR) synchronous dynamic random access memory. Therefore, the disclosed embodiments have higher operation efficiency in batch processing.
[0095] The disclosed embodiments also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the cache processing method described above. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.
[0096] The disclosed embodiments also provide a computer program product comprising computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, which, when executed in a processor of an electronic device, causes the processor in the electronic device to perform the cache processing method described above.
[0097] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, the functions of the modules / units in the systems and devices, can be implemented as software, firmware, hardware, or a suitable combination thereof. In hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, one physical component can have multiple functions, or one function or step can be performed by several physical components in cooperation. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable storage medium, which can include computer storage media (or non-transitory media) and communication media (or transitory media).
[0098] As those skilled in the art will appreciate, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable program instructions, data structures, program modules or other data. Computer storage media include, but are not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM), Electrically Erasable Programmable Read Only Memory (EEPROM), flash memory or other memory technology, Compact Disc Read Only Memory (CD-ROM), Digital Versatile Disc (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Further, as those skilled in the art will appreciate, communication media typically embodies computer readable program instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media. As used herein, the term "exemplary" means serving as an example, instance or illustration. Any implementation described herein as "exemplary" is not necessarily to be construed as preferred or advantageous over other implementations.
[0099] The computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0100] Computer readable program instructions for carrying out operations of the present disclosure can be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state setting data, or any combination of one or more of the above in any combination of one or more programming languages including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0101] The computer program product described herein can be embodied in a specific manner by hardware, software, or a combination thereof. In an optional embodiment, the computer program product is embodied as a computer storage medium. In another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (SDK), and the like.
[0102] The various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer readable program instructions.
[0103] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can be a computer- readable storage medium having no data storage cycles that change state. The instructions can be executed by one or more processors of a computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process such that the instructions which execute via the one or more processors of the computer or other programmable data processing devices create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. The computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0104] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0105] The flow and block diagrams in the drawings show architectural, functional, and operational aspects of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow and block diagrams can represent a module, a segment, or a portion of instructions which comprise one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may
[0106] Example embodiments have been disclosed herein and, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation. In some instances, it will be apparent to those skilled in the art that features, characteristics or aspects described with reference to a particular embodiment can be used alone or in combination with other embodiments, unless expressly stated otherwise. Accordingly, it will be understood by those skilled in the art that various changes in form and details can be made without departing from the scope of the present disclosure as set forth in the appended claims.
Claims
1. A cache system, wherein, The processing core comprises a target storage area: The configuration module is configured to determine multi-level cache configuration information of the first processing core, and divide a first cache area from the target storage area of the plurality of processing cores according to the multi-level cache configuration information, wherein the first processing core and the processing cores corresponding to the divided first cache areas constitute a multi-level cache system of the first processing core, and the first cache areas at different levels are related to position information of the processing cores where the first cache areas are located. The first processing core is configured to, when receiving a data access request, search and acquire data indicated in the data access request from the multi-level cache system of the first processing core.
2. The system of claim 1, wherein, The plurality of processing cores are processing cores on a many-core chip, wherein different processing cores perform different computing tasks in parallel.
3. The system of claim 1, wherein, The configuration module is configured to determine the multi-level cache configuration information of the first processing core according to task information of a task performed by the first processing core, storage resource information of a target storage area in the plurality of processing cores, and position information between the first processing core and the plurality of processing cores. Alternatively, the configuration module is configured to acquire an input configuration instruction, and determine multi-level cache configuration information of the first processing core indicated in the configuration instruction.
4. The system of any one of claims 1-3, wherein, The multi-level cache configuration information comprises a number of levels of the multi-level cache system, positions of a plurality of level cache areas, and sizes of the plurality of level cache areas. The configuration module is configured to divide a first cache area from a target storage area of a processing core corresponding to a position of the plurality of level cache areas according to the number of levels, the position of the plurality of level cache areas, and the sizes of the plurality of level cache areas, and determine link information between a plurality of first cache areas.
5. The system of claim 4, wherein, The first cache area comprises a link area and a cache data area, wherein The cache data area is configured to store target data of a task performed by the first processing core. The link area is configured to store link information of a next level cache area corresponding to a current level cache area, wherein the link information comprises a position of the next level cache area and a data address range of target data stored in the next level cache area.
6. The system of claim 5, wherein, If cache areas at the same level comprise a plurality of sub-cache areas, and the plurality of sub-cache areas are located in target storage areas of different processing cores, then link information between first cache areas at different levels in the multi-level cache system is in a tree structure or an inverse tree structure, and link information of a next level cache area stored in the link information is link information of a sub-cache area of the next level cache area having a connection edge in the tree structure or the inverse tree structure.
7. The system of claim 5, wherein, The target storage area of the processing core comprises one or more first cache areas, and the plurality of first cache areas represent cache areas at different levels in the multi-level cache system of the first processing core. Alternatively, the target storage area of the processing core further comprises second cache areas in the multi-level cache system of another processing core. The link area is further configured to store a cache page table, wherein the cache page table is configured to record a core identifier of a processing core to which a plurality of cache areas included in the target storage area of the processing core belong, and a level identifier.
8. The system of claim 5, wherein, The system further comprises a tag area; The tag area is located in a first cache area of the first processing core, and the tag area is configured to store a data address of target data stored in the first cache area; Alternatively, the tag area is located in a shared area of multiple processing cores, and the tag area is configured to store a data address of target data stored in a first cache area of the multiple processing cores associated with the shared area.
9. The system of claim 4, wherein, The multiple processing cores are divided into m*n grids according to position information, at least one processing core is included in each grid, and the processing cores in the same grid share a cache area of the same level; and the configuration module is configured to: According to the divided m*n grids, a first cache area is divided from a target storage area of a processing core included in a first grid in which the first processing core is located, to serve as a first-level cache area; According to a preset adjacent routing rule, a second grid corresponding to the first grid and meeting the adjacent routing rule is determined, a first cache area is divided from a target storage area of a processing core included in the second grid, to serve as a second-level cache area, and the process is continued with the first grid as the center until the number of levels is met, to obtain a multi-level cache system of the first processing core, the multi-level cache system at least including the first-level cache area and the second-level cache area.
10. The system of claim 1, wherein, The target storage area is located in a private memory of a processing core, or the target storage area is located in a local shared storage area or a global shared storage area of multiple processing cores.
11. A cache processing method, wherein, The method is applied to the cache system of any one of claims 1-10, and the method comprises: determining multi-level cache configuration information of a first processing core; According to the multi-level cache configuration information, a first cache area is divided from a target storage area of multiple processing cores, and the first cache areas corresponding to the divided multiple processing cores constitute a multi-level cache system of the first processing core, and the processing cores in which the first cache areas of different levels are located are related to position information of the first processing core.
12. The method of claim 11, wherein, The method further comprises: receiving a data access request, the data access request at least including a data address, and the data access request being received by the first processing core; According to the data address, the data indicated in the data access request is searched and obtained from the multi-level cache system of the first processing core.
13. An electronic device, comprising: The method comprises: multiple processing cores; and an on-chip network configured to interact data between the multiple processing cores and external data; wherein one or more of the processing cores store one or more instructions, and the one or more instructions are executed by one or more of the processing cores to enable one or more of the processing cores to execute the cache processing method of any one of claims 11-12.
14. A computer readable storage medium having stored thereon a computer program, wherein, The computer program, when executed by a processor, implements the cache processing method of any one of claims 11-12.
15. A computer program product comprising a computer program, which, when executed by a processor, implements the cache processing method of any one of claims 11-12.
Citation Information
Patent Citations
Data access method, shared cache, chip system and electronic equipment
CN114036084A
Cache system, cache processing method, electronic equipment and storage medium
CN119271574A
Allocation of memory space to individual processor cores
US20100268891A1
Allocating processor cores with cache memory associativity
US20110047333A1
Method and system for dynamically power scaling a cache memory of a multi-core processing system
US20130246825A1