Request processing method of multi-core processor, multi-core processor, product and server

By adopting a non-inclusive shared cache architecture and directory maintenance consistency in multi-core processors, the problems of low cache space utilization and large data exchange volume are solved, which reduces the cache consistency maintenance cost and improves request processing efficiency.

CN120336213AActive Publication Date: 2025-07-18SHANDONG BOSUAN ZHIXIN INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510828993.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

The low-level cache space utilization rate, large data exchange volume and high cache consistency maintenance costs in existing multi-core processors.

Method used

A non-inclusive shared cache architecture is adopted, and the request results are obtained through the master node query directory, and different request response policies are adopted to allow data to exist in both private cache and shared cache, and the directory is used to maintain cache consistency.

Benefits of technology

It improves the utilization rate of shared cache, reduces the amount of data exchange and cache consistency maintenance costs, and improves the efficiency and flexibility of request processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336213A_ABST
    Figure CN120336213A_ABST
Patent Text Reader

Abstract

The invention discloses a request processing method of a multi-core processor, the multi-core processor, a product and a server, and relates to the technical field of multi-core processors. In the method, a main node queries a directory according to a request and obtains a query result, and based on the query result, different request response strategies are adopted to respond to the request. According to the method, the consistency of the shared caches in the non-inclusion relation is maintained in a directory mode, the utilization rate of the shared caches is increased, and the chip area occupied by the caches is reduced under the condition that the cache capacity is fixed; secondly, based on the shared cache architecture of the non-inclusion relationship, the data is allowed to be stored in the private cache and the shared cache at the same time, so that the problem of excessive data exchange caused by cache position exchange in an exclusive architecture can be avoided, meanwhile, the private cache is allowed to reserve the data which does not exist in the shared cache, and the problem of reverse failure in an inclusion mode is avoided. According to the method, the utilization rate of the shared cache is improved, and the data exchange amount and the maintenance cost of cache consistency are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-core processors, and particularly to a request processing method for a multi-core processor, a multi-core processor, a product, and a server. Background Art

[0002] With the continuous development of computer technology, in order to reduce the latency of the Central Processing Unit (CPU) accessing memory, multiple levels of caches are set between the processor and the memory. The multi-level cache architecture can be divided into three modes: Inclusive, Exclusive, and Non-Inclusive.

[0003] In the Inclusive mode, the same data is stored between different levels of caches. The data in the upper-level cache must exist in the lower-level cache, but the data in the lower-level cache does not necessarily exist in the upper-level cache, resulting in a low utilization rate of the low-level cache space. In the Exclusive mode, the data in different levels of caches must be different, and the same cache copy can only be stored in a certain level of cache. Since the data cannot coexist in multiple cache levels in the Exclusive mode, when the processor needs to access a certain piece of data, it may cause the movement of data between cache levels, resulting in a large amount of data exchange and a high cost of maintaining cache consistency.

[0004] Therefore, how to improve the utilization rate of the low-level cache space, reduce the amount of data exchange when accessing data, and reduce the cost of maintaining cache consistency are technical problems that need to be solved by those skilled in the art. Summary of the Invention

[0005] The object of the present invention is to provide a request processing method for a multi-core processor, a multi-core processor, a product, and a server to solve the technical problems of low utilization rate of the low-level cache space, large amount of data exchange when accessing data, and high cost of maintaining cache consistency.

[0006] To solve the above technical problems, the present invention provides a request processing method for a multi-core processor, which is applied to a master node and includes: Obtain a request sent by a request node; Query a directory according to the request and obtain a query result; wherein, each directory entry includes at least an address, a data status, a cache level, and the status of the private cache of a slave node; the cache level includes the level at which the data is stored in the shared cache, the level at which the data is stored in the private cache, and the level at which the data is stored in both the shared cache and the private cache; If the query result is a directory miss, obtain the target data corresponding to the request from the memory and transfer the target data to the private cache of the request node after passing through the shared cache to respond to the request; If the query result is a shared cache hit, obtain the target data corresponding to the request from the shared cache, and transfer the target data to the private cache of the request node to respond to the request; If the query result is a hit in other private caches, control the hit other private cache to transfer the target data corresponding to the request to the private cache of the request node to respond to the request; wherein, the other private caches are caches other than the private cache of the request node among all private caches.

[0007] On the one hand, before obtaining the request sent by the request node, it further includes: The request node determines whether the request initiated by itself hits its own private cache; if so, the request node uses the target data corresponding to the request stored in its own private cache to respond to the request and ends; if not, it sends a request to the master node.

[0008] On the other hand, the request is a read request, and the response to the request includes: Starting from receiving the target data from the private cache of the request node, within a preset duration, the request node obtains the target data from its own private cache; uses the target data to respond to the request.

[0009] On the other hand, after obtaining the target data corresponding to the request from the memory and transferring the target data to the private cache of the request node via the shared cache, it further includes: Mark the cache level of the cache line where the address corresponding to the read request is located in the directory as the level where the data is stored in both the shared cache and the private cache; Determine the target state of the private cache of the request node according to the request; Control the state of the private cache of the request node to be the target state.

[0010] On the other hand, after obtaining the target data corresponding to the request from the shared cache and transferring the target data to the private cache of the request node, it further includes: Change the cache level of the cache line where the address corresponding to the read request is located in the directory from the level where the data is stored in the shared cache to the level where the data is stored in both the shared cache and the private cache; Determine the target state of the private cache of the request node according to the request; Control the state of the private cache of the request node to be the target state.

[0011] On the other hand, after controlling the hit other private cache to transfer the target data corresponding to the request to the private cache of the request node, it further includes: Keep the cache level of the cache line where the address corresponding to the read request in the directory remains unchanged; Determine the first target state of the private cache of the request node according to the request, and determine the second target state of other private caches that hit; Control the state of the private cache of the request node to be the first target state, and control the state of other private caches that hit to be the second target state.

[0012] On the other hand, the request is a write request, and the target data is the write address corresponding to the write request; the response to the request includes: Control the state of the private cache of the request node to be the exclusive state; When it is detected that the state of the private cache of the request node is the exclusive state, control the request node to write the data corresponding to the write request into its own private cache of the request node.

[0013] On the other hand, before controlling the request node to write the data corresponding to the write request into its own private cache of the request node, it further includes: When it is detected that the private cache of the request node itself is in the full state, control the request node to evict the shared cache line according to the least recently used policy and the policy of preferentially evicting the data in the shared state in its own private cache of the request node; After evicting the shared state data, enter the step of controlling the request node to write the data corresponding to the write request into its own private cache of the request node.

[0014] On the other hand, after evicting the shared state data and before entering the step of controlling the request node to write the data corresponding to the write request into its own private cache of the request node, it further includes: Obtain the size relationship between the remaining space of the private cache and the space occupied by the data corresponding to the write request; If it is detected that the remaining space of the private cache is greater than or equal to the space occupied by the data corresponding to the write request, enter the step of controlling the request node to write the data corresponding to the write request into its own private cache of the request node; If it is detected that the remaining space of the private cache is less than the space occupied by the data corresponding to the write request, evict the exclusive cache line and return to the step of obtaining the size relationship between the remaining space of the private cache and the space occupied by the data corresponding to the write request.

[0015] On the other hand, before transmitting the target data to the private cache of the request node, it further includes: When it is detected that the private cache of the requesting node itself is full, control the requesting node to evict the shared-state cache lines according to the least recently used policy and the policy of preferentially evicting the data in the shared state in the private cache of the requesting node itself; After evicting the shared-state data, obtain the size relationship between the remaining space in the private cache and the space occupied by the target data; If it is detected that the remaining space in the private cache is greater than or equal to the space occupied by the target data, enter the step of transferring the target data to the private cache of the requesting node; If it is detected that the remaining space in the private cache is less than the space occupied by the target data, evict the exclusive cache lines and return to the step of obtaining the size relationship between the remaining space in the private cache and the space occupied by the target data.

[0016] On the other hand, before transferring the target data to the shared cache, it further includes: When it is detected that the shared cache is full, according to the least recently used policy and the policy of preferentially evicting the cache lines whose cache level is the level where the data is stored in both the shared cache and the private cache and the data state is clean in the shared cache, evict the target cache lines; Change the cache level of the target cache line in the directory from the level where the data is stored in both the shared cache and the private cache to the level where the data is stored in the private cache; Enter the step of transferring the target data to the shared cache.

[0017] On the other hand, when determining multiple target cache lines according to the least recently used policy and the policy of preferentially evicting the cache lines whose cache level is the level where the data is stored in both the shared cache and the private cache and the data state is clean in the shared cache, the evicting the target cache lines includes: Obtain the cache level of each target cache line recorded in the directory; If there is a target cache line with a cache level of the level where the data is stored in both the shared cache and the private cache among all the target cache lines, evict the target cache line with a cache level of the level where the data is stored in both the shared cache and the private cache; If there is no target cache line with a cache level of the level where the data is stored in both the shared cache and the private cache among all the target cache lines, randomly evict the target cache lines.

[0018] On the other hand, it further includes: Determine the address of the next request according to the address corresponding to the request; Control the shared cache to prefetch the data corresponding to the address of the next request, and store the prefetch data corresponding to the address of the next request in the shared cache.

[0019] On the other hand, before storing the prefetch data corresponding to the address of the next request in the shared cache, it further includes: When detecting that the shared cache is in a full state, according to the least recently used policy and the policy of preferentially evicting the cache line whose cache level is the level where the data is stored in both the shared cache and the private cache and the data state is clean, evict the target cache line; Change the cache level of the target cache line in the directory from the level where the data is stored in both the shared cache and the private cache to the level where the data is stored in the shared cache; Enter the step of storing the prefetch data corresponding to the address of the next request in the shared cache.

[0020] On the other hand, before evicting the cache line, it further includes: Obtain the pre-established array for characterizing the data situation in the private cache and the array for characterizing the data situation in the shared cache; wherein, each row in the array includes at least the address, the cache state corresponding to the address, and the usage times corresponding to the address; The evicting the cache line includes: Based on the array for characterizing the data situation in the private cache or the array for characterizing the data situation in the shared cache, and according to the least recently used policy and the preferential eviction policy, evict the cache line.

[0021] On the other hand, when detecting that the shared cache is in a full state, the content of the request is to write data to the memory, and the data corresponding to the write request is in a dirty data state, the method further includes: Pre-set the bypass cache corresponding to the shared cache; Control the request node to transfer the data of the write request to the bypass cache corresponding to the shared cache; Write the data of the write request back to the memory through the bypass cache.

[0022] To solve the above technical problems, the present invention further provides a multi-core processor, including: a main node, slave nodes, a shared cache, a memory, and private caches corresponding to each node, the main node is connected to the slave nodes; the main node is used for: Obtain the request sent by the request node; Query the directory according to the request and obtain the query result; wherein, each directory entry includes at least an address, a data status, a cache level, and the status of the private cache of the slave node; the cache level includes the level at which the data is stored in the shared cache, the level at which the data is stored in the private cache, and the level at which the data is stored in both the shared cache and the private cache; If the query result is a directory miss, obtain the target data corresponding to the request from the memory and transfer the target data to the private cache of the request node via the shared cache to respond to the request; If the query result is a shared cache hit, obtain the target data corresponding to the request from the shared cache and transfer the target data to the private cache of the request node to respond to the request; If the query result is a hit in other private caches, control the hit other private caches to transfer the target data corresponding to the request to the private cache of the request node to respond to the request; wherein, the other private caches are the caches other than the private cache of the request node among all the private caches.

[0023] To solve the above technical problems, the present invention also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above request processing method of the multi-core processor are implemented.

[0024] To solve the above technical problems, the present invention also provides a server, including: A memory for storing a computer program; A processor for implementing the steps of the above request processing method of the multi-core processor when executing the computer program.

[0025] To solve the above technical problems, the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above request processing method of the multi-core processor are implemented.

[0026] The beneficial effects of the present invention are as follows. First, in this method, the cache hierarchy includes the level where data is stored in the shared cache, the level where data is stored in the private cache, and the level where data is stored in both the shared cache and the private cache simultaneously. That is, it allows data to exist in both the private cache and the shared cache at the same time, and also allows data to be stored separately in the private cache or the shared cache. It can be seen that the method provided by the present invention is based on a non-inclusive shared cache architecture. After the master node receives the request sent by the requesting node, it queries the directory according to the request and obtains the query result, and then adopts different request response strategies according to the query result to respond to the request. That is, the directory is used to maintain the consistency of the non-inclusive shared cache. Suppose this multi-core system has N cores, the capacity of the private cache is C1, the capacity of the shared cache is C2, the number of directory entries N1 required by using the directory method, and the ratio S1 of the number of entries N2 required by the inclusive shared cache architecture method is: (N × C1 + C2) / C2; the ratio S2 of the utilization rate of the shared cache is: C2 / (C2 - N × C1); the ratio S3 of the increased overhead of the directory and the increased cache capacity is: (C2 2 - (N × C1) 2 ) / C2 2 . Generally, the capacity of the shared cache is greater than the capacity of the private caches of all cores. Therefore, the value of S3 is greater than 0 and less than 1, which proves that the cost performance of the increased shared cache capacity brought by the increased overhead of the directory is high. That is, compared with maintaining the consistency of the inclusive shared cache, in the method provided by the present invention for maintaining the consistency of the non-inclusive shared cache by using the directory method, the utilization rate of the shared cache is increased, and when the cache capacity is certain, the chip area occupied by the cache is reduced. Second, since the method provided by the present invention is based on a non-inclusive shared cache architecture, in the non-inclusive shared cache architecture, it allows data to exist in both the private cache and the shared cache at the same time, and also allows data to be stored separately in the private cache or the shared cache. Allowing data to be stored in both the private cache and the shared cache at the same time can avoid excessive data exchange problems caused by the swapping of caches in the exclusive architecture. At the same time, allowing the private cache to retain data that does not exist in the shared cache avoids the reverse invalidation problem in the inclusive mode. It can be seen that this method improves the utilization rate of the shared cache and reduces the data exchange volume, thus reducing the maintenance cost of cache consistency. Third, in the process of adopting different request response strategies according to the query result to respond to the request, the flexibility of request processing is realized, and when the query result is a hit in other private caches, the data transfer method from cache to cache is adopted. Compared with the transfer method between two-level caches, in the method provided by the present invention, the transmission delay is reduced and the efficiency of responding to requests is improved.

[0027] In addition, before obtaining the request sent by the requesting node, the requesting node first determines whether the request initiated by itself hits its own private cache; if so, the requesting node directly uses the target data corresponding to the request stored in its own private cache to respond to the request and ends; if not, it sends the request to the master node. Since the private cache of the requesting node is closer to the requesting node than the shared cache and other private caches, when the request sent by the requesting node hits its own private cache, the private cache directly responds to the request, which can improve the efficiency of request response.

[0028] Starting from receiving the target data from the private cache of the requesting node, within a preset duration, the requesting node obtains the target data from its own private cache; and uses the target data to respond to the read request. The response to the read request is achieved through this method.

[0029] After obtaining the target data corresponding to the request from the memory and transmitting the target data to the private cache of the requesting node via the shared cache, mark the cache level of the cache line where the address corresponding to the read request is located in the directory as the level where the data is stored in both the shared cache and the private cache, and update the status of the private cache of the requesting node according to the request; after obtaining the target data corresponding to the request from the shared cache and transmitting the target data to the private cache of the requesting node, update the cache level of the cache line where the address corresponding to the read request is located in the directory and update the status of the private cache of the requesting node according to the request; after controlling other private caches that hit to transmit the target data corresponding to the request to the private cache of the requesting node, keep the cache level of the cache line where the address corresponding to the read request is located in the directory unchanged, and change the private cache status of the request core and the response core, thus realizing the maintenance of cache consistency.

[0030] Control the status of the private cache of the requesting node to the exclusive state; in the case of detecting that the status of the private cache of the requesting node is in the exclusive state, control the requesting node to write the data corresponding to the write request to its own private cache. That is, when ensuring that the private cache of the requesting node is in the exclusive state, write data to the private cache, which can ensure cache consistency as much as possible.

[0031] Before the control request node writes the data corresponding to the write request to its own private cache, if the private cache of the request node is full, then according to the least recently used policy, and the policy of preferentially evicting the shared-state data in the private cache of the request node itself, evict the shared-state cache line, and then write the data to the private cache. That is, in the case where the private cache is full, by evicting the cache line, it is ensured that the data can be written into the private cache; the shared-state data is preferentially evicted because the shared-state data also exists in other private caches, and when a node (i.e., a core) needs this cache line again, the cache copy can still be obtained through cache-to-cache transfer; the data transfer latency between caches is lower than the transfer latency between two levels of caches.

[0032] After evicting the shared-state data, if it is detected that the remaining space in the private cache is less than the space occupied by the data corresponding to the write request, then evict the exclusive cache line. That is, the shared-state data is preferentially evicted, and finally the exclusive cache line is evicted, reserving a higher non-eviction privilege for the exclusive data.

[0033] During the read or write process, before the target data read is transferred to the private cache of the request node, if the private cache is full, then according to the least recently used policy, and the policy of preferentially evicting the shared-state data in the private cache of the request node itself, evict the shared-state cache line, and finally evict the shared-state data. Similarly, the shared-state data is preferentially evicted because the shared-state data also exists in other private caches, and when a node (i.e., a core) needs this cache line again, the cache copy can still be obtained through cache-to-cache transfer; the data transfer latency between caches is lower than the transfer latency between two levels of caches; and a higher non-eviction privilege is reserved for the exclusive data.

[0034] Before transferring the target data to the shared cache, if the shared cache is full, then according to the least recently used policy, and the policy of preferentially evicting the cache line whose cache level is the level where the data is stored in both the shared cache and the private cache and the data state is clean in the shared cache, evict the target cache line; change the cache level of the target cache line in the directory, without the need to evict the data of the same cache line in the private cache at the same time. This method can avoid reverse invalidation. Reverse invalidation will greatly increase the memory access time. If the data invalidated in reverse is in the dirty state, it is necessary to wait for the dirty data to be written back to memory before this cache line can be invalidated and the shared cache can obtain the data copy of the new request. Therefore, this method adopts a multi-level cache maintenance strategy to avoid reverse invalidation, which can reduce the memory access time when the cache misses.

[0035] When determining multiple target cache lines according to the least recently used policy and the policy of preferentially evicting cache lines whose cache level in the shared cache is the level where data is stored in both the shared cache and the private cache and the data status is clean, continue to combine the cache level of the directory record to judge the evicted cache line, and preferentially evict the target cache line whose data is stored in both the shared cache and the private cache, because the shared state data still exists in other private caches, and when a core needs this cache line again, the cache copy can still be obtained through cache-to-cache transfer.

[0036] Determine the address of the next request according to the address corresponding to the request; control the shared cache to prefetch the data corresponding to the address of the next request, and store the prefetch data corresponding to the address of the next request in the shared cache. In this method, the data of the next request is placed in the shared cache in advance according to the address corresponding to the request. By prefetching more cache line copies that meet the requirements, the number of cache misses can be reduced.

[0037] In the scenario where the shared cache actively prefetches data, if the shared cache is full, according to the least recently used policy and the policy of preferentially evicting cache lines whose cache level in the shared cache is the level where data is stored in both the shared cache and the private cache and the data status is clean, evict the target cache line; change the cache level of the target cache line in the directory. Since the evicted data is stored in both the shared cache and the private cache, after evicting the target cache line, the data still remains in the private cache. When a core needs this cache line again, the cache copy can still be obtained through cache-to-cache transfer.

[0038] An array established in advance for characterizing the data situation in the private cache and an array for characterizing the data situation in the shared cache. When evicting a cache line, based on the array for characterizing the data situation in the private cache or the array for characterizing the data situation in the shared cache, and according to the least recently used policy and the preferential eviction policy, evict the cache line. Since each row in the array includes at least the address, the cache status corresponding to the address, and the usage times corresponding to the address, the cache status of each address and the usage times corresponding to the address can be intuitively obtained, improving the efficiency of finding the cache line to be evicted.

[0039] When the shared cache is full and the request is to write dirty data to the memory, write the dirty data back to the memory through the write-back path from the private cache to the bypass cache to the memory. The bypass cache can avoid the passive cache replacement of the shared cache when the shared cache is full due to the active write-back of the private cache, reducing the impact of the private cache on the shared cache.

[0040] In addition, the present invention also provides a multi-core processor, a computer program product, a server, and a computer-readable storage medium, which have the same or corresponding technical features as the request processing method of the multi-core processor mentioned above, and the effects are the same as above. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] To more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0042] Figure 1 Schematic diagram of an eight-core two-level cache system architecture provided by an embodiment of the present invention; Figure 2 Flowchart of a request processing method for a multi-core processor provided by an embodiment of the present invention; Figure 3 Schematic diagram of a directory entry provided by an embodiment of the present invention; Figure 4 Schematic diagram of a private cache tag array provided by an embodiment of the present invention; Figure 5 Schematic diagram of a shared cache tag array provided by an embodiment of the present invention; Figure 6 Schematic diagram of a directory array provided by an embodiment of the present invention; Figure 7 Structure diagram of a server provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0044] The core of the present invention is to provide a request processing method for a multi-core processor, a multi-core processor, a product, and a server to solve the technical problems of low utilization rate of low-level cache space, large data exchange volume when accessing data, and high cache consistency maintenance cost.

[0045] With the continuous development of computer technology, the performance of processors has been significantly improved. However, the growth of the main memory access speed has lagged behind relatively. That is, the computing power of the processor is much greater than the speed of obtaining data from the memory, resulting in an increasingly large performance difference between the processor and the memory, and the "memory wall" problem has emerged. To solve the "memory wall" problem, multi-level cache technology has been introduced. By deploying multiple caches at different levels, the speed and efficiency of data access are improved. Multiple layers of caches are set between the processor and the memory to reduce the latency of the CPU accessing the memory. The closer the cache is to the CPU, the smaller the capacity, the faster the access speed, and the higher the cost. The multi-level cache architecture can be divided into three modes: inclusive mode, exclusive mode, and non-inclusive mode. In the inclusive mode, the same data is stored between different levels of caches. The data in the upper-level cache must exist in the lower-level cache, but the data in the lower-level cache may not exist in the upper-level cache, which may lead to a low utilization rate of the lower-level cache space. In the exclusive mode, the data exchange volume is large, and the cost of maintaining cache consistency is high. Considering the flexibility of the non-inclusive mode, the data in the upper-level cache may or may not exist in the lower-level cache, that is, it does not strictly require whether the data between different levels of caches is inclusive. Therefore, in the present invention, a multi-level cache architecture based on the non-inclusive mode is used to process requests. Taking an eight-core two-level cache system architecture as an example, the multi-core processing system provided by the present invention will be described below. Figure 1 FIG. is a schematic diagram of an eight-core two-level cache system architecture provided by an embodiment of the present invention, as Figure 1 shown. The eight cores are respectively core 0, core 1 to core 7, and the two-level caches are respectively the first-level cache (L1) and the second-level cache (L2). The first-level cache is the cache corresponding to each core, also called a private cache. Each core corresponds to a first-level cache controller, first-level cache instructions, and first-level cache data; the second-level cache is a shared cache. There is a corresponding second-level cache controller for the second-level cache. The memory corresponds to a memory controller, and the second-level cache controller is connected to the memory controller. The first-level cache and the second-level cache are in a non-inclusive mode. It should be noted that the present invention is not limited to the number of levels of the multi-level cache, which can be two levels, three levels, or four levels, as long as it is ensured that the shared last-level cache is a non-inclusive shared architecture, and the cache architecture between other levels can be selected freely. The shared cache described in the present invention is the shared last-level cache (LastLevel Cache, LLC).

[0046] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Figure 2 FIG. is a flowchart of a request processing method for a multi-core processor provided by an embodiment of the present invention. This method is applied to the main node. It should be noted that both the main node and the slave node are processor cores. There is no limitation on the selected main node, which is determined according to the actual situation. This method includes: S10: Obtain the request sent by the requesting node.

[0047] The requesting node can be the master node or the slave node. The request sent can be a read request or a write request. Since the private cache of the requesting node itself is closer to the requesting node compared to the shared cache and other private caches, therefore, in order to improve the efficiency of request response, before obtaining the request sent by the requesting node, it further includes: the requesting node determines whether the request initiated by itself hits its own private cache; if so, the requesting node uses the target data corresponding to the request stored in its own private cache to respond to the request and ends; if not, it sends the request to the master node, and the master node obtains the request.

[0048] It should be noted that when determining whether the request hits the cache, since the requested address is carried in the request, and there is also the address corresponding to the saved data in the cache, if the requested address carried in the request can be found from the address corresponding to the saved data in the cache, it is determined that the request hits the cache; otherwise, it is determined that the request does not hit the cache.

[0049] S11: Query the directory according to the request and obtain the query result.

[0050] Among them, each directory entry includes at least an address, a data status, a cache level, and the status of the private cache of the slave node; the cache level includes the level where the data is stored in the shared cache, the level where the data is stored in the private cache, and the level where the data is stored in both the shared cache and the private cache.

[0051] In a shared last-level cache with a non-inclusive architecture, a directory is used to maintain the consistency between the last-level cache and its upper-level cache. Figure 3 It is a schematic diagram of a directory entry provided by an embodiment of the present invention, as Figure 3As shown in the figure, there are N private caches. Each directory entry contains an address, whether it is exclusive (occupying 1 bit, 1 means exclusive, 0 means shared), cache level (occupying 2 bits, 01 means the data is in the last-level cache (i.e., shared cache), 10 means the data is in the private cache, 11 means the data exists in both the last-level cache and the private cache), whether the last-level cache is dirty (occupying 1 bit, 1 means dirty data, 0 means clean data), and the states of N (N is the number of cores) private caches (occupying 2 bits, and the number of bits is related to the coherence protocol adopted). For example, for the address xxxx000, it is dirty data, the data is in the last-level cache, and the states of all private caches are invalid states (represented as I state); for the address xxxx001, it is clean data, the data is in the private cache, the states of private cache 1 and private cache 2 are both shared clean states (represented as SC state), and the state of private cache N is the invalid state; for the address xxxx111, it is dirty data, the data exists in both the last-level cache and the private cache, the state of private cache 1 is the exclusive state (represented as U state), and the states of other private caches are all invalid states.

[0052] The address is carried in the request, and the directory also records the address, cache level, and the states of private caches. Therefore, the directory can be queried to determine the way to respond to the request.

[0053] S12: Adopt different strategies to respond to the request according to the query result.

[0054] Step S12 specifically includes: S120: If the query result is a directory miss, obtain the target data corresponding to the request from the memory and transfer the target data to the private cache of the request node through the shared cache to respond to the request.

[0055] S121: If the query result is a shared cache hit, obtain the target data corresponding to the request from the shared cache and transfer the target data to the private cache of the request node to respond to the request.

[0056] S122: If the query result is a hit in other private caches, control the hit other private cache to transfer the target data corresponding to the request to the private cache of the request node to respond to the request.

[0057] Among them, other private caches are caches in all private caches except the private cache of the request node.

[0058] It should be noted that when the request is a read request, the target data corresponds to the data; when the request is a write request, the target data corresponds to the address of the write request. If the directory is not hit, the target data corresponding to the request is first fetched from the memory and transferred to the shared cache, and then from the shared cache to the private cache of the requesting node to respond to the request. If the shared cache is hit, the target data corresponding to the request is fetched from the shared cache and transferred to the private cache of the requesting node to respond to the request. If another private cache is hit, the target data is transferred from the hit private cache (the core corresponding to the hit private cache is called the responding core) to the private cache of the requesting node to respond to the request.

[0059] After the target data is received in the private cache of the requesting node, if it is a read request, the response to the request includes: Starting from when the target data is received in the private cache of the requesting node, within a preset duration, the requesting node fetches the target data from its own private cache; and uses the target data to respond to the request.

[0060] There is no limitation on the preset duration, which is determined according to the actual situation.

[0061] To maintain cache coherence, after fetching the target data corresponding to the request from the memory and transferring the target data through the shared cache to the private cache of the requesting node, it also includes: Marking in the directory the cache level of the cache line where the address corresponding to the read request is located as the level where the data is stored in both the shared cache and the private cache; Determining the target state of the private cache of the requesting node according to the request; Controlling the state of the private cache of the requesting node to be the target state.

[0062] If the request is in the SC state, changing the state of the private cache of the requesting node to the SC state.

[0063] After fetching the target data corresponding to the request from the shared cache and transferring the target data to the private cache of the requesting node, it also includes: Changing the cache level of the cache line where the address corresponding to the read request in the directory is located from the level where the data is stored in the shared cache to the level where the data is stored in both the shared cache and the private cache; Determining the target state of the private cache of the requesting node according to the request; Controlling the state of the private cache of the requesting node to be the target state.

[0064] After controlling the hit private cache of another node to transfer the target data corresponding to the request to the private cache of the requesting node, it also includes: Keeping the cache level of the cache line where the address corresponding to the read request in the directory unchanged; Determine the first target state of the private cache of the requesting node according to the request, and determine the second target state of other private caches that hit; Control the state of the private cache of the requesting node to the first target state, and control the states of other private caches that hit to the second target state.

[0065] For example, if the response core is Core 2 and the requesting core is Core 1. If Core 2 is in the U state, the current state of Core 1 is the I state, and Core 1 requests the U state, then change Core 1 from the I state to the U state and change Core 2 from the U state to the I state.

[0066] The above describes the processing process of read requests. Here, the entire process of read request processing is described again. The entire process of read request processing includes: 1. When a read request initiated by a core hits in the private cache, the private cache directly returns a response to the core; 2. When a read request initiated by a core misses in the private cache, the read request is sent to the directory. Query the directory to check if there is a hit in other private caches or the last-level cache. If there is no hit, obtain a data copy from memory, transfer the data copy to the LLC and the private cache, mark the cache level of this cache line in the directory as 11, and set the state of the private cache of the requesting core; if there is a hit in the LLC, transfer the data copy to the private cache corresponding to the requesting core, change the cache level from 01 to 11, and set the state of the private cache of the requesting core; if there is a hit in other private caches, directly transfer the data copy to the core that initiated the request through cache-to-cache transfer. The cache level remains unchanged, and the states of the private caches of the requesting core and the responding core are changed.

[0067] Next, the processing of write requests is described. The request is a write request, and the target data is the write address corresponding to the write request; the response to the request includes: Control the state of the private cache of the requesting node to the exclusive state; In the case where it is detected that the state of the private cache of the requesting node is the exclusive state, control the requesting node to write the data corresponding to the write request into its own private cache.

[0068] During the processing of the write request, after the private cache of the write request core obtains the exclusive state, it directly writes the data into the private cache, adopting the write-back policy to reduce the operation of writing data back to memory.

[0069] The entire process of write request processing includes: 1. If the private cache of the requesting node itself hits, write the data into its own private cache; 2. If there is a cache miss in the directory, retrieve the data from memory. The private cache of the requesting node obtains the exclusive state and writes the data into the cache. If there is a cache hit in the shared cache, read the data corresponding to the write request address in the shared cache into the private cache of the core with the write request, and obtain the exclusive state, then write the data corresponding to the write request into the private cache corresponding to the core with the write request. If there is a cache hit in other private caches, after invalidating other private caches, write it into its own private cache.

[0070] During the process of request processing, the situation may occur that the private cache is full and / or the shared cache is full. In this case, in order to ensure that data can be written into the private cache or the shared cache, by the way of evicting cache lines, new data can be written into the cache.

[0071] First, the processing method in the case of the private cache being full will be described below.

[0072] Specifically, in one implementation, before controlling the requesting node to write the data corresponding to the write request into the private cache of the requesting node itself, it further includes: When it is detected that the private cache of the requesting node itself is in a full state, control the requesting node to evict the shared-state cache lines according to the Least Recently Used (LRU) policy and the policy of preferentially evicting the data in the shared state in the private cache of the requesting node itself; After evicting the shared-state data, enter the step of controlling the requesting node to write the data corresponding to the write request into the private cache of the requesting node itself.

[0073] In practice, it may occur that after evicting the shared-state data, the remaining space in the private cache is less than the space occupied by the data corresponding to the write request. Similarly, in order to ensure that data can be written into the private cache, before entering the step of controlling the requesting node to write the data corresponding to the write request into the private cache of the requesting node itself after evicting the shared-state data, it further includes: Obtain the size relationship between the remaining space in the private cache and the space occupied by the data corresponding to the write request; If it is detected that the remaining space in the private cache is greater than or equal to the space occupied by the data corresponding to the write request, enter the step of controlling the requesting node to write the data corresponding to the write request into the private cache of the requesting node itself; If it is detected that the remaining space in the private cache is less than the space occupied by the data corresponding to the write request, evict the exclusive cache lines and return to the step of obtaining the size relationship between the remaining space in the private cache and the space occupied by the data corresponding to the write request.

[0074] In this method, before the control request node writes the data corresponding to the write request to its own private cache, if the private cache of the request node is full, according to the least recently used policy and the policy of preferentially evicting the data in the shared state in the private cache of the request node itself, the shared state cache line is evicted, and then the data is written to the private cache. That is, when the private cache is full, by evicting the cache line, it is ensured that the data can be written into the private cache; the shared state data is preferentially evicted because the shared state data still exists in other private caches. When a node (i.e., a core) needs this cache line again, the cache copy can still be obtained through cache-to-cache transmission; the cache-to-cache data transmission delay is lower than the transmission delay between the two-level caches. After evicting the shared state data, if it is detected that the remaining space in the private cache is smaller than the space occupied by the data corresponding to the write request, the exclusive cache line is evicted. That is, the shared state data is preferentially evicted, and finally the exclusive cache line is evicted, reserving a higher non-eviction privilege for the exclusive data.

[0075] In another embodiment, before transmitting the target data to the private cache of the request node, it further includes: When it is detected that the private cache of the request node itself is full, control the request node to evict the shared state cache line according to the least recently used policy and the policy of preferentially evicting the data in the shared state in the private cache of the request node itself; After evicting the shared state data, obtain the size relationship between the remaining space in the private cache and the space occupied by the target data; If it is detected that the remaining space in the private cache is greater than or equal to the space occupied by the target data, enter the step of transmitting the target data to the private cache of the request node; If it is detected that the remaining space in the private cache is smaller than the space occupied by the target data, evict the exclusive cache line and return to the step of obtaining the size relationship between the remaining space in the private cache and the space occupied by the target data.

[0076] During the read or write process, before the read target data is transmitted to the private cache of the request node, the effect of evicting the cache line after the private cache is full is the same as that of the above embodiment, which will not be elaborated here.

[0077] To enable those skilled in the art to better understand the above processing method when the private cache is full, the above processing method when the private cache is full will be described again below.

[0078] When the private cache is full and new cache line data needs to be stored, while following the LRU cache replacement algorithm, shared-state data is preferentially evicted because the shared-state data also exists in other private caches. When a core needs the cache line again, the cache copy can still be obtained through cache-to-cache transfer; the cache-to-cache data transfer latency is lower than the transfer latency between two levels of caches; finally, the exclusive cache line is evicted to reserve a higher non-eviction privilege for exclusive data.

[0079] Secondly, the handling method in the case where the shared cache is full will be described.

[0080] Specifically, in one embodiment, before transmitting the target data to the shared cache, it further includes: When it is detected that the shared cache is in a full state, according to the least recently used policy, and in accordance with the policy of preferentially evicting cache lines in the shared cache where the cache level is that the data is stored in both the shared cache and the private cache and the data state is clean, the target cache line is evicted; Change the cache level of the target cache line in the directory from the level where the data is stored in both the shared cache and the private cache to the level where the data is stored in the private cache; Enter the step of transmitting the target data to the shared cache.

[0081] In practice, according to the least recently used policy, and in accordance with the policy of preferentially evicting cache lines in the shared cache where the cache level is that the data is stored in both the shared cache and the private cache and the data state is clean, multiple target cache lines may be determined. To evict the appropriate target cache line, in practice, obtain the cache levels of each target cache line recorded in the directory; If there is a target cache line with a cache level where the data is stored in both the shared cache and the private cache among all the target cache lines, then evict the target cache line with a cache level where the data is stored in both the shared cache and the private cache; If there is no target cache line with a cache level where the data is stored in both the shared cache and the private cache among all the target cache lines, then randomly evict the target cache line.

[0082] When the shared cache is full and new cache line data needs to be stored, while following the LRU cache replacement algorithm, it preferentially evicts the cache line with cache level 11 and clean state, changes the cache level of this cache line in the directory to 10, and does not need to evict the data of the same cache line in the private cache at the same time. This method can avoid reverse invalidation, which greatly increases the memory access time. If the data invalidated in reverse is in a dirty state, it needs to wait for the dirty data to be written back to memory before this cache line can be invalidated and the LLC can obtain the data copy of the new request. Therefore, this architecture adopts a multi-level cache maintenance strategy to avoid reverse invalidation, which can reduce the memory access time when the cache misses.

[0083] To improve the cache hit probability, the request processing method of the multi-core processor further includes: Determine the address of the next request according to the address corresponding to the request; Control the shared cache to prefetch the data corresponding to the address of the next request, and store the prefetched data corresponding to the address of the next request in the shared cache.

[0084] For example, if the current request is for address 01, the address of the next request may be address 02. Then the data corresponding to address 02 is prefetched and placed in the shared cache in advance. In this method, the data of the next request is placed in the shared cache in advance according to the address corresponding to the request. By prefetching more cache line copies that meet the requirements, the number of cache misses can be reduced.

[0085] In the scenario of LLC actively prefetching data, during the process of writing data to the shared cache, it may also encounter the situation where the shared cache is full. To ensure that the data can be written to the shared cache, before storing the prefetched data corresponding to the address of the next request in the shared cache, it further includes: When detecting that the shared cache is in a full state, according to the least recently used strategy, and following the strategy of preferentially evicting the cache line with the cache level where the data is stored in both the shared cache and the private cache and the data state is clean state to evict the target cache line; Change the cache level of the target cache line in the directory from the cache level where the data is stored in both the shared cache and the private cache to the cache level where the data is stored in the shared cache; Enter the step of storing the prefetched data corresponding to the address of the next request in the shared cache.

[0086] In this method, if the LLC actively prefetches data, there is no need to transfer the data to the private cache. It can be stored only in the last-level cache. Mark the cache level as 01. For the data with the evicted cache level of 11, the data still remains in the private cache. By prefetching more cache line copies that meet the requirements, the number of cache misses can be reduced. Therefore, preferentially evicting the data with the cache level of 11 can make full use of the capacity of the multi-level cache, shorten the time of cache misses, and reduce the number of cache misses.

[0087] To improve the efficiency of finding the cache line to be evicted, before evicting the cache line, it further includes: Obtain the pre-established arrays for characterizing the data situation in the private cache and the data situation in the shared cache; wherein, each row in the array includes at least the address, the cache status corresponding to the address, and the usage count corresponding to the address (i.e., the least recently used value (LRU value)).

[0088] Evicting the cache line includes: Based on the array for characterizing the data situation in the private cache or the array for characterizing the data situation in the shared cache, and evict the cache line according to the least recently used policy and the preferential eviction policy.

[0089] The method of evicting the cache line according to the least recently used policy and the preferential eviction policy has been described in detail above and will not be elaborated here.

[0090] When the shared cache is full and a request is made to write dirty data to memory. When the active write-back of the private cache causes the LLC to be full, the LLC performs passive cache replacement, that is, the private cache affects the shared cache. To reduce the impact of the private cache on the shared cache. In implementation, when it is detected that the shared cache is full, the content of the request is to write data to memory, and the data state corresponding to the data of the write request is the dirty data state, the method further includes: Pre-set a bypass cache corresponding to the shared cache; Control the request node to transfer the data of the write request to the bypass cache corresponding to the shared cache; Write the data of the write request back to memory through the bypass cache.

[0091] Such as Figure 1 In, a corresponding bypass cache is set for the secondary cache. Design a bypass cache for the shared last-level cache, which is used when the LLC is full and the private cache initiates a dirty data write-back to memory. The dirty data is written back to memory through the write-back path from the private cache to the bypass cache to memory. The bypass cache can avoid the passive cache replacement of the LLC when the LLC is full due to the active write-back of the private cache, and reduce the impact of the private cache on the shared cache.

[0092] In the method provided by the present invention, the consistency of the non-inclusive architecture multi-level cache is maintained by using a directory. The cache level where the data is located is marked, allowing the data to exist in both the private cache and the shared cache simultaneously, and also allowing the data to be stored separately in the private cache or the shared cache. Allowing the data to be stored in both the private cache and the shared cache simultaneously can avoid excessive data exchange problems caused by the swapping of caches in the exclusive architecture. At the same time, allowing the data that does not exist in the shared cache to be retained in the private cache avoids the reverse invalidation problem in the inclusive mode. In addition, using a directory to maintain the consistency of the non-inclusive shared cache increases the directory overhead compared to maintaining the consistency of the inclusive shared cache, but increases the utilization rate of the shared last-level cache, and reduces the chip area occupied by the cache when the cache capacity is fixed. Suppose the multi-core system has N cores, the private cache capacity is C1, the shared cache capacity is C2, the number of directory entries N1 required by using the directory method, and the ratio S1 of the number of entries N2 required by the inclusive shared cache architecture method is: (N×C1 + C2) / C2; the ratio S2 of the shared cache utilization rate is: C2 / (C2 - N×C1); the ratio S3 of the increased overhead of the directory and the increased cache capacity is: (C2 2 -(N×C1) 2 ) / C2 2 . Generally, the shared cache capacity is greater than the private cache capacity of all cores, so the value of S3 is greater than 0 and less than 1, proving that the cost performance of the increased shared cache capacity brought by the increased overhead of the directory is high. The specific resources occupied by the directory and the increased last-level cache capacity need to be calculated according to the address bit width and the sizes of the private cache and the shared last-level cache.

[0093] Continue below with Figure 1 the eight-core two-level cache system shown as an example to illustrate that using the method provided by the present invention will increase the utilization rate of the shared last-level cache. For the eight-core system, the address bit width is 32 bits, and the addressable range is 4GB. The schematic diagram of the system architecture is as shown in the above Figure 1As shown, in the secondary cache system, the non-inclusive mode is adopted between the L1 and L2 caches. L1 is a private cache, divided into a data cache and an instruction cache, both with a size of 32 KB. L2 is a shared cache with a size of 2 MB, and the size of each cache line is 64 Byte. Using the non-inclusive last-level cache of the present invention, the available space of the last-level cache is 2 MB, and the number of directory entries required is 40960 ((64 KB × 8 + 2 MB) / 64 Byte = 40960). Each directory needs to record a 26-bit address, a 1-bit dirty flag bit, a 2-bit cache layer flag bit, and 2 × 8-bit cache status flag bits. Each entry is 45 bits, and the directory occupies a total of 225 KB for storage. If the secondary cache architecture is in the inclusive mode, the number of directory entries required is 32768 (2 MB / 64 Byte = 32768). Each directory, excluding the cache layer flag bit, requires 43 bits, and the directory requires 172 KB of resources for storage. Compared with the inclusive mode, the non-inclusive mode requires an additional 53 KB of storage resources to store the directory, while the last-level cache can utilize up to 512 KB more resources (assuming that all cache lines in the private cache are different and have data copies in the shared cache). The present invention uses 2.5% of the resources and increases the space utilization rate of the last-level cache by 34%.

[0094] Figure 4 It is a schematic diagram of a private cache tag array provided by an embodiment of the present invention. The tag array is the Tag array. Figure 4Only the status bits required by the present invention are shown, and other status bits are not listed. The address status of xxxx000 is the exclusive dirty state (UD), and the least recently used value (i.e., the LRU value) is 6; the address status of xxxx001 is the shared clean state (SC), and the least recently used value (i.e., the LRU value) is 5; the address status of xxxx010 is the shared dirty state (SD), and the least recently used value (i.e., the LRU value) is 4; the address status of xxxx011 is the exclusive clean state (UC), and the least recently used value (i.e., the LRU value) is 3; the address status of xxxx100 is the shared clean state (SC), and the least recently used value (i.e., the LRU value) is 2; the address status of xxxx101 is the exclusive clean state (UC), and the least recently used value (i.e., the LRU value) is 1; the address status of xxxx110 is the exclusive clean state (UC), and the least recently used value (i.e., the LRU value) is 5; the address status of xxxx111 is the shared clean state (SC), and the least recently used value (i.e., the LRU value) is 1. Assume that the private cache of core 1 is full, and core 1 initiates a read request for the address xxx1000. The L1 cache misses, and the request is sent to the directory. The directory shows that the data copy is in the L2 shared cache. The L2 transfers the data to the L1 cache, but the L1 cache is full and needs to evict a cache line to store the new data. According to the LRU rule, the cache with a smaller LRU value is preferentially evicted. The present invention proposes that under the condition of satisfying the LRU rule, the shared state data is preferentially replaced. In Figure 4 In the shown case, the LRU values of xxxx101 and xxxx111 are both 1. The status of xxxx101 is UC, and the status of xxxx111 is SC. Then, the cache line of xxxx111 is evicted first to store the cache copy of xxx1000. Subsequently, when both the L1 cache and the L2 cache are full, when core 1 initiates a write-back request for the address xxxx000, the data is in the UD state and needs to be written back to the memory. Then, the data is transferred to the bypass cache of the L2 cache and written back to the memory by the bypass cache, avoiding the passive replacement of the shared L2 cache.

[0095] Figure 5 is a schematic diagram of a shared cache tag array provided by an embodiment of the present invention. When the system runs to Figure 5 the shown state, the shared cache is full. Core 1 initiates a read request for the address xxx1111. Both the L1 cache and the L2 cache miss, and a cache copy needs to be obtained from the memory. Then, a cache line in the shared cache needs to be evicted to store the data copy of the address xxxx1111. The LRU values of the addresses xxxx100 and xxxx110 in the shared cache are both 1, and both are clean cache lines. Then, it is necessary to judge by combining the cache hierarchy recorded in the directory. Figure 6 A schematic diagram of a directory array provided by an embodiment of the present invention. As Figure 6As shown, the cache level of address xxxx100 is 01, and the cache level of address xxxx110 is 11. Then, the cache line of address xxxx110 is preferentially evicted to store the data copy of address xxx1111.

[0096] In the method provided by the present invention, the capacity of the multi-level cache is fully utilized, the utilization rate of the shared last-level cache is improved, and the memory access process of the multi-core system is optimized.

[0097] The request processing method of the multi-core processor was described above. This embodiment also provides a multi-core processor, including: a main node, slave nodes, a shared cache, a memory, and private caches corresponding to each node. The main node is connected to the slave nodes. The main node is used for: Obtaining a request sent by a requesting node; Querying a directory according to the request and obtaining a query result; wherein, each directory entry includes at least an address, a data status, a cache level, and the status of the private cache of the slave node; the cache level includes the level at which the data is stored in the shared cache, the level at which the data is stored in the private cache, and the level at which the data is stored in both the shared cache and the private cache; If the query result is a directory miss, obtaining the target data corresponding to the request from the memory and transmitting the target data to the private cache of the requesting node via the shared cache to respond to the request; If the query result is a hit in the shared cache, obtaining the target data corresponding to the request from the shared cache and transmitting the target data to the private cache of the requesting node to respond to the request; If the query result is a hit in other private caches, controlling the hit other private caches to transmit the target data corresponding to the request to the private cache of the requesting node to respond to the request; wherein, the other private caches are the caches other than the private cache of the requesting node among all the private caches.

[0098] The slave nodes are used to send requests and respond to requests based on the target data in their own private caches. When sending a read request, the request is responded to using the target data in its own private cache; when sending a write request, after its own private cache is in the exclusive state, the data is written to its own private cache.

[0099] The multi-core processor provided in this embodiment has the same or corresponding technical features as the request processing method of the multi-core processor described above, which will not be elaborated here, and the effects are the same.

[0100] In the above embodiment, the request processing method of the multi-core processor was described in detail. The present invention also provides corresponding embodiments of a request processing device for the multi-core processor and a computer-readable storage medium. It should be noted that the present invention describes the embodiments of the device part from two perspectives, one is from the perspective of functional modules, and the other is from the perspective of hardware.

[0101] The request processing device of the multi-core processor provided by the embodiment of the present invention, from the perspective of functional modules, includes: The first acquisition module is used to acquire the request sent by the request node; The query and acquisition module is used to query the directory according to the request and obtain the query result; wherein, each directory entry includes at least an address, a data status, a cache level, and the status of the private cache of the slave node; the cache level includes the level at which the data is stored in the shared cache, the level at which the data is stored in the private cache, and the level at which the data is stored in both the shared cache and the private cache; The first acquisition and transmission module is used to, if the query result is a directory miss, acquire the target data corresponding to the request from the memory and transmit the target data to the private cache of the request node via the shared cache to respond to the request; The second acquisition and transmission module is used to, if the query result is a shared cache hit, acquire the target data corresponding to the request from the shared cache and transmit the target data to the private cache of the request node to respond to the request; The first control module is used to, if the query result is a hit in other private caches, control the hit other private caches to transmit the target data corresponding to the request to the private cache of the request node to respond to the request; wherein, the other private caches are the caches in all the private caches except the private cache of the request node.

[0102] In some embodiments, the request is a read request, and the request processing device of the multi-core processor includes a response module for responding to the request.

[0103] The response module specifically includes: The response sub-module is used to, starting from receiving the target data from the private cache of the request node, within a preset duration, the request node acquires the target data from its own private cache; and uses the target data to respond to the request.

[0104] In some embodiments, the request processing device of the multi-core processor further includes: The marking module is used to mark the cache level of the cache line where the address corresponding to the read request is located in the directory as the level at which the data is stored in both the shared cache and the private cache; The first determination module is used to determine the target status of the private cache of the request node according to the request; The second control module is used to control the status of the private cache of the request node to be the target status.

[0105] In some embodiments, the request processing device of the multi-core processor further includes: The first modification module is used to change the cache level of the cache line corresponding to the read request address in the directory from the level where data is stored in the shared cache to the level where data is stored in both the shared cache and the private cache; The second determination module is used to determine the target state of the private cache of the request node according to the request; The third control module is used to control the state of the private cache of the request node to be the target state.

[0106] In some embodiments, the request processing device of the multi-core processor further includes: The retention module is used to keep the cache level of the cache line corresponding to the read request address in the directory unchanged; The third determination module is used to determine the first target state of the private cache of the request node according to the request, and determine the second target state of other private caches that hit; The fourth control module is used to control the state of the private cache of the request node to be the first target state, and control the state of other private caches that hit to be the second target state.

[0107] In some embodiments, when the request is a write request, the response module specifically includes: The fifth control module is used to control the state of the private cache of the request node to be the exclusive state; The sixth control module is used to control the request node to write the data corresponding to the write request into its own private cache when it detects that the state of the private cache of the request node is the exclusive state.

[0108] In some embodiments, the request processing device of the multi-core processor further includes: The seventh control module is used to control the request node to evict the shared-state cache line according to the least recently used policy and the policy of preferentially evicting the data in the shared state in its own private cache when it detects that its own private cache is full; After evicting the shared-state data, trigger the sixth control module.

[0109] In some embodiments, the request processing device of the multi-core processor further includes: The third acquisition module is used to acquire the size relationship between the remaining space of the private cache and the space occupied by the data corresponding to the write request; The first detection and triggering module is used to trigger the sixth control module if it detects that the remaining space of the private cache is greater than or equal to the space occupied by the data corresponding to the write request; The first detection and eviction module is used to evict the exclusive cache line if it detects that the remaining space of the private cache is less than the space occupied by the data corresponding to the write request, and return to trigger the third acquisition module.

[0110] In some embodiments, the request processing device of the multi-core processor further includes: An eighth control module, configured to, when detecting that the private cache of the request node itself is in a full state, control the request node to evict shared-state cache lines according to the least recently used policy and the policy of preferentially evicting data in the shared state in the private cache of the request node itself; A fourth acquisition module, configured to, after evicting the shared-state data, acquire the size relationship between the remaining space in the private cache and the space occupied by the target data; A second detection and triggering module, configured to, if detecting that the remaining space in the private cache is greater than or equal to the space occupied by the target data, trigger the transmission module in the second acquisition and transmission module; A second detection and eviction module, configured to, if detecting that the remaining space in the private cache is less than the space occupied by the target data, evict the exclusive cache line and return to trigger the fourth acquisition module.

[0111] In some embodiments, the request processing device of the multi-core processor further includes: A first eviction module, configured to, when detecting that the shared cache is in a full state, evict the target cache line according to the least recently used policy and the policy of preferentially evicting cache lines whose cache level is the level where data is stored in both the shared cache and the private cache and the data state is a clean state in the shared cache; A second change module, configured to change the cache level of the target cache line in the directory from the level where data is stored in both the shared cache and the private cache to the level where data is stored in the private cache; A triggering module, configured to trigger the transmission module in the first acquisition and transmission module; the first acquisition and transmission module is configured to transmit the target data to the shared cache.

[0112] In some embodiments, the request processing device of the multi-core processor includes a second eviction module, configured to evict the target cache line.

[0113] The second eviction module includes: A fifth acquisition module, configured to acquire the cache levels of the respective target cache lines recorded in the directory; A first eviction sub-module, configured to, if there is a target cache line whose cache level is the level where data is stored in both the shared cache and the private cache among all the target cache lines, evict the target cache line whose cache level is the level where data is stored in both the shared cache and the private cache; A second eviction sub-module, configured to, if there is no target cache line whose cache level is the level where data is stored in both the shared cache and the private cache among all the target cache lines, randomly evict the target cache line.

[0114] In some embodiments, the request processing apparatus of the multi-core processor further includes: A fourth determination module, configured to determine the address of the next request according to the address corresponding to the request; A storage module, configured to control the shared cache to prefetch the data corresponding to the address of the next request, and store the prefetched data corresponding to the address of the next request in the shared cache.

[0115] In some embodiments, the request processing apparatus of the multi-core processor further includes: A third eviction module, configured to, when detecting that the shared cache is in a full state, evict the target cache line according to the least recently used policy and the policy of preferentially evicting the cache lines whose cache levels are the levels where data is stored in both the shared cache and the private cache and the data status is the clean status in the shared cache; A third change module, configured to change the cache level of the target cache line in the directory from the level where data is stored in both the shared cache and the private cache to the level where data is stored in the shared cache; trigger the storage module.

[0116] In some embodiments, the request processing apparatus of the multi-core processor further includes: A sixth acquisition module, configured to acquire an array for characterizing the data situation in the private cache and an array for characterizing the data situation in the shared cache established in advance; wherein, each row in the array includes at least an address, the cache status corresponding to the address, and the usage times corresponding to the address.

[0117] In some embodiments, the request processing apparatus of the multi-core processor further includes: A setting module, configured to preset a bypass cache corresponding to the shared cache; A ninth control module, configured to control the request node to transmit the data of the write request to the bypass cache corresponding to the shared cache; A write-back module, configured to write the data of the write request back to the memory through the bypass cache.

[0118] Since the embodiments of the apparatus part correspond to the embodiments of the method part, for the embodiments of the apparatus part, please refer to the description of the embodiments of the method part, and details are not described herein again.

[0119] Figure 7 The structural diagram of the server provided by the embodiments of the present invention. This embodiment is from a hardware perspective. As Figure 7 shown, the server includes: A memory 20, configured to store a computer program; A processor 21, configured to implement the steps of the request processing method of the multi-core processor as mentioned in the above embodiments when executing the computer program.

[0120] Among them, the processor 21 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 21 may be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), and a programmable logic array. The processor 21 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU; the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 21 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 21 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computational operations related to machine learning.

[0121] The memory 20 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 20 may further include a high-speed random access memory and a non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 20 is at least used to store the following computer program 201. After the computer program is loaded and executed by the processor 21, it can implement the relevant steps of the request processing method of the multi-core processor disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 20 may further include an operating system 202 and data 203, etc., and the storage method may be transient storage or permanent storage. Among them, the operating system 202 may include Windows, Unix, Linux, etc. The data 203 may include, but is not limited to, the data involved in the request processing method of the multi-core processor mentioned above.

[0122] In some embodiments, the server may further include a display screen 22, an input / output interface 23, a communication interface 24, a power supply 25, and a communication bus 26.

[0123] Those skilled in the art can understand that Figure 7 the structure shown in

[0124] does not constitute a limitation on the server, and it may include more or fewer components than shown in the figure.

[0125] An embodiment of the present invention further provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above-mentioned request processing method for a multi-core processor are implemented.

[0126] Finally, the present invention also provides a corresponding embodiment of a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps recorded in the above-mentioned method embodiment are implemented.

[0127] It can be understood that if the method in the above-mentioned embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0128] The computer-readable storage medium provided by the present invention includes the above-mentioned request processing method for a multi-core processor, and the effect is the same.

[0129] The above has introduced in detail the request processing method for a multi-core processor, the multi-core processor, the product, and the server provided by the present invention. The various embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part. It should be noted that for those of ordinary skill in the art in the technical field of the present invention, without departing from the principle of the present invention, several improvements and modifications can be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.

[0130] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the said element.

Claims

1. A request processing method for a multi-core processor, characterized in that, Applied to the master node; including: Obtain the request sent by the requesting node; Query the directory according to the request and obtain the query result; wherein, each directory entry includes at least the address, data status, cache level, and the status of the private cache of the slave node; the cache level includes the level where the data is stored in the shared cache, the level where the data is stored in the private cache, and the level where the data is stored in both the shared cache and the private cache; If the query result is a directory miss, obtain the target data corresponding to the request from the memory and transfer the target data to the private cache of the requesting node via the shared cache to respond to the request; If the query result is a shared cache hit, obtain the target data corresponding to the request from the shared cache and transfer the target data to the private cache of the requesting node to respond to the request; If the query result is a hit in other private caches, control the hit other private cache to transfer the target data corresponding to the request to the private cache of the requesting node to respond to the request; wherein, the other private cache is the cache except the private cache of the requesting node among all private caches.

2. The request processing method of the multi-core processor according to claim 1, wherein Before obtaining the request sent by the requesting node, it further includes: The requesting node determines whether the request initiated by itself hits its own private cache; if so, the requesting node uses the target data stored in its own private cache to respond to the request and ends; if not, it sends a request to the master node.

3. The request processing method of the multi-core processor according to claim 2, wherein The request is a read request, and the response to the request includes: Starting from receiving the target data from the private cache of the requesting node, within a preset duration, the requesting node obtains the target data from its own private cache; uses the target data to respond to the request.

4. The request processing method of the multi-core processor according to claim 3, wherein After obtaining the target data corresponding to the request from the memory and transferring the target data to the private cache of the requesting node via the shared cache, it further includes: Mark the cache level of the cache line where the address corresponding to the read request is located in the directory as the level where the data is stored in both the shared cache and the private cache; Determine the target state of the private cache of the requesting node according to the request; Control the state of the private cache of the requesting node to be the target state.

5. The request processing method of the multi-core processor according to claim 3, characterized in that, After obtaining the target data corresponding to the request from the shared cache and transferring the target data to the private cache of the requesting node, it further includes: Change the cache level of the cache line where the address corresponding to the read request is located in the directory from the level where the data is stored in the shared cache to the level where the data is stored in both the shared cache and the private cache; Determine the target state of the private cache of the requesting node according to the request; Control the state of the private cache of the requesting node to be the target state.

6. The request processing method of the multi-core processor according to claim 3, wherein After controlling the hit other private cache to transfer the target data corresponding to the request to the private cache of the requesting node, it further includes: Keep the cache level of the cache line where the address corresponding to the read request is located in the directory unchanged; Determine the first target state of the private cache of the requesting node according to the request, and determine the second target state of the hit other private cache; Control the state of the private cache of the requesting node to the first target state, and control the states of other hit private caches to the second target state.

7. The request processing method of the multi-core processor according to claim 2, characterized in that, The request is a write request, and the target data is the write address corresponding to the write request; The response to the request includes: Control the state of the private cache of the requesting node to the exclusive state; When it is detected that the state of the private cache of the requesting node is the exclusive state, control the requesting node to write the data corresponding to the write request into its own private cache.

8. The request processing method of the multi-core processor according to claim 7, characterized in that, Before the control of the requesting node to write the data corresponding to the write request into its own private cache, it further includes: When it is detected that the private cache of the requesting node itself is in the full state, control the requesting node to evict the shared-state cache line according to the least recently used policy and the policy of preferentially evicting the data in the shared state in its own private cache; After evicting the shared-state data, enter the step of controlling the requesting node to write the data corresponding to the write request into its own private cache.

9. The request processing method of the multi-core processor according to claim 8, wherein After evicting the shared-state data and before entering the step of controlling the requesting node to write the data corresponding to the write request into its own private cache, it further includes: Obtain the size relationship between the remaining space of the private cache and the space occupied by the data corresponding to the write request; If it is detected that the remaining space of the private cache is greater than or equal to the space occupied by the data corresponding to the write request, enter the step of controlling the requesting node to write the data corresponding to the write request into its own private cache; If it is detected that the remaining space of the private cache is less than the space occupied by the data corresponding to the write request, evict the exclusive cache line and return to the step of obtaining the size relationship between the remaining space of the private cache and the space occupied by the data corresponding to the write request.

10. The request processing method of the multi-core processor according to claim 1, characterized in that Before transmitting the target data to the private cache of the requesting node, it further includes: When it is detected that the private cache of the requesting node itself is in the full state, control the requesting node to evict the shared-state cache line according to the least recently used policy and the policy of preferentially evicting the data in the shared state in its own private cache; After evicting the shared-state data, obtain the size relationship between the remaining space of the private cache and the space occupied by the target data; If it is detected that the remaining space of the private cache is greater than or equal to the space occupied by the target data, enter the step of transmitting the target data to the private cache of the requesting node; If it is detected that the remaining space of the private cache is less than the space occupied by the target data, evict the exclusive cache line and return to the step of obtaining the size relationship between the remaining space of the private cache and the space occupied by the target data.

11. The request processing method of the multi-core processor according to claim 1, characterized in that, Before transmitting the target data to the shared cache, it further includes: When it is detected that the shared cache is in a full state, according to the least recently used policy, and in accordance with the policy of preferentially evicting cache lines in the shared cache where the cache level is that the data is stored in both the shared cache and the private cache and the data state is the clean state, evict the target cache line; Change the cache level of the target cache line in the directory from the level where the data is stored in both the shared cache and the private cache to the level where the data is stored in the private cache; Enter the step of transferring the target data to the shared cache.

12. The request processing method of the multi-core processor according to claim 11, wherein, When multiple target cache lines are determined according to the least recently used policy and in accordance with the policy of preferentially evicting cache lines in the shared cache where the cache level is that the data is stored in both the shared cache and the private cache and the data state is the clean state, the eviction of the target cache line includes: Obtain the cache levels of each target cache line recorded in the directory; If there is a target cache line with a cache level where the data is stored in both the shared cache and the private cache among all the target cache lines, evict the target cache line with a cache level where the data is stored in both the shared cache and the private cache; If there is no target cache line with a cache level where the data is stored in both the shared cache and the private cache among all the target cache lines, randomly evict the target cache line.

13. The request processing method of the multi-core processor according to claim 1, characterized in that, It also includes: Determine the address of the next request according to the address corresponding to the request; Control the shared cache to prefetch the data corresponding to the address of the next request and store the prefetched data corresponding to the address of the next request in the shared cache.

14. The request processing method for a multi-core processor according to claim 13, characterized in that, Before storing the prefetched data corresponding to the address of the next request in the shared cache, it also includes: When it is detected that the shared cache is in a full state, according to the least recently used policy, and in accordance with the policy of preferentially evicting cache lines in the shared cache where the cache level is that the data is stored in both the shared cache and the private cache and the data state is the clean state, evict the target cache line; Change the cache level of the target cache line in the directory from the level where the data is stored in both the shared cache and the private cache to the level where the data is stored in the shared cache; Enter the step of storing the prefetched data corresponding to the address of the next request in the shared cache.

15. The request processing method of the multi-core processor according to any one of claims 8 to 11, characterized in that Before evicting the cache line, it also includes: Obtain an array established in advance for characterizing the data situation in the private cache and an array for characterizing the data situation in the shared cache; where at least the address, the cache state corresponding to the address, and the usage times corresponding to the address are included in each row of the array; The eviction of the cache line includes: Based on the array for characterizing the data situation in the private cache or the array for characterizing the data situation in the shared cache, and according to the least recently used policy and the preferential eviction policy, evict the cache line.

16. The request processing method of the multi-core processor according to claim 1, wherein When it is detected that the shared cache is in a full state, the content of the request is to write data to the memory, and the data state corresponding to the data of the write request is the dirty data state, the method also includes: Pre-set a bypass cache corresponding to the shared cache; Control the request node to transfer the data of the write request to the bypass cache corresponding to the shared cache; Write the data of the write request back to the memory through the bypass cache.

17. A multi-core processor, characterized in that, Comprising: A master node, a slave node, a shared cache, a memory, and a private cache corresponding to each node, where the master node is connected to the slave node; the master node is configured to: Obtain a request sent by a requesting node; Query a directory according to the request and obtain a query result; wherein, each directory entry includes at least an address, a data status, a cache level, and a status of the private cache of the slave node; the cache level includes a level where the data is stored in the shared cache, a level where the data is stored in the private cache, and a level where the data is stored in both the shared cache and the private cache; If the query result is a directory miss, obtain the target data corresponding to the request from the memory and transmit the target data to the private cache of the requesting node after passing through the shared cache to respond to the request; If the query result is a shared cache hit, obtain the target data corresponding to the request from the shared cache and transmit the target data to the private cache of the requesting node to respond to the request; If the query result is a hit in other private caches, control the hit other private cache to transmit the target data corresponding to the request to the private cache of the requesting node to respond to the request; wherein, the other private cache is the cache except the private cache of the requesting node among all the private caches.

18. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the steps of the request processing method of the multi-core processor according to any one of claims 1 to 16 are implemented.

19. A server, characterized in that, Comprising: A memory for storing a computer program; A processor for implementing the steps of the request processing method of the multi-core processor according to any one of claims 1 to 16 when executing the computer program.

20. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the request processing method of the multi-core processor according to any one of claims 1 to 16 are implemented.

Citation Information

Patent Citations

  • Directory cache management method for big data application

    CN104461932A

  • Method and device for accessing data visitor directory in multi-core system

    CN106164874A

  • Method for accessing data in multiprocessor system and multiprocessor system

    CN113342709A

  • Access method of multi-level cache system and data storage method and device

    CN115328820A

  • Cache replacement method and device for multi-core processor

    CN117971718A

Cited By

  • Active maintenance method and device for cache consistency, electronic equipment and storage medium

    CN120523749A

  • Processor, display card, computer equipment and constant reading method

    CN120596427A