Cache data reading method, device, computer equipment and storage medium
By adding a cache controller directory to the computer system, maintaining cache consistency between the L2 cache and the cache additional accelerator cache, the problem of increasing cache consistency complexity after the introduction of the cache additional accelerator is solved, and efficient cache consistency maintenance is achieved.
Patent Information
- Application Number
- CN202510265839.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-03-07
AI Technical Summary
The introduction of cache additional accelerator has significantly increased the complexity of cache consistency, resulting in increased system overhead.
By adding the cache controller directory, maintaining cache consistency between the L2 cache and the cache additional accelerator cache, determine the target cache data when the processor core reads the cache data.
Maintaining efficient cache consistency without significantly increasing system overhead solves the problem of increasing cache consistency complexity after the introduction of cache additional accelerator.
Smart Images

Figure CN119759807B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a cache data reading method, device, computer equipment and storage medium. Background Art
[0002] In some high-performance computing systems, accelerators are integrated into the cache hierarchy as cache-attached accelerators, which can significantly reduce the distance and time of data transmission, thereby improving the overall performance of the system. There is a cache of the cache-attached accelerator in the cache-attached accelerator, and the computing unit has L1 cache, L2 cache, and L3 cache. The cache-attached accelerator shares a unified address space with the CPU (Central Processing Unit) core. If these caches store copies of data at the same memory address at the same time, when one component modifies the data, other components may still use the old copy, resulting in data inconsistency.
[0003] In order to ensure that the data accessed at the same address are consistent when the cache attached accelerator and the CPU core read the cache data, it is necessary to maintain cache consistency between the cache attached accelerator and the cache of the CPU core. Due to the introduction of the cache attached accelerator, the complexity of cache consistency increases significantly, resulting in increased system overhead. Summary of the invention
[0004] In view of this, the present invention provides a cache data reading method, apparatus, computer device and storage medium to solve the problem that the complexity of cache consistency is significantly increased due to the introduction of a cache additional accelerator, resulting in increased system overhead.
[0005] In a first aspect, the present invention provides a cache data reading method, the method comprising:
[0006] Get the first read data request;
[0007] querying first cache data stored in a first cache of a processor core and second cache data stored in a cache of a cache-attached accelerator according to a first read data request;
[0008] Obtaining a first cache state of the first cache and a second cache state of the cache controller directory record, wherein the second cache state is determined based on the first cache state and a cache state of the cache attached accelerator cache;
[0009] Based on the first cache data hitting the third cache data corresponding to the first read data request, the second cache data hitting the third cache data, the first cache state and the second cache state, target cache data corresponding to the first read data request is determined.
[0010] In a second aspect, the present invention provides a cache data reading device, the device comprising:
[0011] A request acquisition module, used for acquiring a first read data request;
[0012] A query module, configured to query first cache data stored in a first cache of a processor core and second cache data stored in a cache of a cache-attached accelerator according to a first read data request;
[0013] A state acquisition module, used to acquire a first cache state of the first cache and a second cache state recorded in the cache controller directory, wherein the second cache state is determined according to the first cache state and the cache state of the cache attached accelerator cache;
[0014] The cache data determination module is used to determine the target cache data corresponding to the first read data request based on the situation that the first cache data hits the third cache data corresponding to the first read data request, the situation that the second cache data hits the third cache data, the first cache state and the second cache state.
[0015] In a third aspect, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, computer instructions being stored in the memory, and the processor executing the cache data reading method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0016] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the cache data reading method of the first aspect or any corresponding embodiment thereof.
[0017] In a fifth aspect, the present invention provides a computer program product, comprising computer instructions for causing a computer to execute the cache data reading method of the first aspect or any corresponding embodiment thereof.
[0018] The cache data reading method provided in this embodiment, when the processor core reads the cache data, first determines the first cache state of the first cache and the second cache state recorded in the cache controller directory; then determines the situation where the first cache data hits the third cache data corresponding to the first read data request and the situation where the second cache data hits the third cache data; finally, determines the target cache data corresponding to the first read data request based on the above information. In this embodiment, the cache consistency maintenance between the cache attached accelerator cache and the first cache is achieved only by adding a cache control directory, and does not add additional states and complex logic to the last-level shared cache, solving the cache consistency problem caused by near data calculations at a minimal cost. It has the effect of maintaining efficient cache consistency without significantly increasing system overhead. It solves the problem that the complexity of cache consistency is significantly increased due to the introduction of a cache attached accelerator, resulting in increased system overhead. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the related technologies, the drawings required for use in the specific embodiments or the related technical descriptions will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0020] Figure 1 is a schematic diagram of a computing unit architecture according to an embodiment of the present invention;
[0021] Figure 2 is a schematic diagram of a cache-attached accelerator composition architecture according to an embodiment of the present invention;
[0022] Figure 3 is a schematic diagram of a cache access relationship in a computing unit according to an embodiment of the present invention;
[0023] Figure 4 is a flow chart of a cache data reading method according to an embodiment of the present invention;
[0024] Figure 5 is a flow chart of a CPU core reading data according to an embodiment of the present invention;
[0025] Figure 6 is a flow chart of a cache-attached accelerator reading data according to an embodiment of the present invention;
[0026] Figure 7 is a flowchart of L2 cache replacement according to an embodiment of the present invention;
[0027] Figure 8 is a flow chart of cache replacement of a cache attachment accelerator according to an embodiment of the present invention;
[0028] Fig. 9 is a flow chart of L3 cache reading data according to an embodiment of the present invention;
[0029] Fig.10 is a structural block diagram of a cache data reading device according to an embodiment of the present invention;
[0030] Fig.11 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0031] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0032] Currently, the performance improvement rate of processors far exceeds that of memories. The bandwidth limitation and access delay between processors and memories lead to bottlenecks in system performance improvement, which is called the "memory wall" problem. Near-Data Computing (NDC) is used to solve the "memory wall" problem. Near-Data Computing refers to reducing the distance and frequency of data transmission by placing computing logic as close as possible to the data storage location, thereby reducing latency and energy consumption. Unlike the traditional architecture where the processor and memory are separated, near-data computing tightly integrates the processor and memory, significantly improving data access efficiency.
[0033] In some high-performance computing systems, a near-data computing strategy is adopted to integrate accelerators into the cache hierarchy, which are called cache-attached accelerators. These cache-attached accelerators share a unified address space with the CPU core, which can significantly reduce the distance and time of data transmission, thereby improving the overall performance of the system. In order to maintain the consistency of the shared cache with the main core, the cache-attached accelerator must implement a cache consistency protocol to effectively manage the consistency and correctness of the data, thereby ensuring the optimization of system performance and resource utilization. However, due to the introduction of cache-attached accelerators, the complexity of cache coherence has increased significantly. The computing unit architecture containing cache-attached accelerators is shown in Figure 1. Figure 1 As shown, Figure 1The computing unit architecture contains an acceleration engine. The computing unit consists of a CPU core, L2 cache, cache controller directory, cache-attached accelerator, and L3 cache. L1 and L2 caches are private caches. The L1 cache is inside the CPU core. The L2 cache and the cache in the cache-attached accelerator need to maintain cache consistency. The L3 cache is a shared cache for multiple computing units in the processor and is also the last level cache. The cache-attached accelerator consists of a dedicated processing unit and a cache-attached accelerator cache, such as Figure 2 As shown in Figure 1, cache-attached accelerators can significantly improve performance in many compute-intensive and data-intensive application scenarios, such as image and video processing, machine learning and artificial intelligence, and big data processing. The dedicated processing unit in the cache-attached accelerator can directly read data from the cache-attached accelerator cache and write the calculation results back to the cache-attached accelerator cache to achieve near data computing. The memory access relationship between caches is shown in Figure 1. Figure 3 As shown, the L2 cache in the computing unit, the cache attached to the accelerator cache, and the L3 cache can access each other, and the caches can read data from each other.
[0034] Based on the above content, an embodiment of the present invention provides a cache data reading method, which uses a cache controller directory to maintain the L2 cache and the cache attached accelerator cache, and records the highest permission status in the L2 cache and the cache attached accelerator cache. When the CPU core and the cache attached accelerator read the cache data, they obtain the required data from the L2 cache, the cache attached accelerator cache, and the L3 cache according to the cache status recorded in the cache controller directory. All modifications of the present invention are limited to the computing unit (tile) where the accelerator is located. The cache consistency between the CPU L2 cache and the cache attached accelerator cache is maintained only by adding a cache controller directory, and no additional states and complex logic are added to the last-level shared cache, thereby solving the cache consistency problem caused by near data calculations at a minimum cost. This achieves the effect of maintaining efficient cache consistency without significantly increasing system overhead.
[0035] According to an embodiment of the present invention, a cache data reading embodiment is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer device with data processing capabilities, such as: a computer, a server, a mobile terminal, etc., and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0036] In this embodiment, a cache data reading method is provided, which can be used in the above-mentioned computer device. Figure 4 is a flow chart of a cache data reading method according to an embodiment of the present invention. Figure 4 As shown, the process includes the following steps:
[0037] Step S401: obtaining a first data read request.
[0038] Specifically, the processor core is, for example, a CPU core. When the processor core reads cache data, the processor core sends a first read data request, first queries the CPU L1 cache, and if the CPU L1 cache hits, the CPU L1 cache directly replies to the cache data read by the processor core. If the CPU L1 cache misses, the target cache data corresponding to the first read data request is read from the first cache of the processor core and the cache-attached accelerator cache. The first cache of the processor core is, for example, an L2 cache.
[0039] The above process is as follows Figure 5 As shown, the CPU core reads data and queries the CPU L1 cache. If the CPU L1 cache hits, the CPU L1 cache replies with data. If the CPU L1 cache misses, the CPU L2 cache and cache controller directory are queried.
[0040] Step S402: query first cache data stored in a first cache of a processor core and second cache data stored in a cache of a cache-attached accelerator according to a first read data request.
[0041] Specifically, the first cache can only be written directly by the processor core, and data generated by the dedicated processing unit in the cache-attached accelerator cannot be directly written into the first cache. Similarly, the cache-attached accelerator cache can only be written directly by the dedicated processing unit, and the processor core cannot directly write into the cache-attached accelerator cache.
[0042] The first cache data stored in the first cache of the processor core and related to the first read data request are queried, and the second cache data stored in the cache of the cache-attached accelerator and related to the first read data request are queried.
[0043] Step S403, obtaining a first cache state of the first cache and a second cache state recorded in the cache controller directory, wherein the second cache state is determined according to the first cache state and the cache state of the cache attached accelerator cache.
[0044] Specifically, the cache controller directory maintains the L2 cache and the cache attached accelerator cache, and records the highest cache state in the L2 cache and the cache attached accelerator cache. The cache states of the cache in this embodiment are, for example: invalid state (Invalid, I), clean data shared state (Shared Clean, SC), dirty data shared state (Shared Dirty, SD), clean data exclusive state (Unique Clean, UC) and dirty data exclusive state (Unique Dirty, UD), where the SC state is a shared state, the cache line data is clean, and there may be cache line copies in other cores; the SD state is a shared state, the cache line data is dirty, and there may be cache line copies in other cores; the UC state is an exclusive state, the cache line data is clean, and there are no cache line copies in other cores; the UD state is an exclusive state, the cache line data is dirty, and there are no cache line copies in other cores. In a computer system, the state of cache data can be divided into two types: dirty (Dirty) and clean (Clean), which is used to describe whether the data in the cache is consistent with the data in the memory. Clean data indicates that the data is consistent with the data in the memory, and dirty data indicates that the data is inconsistent with the data in the memory.
[0045] After receiving the read request, the L2 cache reads the target cache data corresponding to the first read data request from the first cache of the processor core and the cache attached accelerator cache, and first queries the first cache state of the first cache and the second cache state recorded in the cache controller directory. The first cache state can be used to determine whether the data stored in the first cache is the latest data, that is, whether the data stored in the first cache has been processed. Because the cache controller directory maintains the L2 cache and the cache attached accelerator cache, the first cache state and the second cache state can be combined to determine whether the data stored in the cache attached accelerator cache is the latest data.
[0046] Step S404, determining target cache data corresponding to the first read data request based on whether the first cache data hits the third cache data corresponding to the first read data request, whether the second cache data hits the third cache data, the first cache state, and the second cache state.
[0047] Specifically, determining that the first cache data hits the third cache data corresponding to the first read data request can determine whether the first cache contains the third cache data. Determining that the second cache data hits the third cache data can determine whether the cache of the cache add-on accelerator contains the third cache data.
[0048] If the first cache data hits the third cache data, and the first cache state is UD state, UC state or SD state, the third cache data in the first cache is the target cache data and directly replies to the processor core. If the first cache data hits the third cache data, the first cache state is SC state, and the second cache state is SD state, a snoop request is sent to the cache attached accelerator cache, and if the second cache data hits the third cache data, the third cache data in the cache attached accelerator cache is the target cache data, and the cache attached accelerator cache directly replies to the processor core. If the first cache data does not hit the third cache data, and the cache controller directory records the data, the third cache data in the cache attached accelerator cache is the target cache data, and a snoop request is sent to the cache attached accelerator cache, and the cache attached accelerator cache directly replies to the processor core. If neither the cache attached accelerator cache nor the first cache contains the third cache data, the target cache data is read from the L3 cache.
[0049] The cache data reading method provided in this embodiment, when the processor core reads the cache data, first determines the first cache state of the first cache and the second cache state recorded in the cache controller directory; then determines the situation where the first cache data hits the third cache data corresponding to the first read data request and the situation where the second cache data hits the third cache data; finally, determines the target cache data corresponding to the first read data request based on the above information. In this embodiment, the cache consistency maintenance between the cache attached accelerator cache and the first cache is achieved only by adding a cache control directory, and does not add additional states and complex logic to the last-level shared cache, solving the cache consistency problem caused by near data calculations at a minimal cost. It has the effect of maintaining efficient cache consistency without significantly increasing system overhead. It solves the problem that the complexity of cache consistency is significantly increased due to the introduction of a cache attached accelerator, resulting in increased system overhead.
[0050] In some optional embodiments, the method further comprises:
[0051] In a case where there is a second read data request from the processing unit and the cache of the cache-attached accelerator does not contain the first target data corresponding to the second read data request, determining whether the first cache contains the first target data;
[0052] If the first cache contains the first target data, the first target data in the first cache is sent to the processing unit, and the second cache state is modified to a third target state in the cache controller directory, wherein the third target state is determined based on the first cache state and the cache state of the cache-attached accelerator cache.
[0053] Specifically, when the cache-attached accelerator reads cache data, the cache-attached accelerator sends a second read data request, first queries the cache-attached accelerator cache, and if the cache-attached accelerator cache hits, that is, the cache-attached accelerator cache contains the first target data corresponding to the second read data request, since the data cached by the cache-attached accelerator are all dirty data, the cache-attached accelerator directly replies with the data. The cache states recorded in the cache-attached accelerator cache are all dirty states, including data written to the cache-attached accelerator cache through dedicated processing unit operations and data read from the CPU L2 cache. When the cache-attached accelerator cache wants to read data, if the cache does not hit, it will first query the CPU L2 cache. If the cache-attached accelerator cache does not hit, it will read data from the second cache. The second cache is, for example, L3 cache.
[0054] If the cache of the cache-attached accelerator does not hit, a snoop request is sent to the first cache. It is determined whether the first cache contains the first target data. If the first cache contains the first target data, the first cache directly replies with data, sends the first target data in the first cache to the processing unit, and modifies the second cache state to the third target state in the cache controller directory, for example: the first cache state of the first cache is a clean data exclusive state, and the cache state of the cache of the cache-attached accelerator is a dirty data shared state, then the third target state is a dirty data shared state.
[0055] If the first cache misses, a read request is sent to the shared cache L3. After the CPU L3 cache hits, the data is directly sent to the cache attached accelerator cache and the cache controller directory is updated at the same time.
[0056] The above process is as follows Figure 6 As shown, the cache attached accelerator reads data; queries the cache attached accelerator cache, if it hits, the cache attached accelerator cache replies data, if it does not hit, sends a snoop request to the CPU L2 cache; queries the CPU L2 cache for a hit, if it hits, the CPU L2 cache replies data and updates the cache controller directory as needed, if it does not hit, sends a snoop request to the CPU L3 cache.
[0057] In this embodiment, after reading the cache data in the first cache, the processing unit promptly updates the cache status recorded in the cache controller directory to ensure the consistency of data copies between multi-level caches and avoid errors caused by dirty data or status conflicts in subsequent accesses.
[0058] In some optional implementations, modifying the second cache state to a third target state in the cache controller directory includes:
[0059] When the first cache state is the first preset state, the second preset state, or the third preset state, and the second cache state is the first preset state, the second preset state, or the third preset state, modify the cache state of the first cache to a fourth preset state, use the third preset state as a third target state, and modify the second cache state to the third target state in the cache controller directory;
[0060] When the first cache state is the fourth preset state and the second cache state is the fourth preset state, the cache state of the first cache is maintained in the fourth preset state, the third preset state is used as the third target state, and the second cache state is modified to the third target state in the cache controller directory.
[0061] Specifically, the first preset state is, for example, a dirty data exclusive state. The second preset state is, for example, a clean data exclusive state. The third preset state is, for example, a dirty data shared state. The fourth preset state is, for example, a clean data shared state.
[0062] If the first cache state is the first preset state or the second preset state or the third preset state, the cache state changes to the fourth preset state after the first cache sends the second target data to the cache attached accelerator, and the cache state changes to the third preset state after the cache attached accelerator receives the data, and the third preset state is used as the third target state, and the second cache state is modified to the third target state in the cache controller directory.
[0063] If the first cache state is the fourth preset state and the second cache state is the fourth preset state, the cache state remains in the fourth preset state after the first cache sends the second target data to the cache attached accelerator, the third preset state is used as the third target state, and the second cache state is modified to the third target state in the cache controller directory.
[0064] In some optional implementations, after determining whether the first cache contains the first target data, the method further includes:
[0065] In the case that the first cache does not contain the first target data, sending a first read request to the second cache, wherein the first read request is used to determine whether the second cache contains the first target data;
[0066] In the case where the second cache contains the first target data, the first target data in the second cache is sent to the cache-attached accelerator cache, and the cache-attached accelerator cache sends the first target data to the processing unit and updates the cache controller directory;
[0067] In a case where the second cache does not contain the first target data, the first target data is read from the memory and sent to the processing unit.
[0068] Specifically, the second cache is, for example, an L3 cache, which is a shared cache in a computing unit. If the first cache does not contain the first target data, a first read request is sent to the second cache to determine whether the second cache contains the first target data.
[0069] If the first cache does not contain the first target data, a first read request is sent to the second cache to determine whether the second cache contains the first target data. If the data in the second cache hits the first target data, the second cache directly sends the first target data to the cache-attached accelerator cache, which sends the first target data to the processor core and updates the cache controller directory.
[0070] If the second cache does not contain the first target data, the first target data is read from the memory and sent to the processing unit.
[0071] The above process is as follows Figure 6 As shown, the CPU L2 cache is queried for a hit. If not, a snoop request is sent to the CPU L3 cache. The CPU L3 cache is queried for a hit. If it is, the CPU L3 cache replies with data and updates the cache controller directory as required; if not, data is read from the memory.
[0072] In some optional implementations, determining target cache data corresponding to the first read data request based on whether the first cache data hits the third cache data corresponding to the first read data request, whether the second cache data hits the third cache data, the first cache state, and the second cache state includes:
[0073] Determine whether the first cache includes the third cache data according to whether the first cache data hits the cache data corresponding to the first read data request, and determine whether the cache of the cache-attached accelerator includes the third cache data according to whether the second cache data hits the third cache data;
[0074] When the first cache contains the third cache data and the first cache state is the first target state, using the third cache data in the first cache as the target cache data;
[0075] In a case where the first cache contains the third cache data, the first cache state is not the first target state, and the second cache state is the second target state, using the third cache data in the cache of the cache-attached accelerator as the target cache data;
[0076] When the first cache does not contain the third cache data and the cache-attached accelerator cache contains the third cache data, sending a first snoop request to the cache-attached accelerator cache so that the cache-attached accelerator cache uses the third cache data as target cache data and sends it to the processor core;
[0077] In a case where the first cache does not contain the third cache data and the cache-attached accelerator cache does not contain the third cache data, a second read request is sent to the second cache, wherein the second read request is used to use the third cache data as target cache data and send it to the first cache when the second cache contains the third cache data, and the first cache sends the target cache data to the processor core and updates the cache controller directory.
[0078] Specifically, if the first cache data hits the cache data corresponding to the first read data request, it can be determined that the first cache contains the third cache data, otherwise, the first cache does not contain the third cache data. If the second cache data hits the third cache data, it is determined whether the cache of the cache-attached accelerator contains the third cache data.
[0079] The first target state is, for example, a dirty data exclusive state, a clean data exclusive state, or a dirty data shared state. If the first cache contains the third cache data, and the first cache state is the first target state, the third cache data in the first cache is the target cache data and is directly replied to the processor core.
[0080] The second target state is, for example, a dirty data sharing state. If the first cache contains the third cache data, the first cache state is not the first target state, and the second cache state is the second target state, a listening request is sent to the cache-attached accelerator cache, and if the second cache data hits the third cache data, the third cache data in the cache-attached accelerator cache is the target cache data, and the cache-attached accelerator cache directly replies to the processor core.
[0081] If the first cache data does not hit the third cache data, and the cache-attached accelerator cache contains the third cache data, the third cache data in the cache-attached accelerator cache is the target cache data, and a first monitoring request is sent to the cache-attached accelerator cache, which takes the third cache data as the target cache data and sends it to the processor core.
[0082] The second cache is, for example, an L3 cache. If neither the first cache nor the cache attached to the cache accelerator contains the third cache data, a second read request is sent to the second cache to read the target cache data from the second cache. The reading process includes: if the second cache contains the third cache data, the third cache data is used as the target cache data and sent to the first cache, and the first cache sends the target cache data to the processor core and updates the cache controller directory. If the second cache does not contain the third cache data, the target cache data needs to be read from the memory.
[0083] The above process is as follows Figure 5As shown, query the CPU L2 cache and cache controller directory. If the CPU L2 cache and cache controller directory do not hit, the CPU L2 cache does not hit, and the cache controller directory does not hit, and a request is sent to the CPU L3 cache; query the CPU L3 cache. If it hits, the CPU L3 cache replies data and updates the cache controller directory according to the demand. If it does not hit, read the data from the memory. If the CPU L2 cache and cache controller directory hit, if the CPU L2 cache hits and the cache state is UD state, UC state or SD state, the CPU L2 cache replies data; if the CPU L2 cache hits the SC state, the cache controller directory is in the SD state, and the cache attached accelerator cache replies data; the CPU L2 cache does not hit, the cache controller directory hits, and the cache attached accelerator cache replies data.
[0084] In this embodiment, the method of reading cache data is selected according to the first cache state and the second cache state, and cache consistency between the first cache and the cache attached accelerator cache is maintained by adding a cache control directory, thereby achieving cache consistency at a minimum cost.
[0085] In some optional implementations, updating the cache controller directory includes:
[0086] Setting a cache state of the first cache to a fourth target state, wherein the fourth target state is determined according to the first cache state;
[0087] Determining a fifth target state according to a preset high-low rule, a cache state of the first cache, and a cache state of a cache attached to an accelerator cache;
[0088] The second cache state of the cache controller directory is set to a fifth target state.
[0089] Specifically, the second cache sends the third cache data to the first cache, and the first cache sends the target cache data to the processor core, and the processor core processes the third cache data in the first cache. If the first cache state before the first cache is a clean data sharing state, the fourth target state is a clean data exclusive state.
[0090] The preset high and low rules are, for example: dirty data exclusive state > dirty data shared state > clean data exclusive state > clean data shared state. According to the preset high and low rules, the cache state of the first cache and the cache state of the cache attached to the accelerator cache, the fifth target state is determined, for example: the cache state of the first cache is the clean data exclusive state, and the cache state of the cache attached to the accelerator cache is the dirty data exclusive state, then the fifth target state is the dirty data exclusive state.
[0091] The second cache state of the cache controller directory is set to a fifth target state.
[0092] In this implementation, after the cache data is read, the cache controller directory is updated in a timely manner to ensure that the cache status recorded in the cache controller directory can reflect the actual status of the cache.
[0093] In some optional embodiments, the method further comprises:
[0094] In the case where there is a first cache replacement request, determining first data to be replaced in the first cache, determining a first current cache state of the first cache, and acquiring a second current cache state recorded in the cache controller directory;
[0095] When the first current cache state is the first preset state or the third preset state, writing the first data to be replaced back to the second cache, deleting the first data to be replaced in the first cache, and updating the cache controller directory;
[0096] When the first current cache state is the second preset state or the fourth preset state, and the second current cache state is the first preset state or the third preset state, deleting the first to-be-replaced data in the first cache;
[0097] When the first current cache state is the second preset state or the fourth preset state and the second current cache state is the second preset state or the fourth preset state, the first data to be replaced is written back to the second cache and the first data to be replaced is deleted from the first cache.
[0098] Specifically, when the CPU L2 cache or the cache-attached accelerator cache performs a replacement operation, it queries the cache controller directory and determines whether the replaced cache is written back to the L3 cache or directly discarded according to the state of the cache control directory record.
[0099] The first cache is, for example, a CPU L2 cache. The first cache replacement request is for replacing data in the first cache. After receiving the first cache replacement request, the first cache first queries the first cache and the cache controller directory to determine the first current cache state of the first cache, obtains the second current cache state recorded in the cache controller directory, and determines the first data to be replaced in the first cache.
[0100] The first preset state is, for example, a dirty data exclusive state. The second preset state is, for example, a clean data exclusive state. The third preset state is, for example, a dirty data shared state. The fourth preset state is, for example, a clean data shared state.
[0101] When the first current cache state is the first preset state or the third preset state, it means that the first cache has data and the cache state is dirty, the first data to be replaced in the first cache is written back to the second cache, and the cache controller directory is updated at the same time.
[0102] When the first current cache state is the second preset state or the fourth preset state, and the second current cache state is the first preset state or the third preset state, it means that the first cache has data and the cache state is clean, and the cache controller directory record is in a dirty state, then the data in the first cache is directly discarded, that is, the first data to be replaced is deleted in the first cache, and the cache controller directory state is not changed.
[0103] When the first current cache state is the second preset state or the fourth preset state, and the second current cache state is the second preset state or the fourth preset state, it means that the first cache has data and the cache state is clean, and the cache controller directory is recorded as a clean state, the first data to be replaced is written back to the second cache, and the first data to be replaced is deleted in the first cache.
[0104] The above process is as follows Figure 7 As shown, the CPU L2 cache is replaced; the CPU L2 cache is queried whether there is dirty data, if so, it is written back to the L3 cache and the cache controller directory is updated, if not, the cache controller directory is queried whether the dirty state is recorded, if so, the data in the CPU L2 cache is directly discarded, if not, the replacement data is written back to the L3 cache.
[0105] In this embodiment, when performing a replacement operation on the first cache or the cache-attached accelerator cache, the cache controller directory is queried. According to the status of the cache control directory record, it is determined whether the replaced cache is written back to the second cache or directly discarded to avoid data loss.
[0106] In some optional embodiments, the method further comprises:
[0107] In the case that there is a cache replacement request of the cache attachment accelerator, determining second data to be replaced in the cache of the cache attachment accelerator;
[0108] The second data to be replaced is written back to the second cache, and the second data to be replaced is deleted from the cache attached accelerator cache.
[0109] Specifically, the cache addition accelerator cache replacement request is for replacing data in the cache of the cache addition accelerator.
[0110] After receiving the cache additional accelerator cache replacement request, the cache additional accelerator cache directly writes the second data to be replaced back to the second cache because the data in the cache additional accelerator cache are all dirty, and deletes the second data to be replaced in the cache additional accelerator cache.
[0111] The above process is as follows Figure 8 As shown, the cache-attached accelerator performs cache replacement, and the replacement data is written back to the L3 cache.
[0112] In some optional implementations, the specific process of the cache attachment accelerator caching replacement data may include steps A1 to A4.
[0113] Step A1, receiving a data replacement request of a cache attached accelerator; the data replacement request includes an address of a first cache line requested to be replaced.
[0114] Step A2: Obtain a second cache line in other caches that corresponds to the address of the first cache line and has multiple copies.
[0115] Specifically, other caches such as L2 cache, L3 cache, etc., use a second cache line corresponding to the address of the first cache line (the corresponding method is the same index) and having multiple copies as a candidate cache line to be replaced. Since there are multiple copies, replacing the second cache line with multiple copies will not affect the access efficiency of the second cache line, because the copy still exists in the cache system.
[0116] Step A3: selecting a cache line from the acquired second cache lines according to a predetermined selection strategy as a target cache line for replacing the first cache line.
[0117] Specifically, the second cache line that is used the least within a predetermined time period before the current moment is selected as the target cache line, because the cache line that is frequently used recently is more easily accessed in subsequent cache line accesses, thereby improving the access hit rate of the cache line.
[0118] Step A4, reading target cache line information, and writing the first cache line data and state information into the target cache line.
[0119] In this embodiment, the cache is made to replace cache lines with multiple copies first, and in the case of a cache line with only one copy, the cache line that is relatively infrequently accessed is replaced based on age, thereby improving the cache application efficiency and cache hit rate. Cache lines with only one copy are replaced first, while cache lines with multiple copies are retained to avoid affecting the cache hit rate.
[0120] In some optional implementations, after determining the target cache data corresponding to the first read data request, the method further includes:
[0121] Determining a calculation result from the processor core, wherein the calculation result is obtained after the processor core performs a calculation operation based on the target cache data;
[0122] The calculation result is written into the first cache, and the cache state and cache controller directory of the first cache are updated.
[0123] Specifically, the processor core can directly read data from the first cache and perform calculation operations, and finally write the results back to the first cache, and update the cache controller directory at the same time. For example: the processor core obtains the calculation result after performing a calculation operation on the target cache data, determines the calculation result from the processor core, writes the calculation result to the first cache, and updates the cache status of the first cache and the cache controller directory.
[0124] Similarly, the dedicated processing unit in the cache-attached accelerator can directly read data from the cache-attached accelerator cache and perform calculation operations, and finally write the results back to the cache and update the cache controller directory at the same time.
[0125] In this embodiment, the processor core can directly read data from the first cache and perform calculation operations, and finally write the calculation results back to the corresponding cache, and update the cache controller directory at the same time, so as to ensure that the latest cache data can be read later and the cache consistency between caches is guaranteed.
[0126] In some optional embodiments, the method further comprises:
[0127] In the case where there is a third read data request of the second cache, determining whether the first cache contains the second target data corresponding to the third read data request, and obtaining a hit result;
[0128] Determine a third cache state of the first cache, and obtain a fourth cache state recorded in the cache controller directory;
[0129] According to the hit result, the third cache state and the fourth cache state, the second target data is obtained from the first cache or the cache-attached accelerator cache, and the second target data is sent to the second cache.
[0130] Specifically, the second cache is, for example, an L3 cache. The L3 cache needs to read data from the upper level cache, and will query the cache controller directory and the CPU L2 cache at the same time, and select a data reading method according to the cache controller directory record information and the CPU L2 cache query result.
[0131] The second cache sends a third read data request to the first cache, indicating that the first cache must have the required data. After receiving the third read data request, the first cache queries the first cache and the cache controller directory at the same time to determine whether the first cache contains the second target data corresponding to the third read data request, and obtains a hit result. According to the hit result, it can be determined whether the first cache contains the second target data.
[0132] In addition, the third cache state of the first cache and the fourth cache state recorded in the cache controller directory are queried, and the third cache state can be used to determine whether the data stored in the first cache is the latest data, that is, whether the data stored in the first cache has been processed. Combining the third cache state and the fourth cache state can determine whether the data stored in the cache of the cache attachment accelerator is the latest data, that is, the second target data.
[0133] According to the hit result, the third cache state and the fourth cache state, the second target data is obtained from the first cache or the cache-attached accelerator cache, for example: if the hit result is a miss, if the cache-attached accelerator cache contains the second target data, the second target data is obtained from the cache-attached accelerator cache; if the hit result is a hit, and the third cache state is a dirty data exclusive state, the second target data is obtained from the first cache. The second target data is sent to the second cache.
[0134] The above process is as follows Fig. 9 As shown, the CPU L3 cache reads data; queries the CPU L2 cache and the cache controller directory; and obtains the second target data according to the CPU L2 cache hit result, the third cache state, and the fourth cache state.
[0135] In some optional implementations, obtaining the second target data from the first cache or the cache-attached accelerator cache according to the hit result, the third cache state, and the fourth cache state includes:
[0136] When the hit result is a hit and the third cache state is the first preset state, the second preset state, or the third preset state, acquiring the second target data in the first cache;
[0137] When the hit result is a hit, the third cache state is the fourth preset state, and the fourth cache state is the third preset state, a second monitoring request is sent to the cache-attached accelerator cache, and the second target data is obtained from the cache-attached accelerator cache by using the second monitoring request;
[0138] When the hit result is a miss, if the cache-attached accelerator cache contains the second target data, a second snoop request is sent to the cache-attached accelerator cache, and the second target data is acquired from the cache-attached accelerator cache using the second snoop request.
[0139] Specifically, the first preset state is, for example, a dirty data exclusive state. The second preset state is, for example, a clean data exclusive state. The third preset state is, for example, a dirty data shared state. The fourth preset state is, for example, a clean data shared state.
[0140] If the hit result is a hit, and the third cache state is the first preset state, the second preset state, or the third preset state, the first cache directly sends the second target data to the third cache.
[0141] If the hit result is a hit, the third cache state is the fourth preset state, and the fourth cache state is the third preset state, a second listening request is sent to the cache-attached accelerator cache, and the second target data is obtained from the cache-attached accelerator cache using the second listening request, and the cache-attached accelerator cache directly replies the second target data to the second cache.
[0142] When the hit result is a miss, if the cache-attached accelerator cache contains the second target data, a second listening request is sent to the cache-attached accelerator cache, the second target data is obtained from the cache-attached accelerator cache using the second listening request, and the cache-attached accelerator cache directly replies the second target data to the second cache.
[0143] The above process is as follows Fig. 9 As shown, if the CPU L2 cache hits and the cache state is UD state, UC state or SD state, the CPU L2 cache replies data; if the CPU L2 cache hits in SC state, the cache controller directory hits in SD state, the cache attached accelerator cache replies data; if the CPU L2 cache misses, the cache controller directory hits, and the cache attached accelerator cache replies data.
[0144] In some optional implementations, before querying the first cache data stored in the first cache of the processor core and the second cache data stored in the cache of the cache-attached accelerator according to the first read data request, the method further includes:
[0145] determining a third current cache state of the second cache and a fourth current cache state of the cache-attached accelerator cache;
[0146] A higher cache state between the third current cache state and the fourth current cache state is determined according to a preset high-low rule, and the higher cache state is recorded in the cache controller directory, wherein the higher cache state includes the second cache state.
[0147] Specifically, the cache controller directory records the highest status of the CPU L2 cache and the cache-attached accelerator cache.
[0148] Preset high and low rules, for example: dirty data exclusive state > dirty data shared state > clean data exclusive state > clean data shared state. Determine the third current cache state of the second cache and the fourth current cache state of the cache-attached accelerator cache. Determine the higher cache state of the third current cache state and the fourth current cache state according to the preset high and low rules, for example: if the third current cache state is the clean data exclusive state and the fourth current cache state is the dirty data exclusive state, then the higher cache state is the dirty data exclusive state; if the third current cache state is the dirty data shared state and the fourth current cache state is the clean data shared state, then the higher cache state is the dirty data shared state.
[0149] In this embodiment, the cache controller directory is used to record the higher cache states of the first cache and the cache attached accelerator cache, and cache consistency between caches is achieved only by adding a cache control directory, thereby solving the cache consistency problem caused by near data calculation at a minimum cost.
[0150] In this embodiment, a cache data reading device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.
[0151] This embodiment provides a cache data reading device, such as Fig.10 As shown, including:
[0152] A request acquisition module 1001 is used to acquire a first read data request;
[0153] A query module 1002, configured to query first cache data stored in a first cache of a processor core and second cache data stored in a cache of a cache-attached accelerator according to a first read data request;
[0154] A state acquisition module 1003 is used to acquire a first cache state of the first cache and a second cache state recorded in the cache controller directory, wherein the second cache state is determined according to the first cache state and the cache state of the cache attached accelerator cache;
[0155] The cache data determination module 1004 is used to determine the target cache data corresponding to the first read data request based on the situation that the first cache data hits the third cache data corresponding to the first read data request, the situation that the second cache data hits the third cache data, the first cache state and the second cache state.
[0156] In some optional embodiments, the device also includes: a first judgment module, used to judge whether the first cache contains the first target data when there is a second read data request from the processing unit and the cache-attached accelerator cache does not contain the first target data corresponding to the second read data request; a first sending module, used to send the first target data in the first cache to the processing unit if the first cache contains the first target data, and modify the second cache state to a third target state in the cache controller directory, wherein the third target state is determined based on the first cache state and the cache state of the cache-attached accelerator cache.
[0157] In some optional embodiments, the first sending module includes: a first modification unit, which is used to modify the cache state of the first cache to a fourth preset state, using the third preset state as the third target state, and modifying the second cache state to the third target state in the cache controller directory when the first cache state is the first preset state, the second preset state, or the third preset state and the second cache state is the first preset state, the second preset state, or the third preset state; a second modification unit, which is used to keep the cache state of the first cache in the fourth preset state, using the third preset state as the third target state, and modifying the second cache state to the third target state in the cache controller directory when the first cache state is the fourth preset state and the second cache state is the fourth preset state.
[0158] In some optional embodiments, the device also includes: a second sending module, used to send a first read request to the second cache when the first cache does not contain the first target data, wherein the first read request is used to determine whether the second cache contains the first target data; a third sending module, used to send the first target data in the second cache to the cache-attached accelerator cache when the second cache contains the first target data, and the cache-attached accelerator cache sends the first target data to the processing unit and updates the cache controller directory; a reading module, used to read the first target data from the memory and send the first target data to the processing unit when the second cache does not contain the first target data.
[0159] In some optional embodiments, the cache data determination module 1004 includes: a determination unit, configured to determine whether the first cache contains third cache data based on a situation where the first cache data hits the cache data corresponding to the first read data request, and determine whether the cache-attached accelerator cache contains third cache data based on a situation where the second cache data hits the third cache data; a first setting unit, configured to use the third cache data in the first cache as target cache data if the first cache contains the third cache data and the first cache state is the first target state; and a second setting unit, configured to use the third cache data in the cache-attached accelerator cache as target cache data if the first cache contains the third cache data, the first cache state is not the first target state, and the second cache state is the second target state. The third cache data is used as target cache data; a first sending unit is used to send a first listening request to the cache-attached accelerator cache when the first cache does not contain the third cache data and the cache-attached accelerator cache contains the third cache data, so that the cache-attached accelerator cache uses the third cache data as target cache data and sends it to the processor core; a second sending unit is used to send a second read request to the second cache when the first cache does not contain the third cache data and the cache-attached accelerator cache does not contain the third cache data, wherein the second read request is used to use the third cache data as target cache data and send it to the first cache when the second cache contains the third cache data, and the first cache sends the target cache data to the processor core and updates the cache controller directory.
[0160] In some optional embodiments, the second sending unit includes: a first setting submodule, used to set the cache state of the first cache to a fourth target state, wherein the fourth target state is determined based on the first cache state; a determination submodule, used to determine the fifth target state based on preset high and low rules, the cache state of the first cache, and the cache state of the cache attached accelerator cache; and a second setting submodule, used to set the second cache state of the cache controller directory to the fifth target state.
[0161] In some optional embodiments, the device also includes: a first determination module, which is used to determine the first data to be replaced in the first cache, determine the first current cache state of the first cache, and obtain the second current cache state recorded in the cache controller directory when there is a first cache replacement request; a first replacement module, which is used to write the first data to be replaced back to the second cache, delete the first data to be replaced in the first cache, and update the cache controller directory when the first current cache state is the first preset state or the third preset state; a second replacement module, which is used to delete the first data to be replaced in the first cache when the first current cache state is the second preset state or the fourth preset state and the second current cache state is the first preset state or the third preset state; a third replacement module, which is used to write the first data to be replaced back to the second cache and delete the first data to be replaced in the first cache when the first current cache state is the second preset state or the fourth preset state and the second current cache state is the second preset state or the fourth preset state.
[0162] In some optional embodiments, the device also includes: a second determination module, used to determine second data to be replaced in the cache-attached accelerator cache when there is a cache-attached accelerator cache replacement request; a fourth replacement module, used to write the second data to be replaced back to the second cache and delete the second data to be replaced in the cache-attached accelerator cache.
[0163] In some optional embodiments, the device also includes: a third determination module, used to determine a calculation result from a processor core, wherein the calculation result is obtained after the processor core performs a calculation operation based on the target cache data; an update module, used to write the calculation result into the first cache, and update the cache status and cache controller directory of the first cache.
[0164] In some optional embodiments, the device also includes: a second judgment module, used to determine whether the first cache contains second target data corresponding to the third read data request when there is a third read data request for the second cache, and obtain a hit result; a first acquisition module, used to determine the third cache state of the first cache and obtain the fourth cache state recorded in the cache controller directory; a second acquisition module, used to obtain the second target data from the first cache or the cache-attached accelerator cache according to the hit result, the third cache state and the fourth cache state, and send the second target data to the second cache.
[0165] In some optional embodiments, the second acquisition module includes: a first acquisition unit, used to acquire the second target data in the first cache when the hit result is a hit and the third cache state is the first preset state, the second preset state or the third preset state; a second acquisition unit, used to send a second listening request to the cache-attached accelerator cache when the hit result is a hit, the third cache state is the fourth preset state, and the fourth cache state is the third preset state, and to acquire the second target data from the cache-attached accelerator cache using the second listening request; a third acquisition unit, used to send a second listening request to the cache-attached accelerator cache when the hit result is a miss and if the cache-attached accelerator cache contains the second target data, and to acquire the second target data from the cache-attached accelerator cache using the second listening request.
[0166] In some optional embodiments, the device also includes: a fourth determination module, used to determine a third current cache state of the second cache and a fourth current cache state of the cache-attached accelerator cache; a recording module, used to determine a higher cache state between the third current cache state and the fourth current cache state according to preset high and low rules, and record the higher cache state to the cache controller directory, wherein the higher cache state includes the second cache state.
[0167] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0168] The cache data reading device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0169] The embodiment of the present invention also provides a computer device having the above Fig.10 The cache data reading device is shown.
[0170] See also Fig.11 , Fig.11 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present invention, such as Fig.11As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Fig.11 A processor 10 is taken as an example.
[0171] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0172] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.
[0173] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0174] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0175] The computer device further comprises a communication interface 30 for the computer device to communicate with other devices or a communication network.
[0176] The embodiment of the present invention also provides a computer-readable storage medium. The method according to the embodiment of the present invention can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0177] A part of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the existence of the computer program instruction in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc., and accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium accessible to the computer.
[0178] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the present invention.
Claims
1. A cache data reading method, characterized in that: The method comprises: Get the first read data request sent by the processor core; querying first cache data stored in a first cache of a processor core and second cache data stored in a cache of a cache-attached accelerator according to the first read data request, wherein the first cache is an L2 cache of the processor core, and the L2 cache is a private cache; Obtaining a first cache state of the first cache and a second cache state of a cache controller directory record, wherein the second cache state is determined based on the first cache state and a cache state of the cache-attached accelerator cache; Based on whether the first cache data hits the third cache data corresponding to the first read data request, whether the second cache data hits the third cache data, the first cache state and the second cache state, target cache data corresponding to the first read data request is determined.
2. The method according to claim 1, characterized in that The method further comprises: In a case where there is a second read data request of the processing unit and the cache of the cache-attached accelerator does not contain the first target data corresponding to the second read data request, determining whether the first cache contains the first target data, wherein the processing unit is the cache-attached accelerator; If the first cache contains the first target data, the first target data in the first cache is sent to the processing unit, and the second cache state is modified to a third target state in the cache controller directory, wherein the third target state is determined based on the first cache state and the cache state of the cache-attached accelerator cache.
3. The method according to claim 2, characterized in that The modifying the second cache state to a third target state in the cache controller directory includes: When the first cache state is the first preset state, the second preset state, or the third preset state, and the second cache state is the first preset state, the second preset state, or the third preset state, modifying the cache state of the first cache to a fourth preset state, using the third preset state as the third target state, and modifying the second cache state to the third target state in the cache controller directory; When the first cache state is the fourth preset state and the second cache state is the fourth preset state, the cache state of the first cache is maintained at the fourth preset state, the third preset state is used as the third target state, and the second cache state is modified to the third target state in the cache controller directory.
4. The method according to claim 2, characterized in that: After determining whether the first cache contains the first target data, the method further includes: In the case where the first cache does not contain the first target data, sending a first read request to a second cache, wherein the first read request is used to determine whether the second cache contains the first target data, the second cache is a processor core L3 cache, and the L3 cache is a shared cache in a computing unit; In the case where the second cache contains the first target data, sending the first target data in the second cache to the cache-attached accelerator cache, and the cache-attached accelerator cache sends the first target data to the processing unit and updates the cache controller directory; In a case where the second cache does not contain the first target data, the first target data is read from the memory, and the first target data is sent to the processing unit.
5. The method according to claim 1, characterized in that The determining, based on a situation that the first cache data hits the third cache data corresponding to the first read data request, a situation that the second cache data hits the third cache data, the first cache state, and the second cache state, target cache data corresponding to the first read data request includes: Determine whether the first cache includes the third cache data according to whether the first cache data hits the third cache data corresponding to the first read data request, and determine whether the cache-attached accelerator cache includes the third cache data according to whether the second cache data hits the third cache data; In a case where the first cache contains the third cache data and the first cache state is a first target state, using the third cache data in the first cache as the target cache data; In a case where the first cache contains the third cache data, the first cache state is not the first target state, and the second cache state is the second target state, using the third cache data in the cache of the cache-attached accelerator as the target cache data; In a case where the first cache does not include the third cache data and the cache-attached accelerator cache includes the third cache data, sending a first snoop request to the cache-attached accelerator cache so that the cache-attached accelerator cache uses the third cache data as the target cache data and sends it to the processor core; In a case where the first cache does not contain the third cache data and the cache-attached accelerator cache does not contain the third cache data, a second read request is sent to the second cache, wherein the second read request is used to use the third cache data as the target cache data and send it to the first cache when the second cache contains the third cache data, and the first cache sends the target cache data to the processor core and updates the cache controller directory, the second cache is the processor core L3 cache, and the L3 cache is a shared cache in the computing unit.
6. The method according to claim 5, characterized in that The updating of the cache controller directory comprises: Setting a cache state of the first cache to a fourth target state, wherein the fourth target state is determined according to the first cache state; Determining a fifth target state according to a preset high-low rule, a cache state of the first cache, and a cache state of the cache-attached accelerator cache; The second cache state of the cache controller directory is set to the fifth target state.
7. The method according to claim 1, characterized in that The method further comprises: In the case where there is a first cache replacement request, determining first data to be replaced in the first cache, determining a first current cache state of the first cache, and acquiring a second current cache state recorded in the cache controller directory; When the first current cache state is a first preset state or a third preset state, writing the first to-be-replaced data back to a second cache, deleting the first to-be-replaced data in the first cache, and updating the cache controller directory, wherein the second cache is a processor core L3 cache, and the L3 cache is a shared cache in a computing unit; When the first current cache state is the second preset state or the fourth preset state, and the second current cache state is the first preset state or the third preset state, deleting the first to-be-replaced data in the first cache; When the first current cache state is the second preset state or the fourth preset state, and the second current cache state is the second preset state or the fourth preset state, the first data to be replaced is written back to the second cache, and the first data to be replaced is deleted from the first cache.
8. The method according to claim 1, characterized in that The method further comprises: In the case where there is a cache replacement request of the cache attachment accelerator, determining second data to be replaced in the cache of the cache attachment accelerator; The second data to be replaced is written back to the second cache, and the second data to be replaced is deleted from the cache attached accelerator cache.
9. The method according to claim 1, characterized in that: After determining the target cache data corresponding to the first read data request, the method further includes: Determining a calculation result from the processor core, wherein the calculation result is obtained after the processor core performs a calculation operation based on the target cache data; The calculation result is written into the first cache, and the cache state of the first cache and the cache controller directory are updated.
10. The method according to claim 1, characterized in that The method further comprises: In the case where there is a third read data request for the second cache, determining whether the first cache contains second target data corresponding to the third read data request, and obtaining a hit result; Determine a third cache state of the first cache, and obtain a fourth cache state recorded in the cache controller directory; According to the hit result, the third cache state, and the fourth cache state, the second target data is obtained from the first cache or the cache-attached accelerator cache, and the second target data is sent to the second cache.
11. The method according to claim 10, characterized in that The acquiring the second target data from the first cache or the cache-attached accelerator cache according to the hit result, the third cache state, and the fourth cache state comprises: When the hit result is a hit and the third cache state is the first preset state, the second preset state, or the third preset state, acquiring the second target data in the first cache; When the hit result is a hit, the third cache state is a fourth preset state, and the fourth cache state is the third preset state, sending a second listening request to the cache-attached accelerator cache, and obtaining the second target data from the cache-attached accelerator cache using the second listening request; When the hit result is a miss, if the cache-attached accelerator cache contains the second target data, the second snoop request is sent to the cache-attached accelerator cache, and the second target data is acquired from the cache-attached accelerator cache using the second snoop request.
12. The method according to claim 1, characterized in that Before querying the first cache data stored in the first cache of the processor core and the second cache data stored in the cache of the cache-attached accelerator according to the first read data request, the method further includes: Determining a third current cache state of a second cache and a fourth current cache state of the cache-attached accelerator cache, wherein the second cache is a processor core L3 cache, and the L3 cache is a shared cache in a computing unit; A higher cache state between the third current cache state and the fourth current cache state is determined according to a preset high-low rule, and the higher cache state is recorded in the cache controller directory, wherein the higher cache state includes the second cache state.
13. A cache data reading device, characterized in that: The device comprises: A request acquisition module, used for acquiring a first read data request sent by a processor core; A query module, configured to query first cache data stored in a first cache of a processor core and second cache data stored in a cache of a cache-attached accelerator according to the first read data request, wherein the first cache is an L2 cache of the processor core, and the L2 cache is a private cache; a state acquisition module, configured to acquire a first cache state of the first cache and a second cache state recorded in a cache controller directory, wherein the second cache state is determined according to the first cache state and a cache state of the cache attached accelerator cache; A cache data determination module is used to determine target cache data corresponding to the first read data request based on whether the first cache data hits the third cache data corresponding to the first read data request, whether the second cache data hits the third cache data, the first cache state, and the second cache state.
14. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the cache data reading method according to any one of claims 1 to 12 by executing the computer instructions.
15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the cache data reading method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Data processing method of multi-core processor, server, product and medium
CN118838863A
Multi-core processor system
CN119025443A