Method and computer program product for adaptively combining cache coherent directory entries

The adaptive joint method manages cache consistency directory entries, which solves the problem of difficult cache consistency between large numbers of processor cores, and achieves the effect of reducing listening operations and fully utilizing directory capacity.

CN119357084BActive Publication Date: 2025-05-06BEIJING KAPULA SCI&TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411908731.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-24
Publication Date
2025-05-06
Estimated Expiration
2044-12-24

AI Technical Summary

Technical Problem

The prior art is difficult to effectively support cache consistency between large numbers of processor cores, especially when the number of directory table entries is insufficient, cache consistency operations may occur where all processor cores need to be listened to.

Method used

An adaptive joint method for cache consistent directory table entries is proposed. By responding to read or write requests from the processor core, a target directory table entries are obtained from the preset directory, and the numbering information of the processor core is saved according to the free space conditions, or a new table entries are applied to make full use of the directory capacity.

Benefits of technology

This method can reduce the listening operation on all processor cores, make full use of the directory capacity, and improve the search efficiency of directory table entries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119357084B_ABST
    Figure CN119357084B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computers, and discloses an adaptive combination method and computer program product of cache consistency directory entries. The method comprises: in response to a read-in request of a processor core, obtaining a first entry set corresponding to a cache; when there are available entries in the first entry set, saving the numbering information of the processor core in the available entry; when there are no entries in the first entry set or there are no available entries in the first entry set, applying for a new entry and saving the numbering information in the new entry; in response to a write request of the processor core, obtaining a second entry set corresponding to the cache; determining the corresponding processor core according to the numbering information saved in each entry in the second entry set, and controlling each processor core to perform a copy operation of invalidating the cache; making the second entry set contain only one entry, and saving the numbering information in the entry. The capacity of the directory is fully utilized and the monitoring operation on all processor cores can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the computer field, and in particular to an adaptive coalescing method for cache coherence directory entries and a computer program product. Background Art

[0002] Each processor core on a modern multi-core CPU has a cache (i.e., a high-speed cache) to improve data access speed and ensure ease of program writing. When running a parallel program, multiple processes or threads running on multiple processor cores may read and write data in the same shared memory area. In order to ensure the correctness of parallel program operation, cache coherence is required between different processor cores of the same CPU and between different CPUs in the same computing node, so that the data of the same cache block (cache line) remains consistent between the private caches of multiple processor cores in the same computing node.

[0003] In the existing technology, there are mainly two types of cache consistency protocols: a consistency protocol based on listening (referred to as listening protocol) and a consistency protocol based on directory structure (referred to as directory protocol); specifically:

[0004] The implementation of the snooping protocol relies on a bus or bus-like network connection (including mesh networks, etc.); based on this network connection, requests issued by the private cache of a single processor core will be broadcast to the private caches of all other processor cores in the system, and the access requests of all processor cores can also be sequenced on this bus to achieve the requirements for memory access order in the cache consistency model and storage identity model, and handle multiple conflicting requests for the same data block.

[0005] The directory protocol uses a directory structure to manage the shared access to each cache line. In this protocol, the memory access request issued by the private cache of the processor core will first be sent to the directory structure that owns the corresponding cache block. This directory structure records the current sharing status of the cache block. The controller will determine the private cache or memory of the processor core that responds to this request based on the current sharing status.

[0006] The snooping protocol and the directory protocol have their own advantages and disadvantages. Among them, the snooping protocol has low hardware implementation overhead and low operating power consumption, but it is easy to become a bottleneck for parallel performance due to the competitive access and orderly response of all processor cores to the bus. The directory protocol can maintain the consistency of cache blocks with different addresses in parallel, so it can achieve cache consistency efficiently; however, on the one hand, the need to record the shared access status of a large number of cache blocks causes hardware and power consumption overhead, and on the other hand, the process of finding the corresponding record of a cache block in the directory will also introduce a large delay.

[0007] Currently, the number of cores in a CPU has reached hundreds (AMD has released a commercial CPU with 192 cores), and a computing node of a supercomputer usually has at least two CPUs, which requires supporting cache consistency between nearly 400 or even more processor cores. Whether it is a snooping protocol or a directory protocol, it is difficult to support cache consistency between such a large number of processor cores. Therefore, most multi-core CPUs use a hybrid method of directory and snooping, which can be expressed as snoop filtering. On the one hand, this snooping filtering method allows a directory table entry to manage the sharing of multiple cache blocks with consecutive addresses, and on the other hand, it can record the sharing of cache blocks inaccurately: for example, record the processor core number of the owner of one or a few copies of the cache block, the current number of effective sharing, etc. When the sharing of cache blocks is not accurately recorded, a cache consistency operation may require snooping on all processor cores. For example, when a directory entry can only record the numbers of at most two cores that own a cache block, if a cache block is being read and shared by three or more cores, and one core is about to modify the cache block, it is necessary to monitor almost all processor cores on the CPU.

[0008] Whether it is a directory protocol or a hybrid method, it is usually necessary to make the number of directory entries match the total number of private caches in all processor cores. For example, when each directory entry is responsible for recording the sharing of a cache block, the number of directory entries should not be less than the total number of cache blocks in the private caches of all processor cores, otherwise cache blocks in the private caches may be swapped out due to insufficient directory entries. The method of making each directory entry responsible for recording the sharing of multiple cache blocks with consecutive addresses is to reduce the number of directory entries and hardware overhead, but when the memory access behavior of parallel programs is very random, it is likely that there will be far from enough directory entries. Therefore, there are rarely redundant directory entries on the CPU.

[0009] In the prior art, data access by each thread in a parallel program is mostly private variables, and although some methods that can avoid redundant cache consistency operations have been disclosed, these methods will make the access to private variables not go through the cache consistency protocol, resulting in the original insufficient directory table entries becoming free, and the more redundant cache consistency operations are reduced, the more free directory table entries will be.

[0010] Furthermore, a contradiction is likely to occur in the future: there are many free directory entries, but each directory entry cannot accurately record the sharing status of cache blocks. This often requires almost all processor cores to monitor the shared cache blocks, but the directory capacity cannot be fully utilized. Summary of the invention

[0011] In order to solve the defects described in the background art and reduce the snooping operations on all processor cores, the present invention proposes an adaptive association method for cache consistency directory entries, an adaptive association system for cache consistency directory entries, and a computer program product.

[0012] At least one embodiment of the present application provides a method for adaptively combining cache coherence directory entries, the method comprising:

[0013] In response to a request from the current processor core for read ownership of the current cache block, obtaining a first target directory entry set corresponding to the address of the current cache block from a preset directory;

[0014] If there is a target entry that meets a preset free space condition in the first target directory entry set, storing the number information of the current processor core in the target entry;

[0015] When there is no entry in the first target directory entry set, or when there is no target entry in the first target directory entry set with free space for storing the numbering information of the current processor core, a new entry is applied for from the preset directory, and when the new entry is successfully applied for, the numbering information of the current processor core is saved in the new entry, and the new entry is saved in the first target directory entry set.

[0016] At least one embodiment of the present application further provides an adaptive association system for cache consistency directory entries, characterized in that it includes:

[0017] A first target directory entry set determination module, configured to obtain a first target directory entry set corresponding to the address of the current cache block from a preset directory in response to a request from the current processor to check the read ownership of the current cache block;

[0018] A first saving module, configured to save the numbering information of the current processor core in the target entry if there is a target entry satisfying a preset free space condition in the first target directory entry set;

[0019] A second saving module is used to apply for a new entry from the preset directory when there is no entry in the first target directory entry set, or when there is no target entry in the first target directory entry set with free space for saving the numbering information of the current processor core, and when the new entry is successfully applied for, save the numbering information of the current processor core in the new entry and save the new entry in the first target directory entry set.

[0020] At least one embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the method described above.

[0021] At least one embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method described above are implemented.

[0022] At least one embodiment of the present application further provides a computer program product, including a computer program, which implements the steps of the optical fiber detection method described above when executed by a processor.

[0023] The adaptive association method of cache consistency directory table entries, the adaptive association system of cache consistency directory table entries, and the computer program product provided by the embodiments of the present application make full use of the capacity of the directory and reduce the monitoring operation of all processor cores; moreover, since the table entries of the target directory table entry set are arranged in order according to the preset saving strategy, the search efficiency for the table entries in the target directory table entry set is relatively high.

[0024] In some optional embodiments, the method further includes:

[0025] In response to a request by the current processor to verify the write ownership of the current cache block, obtaining a set of second target directory entries corresponding to the address of the current cache block from the preset directory;

[0026] Determine the corresponding processor core according to the number information of the processor core stored in each entry in the second target directory entry set, and issue an invalid copy instruction to enable each processor core to execute an operation of invalidating the copy of the current cache block;

[0027] An entry update instruction is issued so that the second target directory entry set includes only one initial directory entry, and the number information of the current processor core is saved in the initial directory entry.

[0028] In some optional embodiments, the method further includes:

[0029] In response to a swap-out request of the current processor core for the current cache block, acquiring from the preset directory an address of the current cache block and a target directory entry storing serial number information of the current processor core;

[0030] In the case where the target directory entry exists, the number information of the current processor core is deleted from the target directory entry.

[0031] In some optional embodiments, after deleting the numbering information of the current processor core from the target directory entry, the method further includes:

[0032] From the target directory entry set storing the target directory entry, the entry that meets the preset condition is deleted; wherein the preset condition includes: the numbering information of any processor core is not stored in the entry.

[0033] In some optional embodiments, the method further includes:

[0034] In the event that the application for the new entry is unsuccessful, an entry update instruction is issued so that the first target directory entry set contains only one initial directory entry, and the numbering information of the current processor core is saved in the initial directory entry in a preset basic content format, after which the first target directory entry set becomes a state of inexact record sharing; the directory entry set in an inexact record sharing state contains only unique directory entries.

[0035] In some optional embodiments, the target directory entry set includes:

[0036] A homogeneous directory entry set and a heterogeneous directory entry set; wherein:

[0037] For the entries in the same isomorphic directory entry set, the content format of each entry is the same;

[0038] For entries in the same heterogeneous directory entry set, the content formats of the entries include one or more.

[0039] In some optional embodiments, the content format of the table entry includes: a bitmap enumeration method or a number enumeration method; wherein:

[0040] The bitmap enumeration method records the ownership of cache block copies by several processor cores with consecutive processor core numbers in a bitmap manner;

[0041] The number enumeration method records the ownership of cache block copies of several processor cores in the form of processor core number values.

[0042] In some optional embodiments, in the first target directory entry set and the second target directory entry set, each entry is arranged in order according to a preset preservation strategy.

[0043] In some optional embodiments, the preset saving strategy includes:

[0044] For any two adjacent first and second entries in the target directory entry set, the value of the processor core number information stored in the first entry is smaller than or larger than the value of the processor core number information stored in the second entry.

[0045] In some optional embodiments, the system further comprises:

[0046] a second target directory entry set determining module, configured to obtain a second target directory entry set corresponding to the address of the current cache block from the preset directory in response to a request of the current processor to check the write ownership of the current cache block;

[0047] an invalidation module, configured to determine a corresponding processor core according to the number information of the processor core stored in each entry in the second target directory entry set, and issue an invalidation copy instruction so that each processor core executes an operation of invalidating the copy of the current cache block;

[0048] The updating module is used to issue an entry updating instruction so that the second target directory entry set includes only one initial directory entry, and the numbering information of the current processor core is saved in the initial directory entry. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.

[0050] Figure 1 A flowchart of a method for adaptively combining cache coherent directory entries provided by an embodiment of the present disclosure;

[0051] Figure 2 A flowchart of another method for adaptively combining cache coherent directory entries provided by an embodiment of the present disclosure;

[0052] Figure 3 A flowchart of another method for adaptively combining cache consistency directory entries provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical scheme and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, it can be understood by those skilled in the art that in the embodiments of the present invention, many technical details are provided to enable readers to better understand the present invention. However, even without these technical details and various changes and modifications based on the following embodiments, the technical scheme claimed in the present invention can be implemented.

[0054] In view of the shortcomings of the prior art, the purpose of the embodiments of the present invention is to provide a method for adaptively combining cache coherent directory entries, a system for adaptively combining cache coherent directory entries, and a computer program product. Compared with the prior art, the capacity of the directory is fully utilized and the monitoring operations on all processor cores can be reduced; moreover, since the entries of the target directory entry set are arranged in order according to the preset saving strategy, the search efficiency for the entries in the target directory entry set is relatively high.

[0055] Embodiment 1:

[0056] The embodiment of the present invention relates to an adaptive combination method of cache consistency directory entries.

[0057] The following specifically describes the implementation details of the adaptive union method of cache consistency directory entries in this embodiment from the perspective of reading cache blocks. The following content is only provided for easy understanding of the implementation details and is not necessary for implementing this solution.

[0058] The adaptive combination method of cache coherence directory entries of this embodiment can be applied to electronic devices with communication, computing and data storage capabilities. Figure 1 As shown, the adaptive combination method of cache consistency directory entries provided in this embodiment includes the following steps:

[0059] Step 110: In response to a request from the current processor to check the read ownership of the current cache block, a first target directory entry set corresponding to the address of the current cache block is obtained from a preset directory.

[0060] Specifically, the preset directory responds to the current processor core's application for read ownership of the current cache block, and queries the first target directory table entry set corresponding to the address of the current cache block; the target directory table entry set stores table entries, wherein the table entries are used to record the serial numbers of all cores of the current cache block.

[0061] Step 120: When a target entry that meets a preset free space condition exists in the first target directory entry set, the number information of the current processor core is saved in the target entry.

[0062] The preset free space condition includes: the available free space in the table entry is not less than the space required to store the serial number information of the current processor core.

[0063] Specifically, when there is an entry in the first target directory entry set having free space for recording the serial number information of the current processor core, the serial number information of the current processor core is recorded in the entry.

[0064] Step 130: If there is no entry in the first target directory entry set, or if there is no target entry in the first target directory entry set with free space for storing the numbering information of the current processor core, apply for a new entry from the preset directory, and if the new entry is successfully applied for, save the numbering information of the current processor core in the new entry, and save the new entry in the first target directory entry set.

[0065] Specifically, when the first target directory entry set has no entry, or all the entries therein have no free space to record the numbering information of the current processor core, a new entry is applied for from the directory. After successfully applying for the new entry, the new entry is added to the first target directory entry set, and the numbering information of the current processor core is recorded in the new entry.

[0066] Embodiment 2:

[0067] Based on the above embodiments, the implementation details of the adaptive union method of cache consistency directory entries of this embodiment are specifically described below from the perspective of cache block writing. The following content is only provided for the convenience of understanding the implementation details and is not necessary for the implementation of this solution.

[0068] The adaptive combination method of cache coherence directory entries of this embodiment can be applied to electronic devices with communication, computing and data storage capabilities. Figure 2 As shown, the adaptive combination method of cache consistency directory entries provided in this embodiment includes the following steps:

[0069] Step 210: In response to a request from the current processor to check the write ownership of the target cache block, a second target directory entry set corresponding to the address of the target cache block is obtained from a preset directory.

[0070] Specifically, the preset directory queries a set of second target directory entries corresponding to the address of the current cache block in response to the application of the current processor core for the write ownership of the current cache block.

[0071] Step 220: determine the corresponding processor core according to the number information of the processor core stored in each entry in the second target directory entry set, and issue an invalid copy instruction to enable each processor core to execute an operation of invalidating the copy of the target cache block.

[0072] Specifically, an operation of invalidating a valid copy of the current cache block is initiated to each related processor core according to the second target directory entry set.

[0073] Step 230: Issue an entry update instruction so that the second target directory entry set includes only one initial directory entry and removes redundant directory entries, and saves the number information of the current processor core in the initial directory entry.

[0074] Specifically, after the operation of invalidating the valid copy of the current cache block, the second target directory entry set is allowed to retain only one initial directory entry, and the serial number information of the current processor core is recorded in the initial directory entry.

[0075] Embodiment three:

[0076] Based on the above embodiments, an implementation of the present invention relates to an adaptive combination method of cache coherent directory entries.

[0077] The implementation details of the adaptive combination method of cache consistency directory entries in this implementation mode are specifically described below. The following content is only provided for the convenience of understanding the implementation details and is not necessary for implementing this solution.

[0078] The adaptive combination method of cache coherence directory entries of this embodiment can be applied to electronic devices with communication, computing and data storage capabilities. Figure 3 As shown, the adaptive combination method of cache consistency directory entries provided in this embodiment includes the following steps:

[0079] Step 310: In response to the request of the current processor to check the read ownership of the target cache block, obtain a first target directory entry set corresponding to the address of the target cache block from a preset directory;

[0080] Step 320: If there is a target entry that meets a preset free space condition in the first target directory entry set, save the number information of the current processor core in the target entry;

[0081] Step 330: If no entry exists in the first target directory entry set, or if no target entry with free space for storing the serial number information of the current processor core exists in the first target directory entry set, apply for a new entry from the preset directory, and if the new entry is successfully applied for, store the serial number information of the current processor core in the new entry, and store the new entry in the first target directory entry set;

[0082] Step 340: In response to the request of the current processor to check the write ownership of the target cache block, obtain a second target directory entry set corresponding to the address of the target cache block from the preset directory;

[0083] Step 350: determine the corresponding processor core according to the number information of the processor core stored in each entry in the second target directory entry set, and issue an invalid copy instruction to enable each processor core to execute an operation of invalidating the copy of the target cache block;

[0084] Step 360: Issue an entry update instruction so that the second target directory entry set contains only one initial directory entry and removes redundant directory entries, and saves the number information of the current processor core in the initial directory entry.

[0085] As an example, the method disclosed in this embodiment may specifically include the following processing flow:

[0086] The preset directory queries a first target directory entry set corresponding to the address of the current cache block in response to the current processor core's application for read ownership of the current cache block; wherein the entry is used to record the serial numbers of all cores of the current cache block;

[0087] When there is an entry in the first target directory entry set having free space for recording the serial number information of the current processor core, recording the serial number information of the current processor core in the entry;

[0088] When the first target directory entry set has no entry, or all entries therein have no free space for recording the serial number information of the current processor core, apply for a new entry from the directory, and after successfully applying for the new entry, add the new entry to the first target directory entry set, and record the serial number information of the current processor core in the new entry;

[0089] The preset directory responds to the current processor core's application for write ownership of the current cache block, queries a second target directory table entry set corresponding to the address of the current cache block, initiates an operation of invalidating a valid copy of the current cache block to each related processor core based on the second target directory table entry set, and then allows the second target directory table entry set to retain only one initial directory table entry, and records the number information of the current processor core in the initial directory table entry.

[0090] Embodiment 4:

[0091] Based on the above embodiment, this embodiment further explains and illustrates the adaptive combination method of cache coherence directory entries provided in the above embodiment.

[0092] In the related art, no matter how many processor cores have valid copies of the same cache block in their private caches, the address of the cache block can only correspond to at most one entry in the directory (in a two-level directory protocol, an entry in the first-level directory and an entry in the second-level directory of each core group for the same cache block are actually the same entry). One of the main technical features of the present invention is that the address of the same cache block corresponds to a set of entries in the directory, the number of entries in the set changes with the running process of the parallel program, and all the entries in the set are combined to record the sharing of the same cache block among all processor cores, that is, each entry can record partial sharing.

[0093] In step 310: in response to a request from the current processor to check the read ownership of a target cache block, a first target directory entry set corresponding to the address of the target cache block is obtained from a preset directory.

[0094] In some embodiments, the target directory entry set includes:

[0095] A homogeneous directory entry set and a heterogeneous directory entry set; wherein:

[0096] For the entries in the same isomorphic directory entry set, the content format of each entry is the same;

[0097] For entries in the same heterogeneous directory entry set, the content formats of the entries include one or more.

[0098] In some optional embodiments, the content format of the table entry includes: a bitmap enumeration method or a number enumeration method, and conversion between the two methods; wherein:

[0099] The bitmap enumeration method records the ownership of cache block copies by several processor cores with consecutive processor core numbers in a bitmap manner;

[0100] The number enumeration method records the ownership of cache block copies of several processor cores in the form of processor core number values.

[0101] Optionally, the content formats of all entries in the same directory entry set can be exactly the same, which is called a homogeneous directory entry set. The content formats of all entries in the same directory entry set can be multiple, which is called a heterogeneous directory entry set, which means that a directory entry can choose one of multiple modes to store the number information of the processor cores that share the cache block. In the current situation where the number of processor cores in a computing node has reached hundreds, it is difficult for a directory entry to accurately record the sharing of any cache block among all processor cores, and a simplified method is often used.

[0102] For example, a directory entry has 64 bits, of which 34 bits are address mark bits of the cache block, 10 bits record the number of current valid copies, and two 10 bits record the processor core numbers of the owners of the two valid copies. In this application, the above simplified method is referred to as the preset basic content format.

[0103] The following further describes the homogeneous directory entry set and the heterogeneous directory entry set:

[0104] 1) A set of homogeneous directory entries based on the basic content format. If each entry in the set uses the basic content format, the address tags of all entries are the same, and each entry can record the processor core numbers of two valid copy owners; therefore, when there are N entries in the set, the maximum number of valid copy owners that can be accurately recorded is 2*N;

[0105] 2) A heterogeneous set of directory entries based on a linked list structure. All entries in the set are organized into a linked list structure or a link structure or index structure similar to file storage. Taking the linked list structure as an example, only the entry at the head of the linked list records the address tag, and each entry has several bits to mark the next entry. When the linked list structure is a doubly linked list (a doubly linked list helps to sort and remove redundancy in the set), each entry will also have several bits to mark the previous entry. Generally speaking, the directory will also adopt a set-associative mapping strategy similar to cache to speed up the search based on the cache block address, and the number of directory entries in the same set is usually not too many (for example, no more than 64). Since all entries in a set of directory entries come from the same group, only 7 bits are needed to mark the previous or next entry, and the remaining 50 bits of a non-linked list head entry can record the processor core numbers of the 5 valid copy owners (called the number enumeration method). In addition, a non-linked list head entry can also record the ownership of several cores with consecutive processor core numbers in a bitmap manner (called bitmap enumeration mode). In this case, 10 of the remaining 50 bits can record the core number (or core group number) of the starting processor core of the bitmap, and the other 40 bits record the ownership of the corresponding 40 processor cores' cache block copies. In a heterogeneous directory entry set, there can be both number enumeration mode entries and bitmap enumeration mode entries. The number enumeration mode is suitable for the situation where the core numbers between processor cores are sparse, while the bitmap enumeration mode is suitable for the situation where the core numbers between processor cores that share the corresponding cache block are dense. As the shared use of the same cache block between processor cores changes during the running of the parallel program, the entry can be adaptively converted between the number enumeration mode and the bitmap enumeration mode to combine as few entries as possible to achieve accurate recording of the shared situation of the same cache block.

[0106] In step 320: when there is a target entry that meets the preset free space condition in the first target directory entry set, the number information of the current processor core is saved in the target entry.

[0107] Optionally, whether in a homogeneous set of directory entries or a heterogeneous set of directory representations, a directory entry can usually record the sharing of the same cache block by multiple processor cores at the same time. Before recording the number information of the current processor core into the target set of directory entries, it should be checked whether there is an entry in the target set of directory entries that has free space to record the number information of the current processor core (for example, in a number enumeration method or a bitmap enumeration method) to improve the utilization efficiency of the directory entries.

[0108] In step 330: if there is no entry in the first target directory entry set, or if there is no target entry in the first target directory entry set with free space for storing the numbering information of the current processor core, apply for a new entry from the preset directory, and if the new entry is successfully applied for, save the numbering information of the current processor core in the new entry, and save the new entry in the first target directory entry set.

[0109] In some embodiments, the method further comprises:

[0110] In the event that the application for the new entry is unsuccessful, an entry update instruction is issued so that the first target directory entry set contains only one initial directory entry and the redundant directory entries are removed, and the numbering information of the current processor core is saved in the initial directory entry in a preset basic content format, after which the first target directory entry set becomes a state of inexact record sharing; the directory entry set in the inexact record sharing state contains only unique directory entries.

[0111] Optionally, after applying for a new entry, it is added to the target directory entry set, and the number information of the current processor core is recorded in the new entry. This is a natural process. When there are free entries in the target directory entry set (when the directory adopts a group-associative mapping strategy, there are free entries in the corresponding group), the application of the new entry will usually succeed; when there are no free entries in the directory, if there are no entries in the target directory entry set (that is, the data of the cache block is read into the cache for the first time), the application of the new entry must succeed; however, when there are no free entries in the directory, if there are entries in the target directory entry set, the application of the new entry may succeed or fail; and success or failure mainly depends on the relevant priority strategy (similar to the cache replacement algorithm). When the application of the new entry fails, the target directory entry will not be able to accurately record the sharing of the cache block among all processor cores. At this time, the basic content format can be used for recording, and often only one entry needs to be retained in the target directory entry set (the remaining entries are released). When the application of a new entry causes an active entry to be preempted, the set of directory entries to which the active entry belongs can no longer accurately record the sharing of cache blocks among all processor cores.

[0112] In step 340: in response to the request of the current processor to check the write ownership of the target cache block, a second target directory entry set corresponding to the address of the target cache block is obtained from the preset directory.

[0113] In step 350: determine the corresponding processor core according to the processor core number information stored in each entry in the second target directory entry set, and issue an invalid copy instruction to enable each processor core to execute an operation of invalidating the copy of the target cache block.

[0114] In step 360: an entry update instruction is issued so that the second target directory entry set includes only one initial directory entry and removes redundant directory entries, and the number information of the current processor core is saved in the initial directory entry.

[0115] As an example, the specific process of the method disclosed in this embodiment includes:

[0116] The preset directory queries a first target directory entry set corresponding to the address of the current cache block in response to the current processor core's application for read ownership of the current cache block; wherein the entry is used to record the serial numbers of all cores of the current cache block;

[0117] When there is an entry in the first target directory entry set having free space for recording the serial number information of the current processor core, recording the serial number information of the current processor core in the entry;

[0118] When the first target directory entry set has no entry, or all entries therein have no free space for recording the serial number information of the current processor core, apply for a new entry from the directory, and after successfully applying for the new entry, add the new entry to the first target directory entry set, and record the serial number information of the current processor core in the new entry;

[0119] The preset directory responds to the current processor core's application for write ownership of the current cache block, queries a second target directory table entry set corresponding to the address of the current cache block, initiates an operation of invalidating a valid copy of the current cache block to each related processor core based on the second target directory table entry set, and then allows the second target directory table entry set to retain only one initial directory table entry, and records the number information of the current processor core in the initial directory table entry.

[0120] When the second target directory entry set accurately records all relevant processor cores that share the current cache block, an operation of invalidating the valid copy of the current cache block is initiated to each of these relevant processor cores (corresponding to the classic implementation of the cache consistency protocol, while invalidating, the current processor core will also obtain a latest copy of the cache block). When the second target directory entry set does not accurately record all relevant processor cores that share the current cache block, an operation of invalidating all valid copies of the current cache block needs to be initiated to almost all processor cores. If the second target directory entry set is an empty set, no invalidation operation needs to be initiated, and then a new directory entry needs to be applied for. A write request will result in only the current processor core having a valid copy of the current cache block among all processor cores. Therefore, the second target directory entry set only needs to retain one entry in the end.

[0121] In some embodiments, the method further comprises:

[0122] In response to a swap-out request of the current processor core for the target cache block, acquiring from the preset directory an address of the current cache block and a target directory entry storing serial number information of the current processor core;

[0123] In the case where the target directory entry exists, the number information of the current processor core is deleted from the target directory entry.

[0124] In some embodiments, after deleting the numbering information of the current processor core from the target directory entry, the method further includes:

[0125] Delete the entries that meet the preset conditions from the target directory entry set storing the target directory entries.

[0126] In some embodiments, the preset conditions include:

[0127] The table entry does not store any processor core number information.

[0128] Optionally, regardless of whether the target directory entry set accurately records the sharing status of the current cache block, it is necessary to update the relevant content in the target directory entry set before swapping out the current cache block. Among them, when a target directory entry exists (corresponding to the situation of accurate recording), it is necessary to remove the numbering information of the current processor core from the content of the target directory entry. Further, the target directory entry set can be optimized and adjusted, for example, so that at most one entry in the set has free space, removing entries without substantial content, or converting a certain entry between bitmap enumeration mode and number enumeration mode, etc.

[0129] In some embodiments, in the first target directory entry set and the second target directory entry set, each entry is arranged in order according to a preset preservation strategy.

[0130] In some embodiments, the preset saving strategy includes:

[0131] For any two adjacent first and second entries in the target directory entry set, the value of the processor core number information stored in the first entry is smaller than or larger than the value of the processor core number information stored in the second entry.

[0132] Optionally, if there are multiple entries in the set, it is necessary to find the entry related to the number information of the current processor core from all the entries in the set. If there is no order relationship between the entries, it is usually necessary to traverse all the entries. In order to speed up the search process, the entries can always meet a certain order, such as arranging the processor core numbers in ascending or descending order, and when the content of the set changes, the order is maintained by adjusting the content between the entries.

[0133] Embodiment five:

[0134] Another embodiment of the present application relates to an adaptive union system for cache coherent directory entries.

[0135] The following is a specific description of the implementation details of the adaptive union system of cache coherence directory entries in this embodiment from the perspective of reading cache blocks. The following content is only provided for the convenience of understanding the implementation details and is not necessary for implementing this solution. The adaptive union system of cache coherence directory entries provided in this embodiment includes:

[0136] A first target directory entry set determination module, configured to obtain a first target directory entry set corresponding to an address of the target cache block from a preset directory in response to a request by the current processor to check the read ownership of the target cache block;

[0137] A first saving module, configured to save the numbering information of the current processor core in the target entry if there is a target entry satisfying a preset free space condition in the first target directory entry set;

[0138] A second saving module is used to apply for a new entry from the preset directory when there is no entry in the first target directory entry set, or when there is no target entry in the first target directory entry set with free space for saving the numbering information of the current processor core, and when the new entry is successfully applied for, save the numbering information of the current processor core in the new entry and save the new entry in the first target directory entry set.

[0139] Technical personnel in the relevant field can clearly understand that, for the convenience and brevity of description, the specific working process of each module in the adaptive joint system of the cache consistency directory table entries can refer to the corresponding process in the aforementioned method embodiment, and this embodiment will not be repeated here.

[0140] It is worth mentioning that all modules involved in this embodiment are logic modules. In practical applications, a logic unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed by this application, but this does not mean that there are no other units in this embodiment.

[0141] Embodiment six:

[0142] Based on the above embodiments, another embodiment of the present application relates to an adaptive union system for cache coherent directory entries.

[0143] The following is a specific description of the implementation details of the adaptive union system of cache coherence directory entries in this embodiment from the perspective of cache block writing. The following content is only provided for the convenience of understanding the implementation details and is not necessary for implementing this solution. The adaptive union system of cache coherence directory entries provided in this embodiment includes:

[0144] A second target directory entry set determination module, configured to obtain a second target directory entry set corresponding to the address of the target cache block from a preset directory in response to a request from the current processor core for writing ownership of the target cache block;

[0145] an invalidation module, configured to determine a corresponding processor core according to the number information of the processor core stored in each entry in the second target directory entry set, and issue an invalidation copy instruction so that each processor core executes an operation of invalidating the copy of the target cache block;

[0146] The updating module is used to issue an entry updating instruction so that the second target directory entry set only includes one initial directory entry and removes redundant directory entries, and saves the serial number information of the current processor core in the initial directory entry.

[0147] Technical personnel in the relevant field can clearly understand that, for the convenience and brevity of description, the specific working process of each module in the adaptive joint system of the cache consistency directory table entries can refer to the corresponding process in the aforementioned method embodiment, and this embodiment will not be repeated here.

[0148] It is worth mentioning that all modules involved in this embodiment are logic modules. In practical applications, a logic unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed by this application, but this does not mean that there are no other units in this embodiment.

[0149] Embodiment seven:

[0150] Based on the above embodiments, another embodiment of the present application relates to an adaptive union system for cache coherent directory entries.

[0151] The following is a detailed description of the implementation details of the cache coherence directory entry adaptive union system of this embodiment. The following content is only for the convenience of understanding the implementation details provided, and is not necessary for implementing this solution. The cache coherence directory entry adaptive union system provided in this embodiment includes:

[0152] A first target directory entry set determination module, configured to obtain a first target directory entry set corresponding to an address of the target cache block from a preset directory in response to a request by the current processor to check the read ownership of the target cache block;

[0153] A first saving module, configured to save the numbering information of the current processor core in the target entry if there is a target entry satisfying a preset free space condition in the first target directory entry set;

[0154] a second saving module, configured to apply for a new entry from the preset directory if no entry exists in the first target directory entry set, or if no target entry with free space for saving the serial number information of the current processor core exists in the first target directory entry set, and if the new entry is successfully applied for, save the serial number information of the current processor core in the new entry, and save the new entry in the first target directory entry set;

[0155] a second target directory entry set determining module, configured to obtain a second target directory entry set corresponding to the address of the target cache block from the preset directory in response to a request by the current processor to check the write ownership of the target cache block;

[0156] an invalidation module, configured to determine a corresponding processor core according to the number information of the processor core stored in each entry in the second target directory entry set, and issue an invalidation copy instruction so that each processor core executes an operation of invalidating the copy of the target cache block;

[0157] The updating module is used to issue an entry updating instruction so that the second target directory entry set only includes one initial directory entry and removes redundant directory entries, and saves the serial number information of the current processor core in the initial directory entry.

[0158] Technical personnel in the relevant field can clearly understand that, for the convenience and brevity of description, the specific working process of each module in the adaptive joint system of the cache consistency directory table entries can refer to the corresponding process in the aforementioned method embodiment, and this embodiment will not be repeated here.

[0159] It is worth mentioning that all modules involved in this embodiment are logic modules. In practical applications, a logic unit can be a physical unit, a part of a physical unit, or a combination of multiple physical units. In addition, in order to highlight the innovative part of this application, this embodiment does not introduce units that are not closely related to solving the technical problems proposed by this application, but this does not mean that there are no other units in this embodiment.

[0160] Embodiment eight:

[0161] Another embodiment of the present application relates to an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the methods in the above-mentioned embodiments.

[0162] Among them, the memory and the processor are connected in a bus manner, and the bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors and memories together. The bus can also connect various other circuits such as peripherals, voltage regulators, and power management circuits, which are well known in the art and are therefore not further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver can be one element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices on a transmission medium. The data processed by the processor is transmitted on a wireless medium via an antenna, and further, the antenna also receives data and transmits the data to the processor.

[0163] The processor is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interfaces, voltage regulation, power management, and other control functions. Memory can be used to store data used by the processor when performing operations.

[0164] Embodiment nine:

[0165] Another embodiment of the present application relates to a computer-readable storage medium storing a computer program, which implements the above method embodiment when executed by a processor.

[0166] That is, those skilled in the art can understand that all or part of the steps in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a program, and the program is stored in a storage medium, including a number of instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as: ROM), random access memory (Random Access Memory, referred to as: RAM), disk or optical disk and other media that can store program codes.

[0167] In some embodiments of the present application, a computer program product is also provided, including a computer program, which implements the steps of the method described in the above embodiment when executed by a processor.

[0168] Those skilled in the art will appreciate that the above embodiments are specific embodiments for implementing the present application, and in actual applications, various changes may be made thereto in form and detail without departing from the spirit and scope of the present application.

Claims

1. A method for adaptively combining cache coherent directory entries, characterized in that: The method comprises: In response to a request by the current processor core for read ownership of the current cache block, a first target directory entry set corresponding to the address of the current cache block is obtained from a preset directory; wherein all entries in the first target directory entry set are combined to record the sharing status of the current cache block among all processor cores; If there is a target entry that meets a preset free space condition in the first target directory entry set, storing the number information of the current processor core in the target entry; When there is no entry in the first target directory entry set, or when there is no target entry in the first target directory entry set with free space for storing the numbering information of the current processor core, a new entry is applied for from the preset directory, and when the new entry is successfully applied for, the numbering information of the current processor core is saved in the new entry, and the new entry is saved in the first target directory entry set.

2. The method according to claim 1, characterized in that The method further comprises: In response to a request by the current processor to verify the write ownership of the current cache block, obtaining a set of second target directory entries corresponding to the address of the current cache block from the preset directory; Determine the corresponding processor core according to the number information of the processor core stored in each entry in the second target directory entry set, and issue an invalid copy instruction to enable each processor core to execute an operation of invalidating the copy of the current cache block; An entry update instruction is issued so that the second target directory entry set includes only one initial directory entry, and the number information of the current processor core is saved in the initial directory entry.

3. The method according to claim 1, characterized in that The method further comprises: In response to a swap-out request of the current processor core for the current cache block, acquiring from the preset directory an address of the current cache block and a target directory entry storing serial number information of the current processor core; In the case where the target directory entry exists, the number information of the current processor core is deleted from the target directory entry.

4. The method according to claim 3, characterized in that After deleting the numbering information of the current processor core from the target directory entry, the method further includes: From the target directory entry set storing the target directory entry, the entry that meets the preset condition is deleted; wherein the preset condition includes: the numbering information of any processor core is not stored in the entry.

5. The method according to claim 1, characterized in that The method further comprises: If the application for the new entry is unsuccessful, an entry update instruction is issued so that the first target directory entry set contains only one initial directory entry and removes redundant directory entries, and the numbering information of the current processor core is saved in the initial directory entry in a preset basic content format.

6. The method according to any one of claims 1 to 5, characterized in that The target directory entry set includes: A homogeneous directory entry set and a heterogeneous directory entry set; wherein: For the entries in the same isomorphic directory entry set, the content format of each entry is the same; For entries in the same heterogeneous directory entry set, the content formats of the entries include one or more.

7. The method according to claim 6, characterized in that The content format of the table entry includes: bitmap enumeration mode or number enumeration mode; wherein: The bitmap enumeration method records the ownership of cache block copies by several processor cores with consecutive processor core numbers in a bitmap manner; The number enumeration method records the ownership of cache block copies of several processor cores in the form of processor core number values.

8. The method according to claim 2, characterized in that: In the first target directory entry set and the second target directory entry set, each entry is arranged in order according to a preset saving strategy.

9. The method according to claim 8, characterized in that The preset saving strategy includes: For any two adjacent first and second entries in the target directory entry set, the value of the processor core number information stored in the first entry is smaller than or larger than the value of the processor core number information stored in the second entry.

10. An adaptive federation system for cache coherent directory entries, characterized in that: include: A first target directory entry set determination module is used to obtain a first target directory entry set corresponding to the address of the current cache block from a preset directory in response to a request from the current processor core for read ownership of the current cache block; wherein all entries in the first target directory entry set are combined to record the sharing status of the current cache block among all processor cores; A first saving module, configured to save the numbering information of the current processor core in the target entry if there is a target entry satisfying a preset free space condition in the first target directory entry set; A second saving module is used to apply for a new entry from the preset directory when there is no entry in the first target directory entry set, or when there is no target entry in the first target directory entry set with free space for saving the numbering information of the current processor core, and when the new entry is successfully applied for, save the numbering information of the current processor core in the new entry and save the new entry in the first target directory entry set.

11. The system according to claim 10, characterized in that: Also includes: a second target directory entry set determining module, configured to obtain a second target directory entry set corresponding to the address of the current cache block from the preset directory in response to a request of the current processor to check the write ownership of the current cache block; an invalidation module, configured to determine a corresponding processor core according to the number information of the processor core stored in each entry in the second target directory entry set, and issue an invalidation copy instruction so that each processor core executes an operation of invalidating the copy of the current cache block; The updating module is used to issue an entry updating instruction so that the second target directory entry set includes only one initial directory entry, and the numbering information of the current processor core is saved in the initial directory entry.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.

Citation Information

Patent Citations

  • Cache directory handling methods and directory controllers for multi-core processor systems

    CN105659216B