A data access control architecture and control method for multi-processors
By introducing the snoop filtering processing component and shared cache memory in the multiprocessor architecture, the data inconsistency problem between the main processor and the accelerator when accessing the main memory is solved, and the consistency of data access and the reliability of transaction processing is achieved.
Patent Information
- Application Number
- CN202411833808.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-13
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2044-12-13
AI Technical Summary
In a heterogeneous multi-core multi-processor architecture, the main processor and the accelerator are prone to inconsistent cache data when accessing the main memory, especially after the main processor has taken away the data copy and the accelerator updates the data, the data cached by the main processor expires or is invalid, resulting in inconsistent data access.
A multiprocessor data access control architecture is designed, including main processor, accelerator, main memory and snoop filtering processing components. The snoop filtering processing component stores the snoop tag in the local cache of the main processor and the shared cache storage. The snoop filtering processing component receives the accelerator's data access request, querys the snoop tag and the shared tag, and if hit, sends the snoop request to the main processor. The main processor invalidates the data copy in the local cache memory, merges the dirty data with the accelerator's updated data, and writes it to the shared cache memory.
It effectively solves the problem of inconsistent data access in the multiprocessor architecture, ensures the consistency of data access between the main processor and accelerator, improves the accuracy and reliability of transaction processing, and reduces the read and write pressure on the main memory.
Smart Images

Figure CN119292963B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a data access control architecture and control method for multi-processors. Background Art
[0002] A heterogeneous multi-core multi-processor architecture is a design that integrates different types of processing units on the same chip. A multi-processor usually includes a main processor and one or more accelerators. The heterogeneous multi-core multi-processor architecture can make full use of the high-performance computing power of the accelerators to achieve parallel and fast computing. Therefore, the heterogeneous multi-core multi-processor architecture is widely used in the fields of high-performance computing (HPC) and artificial intelligence (AI). It also plays an important role in scientific research in disciplines such as astronomy, physics, and biology.
[0003] In a heterogeneous multi-core multi-processor architecture, in addition to the main processor needing to access the main memory, other accelerators also need to read and write the main memory to obtain or cache data during the computing process. That is, both the main processor and each accelerator can read or write to a single shared address. This access mechanism will cause the problem of inconsistent cache data in the heterogeneous multi-core multi-processor architecture.
[0004] For example, after the data in the main memory has been fetched by the main processor, when the accelerator needs to update the data, the data cached in the main processor is expired or invalid at this time. Summary of the Invention
[0005] The present invention provides a data access control architecture and control method for multi-processors to ensure the consistency of data access in a heterogeneous multi-core multi-processor architecture.
[0006] According to one aspect of the present invention, there is provided a data access control architecture for multi-processors, the data access control architecture for multi-processors including: a main processor, at least one accelerator, a main memory, and a snooping filtering processing component, the main processor and the accelerator access a shared cache memory, and a local cache memory is provided in the main processor; wherein:
[0007] The snooping filtering processing component stores snooping tags corresponding to all memory addresses for which there are copies in the local cache memory of the main processor;
[0008] The shared cache memory stores shared tags corresponding to all memory addresses for which there are copies in the shared cache memory by the main processor and / or the accelerator;
[0009] The shared cache memory and the local cache memory are non-inclusive memories;
[0010] The snooping filter processing component receives the data access requests of each accelerator, and queries the snooper tag and the shared tag according to the current memory address accessed by the accelerator;
[0011] If the snooper tag is hit, the snooping filter processing component sends a snooping request to the main processor;
[0012] When the main processor determines that the data access request of the accelerator is a write operation according to the snooping request, it invalidates the data copy in the local cache memory corresponding to the current memory address accessed by the accelerator;
[0013] When the data in the main processor is dirty data, the dirty data is merged with the updated data of the accelerator and written into the shared cache memory.
[0014] According to another aspect of the present invention, there is provided a data access control method for a multi-processor, which is applied to the data access control architecture of a multi-processor provided in any embodiment of the present invention. The data access control method for the multi-processor includes:
[0015] Receive the data access requests of each accelerator through the snooping filter processing component, and query the snooper tag and the shared tag according to the current memory address accessed by the accelerator;
[0016] If the snooper tag is hit, send a snooping request to the main processor through the snooping filter processing component;
[0017] When it is determined according to the snooping request that the data access request of the accelerator is a write operation, invalidate the data copy in the local cache memory corresponding to the current memory address accessed by the accelerator through the main processor;
[0018] When the data in the main processor is dirty data, merge the dirty data with the updated data of the accelerator and write it into the shared cache memory.
[0019] According to another aspect of the present invention, there is provided a chip, which includes: the data access control architecture of a multi-processor provided in any embodiment of the present invention.
[0020] According to another aspect of the present invention, there is provided an electronic device, which includes: the chip provided in any embodiment of the present invention.
[0021] The technical solution of the embodiment of the present invention is to set up a data access control architecture for a multi-processor including a main processor, at least one accelerator, a main memory, and a snooping filter processing component. Among them, the main processor and the accelerator access a shared cache memory, and a local cache memory is set in the main processor; the snooping filter processing component stores snooper tags corresponding to all memory addresses with copies in the local cache memory of the main processor; the shared cache memory stores shared tags corresponding to all memory addresses with copies in the shared cache memory by the main processor and / or the accelerator; the shared cache memory and the local cache memory are non-inclusive memories; the snooping filter processing component receives data access requests from each accelerator, and queries the snooper tags and shared tags according to the memory address currently accessed by the accelerator; if the snooper tag is hit, the snooping filter processing component sends a snooping request to the main processor; when the main processor determines that the data access request of the accelerator is a write operation according to the snooping request, it invalidates the data copy corresponding to the memory address currently accessed by the accelerator in the local cache memory; when the data in the main processor is dirty data, it merges the dirty data with the updated data of the accelerator and writes it into the shared cache memory, solving the data inconsistency problem during data access by the multi-processor, especially solving the data access inconsistency problem caused when the accelerator needs to update the data after the data in the main memory has been fetched by the main processor and the data cached in the main processor is expired or invalid, ensuring that the data accessed by the main processor and the accelerator is consistent, and improving the correctness and reliability of transaction processing.
[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 FIG. is a schematic structural diagram of a data access control architecture for a multi-processor according to an embodiment of the present invention.
[0025] Figure 2 FIG. is a schematic structural diagram of another data access control architecture for a multi-processor according to an embodiment of the present invention.
[0026] Figure 3It is a schematic flowchart of a data access control method for a multi-processor according to an embodiment of the present invention. Detailed implementation manners
[0027] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0028] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0029] Figure 1 It is a schematic structural diagram of a data access control architecture for a multi-processor according to an embodiment of the present invention. This embodiment is applicable to the situation of ensuring data access consistency in a heterogeneous multi-core multi-processor architecture. As Figure 1 shown, the data access control architecture of the multi-processor includes: a main processor, at least one accelerator, a main memory, and a snooping filter processing component. The main processor and the accelerator access a shared cache memory, and a local cache memory is provided in the main processor. The main processor and the accelerator can also access the main memory.
[0030] In the embodiment of the present invention, a cache memory may not be provided in the accelerator. Alternatively, a cache memory may be provided in the accelerator, but its data consistency is not maintained through a hardware structure, but through software.
[0031] Among them, the cache memory may be a first-level cache (L1 Cache), a second-level cache (L2 Cache), and a third-level cache (L3 Cache). In an application scenario of the embodiment of the present invention, an L1 Cache and an L2 Cache may be provided in the main processor.
[0032] The snooping filter processing component stores snoop tags corresponding to all memory addresses for which copies are stored in the local cache memory of the main processor. The shared cache memory stores shared tags corresponding to all memory addresses for which copies are stored in the shared cache memory by the main processor and / or the accelerator.
[0033] Combined with the specific processing method in the multi-processor data access control architecture provided in the embodiments of the present invention, in the snooping filter processing component and the shared cache memory, only storing the tag values of the addresses can ensure data access consistency. In the snooping filter processing component and the shared cache memory, there is no need to store information such as cache line state information and transaction processing source identifiers. The information that needs to be stored in the snooping filter processing component and the shared cache memory can be simplified, greatly reducing the hardware overhead.
[0034] In the embodiments of the present invention, the shared cache memory and the local cache memory are non-inclusive memories. That is, the data corresponding to the memory address stored in the local cache memory is not stored in the shared cache memory. The data corresponding to the memory address stored in the shared cache memory is not stored in the local cache memory.
[0035] When the accelerator needs to access data, it can first send a data access request to the snooping filter processing component. The snooping filter processing component receives the data access requests from each accelerator, and queries the snoop tags and shared tags according to the memory address currently accessed by the accelerator. If the snoop tag is hit, the snooping filter processing component sends a snooping request to the main processor. When the main processor determines that the data access request of the accelerator is a write operation according to the snooping request, it invalidates the data copy corresponding to the memory address currently accessed by the accelerator in the local cache memory. When the data in the main processor is dirty data, the dirty data is merged with the updated data of the accelerator and written into the shared cache memory.
[0036] Among them, when the memory address currently accessed by the accelerator hits the snoop tag, it indicates that the main processor has fetched the data copy from the main memory and stored it in the local cache memory of the main processor. At this time, when the accelerator needs to perform a write operation on the data at this memory address, the data in the local cache memory of the main processor is expired data. Therefore, in the embodiments of the present invention, when the main processor determines that the data access request of the accelerator is a write operation according to the snooping request, it invalidates the data copy corresponding to the memory address currently accessed by the accelerator in the local cache memory, thereby avoiding the main processor from processing expired data and causing data inconsistency between the main processor and the accelerator.
[0037] In the local cache memory of the main processor, the data can be either dirty data or clean data. Among them, dirty data is data that is inconsistent with the corresponding data in the main memory, that is, the main processor has updated the data at this address. Clean data is data that is consistent with the corresponding data in the main memory, that is, the main processor has not updated the data at this address. To ensure data consistency, when the data in the main processor is dirty data, the dirty data copy is returned, and the dirty data is merged with the updated data of the accelerator and written into the shared cache memory, so that the main processing can continue to perform consistency processing on the data. When the data in the main processor is clean data, the accelerator updates the data in the main memory.
[0038] By returning the dirty data, merging it with the part of the data that the accelerator hopes to update, and storing it in the shared cache memory, the network pressure and performance loss caused by reading and writing the main memory can be reduced. When the data in the main processor is clean data, the data copy does not need to be returned, and the accelerator can directly update the data in the main memory. At this time, the data copy in the local cache memory of the main processor has been deleted, thus ensuring data consistency. Through the above operations, the problem of data access inconsistency caused by the data in the main memory being updated by another accelerator after the main processor has taken the copy, and the data in the main processor cache being expired or invalid, is solved.
[0039] In the embodiment of the present invention, when the main processor determines that the data access request of the accelerator is a read operation according to the snooping request, the data copy corresponding to the current memory address accessed by the accelerator in the local cache memory is returned, and the data copy is kept valid.
[0040] Among them, when the current memory address accessed by the accelerator hits the snooper tag, it indicates that the main processor has taken the data copy from the main memory and saved it in the local cache memory of the main processor. At this time, when the accelerator needs to perform a read operation on the data at this memory address, the main memory will no longer contain the latest data. Therefore, when the accelerator reads the data corresponding to this memory address in the main memory, it will obtain expired data. In the embodiment of the present invention, by the main processor returning the data copy corresponding to the current memory address accessed by the accelerator in the local cache memory, the problem that the accelerator obtains expired data in the main memory can be solved, and data access consistency can be ensured.
[0041] In the embodiment of the present invention, if the shared tag is hit, the data access request of the accelerator is processed in the shared cache memory; if the snooper tag and the shared tag are not hit, the data access request of the accelerator is processed in the main memory.
[0042] By storing the snooper tag in the snooping filter processing component and storing the shared tag in the shared cache memory; when the accelerator needs to access data, comparing the current memory address accessed by the accelerator with the snooper tag and the shared tag; determining the memory for the accelerator data access according to the hit situation, it is possible to solve the problem of data access inconsistency caused by the data in the main memory being updated by another accelerator after the main processor has taken a copy, and at this time, the data in the main processor cache is expired or invalid; and when the main processor takes a copy of the data and stores it in the local cache memory of the main processor, when the accelerator needs to perform a read operation on the data at this memory address, the main memory will no longer contain the latest data, resulting in the problem of data access inconsistency.
[0043] By storing the snooper tag in the snooping filter processing component and storing the shared tag in the shared cache memory, the stored information can be simplified, without storing cache line state information and transaction processing source identifiers, reducing the hardware overhead; when the accelerator is performing a write operation and the main processor returns dirty data, the data is written through the shared cache memory instead of writing back to the main memory, alleviating the transmission pressure on the system bus and network, and also reducing the impact of consistency operations on the performance of the main memory.
[0044] Based on the above embodiments, optionally, when the main processor receives a data access request for the main processor, it queries the shared tag in the shared cache memory according to the current memory address accessed by the main processor to determine whether there is a hit; if the current memory address accessed by the main processor hits the shared tag, the data access to the main processor is performed in the shared cache memory; otherwise, the data access to the main processor is performed through the main memory. Through this processing, it is possible to solve the problem of data access inconsistency caused when the accelerator has taken a copy of the data and the main processor also needs to access the data at this address when accessing the data through the shared cache memory.
[0045] Figure 2 It is a schematic structural diagram of another data access control architecture for multi-processors provided according to an embodiment of the present invention. As Figure 2 shown, the structure of the data access control architecture for multi-processors includes a main processor, at least one accelerator, a main memory, and a snooping filter processing component; the main processor and the accelerator access the shared cache memory, and a local cache memory is set in the main processor ( Figure 2(not shown). The snooping filter processing component stores snooper tags corresponding to all memory addresses that have copies stored in the local cache memory of the main processor. The shared cache memory stores shared tags corresponding to all memory addresses that have copies stored in the shared cache memory of the main processor and / or the accelerator. The shared cache memory and the local cache memory are non-inclusive memories. The accelerator may not be provided with a cache memory. Alternatively, the accelerator may be provided with a cache memory, but its data consistency is not maintained through a hardware structure, but through software.
[0046] As Figure 2 shown, optionally, the data access control architecture of the multi-processor further includes: a coherence processing component; a snooping filter processing component, including: a pre-filter controller, a post-filter controller, a snooping filter controller, a snooping filter array, and a snooping request processing component. The snooping filter array stores snooper tags corresponding to all memory addresses that have copies stored in the local cache memory of the main processor.
[0047] Optionally, the main processor includes: a main processor request processing component. The data access control architecture of the multi-processor further includes: a main memory access network and a main memory access controller.
[0048] In a specific embodiment, when the main processor receives a data access request for the main processor, it transmits the current access memory address of the main processor to the coherence processing component through the main processor request processing component; the coherence processing component queries the shared tags in the shared cache memory to determine whether there is a hit; if the shared tag is hit for the current access memory address of the main processor, the data access to the main processor is performed in the shared cache memory; otherwise, the coherence processing component accesses the main memory access network through the main memory access controller, and the data access to the main processor is performed in the main memory.
[0049] Optionally, the pre-filter controller receives the data access requests of each accelerator and sends the data access requests of each accelerator to the coherence processing component; the coherence processing component queries the snooper tags stored by the snooping filter controller through the snooping filter array and the shared tags in the shared cache memory according to the current access memory address of the accelerator; when the coherence processing component determines that the current access memory address of the accelerator hits the snooper tag, it sends the hit information to the snooping request processing component; the snooping request processing component sends a snooping request to the main processor.
[0050] Optionally, when the main processor determines that the data access request of the accelerator is a write operation according to the snooping request, it invalidates the data copy in the local cache corresponding to the current memory address accessed by the accelerator; the main processor returns the corresponding dirty data to the coherence processing component through the snooping request processing component; the coherence processing component obtains the updated data of the accelerator transmitted by the accelerator through the pre-filter controller, and merges the dirty data with the updated data, and writes the merged result into the shared cache. When the data access request of the accelerator is a write operation and the data in the main processor is clean data, a reply response is made through the snooping request processing component, and the reply response does not contain data; the snooping request processing component transmits the reply response to the coherence processing component; the coherence processing component accesses the main memory access network through the main memory access controller and performs the write operation of the updated data in the accelerator data access request in the main memory.
[0051] When the main processor determines that the data access request of the accelerator is a read operation according to the snooping request, it returns the data copy in the local cache corresponding to the current memory address accessed by the accelerator to the shared cache through the snooping request processing component and the coherence processing component, so that the accelerator can perform data reading and keep the data copy in the local cache valid.
[0052] When the coherence processing component determines that the current memory address accessed by the accelerator misses the snooper tag and the shared tag, it sends the miss information to the post-filter controller; the post-filter controller accesses the main memory access network and processes the data access request of the accelerator in the main memory.
[0053] Through the above multi-processor data access control architecture and hardware data access coherence maintenance, on the basis of simplifying the snooping filter processing component and the information stored in the shared cache and reducing the hardware overhead, when it is caused by a coherence operation rather than the main processor actively clearing the storage entry, the storage entry that needs to be cleared can be stored in the shared cache instead of being written back to the main memory, alleviating the transmission pressure on the system bus and network, and also reducing the impact of coherence operations on the performance of the main memory.
[0054] In addition, by in Figure 2In the data access control architecture of the multi-processor shown, a first interface formed by a main memory access controller and a main memory access network, and a second interface formed by a filtered controller and the main memory access network are provided. The access to the main memory can be split, the transmission pressure on the system bus and network can be alleviated, and the impact of coherence operations on the performance of the main processor core can be reduced. Among them, the first interface can be used for the read and write access of the main processor, and the read and write access forwarding when valid data cannot be obtained by the snooping command. The second interface can be used for read and write operations from the accelerator and where the memory addresses accessed do not hit the tags in the snooping filter processing component and the shared cache memory, which can filter out useless address accesses and reduce the impact of the accelerator's read and write data on the performance of the main processor.
[0055] Based on the above implementation, when the number of snooper tags stored in the snooping filter array reaches the upper limit, the stored tags can be removed. For example, the victim entry can be evicted in a first-in, first-out manner, and an invalid request can be sent to the main processor. The data copy in the local cache memory of the main processor can be invalidated. The invalidated data copy can be taken out from the local cache memory of the main processor and stored in the shared cache memory. By writing the cleared entry into the shared cache memory instead of writing it back to the main memory, both the transmission pressure on the system bus and network can be alleviated, and the impact of coherence operations on the performance of the main memory can be reduced.
[0056] In the Figure 2 data access control architecture of the multi-processor shown, a shared cache memory controller can also be included. The shared cache memory controller is used to control the shared cache memory.
[0057] The technical solution of this embodiment is to listen to the listener tags corresponding to all memory addresses with copies stored in the local high-speed cache memory of the main processor in the snoop filtering processing component; share the shared tags corresponding to all memory addresses with copies stored in the shared high-speed cache memory by the main processor and / or the accelerator in the shared high-speed cache memory; the shared high-speed cache memory and the local high-speed cache memory are non-inclusive memories; the shared high-speed cache memory and the local high-speed cache memory are non-inclusive memories; if the listener tag is hit, the snoop filtering processing component sends a snoop request to the main processor; when the main processor determines that the data access request of the accelerator is a write operation according to the snoop request, invalidate the data copy corresponding to the current memory address accessed by the accelerator in the local high-speed cache memory; when the data in the main processor is dirty data, merge the dirty data with the updated data of the accelerator and write it into the shared high-speed cache memory; when the data in the main processor is clean data, the accelerator updates the data in the main memory; when the main processor determines that the data access request of the accelerator is a read operation according to the snoop request, return the data copy corresponding to the current memory address accessed by the accelerator in the local high-speed cache memory and keep the data copy valid; if the shared tag is hit, process the data access request of the accelerator in the shared high-speed cache memory; if the listener tag and the shared tag are not hit, process the data access request of the accelerator in the main memory, which solves the problem that the data in the main memory is updated by another accelerator after the main processor has taken a copy, and at this time the data in the main processor cache will be expired or invalid; and when the main processor writes to the copy in its cache, the main memory will no longer contain the latest data. If the accelerator reads the same address in the main memory at this time, it will see expired data, and the problem of inconsistent data access. On the basis of simplifying the stored information and reducing the hardware overhead, data access consistency can be achieved, and the transmission pressure on the system bus and network can be relieved, and the impact of consistency operations on the performance of the main processor core can also be reduced.
[0058] Figure 3 It is a flowchart of a data access control method for a multi-processor provided according to an embodiment of the present invention. The data access control method for a multi-processor is applied to the data access control architecture of a multi-processor provided in any embodiment of the present invention. As Figure 3 shown, the data access control method for a multi-processor includes:
[0059] Step 310, receive the data access requests of each accelerator through the snoop filtering processing component, and query the listener tag and the shared tag according to the current memory address accessed by the accelerator.
[0060] Specifically, the snooping filter processing component includes: a pre-filter controller, a post-filter controller, a snooping filter controller, a snooping filter array, and a snooping request processing component. The snooping filter array stores the snooper tags corresponding to all memory addresses that have copies stored in the main processor's local cache memory.
[0061] Optionally, receive data access requests from each accelerator through the snooping filter processing component, and query the snooper tags and shared tags according to the memory address currently accessed by the accelerator, including: receiving data access requests from each accelerator through the pre-filter controller, and sending the data access requests from each accelerator to the coherence processing component; querying the snooper tags stored by the snooping filter controller through the snooping filter array and the shared tags in the shared cache memory according to the memory address currently accessed by the accelerator through the coherence processing component.
[0062] According to the specific situation, execute step 320, step 330, and step 340; or, execute step 320, step 330, and step 350; or, execute step 320, step 360; or, execute step 370; or, execute step 380.
[0063] Step 320: If the snooper tag is hit, send a snooping request to the main processor through the snooping filter processing component.
[0064] Optionally, if the snooper tag is hit, send a snooping request to the main processor through the snooping filter processing component, including: when the coherence processing component determines that the memory address currently accessed by the accelerator hits the snooper tag, send the hit information to the snooping request processing component; send a snooping request to the main processor through the snooping request processing component.
[0065] Step 330: When it is determined according to the snooping request that the data access request of the accelerator is a write operation, invalidate the data copy in the local cache memory of the main processor corresponding to the memory address currently accessed by the accelerator through the main processor.
[0066] Optionally, the main processor includes: a main processor request processing component.
[0067] Step 340: When the data in the main processor is dirty data, merge the dirty data with the updated data of the accelerator and write it into the shared cache memory.
[0068] Optionally, when the data in the main processor is dirty data, the dirty data is merged with the updated data of the accelerator and written into the shared cache memory, including: returning the corresponding dirty data to the coherence processing component through the snooping request processing component by the main processor; obtaining the updated data of the accelerator transmitted by the accelerator through the pre-filter controller by the coherence processing component, merging the dirty data with the updated data, and writing the merged result into the shared cache memory.
[0069] Step 350, when the data in the main processor is clean data, update the data accessed by the accelerator in the main memory through the snooping filtering processing component.
[0070] Optionally, when the data in the main processor is clean data, update the data accessed by the accelerator in the main memory through the snooping filtering processing component, including: when the data access request of the accelerator is a write operation and the data in the main processor is clean data, perform a reply response through the snooping request processing component, and the reply response does not include data; transmitting the reply response to the coherence processing component through the snooping request processing component; accessing the main memory access network through the main memory access controller by the coherence processing component, and performing the write operation of the updated data in the accelerator data access request in the main memory.
[0071] Step 360, when it is determined according to the snooping request that the data access request of the accelerator is a read operation, return the data copy corresponding to the current memory address accessed by the accelerator in the local cache memory of the main processor and keep the data copy valid.
[0072] Specifically, when it is determined according to the snooping request that the data access request of the accelerator is a read operation, return the data copy corresponding to the current memory address accessed by the accelerator in the local cache memory of the main processor to the shared cache memory through the snooping request processing component and the coherence processing component, and keep the data copy in the local cache memory of the main processor valid.
[0073] Step 370, if the shared tag is hit, process the data access request of the accelerator in the shared cache memory.
[0074] Specifically, when the coherence processing component determines that the current memory address accessed by the accelerator hits the shared tag, send the hit information to the shared cache memory controller. Control the processing of the data access request of the accelerator in the shared cache memory through the shared cache memory controller.
[0075] Step 380, if the snooper tag and the shared tag are not hit, process the data access request of the accelerator in the main memory.
[0076] Optionally, if the listener tag and the shared tag are not hit, the data access request processing for the accelerator is performed in the main memory, including: when the coherence processing component determines that the current memory address accessed by the accelerator does not hit the listener tag and the shared tag, sending the miss information to the filtered controller; and the filtered controller accessing the main memory access network to perform the data access request processing for the accelerator in the main memory.
[0077] Before step 310, or after step 380, it may further include: when a data access request for the main processor is received through the main processor, querying the shared cache for the shared tag according to the current memory address accessed by the main processor to determine whether there is a hit; if the current memory address accessed by the main processor hits the shared tag, performing the data access for the main processor in the shared cache; otherwise, performing the data access for the main processor through the main memory.
[0078] Specifically, when the main processor receives a data access request for the main processor, the current memory address of the main processor is transmitted to the coherence processing component through the main processor request processing component; the coherence processing component queries the shared cache for the shared tag to determine whether there is a hit; if the current memory address accessed by the main processor hits the shared tag, performing the data access for the main processor in the shared cache; otherwise, the coherence processing component accesses the main memory access network through the main memory controller to perform the data access for the main processor in the main memory.
[0079] The data access control method for multiple processors provided by the embodiments of the present invention can achieve the same beneficial effects as the data access control architecture for multiple processors.
[0080] It should be understood that the various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and this is not limited herein.
[0081] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data access control architecture for a multiprocessor, characterized in that: The data access control architecture of the multiprocessor includes: a main processor, at least one accelerator, a main memory, and a snoop filtering processing component, the main processor and the accelerator access a shared cache memory, and the main processor is provided with a local cache memory; wherein: The snoop filter processing component stores snooper tags corresponding to all memory addresses having copies stored in the local cache memory of the master processor; The shared cache memory stores shared tags corresponding to all memory addresses of which the host processor and / or accelerator has copies stored in the shared cache memory; The shared cache memory and the local cache memory are non-inclusive memories; The snoop filter processing component receives data access requests from each accelerator and queries the snoop tag and the shared tag according to the memory address currently accessed by the accelerator; If the snoop tag is hit, the snoop filter processing component sends a snoop request to the main processor; When the main processor determines that the data access request of the accelerator is a write operation according to the snoop request, the data copy corresponding to the memory address currently accessed by the accelerator in the local cache memory is invalidated; When the data of the main processor is dirty data, the dirty data is merged with the updated data of the accelerator and written into the shared cache memory; The data access control architecture of the multiprocessor further includes: a consistency processing component; A snoop filtering processing component includes: a pre-filtering controller, a post-filtering controller, a snoop filter controller, a snoop filter array, and a snoop request processing component; wherein: The snoop filter array stores snoop tags corresponding to all memory addresses having copies stored in the local cache memory of the master processor; The pre-filter controller receives the data access request of each accelerator and sends the data access request of each accelerator to the consistency processing component; The consistency processing component queries the snoop tag stored by the snoop filter controller through the snoop filter array and the shared tag in the shared cache memory according to the memory address currently accessed by the accelerator; The consistency processing component sends hit information to the snoop request processing component when determining that the memory address currently accessed by the accelerator hits the snoop tag; The snoop request processing component sends a snoop request to the main processor.
2. The multiprocessor data access control architecture according to claim 1, characterized in that: Also includes: When the main processor determines that the data access request of the accelerator is a read operation according to the snoop request, the main processor returns the data copy corresponding to the memory address currently accessed by the accelerator in the local cache memory and keeps the data copy valid.
3. The multiprocessor data access control architecture according to claim 1, characterized in that: Also includes: When the data in the main processor is clean data, the accelerator updates the data in the main memory.
4. The multiprocessor data access control architecture according to claim 1, characterized in that: Also includes: If the shared tag is hit, the data access request of the accelerator is processed in the shared cache memory; If the snooper tag and the shared tag are not hit, the data access request processing of the accelerator is performed in the main memory.
5. The multiprocessor data access control architecture according to claim 1, characterized in that: Also includes: When receiving a data access request to the main processor, the main processor queries the shared tag in the shared cache memory according to the current access memory address of the main processor to determine whether it hits; If the current access memory address of the main processor hits the shared tag, data access to the main processor is performed in the shared cache memory; otherwise, data access to the main processor is performed through the main memory.
6. The multiprocessor data access control architecture according to claim 1, characterized in that: The main processor includes: A main processor request processing unit; When the main processor determines that the data access request of the accelerator is a write operation according to the snoop request, the data copy corresponding to the memory address currently accessed by the accelerator in the local cache memory is invalidated; The main processor returns the corresponding dirty data to the consistency processing component through the snoop request processing component; The consistency processing component obtains the update data of the accelerator transmitted by the accelerator through the pre-filtering controller, merges the dirty data with the update data, and writes the merged result into the shared high-speed cache memory.
7. The multiprocessor data access control architecture according to claim 6, characterized in that: The data access control architecture of the multiprocessor also includes: a main memory access network; When the data access request of the accelerator is a write operation and the data in the main processor is clean, a reply response is made through the snoop request processing component, and the reply response does not contain data; The snoop request processing component transmits the reply response to the consistency processing component; The consistency processing component accesses the main memory access network through the main memory access controller, and executes the update data write operation in the accelerator data access request in the main memory; or, The consistency processing component sends miss information to the post-filtering controller when determining that the memory address currently accessed by the accelerator misses the snooper tag and the shared tag; After filtering, the controller accesses the main memory access network and processes the data access request of the accelerator in the main memory.
8. The multiprocessor data access control architecture according to claim 7, characterized in that: The data access control architecture of the multiprocessor further includes: a main memory access controller; When receiving a data access request to the main processor, the main processor transmits the current access memory address of the main processor to the consistency processing component through the main processor request processing component; The consistency processing component queries the shared tag in the shared cache memory to determine whether it hits; If the current access memory address of the main processor hits the shared tag, data access to the main processor is performed in the shared cache memory; Otherwise, the consistency processing component accesses the main memory access network through the main memory access controller, and performs data access to the main processor in the main memory.
9. A data access control method for a multiprocessor, characterized in that: The data access control method for a multiprocessor is applied to the data access control architecture for a multiprocessor as claimed in any one of claims 1 to 8, and the data access control method for a multiprocessor comprises: Receive data access requests from each accelerator through a snoop filter processing component, and query the snoop tag and the shared tag according to the current memory address accessed by the accelerator; If the snooper tag is hit, a snoop request is sent to the main processor through the snoop filter processing component; When it is determined according to the snoop request that the data access request of the accelerator is a write operation, the data copy corresponding to the memory address currently accessed by the accelerator in the local cache memory is invalidated by the main processor; When the data in the main processor is dirty data, the dirty data is merged with the updated data of the accelerator and written into the shared cache memory; Receiving data access requests from each accelerator through a snoop filter processing component, and querying a snoop tag and a shared tag according to a current memory access address of the accelerator, including: receiving data access requests from each accelerator through a pre-filter controller, and sending the data access requests from each accelerator to a consistency processing component; querying a snoop tag stored by a snoop filter controller through a snoop filter array and a shared tag in a shared cache memory according to a current memory access address of the accelerator through the consistency processing component; If the snoop tag is hit, a snoop request is sent to the main processor through the snoop filter processing component, including: when the consistency processing component determines that the accelerator's current access memory address hits the snoop tag, the hit information is sent to the snoop request processing component; and the snoop request processing component sends a snoop request to the main processor.
Citation Information
Patent Citations
Snoop filter and non-inclusive shared cache memory
CN103136117A
Coherent interconnect for managing snoop operation and data processing apparatus including the same
CN107015923A