Data exchange architecture, system and method capable of realizing cache consistency

By configuring shared storage space and cache state array list in the data exchange device, the management module controls access permissions, solving the cache consistency problem in parallel computing of multiple devices, realizing cache consistency of the data exchange architecture, and is suitable for systems such as deep learning training and graphics rendering.

CN120508413AActive Publication Date: 2025-08-19SHANGHAI XINLIJI SEMICON CO LTD

Patent Information

Application Number
CN202511006599.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-08-19
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

In multi-device parallel computing environment, especially in deep learning training and graphics rendering, traditional PCIe architectures face cache consistency problems, resulting in data inconsistency and calculation errors.

Method used

By configuring the data exchange device, including shared storage space, cache state array list and management module, the management module controls the device's access rights to memory units according to the cache state array to achieve cache consistency, including management of exclusive writes, shared reads, data failure and uncached states.

Benefits of technology

It effectively resolves cache consistency conflicts between multiple devices, avoids the risks of data failure and parallel write and read errors, and is suitable for large-scale parallel systems such as deep learning training and graphics rendering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120508413A_ABST
    Figure CN120508413A_ABST
Patent Text Reader

Abstract

The invention discloses a data exchange architecture, system and method capable of realizing cache consistency. The data exchange architecture comprises a data exchange device and a plurality of devices electrically connected with the data exchange device, the data exchange device comprises a shared storage space, a cache state array list and a management module, wherein the shared storage space comprises memory units in one-to-one correspondence with device IDs, the cache state array list comprises the device IDs in one-to-one correspondence with the device IDs and cache state arrays, and the cache state arrays comprise cache states in one-to-one correspondence with the device IDs. The memory unit of which the cache state is exclusively written cannot be written or read by other equipment except the write-in equipment; the memory unit which is in an exclusive writing state in a non-cache state can be written or read by any equipment; other equipment cannot cache the data of the memory unit through the equipment corresponding to the data failure / non-cache; other devices may cache the data of the memory unit for the shared read device through the cache state. According to the invention, the consistency of caches among multiple devices can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer and communication technologies, and in particular to a data exchange architecture, system and method capable of achieving cache consistency. Background Art

[0002] With the rapid development of artificial intelligence (AI), deep learning (DL), high-performance computing (HPC), and big data processing technologies, data transmission demands are exploding. In particular, in high-bandwidth, low-latency data exchange scenarios, traditional PCIe architectures face numerous challenges, including insufficient bandwidth, latency, and lengthy data transmission paths. To address these challenges, shared memory is a viable technical solution.

[0003] However, the complexity and frequency of data transmission are exploding. Especially in multi-device data exchange scenarios, such as in massively parallel systems like deep learning training, graphics rendering, and high-performance scientific computing, shared memory-based data transmission methods are particularly prone to cache coherence issues caused by data modifications and concurrent access.

[0004] For example, when a GPU's memory is frequently accessed during multi-device parallel computing, cache inconsistency can occur. This means that when data in one GPU's memory is modified, the caches of other GPUs may not be updated in a timely manner, leading to inconsistent or incorrect data. This is a typical synchronization issue in parallel computing environments, and cache coherence becomes particularly prominent when multiple devices (such as GPUs) need to read and write shared data simultaneously.

[0005] The disclosure of the above background technology content is only used to assist in understanding the inventive concept and technical solution of the present invention. It does not necessarily belong to the prior art of the present application, nor does it necessarily provide technical guidance. In the absence of clear evidence that the above content has been disclosed before the filing date of the present application, the above background technology should not be used to evaluate the novelty and creativity of the present application. Summary of the Invention

[0006] The purpose of the present invention is to provide a data exchange architecture, system and method capable of achieving cache consistency, which can achieve cache consistency among multiple devices.

[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows: A data exchange architecture capable of achieving cache consistency, comprising a data exchange device and a plurality of devices to be communicated electrically connected to the data exchange device, each of the devices being configured with a unique device ID; The data exchange device includes a shared storage space, a cache status array table, and a management module. The shared storage space includes memory cells corresponding to the device IDs one by one. The cache status array table includes a cache status array corresponding to the device IDs one by one. The cache status array includes cache status corresponding to each device ID one by one. The cache status includes exclusive write, shared read, data invalid, and uncached. The exclusive write indicates that the memory cell is being written to and can only be written to by one device. The shared read indicates that the device only reads data from the memory cell and does not modify the data. The data invalid indicates that the data read from the memory cell by the device is invalid. The uncached indicates that the device does not read data from the memory cell. For each memory unit, the management module is configured to determine the access rights of each device to the memory unit according to the cache status array corresponding to the memory unit, including: When a cache state in the cache state array is exclusive write, the memory unit cannot be written / read by other devices except the current write device, and the device ID corresponding to the current write device is consistent with the device ID corresponding to the exclusive write; When the cache state array does not contain a cache state indicating exclusive write, the memory unit can be written to / read from by one device; The management module is further configured to determine, based on the cache status array, access rights of other devices to data in a memory unit cached by a device, including: For a cache status array corresponding to a memory unit, other devices cannot cache data in the memory unit through a device corresponding to a cache status of data invalid / uncached; For a cache status array corresponding to a memory unit, other devices can cache data in the memory unit corresponding to the device for shared reading through the cache status.

[0008] Further, based on any one of the above technical solutions or a combination of multiple technical solutions, the management module is further configured to update the cache status array according to the current access status of the memory unit and the current cache status of its corresponding cache status array; The access status includes whether the accessed object reads data or writes data, and the accessed object is one or more of the multiple devices.

[0009] Further, based on any one of the above technical solutions or a combination of multiple technical solutions, for each of the memory units, the management module updates the cache status array in the following manner: If the memory unit is written with data by a requesting device, the management module modifies the cache state corresponding to the requesting device ID in the cache state array to exclusive write, and the requesting device is one of the devices; Update the cache status of other device IDs corresponding to shared read to invalid status; The cache status corresponding to other device IDs remains unchanged as uncached / data invalid.

[0010] Further, based on any one of the above technical solutions or a combination of multiple technical solutions, for each of the memory units, the management module updates the cache status array in the following manner: If the requesting device completes writing data to the memory unit, the management module changes the cache status corresponding to the requesting device ID in the cache status array from exclusive writing to shared reading.

[0011] Further, based on any one of the above technical solutions or a combination of multiple technical solutions, for each of the memory units, the management module updates the cache status array in the following manner: If the requesting device completes writing data to the memory unit, the management module broadcasts a data update notification to other devices; In response to receiving the data update notification, the other device may selectively cache the data in the memory unit; If the device caches the data in the memory unit, the management module updates the cache state corresponding to the device ID to data sharing; If the device does not cache the data in the memory unit, the management module maintains the cache state corresponding to the device ID unchanged.

[0012] Furthermore, based on any one of the technical solutions or a combination of multiple technical solutions described above, for each of the memory units, the management module responds to the requesting device completing writing data to the memory unit, then for the cache status array corresponding to the memory unit, determines that the device ID corresponding to the cache status of data sharing is the target device ID, and the management module broadcasts a data update notification to the device corresponding to the target device ID.

[0013] Further, based on any one of the above technical solutions or a combination of multiple technical solutions, for each of the memory units, the management module updates the cache status array in the following manner: If the memory unit is requested to read data by a device, the management module modifies the cache state corresponding to the requesting device ID in the cache state array to shared read, and the requesting device is one of the devices; The cache status corresponding to other device IDs remains unchanged as data sharing / uncached / data invalid.

[0014] Furthermore, based on any one of the above technical solutions or a combination of multiple technical solutions, the shared storage space is a DMA-based global address space, the memory unit is a DMA address, and the DMA address corresponding to the same device ID directly communicates with the physical memory address of the device; The management module is further configured to transfer data in one DMA address to another DMA address.

[0015] Furthermore, based on any one of the above technical solutions or a combination of multiple technical solutions, it further includes an uplink switch, wherein the uplink switch is electrically connected to each of the devices, and the uplink switch is electrically connected to the CPU; When the two devices need to communicate, the two devices communicate with the CPU through the data transmission unit instead of through the uplink switch.

[0016] Furthermore, any one of the above technical solutions or a combination of multiple technical solutions includes a first device group and a second device group, wherein the first device group and the second device group each include a plurality of the devices; The plurality of devices in the first device group are electrically connected to a first uplink switch, respectively; the plurality of devices in the second device group are electrically connected to a second uplink switch, respectively; and the first uplink switch and the second uplink switch are electrically connected to a CPU, respectively; Each device in the first device group and the second device group is electrically connected to the data exchange device, and any two devices communicate through the data transmission unit instead of through the first uplink switch, the second uplink switch and the CPU.

[0017] Furthermore, according to any one of the above technical solutions or a combination of multiple technical solutions, the DMA addresses do not overlap; and / or, There is no communication between the DMA addresses corresponding to different device IDs and the physical memory addresses of the devices; and / or, The device is a PCIe device.

[0018] According to another aspect of the present invention, a communication system is provided, which includes a data exchange architecture capable of achieving cache consistency as described in any one of the above technical solutions or a combination of multiple technical solutions.

[0019] According to another aspect of the present invention, a data exchange method capable of achieving cache consistency is provided, comprising the following steps: A data exchange device is configured to be electrically connected to a plurality of devices to be communicated, each of the devices being configured with a unique device ID, the data exchange device comprising a shared storage space, a cache status array table, and a management module, the shared storage space comprising memory cells corresponding one-to-one to the device IDs, the cache status array comprising a cache status corresponding one-to-one to the device IDs, the cache status comprising exclusive write, shared read, data invalidation, and uncached; For each memory unit, the management module is configured to determine the access rights of each device to the memory unit according to the cache status array corresponding to the memory unit, including: When a cache state in the cache state array is exclusive write, the memory unit cannot be written / read by other devices except the current write device, and the device ID corresponding to the current write device is consistent with the device ID corresponding to the exclusive write; When there is no cache state of exclusive write in the cache state array, the memory unit can be written / read by any device; The management module is further configured to determine, based on the cache status array, access rights of other devices to data in a memory unit cached by a device, including: For a cache status array corresponding to a memory unit, other devices cannot cache data in the memory unit through a device corresponding to a cache status of data invalid / uncached; For a cache status array corresponding to a memory unit, other devices can cache data in the memory unit corresponding to the device for shared reading through the cache status.

[0020] Furthermore, based on any one of the above technical solutions or a combination of multiple technical solutions, the cache status array table is updated according to the data exchange process by using the management module, including: For the cache status array corresponding to each of the memory units, the management module updates the cache status array according to the current access status of the memory unit and the current cache status of each cache status in the cache status array.

[0021] Further, based on any one of the above technical solutions or a combination of multiple technical solutions, for each of the memory units, the management module updates the cache status array in the following manner: If the memory unit is written with data by a requesting device, the management module modifies the cache state corresponding to the requesting device ID in the cache state array to exclusive write, and the requesting device is one of the devices; Update the cache status of other device IDs corresponding to shared read to invalid status; The cache status corresponding to other device IDs remains unchanged as uncached / data invalid.

[0022] Further, based on any one of the above technical solutions or a combination of multiple technical solutions, for each of the memory units, the management module updates the cache status array in the following manner: If the requesting device completes writing data to the memory unit, the management module sends a data update notification to other devices; In response to receiving the data update notification, the other device may selectively cache the data in the memory unit; If the device caches the data in the memory unit, the management module updates the cache state corresponding to the device ID to data sharing; If the device does not cache the data in the memory unit, the management module maintains the cache state corresponding to the device ID unchanged.

[0023] The beneficial effects brought about by the technical solution provided by the present invention are as follows: a. The present invention configures a shared storage space in a data exchange device that includes memory cells corresponding one-to-one to devices, a cache status array table including a cache status array corresponding one-to-one to each memory cell, the cache status array including cache states corresponding one-to-one to each device, and multiple cache states including exclusive write, shared read, data invalidation, and uncached. A management module controls each device's access rights to each memory cell and the access rights of a device's previously cached memory cell based on the cache status array. This avoids the risk of errors such as data being accessed after invalidation and multiple devices competing for the same memory cell, and facilitates consistent operations such as write synchronization and read redirection. b. The present invention uses a management module to update the cache status of a memory cell based on its access status and its corresponding cache status array. This allows for the effective and rational adjustment of the corresponding cache status array based on the access status of each memory cell, thereby effectively adjusting the access rights of other devices to the memory cell and effectively avoiding the risk of parallel write and read errors caused by caching old values. c. The present invention not only timely adjusts the cache status array corresponding to a memory unit through the management module while the memory unit is being accessed, but also, after the memory unit is accessed, the management module broadcasts the updated information of the memory unit to other devices, and updates the cache status array corresponding to the memory unit again based on whether the other devices synchronize the data of the memory unit. This can completely resolve cache consistency conflicts caused by data modification and concurrent access, and is particularly suitable for large-scale parallel systems such as deep learning training, graphics rendering, and high-performance scientific computing. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0025] Figure 1 A schematic diagram of a data exchange architecture according to an exemplary embodiment of the present invention; Figure 2 A schematic diagram illustrating a configuration principle of a cache status array table provided for an exemplary embodiment of the present invention; Figure 3 A schematic diagram of a principle for updating a cache status array 1 when data is written to a memory unit 1 by a device 1 according to an exemplary embodiment of the present invention is provided; Figure 4 A schematic diagram of a principle for updating a cache status array 1 when a device 3 reads data from a memory unit 1 according to an exemplary embodiment of the present invention is provided; Figure 5 A schematic diagram of a first principle of implementing data transmission between devices based on shared memory is provided as an exemplary embodiment of the present invention; Figure 6 A schematic diagram of a second principle of implementing data transmission between devices based on shared memory, provided as an exemplary embodiment of the present invention; Figure 7 A flowchart of a data exchange architecture provided for an exemplary embodiment of the present invention when device 1 needs to write data to memory unit 1. DETAILED DESCRIPTION

[0026] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0027] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0028] In a multi-device data exchange architecture based on shared memory, the shared memory will be frequently accessed and modified, and each GPU will read data from the shared memory into its own cache, which will lead to cache consistency.

[0029] Specifically, if multiple GPUs share the same memory area and each GPU caches its own copy of the data, cache consistency issues will occur, including at least the following aspects.

[0030] Data modification asynchrony: Suppose GPU 1 modifies data in shared memory (for example, by updating the weight of a model parameter), but this modification only exists in the newly updated shared memory. If GPU 2 continues to use the old data in its local cache instead of retrieving the updated data from shared memory, data inconsistency will occur, potentially leading to incorrect calculation results.

[0031] Cache invalidation: When GPU1 modifies data, GPU2 needs to know that the data in its cache has been invalidated so that it can reload the latest data from shared memory. If GPU2 is not notified in time to update its cache, GPU2 will continue to use outdated data, causing calculation problems.

[0032] Based on the above problems, in one embodiment of the present invention, a data exchange architecture capable of achieving cache consistency is provided, such as Figure 1 、 Figure 2 、 Figure 7As shown in Table 1, it includes a data exchange device and a plurality of devices to be communicated electrically connected to the data exchange device, each of the devices is configured with a unique device ID, and optionally, the device is a PCIe device such as a GPU, an ASIC, etc.; The data exchange device includes a shared storage space, a cache status array table (CSA), and a cache coherence monitor (CCM module). The shared storage space includes memory cells corresponding to the device IDs one by one. The cache status array table includes a cache status array corresponding to the device IDs one by one. The cache status array includes cache status corresponding to each device ID one by one. The cache status includes exclusive write (Write-Owner), shared read (R), data expired (X), and not cached (N). Exclusive write indicates that the memory cell is being written by a device and can only be written by the device. Shared read indicates that the device only reads data from the memory cell and does not modify the data. Data expire indicates that data read from the memory cell by the device is invalid. Not cached indicates that the device does not read data from the memory cell. For each memory unit, the management module is configured to determine the access rights of each device to the memory unit according to the cache status array corresponding to the memory unit, including: When a cache state in the cache state array is exclusive write, the memory unit cannot be written / read by other devices except the current write device, and the device ID corresponding to the current write device is consistent with the device ID corresponding to the exclusive write; When the cache state array does not contain a cache state indicating exclusive write, the memory unit can be written to / read from by one device; The management module is further configured to determine, based on the cache status array, access rights of other devices to data in a memory unit cached by a device, including: For a cache status array corresponding to a memory unit, other devices cannot cache data in the memory unit through a device corresponding to a cache status of data invalid / uncached; For a cache status array corresponding to a memory unit, other devices can cache data in the memory unit corresponding to the device for shared reading through the cache status.

[0033] See also Figure 3For the current cache status array of memory unit 1 corresponding to device 1's ID, if the cache status corresponding to device 2's ID is R, this indicates that device 2 has previously shared read data from memory unit 1, and the data cached by device 2 in memory unit 1 can be read by other devices. In other words, device 2 can share the data it read from memory unit 1 with other devices. If the cache status corresponding to device 3's ID is X, this indicates that the data previously read from memory unit 1 by device 3 has expired (if there is a need to use the data, the data must be updated first; the expired data cannot be used). The data previously cached by device 3 in memory unit 1 cannot be read by other devices. In other words, device 3 cannot share the data it read from memory unit 1 with other devices. If the cache status corresponding to device 4's ID is N, this indicates that device 4 has not cached the data read from memory unit 1, and other devices cannot access the data cached by device 4 in memory unit 1. In other words, device 4 cannot share the data in memory unit 1 with other devices. Data corresponding to shared reads can be used and shared, but data corresponding to expired data cannot be used or shared. If the data in the memory unit is not cached, it does not exist and therefore cannot be used or shared.

[0034] In this embodiment, during the configuration of the data exchange device, an empty CSA table is initialized (indexed by device ID, i.e., block address). After each device is booted, the cache controller of each device actively registers with the CCM and establishes a status reporting channel. Each row in the cache status array in the CSA table consists of a block address (device ID) and a status bit field for devices 1 through n (GPU0 through GPUn).

[0035] When accessing the memory unit, the cache controller of each device uploads the cache block status of the current operation, i.e., the access request, to the management module via a corresponding transaction layer, such as a status reporting transaction layer, during a read / write operation. The management module queries the cache status array table and determines whether to allow the current access request operation, including triggering an invalidation / rejection notification, allowing the read / write operation, or allowing the read / write operation or read forwarding.

[0036] Specifically, when any GPU (such as GPU1) makes an access request to write data to memory unit 1 of the shared memory, the data exchange device first sends GPU1's access request to the management module; the management module queries each cache state in the cache state array 1 corresponding to the GPU1 ID in the cache state array table. If the cache state corresponding to other GPU IDs in each cache state is exclusive write, the management module returns a write rejection signal to GPU1; if the cache state corresponding to no other GPU ID is exclusive write, the management module returns a write approval signal to GPU1.

[0037] Specifically, when any GPU (such as GPU1) makes an access request to read data from memory unit 1 of the shared memory, the data exchange device first sends the access request of GPU1 to the management module in the management module; the management module queries the cache status in the cache status array 1 corresponding to the GPU1 ID in the cache status array table. If the cache status corresponding to one GPU ID among the cache statuses is exclusive write, the management module returns a read rejection signal to GPU1; if the cache status corresponding to no other GPU ID is exclusive write, the management module returns a read approval signal to GPU1.

[0038]

[0039] In this embodiment, the management module is also configured to update the cache status array based on the current access status of the memory unit and the current cache status of its corresponding cache status array; the access status includes reading data or writing data by the accessed object, and the access object is one or more of the multiple devices.

[0040] Specifically, for each of the memory units, the management module updates the cache status array in the following manner.

[0041] If the memory cell is being written to by the requesting device, the management module modifies the cache status corresponding to the requesting device ID in the cache status array to exclusive write, where the requesting device is one of the devices. It should be noted that if the memory cell is currently being written to by another device, then according to Table 1, the memory cell cannot be written to by the requesting device. That is, for the same memory cell, there can only be one cache status in the cache status array that is exclusive write. This prevents data errors caused by two devices modifying the same memory cell at the same time.

[0042] like Figure 3 As shown, in the cache status array 1 corresponding to memory unit 1, cache status 1 to cache status n are the access statuses of devices 1 to n to memory unit 1, respectively. The current access status of memory unit 1 is data being written by device 1. It should be noted that, for cache status array 1 corresponding to memory unit 1, when the current cache status corresponding to device 1 ID is shared read / uncached / data invalid (R / N / X) and the cache status corresponding to other device IDs does not have exclusive write (W), then device 1 can write data to memory unit 1. Based on the current access status of memory unit 1, the management module sets cache status 1 in the cache status array 1 corresponding to memory unit 1 to exclusive write (W) and updates the cache status corresponding to other devices in cache status array 1.

[0043] exist Figure 3 In the example, if cache status 2 is shared read (R), it means that device 2 has previously read data from memory unit 1 and the data is shared. Since new data is currently being written into memory unit 1 by device 1, in order to ensure cache data consistency between multiple devices, including that the data read by device 2 from memory unit 1 is not used by itself and is not further read by other devices, cache status 2 needs to be changed from shared read (R) status to data invalidation (X) status.

[0044] After device 1 completes writing to memory unit 1, the management module changes the cache state 1 corresponding to the requesting device ID in the cache state array from exclusive write to shared read, and the management module sends a data update notification to other devices. In response to receiving the data update notification, the other devices may selectively cache the data in the memory unit.

[0045] For example, if device 2 receives the data update notification and caches the data in memory unit 1, the management module updates the cache state 2 corresponding to device 2 to data sharing. In this way, when the data in memory unit 1 is updated, other devices that previously cached the data in memory unit 1 can also update the data synchronously.

[0046] Alternatively, device 2 responds to receiving the data update notification but does not cache the data in memory unit 1. For example, if device 2 no longer needs to use the data in memory unit 1, it may choose not to cache it. In this case, cache status 2 corresponding to device 2 maintains the data invalid (X). This ensures data cache consistency without adding additional system overhead.

[0047] exist Figure 3 In the example, cache status 3 in cache status array 1 is "Data Stale" (X), indicating that device 3 previously cached data in memory unit 1. However, after the data in memory unit 1 was updated, device 3 did not synchronize its cached data from memory unit 1 (as described above for device 2). While in this "Data Stale" state, the previously cached data in memory unit 1 cannot be read by other devices. Furthermore, because memory unit 1 is currently being written to by device 1, meaning the previously cached data is being updated, the management module maintains cache status 3 in cache status array 1 as "Data Stale" (X) in either case to prevent other devices from caching outdated, erroneous data.

[0048] Similarly, cache state 4 in cache state array 1 is not cached (N), indicating that device 4 has not previously cached the data in memory unit 1. Since device 4 has not previously cached the data in memory unit 1, the management module maintains cache state 4 in cache state array 1 as not cached (N) to prevent other devices from caching delayed or erroneous data.

[0049] For each of the memory units, the management module also updates the cache status array in the following manner: if the memory unit is read by a requesting device, the management module modifies the cache status corresponding to the requesting device ID in the cache status array to shared read, and the requesting device is one of the devices; and maintains the cache status corresponding to other device IDs as data sharing / uncached / data invalid.

[0050] like Figure 4 As shown, in cache status array 1 corresponding to memory unit 1, cache status 1 to cache status n represent the access statuses of devices 1 to n to memory unit 1, respectively. The current access status of memory unit 1 is data being read by device 3. It should be noted that for cache status array 1 corresponding to memory unit 1, if the cache status corresponding to each device ID does not have exclusive write (W), device 3 can read data from memory unit 1. The management module changes cache status 1 in the cache status array from shared read / data invalidated / not cached (R / X / N) to shared read (R).

[0051] For the cache status corresponding to other device IDs, there are only three possible states: shared read, data invalidation, and uncached. For example Figure 4 As shown, the management module maintains the cache status corresponding to other device IDs as data sharing / uncached / data invalid.

[0052] As described above, in the present invention, the management module serves as the control core of the data exchange device, responsible for coordinating the cache status of each GPU. The cache status array table (CSA table) is a database of the real-time cache status of each memory unit in each device. During each data read / write step, the management module performs table lookup, update, broadcast, and state transfer operations on the cache status array table, thus efficiently achieving data cache consistency across multiple devices.

[0053] Of course, in another embodiment of the present invention, for each of the memory units, the management module responds to the requesting device completing the write operation on the corresponding memory unit, then for the cache status array corresponding to the memory unit, it determines that the device ID corresponding to the cache status of data sharing is the target device ID, and the management module broadcasts a data update notification to the device corresponding to the target device ID. In response to receiving the data update notification, the device reads the latest data in the memory unit, that is, receives the broadcast data update notification and its corresponding previous cache data. This embodiment is different from the above embodiment in that, in this embodiment, after the data in a memory unit is updated, the management module does not broadcast the data update notification to all other devices, but only broadcasts to the devices that have cached the memory unit. Such a broadcast operation is more targeted, and the device that receives the broadcast data update notification will definitely update its corresponding previous cache to ensure the consistency of the data cache.

[0054] With respect to the data exchange architecture capable of achieving cache coherence described in any of the above embodiments, the shared memory space is preferably a DMA-based global address space, the memory units are DMA addresses, and the DMA addresses corresponding to the same device ID directly communicate with the physical memory addresses of the device. Preferably, the DMA addresses do not overlap.

[0055] In this embodiment, the data exchange device further includes a global address mapping table, which includes one-to-one correspondences between device IDs and DMA addresses. The DMA addresses corresponding to the same device ID directly communicate with the physical memory address.

[0056] When a requesting device sends an access request to the data exchange apparatus, the access request including the requesting device ID, the target device ID, and the access request type, the management module determines a DMA address in the global address mapping table that matches the target device ID as the target DMA address, and implements a data transfer operation from the requesting device to the target device via the target DMA address.

[0057] In this embodiment, the data exchange architecture further includes an uplink switch, which is electrically connected to each of the devices, and the uplink switch is electrically connected to the CPU; when two of the devices need to communicate, the two devices do not communicate with the CPU through the uplink switch, but communicate through the data transmission unit.

[0058] There are two ways to implement the data transfer operation from the requesting device to the target device through the target DMA address, including reading / writing. One is the memory sharing implementation method such as Figure 5As shown, each DMA address serves as a memory unit. The DMA address corresponding to the same device ID communicates directly with the physical memory address, and the physical memory address corresponding to different device IDs cannot communicate directly with the DMA address and needs to be forwarded through the management module.

[0059] For example, when device 1 needs to send data to device 4, it sends an access request to the management module, i.e., the access request type is write. The management module determines that the target DMA address is DMA address 4 based on the target device ID in the access request, and searches the cache status array 4 corresponding to DMA address 4 in the CAS table to find that each cache status does not have exclusive write access. (It should be noted that because DMA address 1 can only be written by device 1, it is not necessary to query each cache status in the cache status array 1 corresponding to DMA address 1 when determining whether to grant the access request.) The management module then returns an access approval signal to device 1 and sends the access request to device 4. Otherwise, the management module returns an access rejection signal to device 1.

[0060] In response to receiving the access consent signal, device 1 directly moves the data in the physical memory address of device 1 to DMA address 1 (equivalent to memory unit 1) through the DMA controller corresponding to device 1. During this process, as in the data cache consistency control method described in the above embodiment, the management module updates its cache status array according to the access status of DMA address 1 and its current cache status array; the management module moves the data in DMA address 1 to DMA address 4. During this process, as in the data cache consistency control method described in the above embodiment, the management module updates its cache status array according to the access status of DMA address 4 and its current cache status array; the DMA controller corresponding to device 4 directly moves the data in DMA address 4 to the physical memory address of device 4. During this process, as in the data cache consistency control method described in the above embodiment, the management module updates its cache status array according to the access status of DMA address 4 and its current cache status array.

[0061] For example, when device 1 needs to read data from device 4, it sends an access request (read type) to the management module. Based on the target device ID in the access request, the management module determines that the target DMA address is DMA address 4. It then searches the CAS table for cache status array 4 corresponding to DMA address 4 to verify that each cache status in the cache status array 1 corresponding to DMA address 4 is not exclusive write-capable. (It should be noted that because DMA address 1 can only be written by device 1, it is not necessary to query each cache status in the cache status array 1 corresponding to DMA address 1 when determining whether the access request is approved.) The management module then returns an access approval signal to device 1 and sends the access request to device 4. Otherwise, the management module returns an access denial signal to device 1.

[0062] In response to receiving the access request, the device 4's corresponding DMA controller directly moves the data in the physical memory address of the device 4 to the DMA address 4. During this process, as in the data cache consistency control method described in the above embodiment, the management module updates its cache status array based on the access status of the DMA address 4 and its current cache status array. The management module moves the data in the DMA address 4 to the DMA address 1. During this process, as in the data cache consistency control method described in the above embodiment, the management module updates its cache status array based on the access status of the DMA address 1 and its current cache status array. In response to receiving the access consent signal, the device 1 directly moves the data in the DMA address 1 (equivalent to the memory unit 1) to the physical memory address of the device 1 through the DMA controller corresponding to the device 1. During this process, as in the data cache consistency control method described in the above embodiment, the management module updates its cache status array based on the access status of the DMA address 1 and its current cache status array.

[0063] Another way to implement memory sharing is as follows Figure 6 As shown, each of the DMA addresses serves as a memory unit, and each of the devices can perform read / write operations on the DMA address corresponding to any device ID. In this manner, data forwarding in different DMA addresses does not need to be performed through a management module.

[0064] For example, when device 1 needs to send data to device 4, device 1 sends an access request to the management module, i.e., the access request type is write. The management module determines that the target DMA address is DMA address 4 based on the target device ID in the access request, and searches the cache status array 4 corresponding to DMA address 4 in the CAS table to find out whether each cache status does not have exclusive write (it should be noted that because data in the physical memory address of device 1 can be written directly to DMA address 4 without being transferred through DMA address 1, it is not necessary to query the cache status array 1 corresponding to DMA address 1 when determining whether the access request can be approved). The management module then returns an access approval signal to device 1 and sends the access request to device 4. Otherwise, the management module returns an access rejection signal to device 1.

[0065] In response to receiving the access approval signal, device 1 directly moves the data in the physical memory address of device 1 to DMA address 4 (equivalent to memory unit 4) through the DMA controller corresponding to device 1. During this process, as in the data cache consistency control method described in the above embodiment, the management module updates its cache status array based on the access status of DMA address 4 and its current cache status array. In response to receiving the access request, the DMA controller corresponding to device 4 directly moves the data in DMA address 4 to the physical memory address of device 4. During this process, as in the data cache consistency control method described in the above embodiment, the management module again updates its cache status array 4 based on the access status of DMA address 4 and its current cache status array.

[0066] For example, when device 1 needs to read data from device 4, it sends an access request (i.e., the access request type is read) to the management module. The management module determines that the target DMA address is DMA address 4 (equivalent to memory unit 4) based on the target device ID in the access request. It then searches the CAS table for each cache state in cache state array 4 corresponding to DMA address 4 to verify that there is no exclusive write. (It should be noted that because the physical memory address of device 1 can directly read data from DMA address 4 without being transferred through DMA address 1, it is not necessary to query each cache state in cache state array 1 corresponding to DMA address 1 when determining whether the access request can be granted.) The management module then returns an access approval signal to device 1 and sends the access request to device 4. Otherwise, the management module returns an access denial signal to device 1.

[0067] In response to receiving the access request, the DMA controller corresponding to device 4 directly moves the data in the physical memory address of device 4 to DMA address 4. During this process, as in the data cache consistency control method described in the above embodiment, the management module updates its cache status array based on the access status of DMA address 4 and its current cache status array. In response to receiving the access consent signal, the DMA controller corresponding to device 1 directly moves the data in DMA address 4 (equivalent to memory unit 1) to the physical memory address of device 1 through the DMA controller corresponding to device 1. During this process, as in the data cache consistency control method described in the above embodiment, the management module updates the cache status array 4 based on the access status of DMA address 1 and its current cache status array.

[0068] It should be noted that, in the above two implementation examples, when device 1 needs to read data from device 4, if the cache status corresponding to the device 1 ID in the cache status array 4 corresponding to the DMA address 4 is data invalid / uncached, the management module can also directly return an access denial signal to device 1.

[0069] In another embodiment of the present invention, the data exchange architecture includes a first device group and a second device group, each of which includes multiple devices. The multiple devices in the first device group are electrically connected to a first upstream switch, and the multiple devices in the second device group are electrically connected to a second upstream switch. The first upstream switch and the second upstream switch are electrically connected to a CPU. Each device in the first device group and the second device group is electrically connected to the data exchange device, and any two devices communicate not through the first upstream switch, the second upstream switch, and the CPU, but through the data transmission unit. The first upstream switch and the second upstream switch can be electrically connected to the same CPU or to different CPUs.

[0070] In one embodiment of the present invention, a communication system is further provided, which includes the data exchange architecture capable of achieving cache consistency as described in any one of the above embodiments or a combination of multiple embodiments.

[0071] In one embodiment of the present invention, a data exchange method capable of achieving cache consistency is also provided. Figure 7 As shown, the data exchange method includes the following steps: A data exchange device is configured to be electrically connected to a plurality of devices to be communicated, each of the devices being configured with a unique device ID, the data exchange device comprising a shared storage space, a cache status array table, and a management module, the shared storage space comprising memory cells corresponding one-to-one to the device IDs, the cache status array comprising a cache status corresponding one-to-one to the device IDs, the cache status comprising exclusive write, shared read, data invalidation, and uncached; For each memory unit, the management module is configured to determine the access rights of each device to the memory unit according to the cache status array corresponding to the memory unit, including: When a cache state in the cache state array is exclusive write, the memory unit cannot be written / read by other devices except the current write device, and the device ID corresponding to the current write device is consistent with the device ID corresponding to the exclusive write; When there is no cache state of exclusive write in the cache state array, the memory unit can be written / read by any device; The management module is further configured to determine, based on the cache status array, access rights of other devices to data in a memory unit cached by a device, including: For a cache status array corresponding to a memory unit, other devices cannot cache data in the memory unit through a device corresponding to a cache status of data invalid / uncached; For a cache status array corresponding to a memory unit, other devices can cache data in the memory unit corresponding to the device for shared reading through the cache status.

[0072] In this embodiment, the management module is used to update the cache status array table according to the data exchange process, including: for the cache status array corresponding to each memory unit, the management module updates the cache status array according to the current access status of the memory unit and the current cache status of each cache status in the cache status array.

[0073] Specifically, for each of the memory units, the management module updates the cache status array in the following manner: if the memory unit is written by a requesting device, the management module modifies the cache status corresponding to the requesting device ID in the cache status array to exclusive write, and the requesting device is one of the devices; the management module updates the cache status corresponding to other device IDs from shared read to invalid state; the management module maintains the cache status corresponding to other device IDs as uncached / data invalid.

[0074] For each of the memory units, the management module updates the cache status array in the following manner: if the requesting device completes writing to the memory unit, the management module sends a data update notification to other devices; in response to receiving the data update notification, the other devices may selectively cache the data in the memory unit; if the device caches the data in the memory unit, the management module updates the cache status corresponding to the device ID to data sharing; if the device does not cache the data in the memory unit, the management module maintains the cache status corresponding to the device ID unchanged.

[0075] In one embodiment of the present invention, a data exchange device capable of achieving cache consistency is further provided, the data exchange device comprising the shared storage space, cache status array table, and management module as described in the above embodiments. The shared storage space comprises memory cells corresponding one-to-one with the device IDs, the data exchange device comprises a plurality of ports, each of which is a high-speed communication port supporting high-bandwidth parallel communication, the management module being electrically connected to each of the ports, each of which is configured to electrically connect to a device to be communicated with, such as a PCIe device.

[0076] The management module controls the access rights of each device to the data in each memory unit based on the access request made by each device and the cache status array table. The working principle and working method of the data exchange device are the same as the above-mentioned data exchange architecture embodiment capable of achieving cache consistency, and will not be repeated here.

[0077] It should be noted that the communication system, the data exchange method capable of achieving cache consistency and the data exchange device capable of achieving cache consistency provided by the present invention are the same as the inventive concept of the above-mentioned data exchange architecture embodiment capable of achieving cache consistency. The entire content of the data exchange architecture embodiment capable of achieving cache consistency is incorporated into the communication system, the data exchange method capable of achieving cache consistency and the data exchange device embodiment capable of achieving cache consistency by introduction.

[0078] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0079] The above is only a specific implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A data exchange architecture capable of achieving cache consistency, characterized in that: It includes a data exchange device and a plurality of devices to be communicated with which are electrically connected, each of the devices being configured with a unique device ID; The data exchange device includes a shared storage space, a cache status array table and a management module, wherein the shared storage space includes memory units corresponding to the device IDs one by one, and the cache status array table includes cache status arrays corresponding to the device IDs one by one; The cache status array includes cache status corresponding to each device ID, and the cache status includes exclusive write, shared read, data invalidation and uncached; For each memory unit, the management module is configured to determine the access rights of each device to the memory unit according to the cache status array corresponding to the memory unit, including: When a cache state in the cache state array is exclusive write, the memory unit cannot be written / read by other devices except the current write device, and the device ID corresponding to the current write device is consistent with the device ID corresponding to the exclusive write; When there is no cache state of exclusive write in the cache state array, the memory unit can be written / read by any device; The management module is further configured to determine, based on the cache status array, access rights of other devices to data in a memory unit cached by a device, including: For a cache status array corresponding to a memory unit, other devices cannot cache data in the memory unit through a device corresponding to a cache status of data invalid / uncached; For a cache status array corresponding to a memory unit, other devices can cache data in the memory unit corresponding to the device for shared reading through the cache status.

2. The data exchange architecture capable of achieving cache consistency according to claim 1, wherein: The management module is further configured to update the cache status array according to the current access status of the memory unit and the current cache status of the corresponding cache status array; The access status includes whether the accessed object reads data or writes data, and the accessed object is one or more of the multiple devices.

3. The data exchange architecture capable of achieving cache consistency according to claim 2, wherein: For each of the memory units, the management module updates the cache status array in the following manner: If the memory unit is written with data by a requesting device, the management module modifies the cache state corresponding to the requesting device ID in the cache state array to exclusive write, and the requesting device is one of the devices; Update the cache status of other device IDs corresponding to shared read to invalid status; The cache status corresponding to other device IDs remains unchanged as uncached / data invalid.

4. The data exchange architecture capable of achieving cache consistency according to claim 3, wherein: For each of the memory units, the management module updates the cache status array in the following manner: If the requesting device completes writing data to the memory unit, the management module changes the cache status corresponding to the requesting device ID in the cache status array from exclusive writing to shared reading.

5. The data exchange architecture capable of achieving cache consistency according to claim 3, wherein: For each of the memory units, the management module updates the cache status array in the following manner: If the requesting device completes writing data to the memory unit, the management module broadcasts a data update notification to other devices; In response to receiving the data update notification, the other device may selectively cache the data in the memory unit; If the device caches the data in the memory unit, the management module updates the cache state corresponding to the device ID to data sharing; If the device does not cache the data in the memory unit, the management module maintains the cache state corresponding to the device ID unchanged.

6. The data exchange architecture capable of achieving cache consistency according to claim 3, wherein: For each of the memory units, the management module responds to the requesting device completing writing data to the memory unit, then for the cache status array corresponding to the memory unit, determines that the device ID corresponding to the cache status of data sharing is the target device ID, and the management module broadcasts a data update notification to the device corresponding to the target device ID.

7. The data exchange architecture capable of achieving cache consistency according to claim 2, wherein: For each of the memory units, the management module updates the cache status array in the following manner: If the memory unit is requested to read data by a device, the management module modifies the cache state corresponding to the requesting device ID in the cache state array to shared read, and the requesting device is one of the devices; The cache status corresponding to other device IDs remains unchanged as data sharing / uncached / data invalid.

8. The data exchange architecture capable of achieving cache consistency according to claim 1, wherein: The shared memory space is a global address space based on DMA, the memory unit is a DMA address, and the DMA address corresponding to the same device ID directly communicates with the physical memory address of the device; The management module is further configured to transfer data in one DMA address to another DMA address.

9. The data exchange architecture capable of achieving cache consistency according to claim 8, characterized in that: It also includes an uplink switch, wherein the uplink switch is electrically connected to each of the devices, and the uplink switch is electrically connected to the CPU; When the two devices need to communicate, the two devices communicate with the CPU through the data transmission unit instead of through the uplink switch.

10. The data exchange architecture capable of achieving cache consistency according to claim 9, characterized in that: comprising a first device group and a second device group, wherein the first device group and the second device group each comprise a plurality of the devices; The plurality of devices in the first device group are electrically connected to a first uplink switch, respectively; the plurality of devices in the second device group are electrically connected to a second uplink switch, respectively; and the first uplink switch and the second uplink switch are electrically connected to a CPU, respectively; Each device in the first device group and the second device group is electrically connected to the data exchange device, and any two devices communicate through the data transmission unit instead of through the first uplink switch, the second uplink switch and the CPU.

11. The data exchange architecture capable of achieving cache consistency according to claim 8, wherein: The DMA addresses do not overlap; and / or, There is no communication between the DMA addresses corresponding to different device IDs and the physical memory addresses of the devices; and / or, The device is a PCIe device.

12. A communication system, characterized in that: The communication system includes the data exchange architecture capable of achieving cache consistency as claimed in any one of claims 1 to 11.

13. A data exchange method capable of achieving cache consistency, characterized in that: The following steps are involved: A data exchange device is configured to be electrically connected to a plurality of devices to be communicated, each of the devices being configured with a unique device ID, the data exchange device comprising a shared storage space, a cache status array table, and a management module, the shared storage space comprising memory cells corresponding one-to-one to the device IDs, the cache status array comprising a cache status corresponding one-to-one to the device IDs, the cache status comprising exclusive write, shared read, data invalidation, and uncached; For each memory unit, the management module is configured to determine the access rights of each device to the memory unit according to the cache status array corresponding to the memory unit, including: When a cache state in the cache state array is exclusive write, the memory unit cannot be written / read by other devices except the current write device, and the device ID corresponding to the current write device is consistent with the device ID corresponding to the exclusive write; When there is no cache state of exclusive write in the cache state array, the memory unit can be written / read by any device; The management module is further configured to determine, based on the cache status array, access rights of other devices to data in a memory unit cached by a device, including: For a cache status array corresponding to a memory unit, other devices cannot cache data in the memory unit through a device corresponding to a cache status of data invalid / uncached; For a cache status array corresponding to a memory unit, other devices can cache data in the memory unit corresponding to the device for shared reading through the cache status.

14. The data exchange method capable of achieving cache consistency according to claim 13, characterized in that: Utilizing the management module to update the cache status array table according to the data exchange process includes: For the cache status array corresponding to each of the memory units, the management module updates the cache status array according to the current access status of the memory unit and the current cache status of each cache status in the cache status array.

15. The data exchange method capable of achieving cache consistency according to claim 14, characterized in that: For each of the memory units, the management module updates the cache status array in the following manner: If the memory unit is written with data by a requesting device, the management module modifies the cache state corresponding to the requesting device ID in the cache state array to exclusive write, and the requesting device is one of the devices; Update the cache status of other device IDs corresponding to shared read to invalid status; The cache status corresponding to other device IDs remains unchanged as uncached / data invalid.

16. The data exchange method capable of achieving cache consistency according to claim 14, characterized in that: For each of the memory units, the management module updates the cache status array in the following manner: If the requesting device completes writing data to the memory unit, the management module sends a data update notification to other devices; In response to receiving the data update notification, the other device may selectively cache the data in the memory unit; If the device caches the data in the memory unit, the management module updates the cache state corresponding to the device ID to data sharing; If the device does not cache the data in the memory unit, the management module maintains the cache state corresponding to the device ID unchanged.

Citation Information

Patent Citations

  • Multi-core processor supporting cache consistency, reading and writing methods and apparatuses as well as device

    CN105740164A

  • Implementation method and device for cache consistency of multi-core processor, the multi-core processor and storage medium

    CN112416615A

  • Method and system for achieving CACHE COHERENCY FOR HOST-DEVICE SYSTEMS

    CN113495854A

  • Multi-source heterogeneous distributed system, memory access method and storage medium

    CN117806553A

  • Data processing system, cache system and method for updating an invalid coherency state in response to snooping an operation

    US20070226427A1

Cited By

  • Access method, device and equipment and computer readable storage medium

    CN120950278A